Create virtual networks across multiple public clouds

By establishing a virtual network on the public cloud data center and optimizing data message routing using controller clusters and software forwarding components, the reliability and security issues of company network traffic in the interconnection of public cloud data centers are solved, and efficient and low-cost cross-cloud data center communication is achieved.

CN115051869BActive Publication Date: 2025-08-29VMWARE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210719316.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-04
Filing Date
2018-10-01
Publication Date
2025-08-29
Estimated Expiration
2038-10-01

AI Technical Summary

Technical Problem

Most of the company's network traffic is carried on expensive rental lines and consumer Internet, resulting in slowing communication speeds, and the existing technology cannot effectively optimize data center interconnections across public clouds, lacking reliability and security.

Method used

Establish a virtual network on several public cloud data centers of one or more public cloud providers, configure software forwarding components and intermediate box service machines through a logically centralized controller cluster, optimize data message routing, achieve end-to-end performance, reliability, and security, while reducing traffic routing through the Internet.

Benefits of technology

It realizes high-speed, reliable private network interconnection across public clouds, optimizes the routing of data messages, improves end-to-end performance and security, reduces Internet traffic, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115051869B_ABST
    Figure CN115051869B_ABST
Patent Text Reader

Abstract

The present disclosure relates to creating a virtual network across multiple public clouds. Some embodiments establish a virtual network for an entity on several public cloud providers and / or on several public clouds in several regions. In some embodiments, the virtual network is an overlay network across several public clouds to interconnect private networks (e.g., networks within the entity's branches, departments, divisions, or their associated data centers), mobile users, and SaaS (Software as a Service) provider machines, as well as one or more of the entity's other web applications. In some embodiments, the virtual network can be configured to optimize the routing of the entity's data messages to their destinations for optimal end-to-end performance, reliability, and security, while attempting to minimize the routing of such traffic through the Internet. Furthermore, in some embodiments, the virtual network can be configured to optimize layer 4 processing of data message flows through the network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of October 1, 2018, application number 201880058539.9, and name “Creating a virtual network across multiple public clouds”. Background Art

[0002] Today, a company's enterprise network is the communications backbone that securely connects a company's various offices and departments. This network is typically a wide area network (WAN) that connects (1) users in branch offices and regional campuses, (2) the company's data center hosting business applications, intranets, and their corresponding data, and (3) the global Internet through the company's firewall and DMZ (demilitarized zone). Enterprise networks consist of specialized hardware such as switches, routers, and middlebox devices interconnected by expensive leased lines such as Frame Relay and MPLS (Multi-Protocol Label Switching).

[0003] In recent years, there has been a paradigm shift in how companies access and consume communications services. First, the mobility revolution has enabled users to access services anytime, anywhere using mobile devices (primarily smartphones). These users access business services over the public internet and cellular networks. Simultaneously, third-party SaaS (Software as a Service) providers (e.g., Salesforce, Workday, Zendesk) have replaced traditional on-premises applications, while other applications hosted in private data centers have been relocated to the public cloud. While this traffic still travels within the corporate network, a significant portion originates and terminates outside the corporate network perimeter and must traverse both the public internet (once or twice) and the corporate network. Recent research indicates that 40% of corporate networks report that backhaul traffic (i.e., internet traffic observed within the corporate network) accounts for over 80%. This means that the majority of corporate traffic is carried over expensive leased lines and the consumer internet.

[0004] As a consumer-centric service, the internet is inherently a poor medium for business traffic. It lacks the reliability, QoS (Quality of Service) guarantees, and security expected by critical business applications. Furthermore, growing consumer traffic demands, net neutrality regulations, and internet-bypassing technologies created by major players (e.g., Netflix, Google, public clouds) have reduced the monetary returns per unit of traffic. These trends have reduced the incentive for service providers to quickly catch up to consumer demand and provide appropriate business services.

[0005] Given the growth of public cloud computing, companies are migrating more of their computing infrastructure to public cloud data centers. Public cloud providers have been at the forefront of computing and networking infrastructure investment. These cloud services have established numerous data centers worldwide, with Azure, AWS, IBM, and Google expanding to 38, 16, 25, and 14 global regions, respectively, in 2016. Each public cloud provider interconnects its data centers using expensive high-speed networks, including dark fiber and submarine cables deployed by submarines.

[0006] Today, despite these changes, corporate network policies often mandate that all corporate traffic traverse their secure WAN gateways. As users become more mobile and applications migrate to SaaS and public clouds, the cost of circumventing the corporate WAN becomes high, slowing down all corporate communications. Most corporate WAN traffic is either sourced from or destined for the internet. Alternative security solutions for routing this traffic over the internet are inadequate due to poor performance and unreliability. Summary of the Invention

[0007] Some embodiments establish a virtual network for an entity across public cloud data centers of one or more public cloud providers in one or more regions (e.g., cities, states, countries, etc.). Examples of entities for which such a virtual network may be established include commercial entities (e.g., companies), non-profit entities (e.g., hospitals, research institutions, etc.), and educational entities (e.g., universities, colleges, etc.), or any other type of entity. Examples of public cloud providers include Amazon Web Services (AWS), Google Cloud Platform (GCP), Microsoft Azure, etc.

[0008] In some embodiments, a high-speed, reliable private network interconnects two or more public cloud data centers (public clouds). Some embodiments define a virtual network as an overlay network that spans several public clouds to interconnect private networks (e.g., networks within an entity's branches, divisions, departments, or their associated data centers), mobile users, machines of SaaS (Software as a Service) providers, machines and / or services in one or more public clouds, and other web applications.

[0009] In some embodiments, the virtual network can be configured to optimize the routing of an entity's data messages to their destinations for optimal end-to-end performance, reliability, and security, while attempting to minimize the routing of such traffic through the Internet. Furthermore, in some embodiments, the virtual network can be configured to optimize layer 4 processing of data message flows through the network. For example, in some embodiments, the virtual network optimizes the end-to-end rate of a TCP (Transmission Control Protocol) connection by splitting the rate control mechanism across the connection path.

[0010] Some embodiments establish a virtual network by configuring several components deployed in several public clouds. In some embodiments, these components include software-based measurement agents, software forwarding elements (e.g., software routers, switches, gateways, etc.), layer 4 connection agents, and middlebox service machines (e.g., appliances, VMs, containers, etc.). In some embodiments, one or more of these components use standardized or commonly available solutions such as Open vSwitch, Open VPN, strong Swan, and Ryu.

[0011] Some embodiments utilize a logically centralized controller cluster (e.g., a collection of one or more controller servers) that configures public cloud components to implement virtual networks across several public clouds. In some embodiments, the controllers in this cluster are located in various locations (e.g., in different public cloud data centers) to improve redundancy and high availability. In some embodiments, the controller cluster scales up or down the number of public cloud components used to build a virtual network or the amount of computing or network resources allocated to these components.

[0012] Some embodiments establish different virtual networks for different entities on the same set of public clouds of the same public cloud provider and / or on different sets of public clouds of the same or different public cloud providers. In some embodiments, the virtual network provider provides software and services that allow different tenants to define different virtual networks on the same or different public clouds. In some embodiments, the same controller cluster or different controller clusters can be used to configure public cloud components to implement different virtual networks for several different entities on the same or different sets of public clouds.

[0013] To deploy a virtual network for a tenant on one or more public clouds, the controller cluster (1) identifies possible ingress and egress routers for entering and exiting the tenant's virtual network based on the location of the tenant's branch offices, data centers, mobile users, and SaaS providers, and (2) identifies routes that traverse from the identified ingress routers to the identified egress routers through other intermediate public cloud routers that implement the virtual network. After identifying these routes, the controller cluster propagates these routes to the forwarding tables of the virtual network routers in the public cloud(s). In an embodiment using an OVS-based virtual network router, the controller distributes the routes using OpenFlow.

[0014] The preceding summary is intended to serve as a brief introduction to some embodiments of the present invention. It is not intended to be an introduction or overview of all inventive subject matter disclosed in this document. The following detailed description and accompanying figures referenced in the detailed description further describe the embodiments described in the summary, as well as other embodiments. Therefore, a comprehensive review of the summary, detailed description, accompanying figures, and claims is required to understand all embodiments described in this document. Furthermore, the claimed subject matter is not limited by the illustrative details in the summary, detailed description, and accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The novel features of the invention are set forth in the appended claims.For purposes of illustration, however, several embodiments of the invention are set forth in the following drawings.

[0016] Figure 1A Given a virtual network defined for a company across several public cloud data centers of two public cloud providers.

[0017] Figure 1B The diagram shows an example of two virtual networks for two corporate tenants deployed on a public cloud.

[0018] Figure 1C An example of two virtual networks is alternatively illustrated, where one network is deployed on a public cloud and the other virtual network is deployed on another pair of public clouds.

[0019] Figure 2 An example of a managed cluster of forwarding nodes and controllers according to some embodiments of the present invention is illustrated.

[0020] Figure 3 An example of a measurement map generated by a controller measurement processing layer in some embodiments is illustrated.

[0021] Figure 4A An example of a routing map generated by a controller path identification layer from a measurement map in some embodiments is illustrated.

[0022] Figure 4B An example is illustrated in which known IPs of two SaaS providers are added to two nodes in a data center that are closest to the data centers of these SaaS providers in a routing map.

[0023] Figure 4C The routing graph generated by adding two nodes to represent two SaaS providers is shown.

[0024] Figure 4D A routing graph is illustrated with additional nodes added to represent branch offices and data centers with known IP addresses connected to two public clouds, respectively.

[0025] Figure 5 The diagram illustrates the process by which the controller path identification layer generates a routing map based on a measurement map received from the controller measurement layer.

[0026] Figure 6 The IPsec data message format of some embodiments is illustrated.

[0027] Figure 7 illustrates an example of two encapsulation headers of some embodiments, and Figure 8 An example is given illustrating how these two headers are used in some embodiments.

[0028] Figures 9-11 The diagram illustrates message handling processes performed by the ingress, intermediate, and egress MFNs, respectively, when the ingress, intermediate, and egress MFNs receive messages sent between two computing devices in two different branch offices.

[0029] Figure 12 An example is illustrated in which no intermediate MFN is involved between the ingress and egress MFNs.

[0030] Figure 13 The diagram illustrates the message handling process performed by the CFE of the ingress MFN when the ingress MFN receives a message sent from a corporate computing device in a branch office to another device in another branch office or a SaaS provider data center.

[0031] Figure 14 The diagram shows the NAT operation performed at the egress router.

[0032] Figure 15 Illustrated is the message handling process performed by an ingress router that receives a message sent from a SaaS provider machine to a tenant machine.

[0033] Figure 16 Such a TM engine is illustrated as being placed in each virtual network gateway on the virtual network's egress path to the Internet.

[0034] Figure 17 The double NAT method used in some embodiments is illustrated instead of Figure 16 The single NAT method shown in .

[0035] Figure 18 An example is provided illustrating source port translation for an ingress NAT engine.

[0036] Figure 19 The diagram shows the SaaS machine responding to its Figure 18 The processing of the data message and the processing of the reply message sent.

[0037] Figure 20 An example is given showing M virtual corporate WANs for M tenants of a virtual network provider having network infrastructure and controller cluster(s) in N public clouds of one or more public cloud providers.

[0038] Figure 21 The processing performed by a virtual network provider's controller cluster to deploy and manage a virtual WAN for a particular tenant is conceptually illustrated.

[0039] Figure 22 A computer system is conceptually illustrated with which some embodiments of the present invention are implemented. DETAILED DESCRIPTION

[0040] In the following detailed description of the present invention, many details, examples and embodiments of the present invention are set forth and described. However, it will be apparent to those skilled in the art that the present invention is not limited to the embodiments set forth, and that the present invention can be practiced without discussing some of these specific details and examples.

[0041] Some embodiments establish a virtual network for an entity across public cloud data centers of one or more public cloud providers in one or more regions (e.g., cities, states, countries, etc.). Examples of entities for which such a virtual network may be established include commercial entities (e.g., companies), non-profit entities (e.g., hospitals, research institutions, etc.), and educational entities (e.g., universities, colleges, etc.), or any other type of entity. Examples of public cloud providers include Amazon Web Services (AWS), Google Cloud Platform (GCP), Microsoft Azure, etc.

[0042] Some embodiments define a virtual network as an overlay network that spans several public cloud data centers (public clouds) to interconnect private networks (e.g., networks within an entity's branches, departments, divisions, or their associated data centers), mobile users, machines of SaaS (software as a service) providers, machines and / or services in (one or more) public clouds, and one or more other web applications. In some embodiments, a high-speed, reliable private network interconnects two or more of the public cloud data centers (public clouds).

[0043] In some embodiments, the virtual network can be configured to optimize the routing of an entity's data messages to their destinations for optimal end-to-end performance, reliability, and security, while attempting to minimize the routing of such traffic through the Internet. Furthermore, in some embodiments, the virtual network can be configured to optimize layer 4 processing of data message flows through the network. For example, in some embodiments, the virtual network optimizes the end-to-end rate of a TCP (Transmission Control Protocol) connection by splitting the rate control mechanism across the connection path.

[0044] Some embodiments establish a virtual network by configuring several components deployed in several public clouds. In some embodiments, these components include software-based measurement agents, software forwarding elements (e.g., software routers, switches, gateways, etc.), layer 4 connection agents, and middlebox service machines (e.g., appliances, VMs, containers, etc.).

[0045] Some embodiments utilize a logically centralized controller cluster (e.g., a collection of one or more controller servers) that configures public cloud components to implement virtual networks across several public clouds. In some embodiments, the controllers in this cluster are located in various locations (e.g., in different public cloud data centers) to improve redundancy and high availability. When different controllers in a controller cluster are located in different public cloud data centers, in some embodiments, the controllers share their state (e.g., configuration data they generate to identify tenants, routes through virtual networks, etc.). In some embodiments, the controller cluster scales up or down the number of public cloud components used to establish a virtual network or the amount of computing or network resources allocated to these components.

[0046] Some embodiments establish different virtual networks for different entities on the same set of public clouds of the same public cloud provider and / or on different sets of public clouds of the same or different public cloud providers. In some embodiments, the virtual network provider provides software and services that allow different tenants to define different virtual networks on the same or different public clouds. In some embodiments, the same controller cluster or different controller clusters can be used to configure public cloud components to implement different virtual networks for several different entities on the same or different sets of public clouds.

[0047] Several examples of corporate virtual networks are provided in the discussion below. However, those of ordinary skill in the art will recognize that some embodiments define virtual networks for other types of entities (such as other business entities, non-profit organizations, educational entities, etc.). Furthermore, as used in this document, a data message refers to a collection of bits of a particular format sent across a network. Those of ordinary skill in the art will recognize that the term "data message" is used in this document to refer to a collection of bits of various formats sent across a network. The format of these bits can be specified by a standardized protocol or a non-standardized protocol. Examples of data messages that follow standardized protocols include Ethernet frames, IP packets, TCP segments, UDP datagrams, etc. Furthermore, as used in this document, references to L2, L3, L4, and L7 layers (or Layer 2, Layer 3, Layer 4, or Layer 7) are references to the second data link layer, the third network layer, the fourth transport layer, and the seventh application layer, respectively, of the OSI (Open Systems Interconnection) layer model.

[0048] Figure 1A A virtual network 100 is shown, which is defined for a company across several public cloud data centers 105 and 110 of two public cloud providers, A and B. As shown, virtual network 100 is a secure overlay network established by deploying different managed forwarding nodes 150 in different public clouds and connecting the managed forwarding nodes (MFNs) to each other via overlay tunnels 152. In some embodiments, an MFN is a conceptual grouping of several different components in a public cloud data center that, together with other MFNs (and other groups of components) in other public cloud data centers, establishes one or more overlay virtual networks for one or more entities.

[0049] As further described below, in some embodiments, the set of components forming the MFN includes (1) one or more VPN gateways for establishing VPN connections with compute nodes that are entities at external machine locations outside the public cloud data center (e.g., an office, a private data center, a remote user, etc.), (2) one or more forwarding elements for forwarding encapsulated data messages between each other to define an overlay virtual network on the shared public cloud network fabric, (3) one or more service machines for performing middlebox service operations and L4-L7 optimization, and (4) one or more measurement agents for obtaining measurements regarding the quality of network connections between the public cloud data centers to identify desired paths through the public cloud data centers. In some embodiments, for reasons of redundancy and scalability, different MFNs may have different arrangements and different numbers of such components, and one MFN may have different numbers of such components.

[0050] Furthermore, in some embodiments, each MFN component group executes on a different computer in the MFN's public cloud data center. In some embodiments, some or all of the MFN components may execute on a single computer in the public cloud data center. In some embodiments, the MFN components execute on hosts that also execute other machines of other tenants. These other machines may be other machines of other MFNs of other tenants, or they may be unrelated machines (e.g., compute VMs or containers) of other tenants.

[0051] In some embodiments, the virtual network 100 is deployed by a virtual network provider (VNP), which deploys different virtual networks for different entities (e.g., different corporate customers / tenants of the virtual network provider) on the same or different public cloud data centers. In some embodiments, the virtual network provider is the entity that deploys the MFNs and provides a controller cluster for configuring and managing these MFNs.

[0052] Virtual network 100 connects corporate computing endpoints (such as data centers, branch offices, and mobile users) to each other and to external services (e.g., public web services or SaaS services such as Office 365 or Salesforce) residing in public clouds or in private data centers accessible via the Internet. This virtual network takes advantage of the different locations of different public clouds to connect different corporate computing endpoints (e.g., different private networks of a company and / or different mobile users) to their nearby public clouds. In the following discussion, corporate computing endpoints are also referred to as corporate computing nodes.

[0053] In some embodiments, virtual network 100 also leverages the high-speed networks that interconnect these public clouds to forward data messages through the public clouds to their destinations, or as close to their destinations as possible, while minimizing the amount of time they traverse the Internet. When corporate computing endpoints are outside the public cloud data centers spanned by the virtual network, these endpoints are referred to as external machine locations. This is the case for corporate branch offices, private data centers, and remote users' devices.

[0054] exist Figure 1AIn the example shown, virtual network 100 spans six data centers 105a-105f of public cloud provider A and four data centers 110a-110d of public cloud provider B. While spanning these public clouds, this virtual network connects several of the branch offices, corporate data centers, SaaS providers, and mobile users of corporate tenants located in different geographic regions. Specifically, virtual network 100 connects two branch offices 130a and 130b in two different cities (e.g., San Francisco, California and Pune, India), a corporate data center 134 in another city (e.g., Seattle, Washington), two SaaS provider data centers 136a and 136b in two other cities (Redmond, Washington and Paris, France), and mobile users 140 around the world. As such, this virtual network can be considered a virtual corporate WAN.

[0055] In some embodiments, branch offices 130a and 130b have their own private networks (e.g., local area networks) that connect computers at the branch locations to the branch's private data centers outside of the public cloud. Similarly, in some embodiments, corporate data center 134 has its own private network and resides outside of any public cloud data centers. However, in other embodiments, corporate data center 134 or the data centers of branches 130a and 130b may be within a public cloud, but the virtual network does not span this public cloud because the corporate or branch data centers are connected to the edge of virtual network 100.

[0056] As mentioned above, virtual network 100 is established by connecting managed forwarding nodes 150 deployed in different public clouds via overlay tunnels 152. Each managed forwarding node 150 includes several configurable components. As further described above and below, in some embodiments, the MFN components include software-based measurement agents, software forwarding elements (e.g., software routers, switches, gateways, etc.), layer 4 agents (e.g., TCP agents), and middlebox service machines (e.g., VMs, containers, etc.). In some embodiments, one or more of these components utilize standardized or commonly available solutions, such as Open vSwitch, OpenVPN, strongSwan, etc.

[0057] In some embodiments, each MFN (i.e., the conceptual group of components that form the MFN) can be shared by different tenants of a virtual network provider that deploys and configures the MFN in a public cloud data center. In conjunction with or alternatively, in some embodiments, a virtual network provider can deploy a unique set of MFNs for a specific tenant in one or more public cloud data centers. For example, for security or quality of service reasons, a particular tenant may not wish to share MFN resources with another tenant. For such tenants, the virtual network provider can deploy its own set of MFNs across several public cloud data centers.

[0058] In some embodiments, a logically centralized controller cluster 160 (e.g., a collection of one or more controller servers) operates within or outside of one or more of public clouds 105 and 110 and configures the public cloud components of the managed forwarding nodes 150 to implement virtual networks on public clouds 105 and 110. In some embodiments, the controllers in this cluster are located in various locations (e.g., in different public cloud data centers) to improve redundancy and high availability. In some embodiments, the controller cluster scales up or down the number of public cloud components used to establish a virtual network or the amount of computing or network resources allocated to these components.

[0059] In some embodiments, controller cluster 160 or another controller cluster of the virtual network provider establishes a different virtual network for another corporate tenant on the same public cloud 105 and 110 and / or on a different public cloud of a different public cloud provider. In addition to the controller cluster(s), in other embodiments, the virtual network provider also deploys forwarding elements and service machines in the public cloud, which allows different tenants to deploy different virtual networks on the same or different public clouds. Figure 1B An example of two virtual networks 100 and 180 for two corporate tenants deployed on public clouds 105 and 110 is illustrated. Figure 1C An example of two virtual networks 100 and 182 is alternatively illustrated, where one network 100 is deployed on public clouds 105 and 110 , and the other virtual network 182 is deployed on another pair of public clouds 110 and 115 .

[0060] By configuring components of MFN, Figure 1AThe virtual network 100 allows different private networks and / or different mobile users of a company tenant to connect to different public clouds that are optimally located (e.g., as measured in terms of physical distance, in terms of connection speed, loss, latency and / or cost, and / or in terms of network connection reliability, etc.) relative to these private networks and / or mobile users. In some embodiments, these components also allow the virtual network 100 to use the high-speed networks interconnecting the public clouds to forward data messages through the public clouds to their destinations while reducing the number of times they traverse the Internet.

[0061] In some embodiments, the MFN components are also configured to run novel processes at the network, transport, and application layers to optimize end-to-end performance, reliability, and security. In some embodiments, one or more of these processes implement proprietary, high-performance networking protocols without the rigidity of current network protocols. As such, in some embodiments, virtual network 100 is not constrained by the Internet's autonomous system, routing protocols, or even end-to-end transport mechanisms.

[0062] For example, in some embodiments, components of the MFN 150 (1) create optimized multi-path and adaptive centralized routing, (2) provide strong QoS (Quality of Service) guarantees, (3) optimize end-to-end TCP rates through intermediate TCP splitting and / or termination, and (4) relocate scalable application-level middlebox services (e.g., firewalls, intrusion detection systems (IDS), intrusion prevention systems (IPS), WAN optimization, etc.) to the computing portion of the cloud within a global network function virtualization (NFV). Thus, the virtual network can be optimized to adapt to the customized and ever-changing needs of a company without being tied to existing network protocols. Moreover, in some embodiments, the virtual network can be configured as a "pay-as-you-go" infrastructure that can dynamically and elastically expand and contract in both performance capabilities and geographic reach as demand continuously changes.

[0063] To implement the virtual network 100 , at least one managed forwarding node 150 in each public cloud data center 105 a - 105 f and 110 a - 110 d that the virtual network spans must be configured by a set of controllers. Figure 2The diagram illustrates examples of managed forwarding nodes 150 and controller clusters 160 according to some embodiments of the present invention. In some embodiments, each managed forwarding node 150 is a machine (e.g., a VM or container) executing on a host computer in a public cloud data center. In other embodiments, each managed forwarding node 150 is implemented by multiple machines (e.g., multiple VMs or containers) executing on the same host computer in a public cloud data center. In still other embodiments, two or more components of an MFN may be implemented by two or more machines executing on two or more host computers in one or more public cloud data centers.

[0064] As shown, the managed forwarding node 150 includes a measurement agent 205, firewall and NAT middlebox service engines 210 and 215, one or more optimization engines 220, edge gateways 225 and 230, and a cloud forwarding element 235 (e.g., a cloud router). In some embodiments, each of these components 205-235 can be implemented as a cluster of two or more components.

[0065] In some embodiments, the controller cluster 160 can dynamically scale each component cluster (1) to add or remove machines (e.g., VMs or containers) to implement the functionality of each component, and / or (2) to add or remove computing and / or network resources to previously deployed computers that implement the components of that cluster. As such, each deployed MFN 150 in a public cloud data center can be viewed as a cluster of MFNs, or it can be viewed as a node comprising multiple different component clusters that perform different operations of the MFN.

[0066] Furthermore, in some embodiments, the controller cluster deploys different sets of MFNs for different tenants in a public cloud data center, defining virtual networks for the tenants in the public cloud data center. In this approach, no two tenants' virtual networks share any MFNs. However, in the embodiments described below, each MFN can be used to implement different virtual networks for different tenants. Those skilled in the art will appreciate that in other embodiments, the controller cluster 160 can utilize its own dedicated set of deployed MFNs to implement the virtual network of each tenant in a first set of tenants, while simultaneously utilizing a shared set of deployed MFNs to implement the virtual network of each tenant in a second set of tenants.

[0067] In some embodiments, the branch gateway 225 and the remote device gateway 230 establish secure VPN connections with one or more branch offices 130 and remote devices (e.g., mobile devices 140) connected to the MFN 150, respectively. Figure 2. An example of such a VPN connection is an IPsec connection, which will be described further below. However, those skilled in the art will recognize that in other embodiments, such gateways 225 and / or 230 establish different types of VPN connections.

[0068] In some embodiments, the MFN 150 includes one or more middlebox engines that perform one or more middlebox service operations, such as firewall operations, NAT operations, IPS operations, IDS operations, load balancing operations, WAN optimization operations, etc. By incorporating these middlebox operations (e.g., firewall operations, WAN optimization operations, etc.) in the MFN deployed in the public cloud, the virtual network 100 implements many functions in the public cloud that are traditionally performed by the company's WAN infrastructure at the company's data center(s) and / or branch office(s).

[0069] Thus, for many middlebox services, corporate computing nodes (e.g., remote devices, branch offices, and data centers) no longer have to access the company's WAN infrastructure in private data centers or branches, as many of these services are now deployed in the public cloud. This approach speeds up access to these services by corporate computing nodes (e.g., remote devices, branch offices, and data centers) and avoids expensive, congested network bottlenecks at private data centers that were originally dedicated to providing such services.

[0070] This approach effectively distributes WAN gateway functionality to various MFNs within a public cloud data center. For example, in some embodiments of the virtual network 100, most or all traditional corporate WAN gateway security functions (e.g., firewall operations, intrusion detection operations, intrusion prevention operations, etc.) are moved to a public cloud MFN (e.g., an ingress MFN that receives data from computing endpoints into the virtual network). This effectively allows the virtual network 100 to have distributed WAN gateways implemented across the many different MFNs that implement the virtual network 100.

[0071] exist Figure 2In the illustrated example, the MFN 150 is shown as including a firewall engine 210, a NAT engine 215, and one or more L4-L7 optimization engines. Those of ordinary skill in the art will appreciate that in other embodiments, the MFN 150 includes other middlebox engines for performing other middlebox operations. In some embodiments, the firewall engine 210 enforces firewall rules on: (1) data message flows on their ingress path into the virtual network (e.g., data message flows received and processed by gateways 225 and 230 from branch offices 130 and mobile devices 140) and (2) data message flows on their egress path out of the virtual network (e.g., data message flows sent through the NAT engine 215 and the Internet 202 to a SaaS provider data center).

[0072] In some embodiments, firewall engine 210 of MFN 150 also enforces firewall rules when the firewall engine belongs to an MFN that is an intermediate hop between the ingress MFN where a data message flow enters the virtual network and the egress MFN where the data message flow exits the virtual network. In other embodiments, firewall engine 210 enforces firewall rules only when it is part of the ingress MFN and / or egress MFN for a data message flow.

[0073] In some embodiments, the NAT engine 215 performs network address translation to change the source network address of a data message stream as it exits the virtual network, passes through the Internet 202, and reaches a third-party device (e.g., a SaaS provider machine). This network address translation ensures that the third-party machine (e.g., a SaaS machine) can be correctly configured to handle the data message stream, which, without address translation, might specify the private network address of a tenant and / or public cloud provider. This is particularly problematic because the private network addresses of different tenants and / or cloud providers may overlap. Address translation also ensures that reply messages from the third-party device (e.g., a SaaS machine) can be correctly received by the virtual network (e.g., by the MFN NAT engine from which the message exited the virtual network).

[0074] In some embodiments, the NAT engine 215 of the MFN performs a double NAT operation on each data message flow that leaves the virtual network and reaches a third-party machine or enters the virtual network from a third-party machine. As further described below, one NAT operation is performed on the data message flow at its ingress MFN when the data message flow enters the virtual network, and the other NAT operation is performed on the data message flow at its egress MFN when the data message flow leaves the virtual network.

[0075] This double NAT approach allows more tenant private networks to be mapped to the public cloud provider's network. It also reduces the burden of distributing data regarding changes to tenant private networks to the MFN. Prior to ingress or egress NAT operations, some embodiments perform a tenant mapping operation that uses a tenant identifier to first map the tenant's source network address to another source network address, and then maps it to yet another source network address through a NAT operation. Performing double NAT operations reduces the burden of distributing data regarding changes to tenant private networks.

[0076] The optimization engine 220 performs novel processing that optimizes the forwarding of an entity's data messages to their destinations to achieve optimal end-to-end performance and reliability. Some of this processing implements proprietary high-performance networking protocols without the rigidity of current network protocols. For example, in some embodiments, the optimization engine 220 optimizes end-to-end TCP rates through intermediate TCP splitting and / or termination.

[0077] Cloud forwarding element 235 is an MFN engine that forwards data message flows to a cloud forwarding element (CFE) in the next-hop MFN when the data message flow must traverse another public cloud to reach its destination, or to an egress router in the same public cloud when the data message flow can reach its destination through the same public cloud. In some embodiments, CFE 235 of MFN 150 is a software router.

[0078] To forward data messages, the CFE encapsulates the message with a tunnel header. Different embodiments use different methods to encapsulate data messages with tunnel headers. When a data message must traverse one or more intermediate MFNs to reach the egress MFN, some embodiments described below use a tunnel header to identify the network ingress / egress addresses for entering and exiting the virtual network, and another tunnel header to identify the next-hop MFN.

[0079] Specifically, in some embodiments, the CFE sends a data message with two tunnel headers: (1) an inner header identifying the ingress and egress CFEs entering and exiting the virtual network, and (2) an outer header identifying the next-hop CFE. In some embodiments, the inner tunnel header also includes a tenant identifier (TID) to allow multiple different tenants of a virtual network provider to use a common set of MFN CFEs of the virtual network provider. Other embodiments define the tunnel header differently to define an overlay virtual network.

[0080] To deploy a virtual network for a tenant on one or more public clouds, the controller cluster (1) identifies possible ingress and egress routers for entering and exiting the tenant's virtual network based on the location of the tenant's corporate computing nodes (e.g., branch offices, data centers, mobile users, and SaaS providers), and (2) identifies routes that traverse from the identified ingress routers to the identified egress routers through other intermediate public cloud routers that implement the virtual network. After identifying these routes, the controller cluster propagates these routes to the forwarding tables of the MFN CFEs 235 in the public cloud(s). In an embodiment using an OVS-based virtual network router, the controller distributes the routes using OpenFlow.

[0081] In some embodiments, the controller cluster 160 can also be configured to implement the virtual network to optimize the components 205-235 of each MFN 150 at several network processing layers to achieve optimal end-to-end performance, reliability, and security. For example, in some embodiments, these components are configured to (1) optimize layer 3 traffic routing (e.g., shortest path, packet duplication), (2) optimize layer 4 TCP congestion control (e.g., segmentation, rate control), (3) implement security features (e.g., encryption, deep packet inspection, firewall), and (4) implement application layer compression features (e.g., deduplication, caching). Within the virtual network, corporate traffic is secured, inspected, and logged.

[0082] In some embodiments, a measurement agent is deployed for each MFN in a public cloud data center. In other embodiments, a measurement agent is shared by multiple MFNs in a public cloud data center or a collection of data centers (e.g., a collection of nearby, associated data centers, such as data centers in an availability zone). To optimize Layer 3 and Layer 4 processing, the measurement agent 205 associated with each managed forwarding node 150 repeatedly generates measurements that quantify the quality of the network connection between that node and each of several other "neighboring" nodes.

[0083] Different embodiments define neighboring nodes in different ways. For a particular MFN in one public cloud data center of a particular public cloud provider, in some embodiments, neighboring nodes include (1) any other MFN operating in any public cloud data center of the particular public cloud provider, and (2) any other MFN operating in a data center of another public cloud provider within the same "region" as the particular MFN.

[0084] Different embodiments define the same region in different ways. For example, some embodiments define regions by distance, which specifies the shape of the boundary around a particular managed forwarding node. Other embodiments define regions by city, state, or region, such as Northern California, Southern California, etc. This approach assumes that different data centers of the same public cloud provider are connected by very high-speed network connections, and that network connections between data centers of different public cloud providers may be fast when the data centers are located in the same region, but may not be as fast when the data centers are located in different regions. When the data centers are located in different regions, connections between data centers of different public cloud providers may have to traverse long distances over the public Internet.

[0085] In different embodiments, the measurement agent 205 generates the measurements in different ways. In some embodiments, the measurement agent periodically (e.g., once per second, every N seconds, every minute, every M minutes, etc.) sends a ping message (e.g., a UDP echo message) to each measurement agent of its neighboring managed forwarding nodes. Given the small size of ping messages, they do not incur significant network connection costs. For example, for 100 nodes, each node sending a ping to every other node every 10 seconds would generate approximately 10 kb / s of ingress and egress measurement traffic for each node, and given current public cloud prices, this would result in a few dollars in network consumption per node per year (e.g., $5).

[0086] Based on the rate at which it receives reply messages, measurement agent 205 calculates and updates measurement metrics (such as network connection throughput, latency, loss, and link reliability). By repeating these operations, measurement agent 205 defines and updates a measurement matrix that represents the quality of the network connections with its neighboring nodes. When agent 205 interacts with the measurement agents of its neighboring nodes, its measurement matrix only quantifies the quality of the connections with its local node clique.

[0087] The measurement agents of the different managed forwarding nodes send their measurement matrices to the controller cluster 160, which then aggregates all the different clique connection data to obtain an aggregated mesh view of the connections between different pairs of managed forwarding nodes. When the controller cluster 160 collects different measurements for the links between two pairs of forwarding nodes (e.g., measurements taken by a node at different times), the controller cluster generates a blended value based on the different measurements (e.g., an average or weighted average of the measurements). In some embodiments, the aggregated mesh view is a full mesh view of all network connections between each pair of managed forwarding nodes, while in other embodiments, it is a more complete view than the views generated by the measurement agents of the individual managed forwarding nodes.

[0088] like Figure 2 As shown in FIG, controller cluster 160 includes a cluster of one or more measurement processing engines 280, one or more path identification engines 282, and one or more management interfaces 284. To avoid obscuring the description with unnecessary detail, each of these clusters will be referred to below in terms of a single engine or interface layer (i.e., in terms of measurement processing layer 280, path identification layer 282, and management interface layer 284).

[0089] The measurement processing layer 280 receives measurement matrices from the measurement agents 205 of the managed forwarding nodes and processes these measurement matrices to generate an aggregated mesh matrix that describes the quality of connections between different pairs of managed forwarding nodes. The measurement processing layer 280 provides the aggregated mesh matrix to the path identification layer 282. Based on the aggregated mesh matrix, the path identification layer 282 identifies different desired routing paths through the virtual network for connecting different corporate data endpoints (e.g., different branch offices, corporate data centers, SaaS provider data centers, and / or remote devices). This layer 282 then provides these routing paths in a routing table, which is distributed to the cloud forwarding elements 235 of the managed forwarding nodes 150.

[0090] In some embodiments, the routing path identified for each pair of data message endpoints is the routing path that is considered optimal based on a set of optimization criteria, for example, whether it is the fastest routing path, the shortest routing path, or the path that uses the least amount of Internet. In other embodiments, the path identification engine may identify and provide (in a routing table) multiple different routing paths between the same two endpoints. In these embodiments, the cloud forwarding elements 235 of the managed forwarding nodes 150 then select one of the paths based on the QoS criteria or other runtime criteria they are enforcing. In some embodiments, each CFE 235 does not receive the entire routing path from the CFE to the egress point of the virtual network, but rather receives the next hop of the path.

[0091] In some embodiments, the path identification layer 282 uses the measurements in the aggregated grid matrix as input to the routing algorithm it executes to construct a global routing map. In some embodiments, this global routing map is an aggregated and optimized version of the measurement map produced by the measurement processing layer 280. Figure 3 An example of a measurement graph 300 generated by the controller measurement processing layer 280 in some embodiments is illustrated. This graph depicts the network connections between various managed forwarding nodes 150 in the AWS and GCP public clouds 310 and 320 (ie, in AWS and GCP's data centers). Figure 4A An example of a routing map 400 generated by the controller path identification layer 282 based on the measurement map 300 in some embodiments is illustrated.

[0092] Figure 5The diagram illustrates a process 500 by which the controller path identification layer generates a routing map based on a measurement map received from the controller measurement layer. As the path identification layer 282 repeatedly receives updated measurement maps from the controller measurement layer, the path identification layer 282 repeatedly performs this process 500 (e.g., each time a new measurement map is received or every Nth time a new measurement map is received). In other embodiments, the path identification layer 282 performs this process periodically (e.g., once every 12 hours or once every 24 hours).

[0093] As shown, the path identification layer initially defines (at 505) the routing graph as being identical to the measurement graph (i.e., having identical links between identical pairs of managed forwarding nodes). At 510, the process removes bad links from the measurement graph 300. Examples of bad links include links with excessive message loss or poor reliability (e.g., links with message loss greater than 2% in the last 15 minutes, or links with message loss greater than 10% in the last 2 minutes). Figure 4A The diagram illustrates that links 302, 304, and 306 in the measurement map 300 are excluded from the routing map 400. This diagram illustrates the exclusion of these links by depicting these links with dashed lines.

[0094] Next, at 515, process 500 calculates a link weight score (cost score) as a weighted combination of several calculated and provider-specific values. In some embodiments, the weight score is a weighted combination of the following for the link: (1) a calculated delay value, (2) a calculated loss value, (3) the provider network connection cost, and (4) the provider computational cost. In some embodiments, the provider's computational cost is taken into account because the managed forwarding nodes connected by the link are machines (e.g., VMs or containers) executing on host computers in (one or more) public cloud data centers.

[0095] At 520, the process adds known source and destination IP addresses for data message flows in the virtual network (e.g., known IP addresses of SaaS providers used by corporate entities) to the routing map. In some embodiments, the process adds each known IP address of a possible message flow endpoint to the node in the routing map closest to that endpoint (e.g., to the node representing the MFN). In doing so, in some embodiments, the process assumes that each such endpoint is connected to the virtual network via a link with zero delay cost and zero loss cost. Figure 4B The figure shows an example of adding two nodes 402 and 404 (representing two MFNs) of the known IP addresses of two SaaS providers to the routing map. The two nodes are located in data centers closest to the data centers of these SaaS providers. In this example, one node is in the AWS public cloud and the other node is in the GCP public cloud.

[0096] Alternatively, or in combination, in some embodiments, process 500 adds known source and destination IP addresses to the routing graph by adding nodes to the graph to represent source and destination endpoints, assigning IP addresses to these nodes, and assigning weight values ​​to connect these added nodes to other nodes in the routing graph (e.g., to nodes in the routing graph representing MFNs in a public cloud). When the source and destination endpoints of a flow are added as nodes, path identification engine 282 can consider the cost of reaching these nodes (e.g., distance cost, latency cost, and / or financial cost, etc.) when identifying different routes through virtual networks between different source and destination endpoints.

[0097] Figure 4C The diagram shows that by adding two nodes 412 and 414 to Figure 4A A routing graph 410 is generated by replacing the node graph 400 with a node graph 412 to represent two SaaS providers. In this example, known IP addresses are assigned to nodes 412 and 414, and these nodes are connected to nodes 402 and 404 (representing two MFNs) via links 416 and 418, which are assigned weights W1 and W2. This approach is used to add the known IP addresses of the two SaaS providers to the routing graph 410. Figure 4B An alternative to the method shown in .

[0098] Figure 4D A more detailed routing graph 415 is shown. In this more detailed routing graph, additional nodes 422 and 424 have been added to represent external corporate computing nodes (e.g., branch offices and data centers) with known IP addresses connected to AWS and GCP public clouds 310 and 320, respectively. Each of these nodes 422 / 424 is connected to at least one routing graph node representing an MFN in the routing graph nodes via at least one link 426 having an associated weight value Wi. Some of these nodes (e.g., some of the branch offices) are connected to the same MFN or to different MFNs via multiple links.

[0099] Next, at 525, process 500 calculates the lowest cost path (e.g., shortest path, etc.) between each MFN and each other MFN that can serve as a virtual network egress location for the corporate entity's data message flows. In some embodiments, egress MFNs include MFNs connected to external corporate computing nodes (e.g., branch offices, corporate data centers, and SaaS provider data centers), as well as MFNs that are candidate locations for mobile device connections and egress Internet connections. In some embodiments, this calculation utilizes a conventional lowest cost (e.g., shortest path) identification process that identifies the shortest path between different pairs of MFNs.

[0100] For each candidate MFN pair, when multiple such paths exist between the MFN pairs, the lowest cost identification process uses the calculated weight scores (i.e., the scores calculated at 510) to identify the path with the lowest score. Several approaches for calculating the lowest cost path are described further below. As mentioned above, in some embodiments, the path identification layer 282 identifies multiple paths between two MFN pairs. This is to allow the cloud forwarding element 235 to use different paths in different situations. Thus, in these embodiments, the process 500 can identify multiple paths between two MFN pairs.

[0101] At 530, the process removes from the routing map the links between the MFN pairs that are not used by any of the lowest-cost paths identified at 525. Next, at 535, the process generates routing tables for the cloud forwarding elements 235 based on the routing map. At 535, the process distributes these routing tables to the cloud forwarding elements 235 of the managed forwarding nodes. After 535, the process ends.

[0102] In some embodiments, a virtual network has two types of external connections, which are:

[0103] (1) external secure connections to an entity's computing nodes (e.g., a branch office, a data center, a mobile user, etc.), and (2) external connections to third-party computers (e.g., a SaaS provider server) over the Internet. Some embodiments optimize a virtual network by finding the optimal virtual network entry and exit locations for each data path that terminates at a source node and a destination node outside the virtual network. For example, to connect a branch office to a SaaS provider server (e.g., a salesforce.com server), some embodiments connect the branch office to an optimal edge MFN (e.g., an MFN with the fastest network connection to the branch office or a network connection closest to the branch office) and identify the optimal edge MFN to the optimally located SaaS provider server (e.g., the SaaS closest to the edge MFN for the branch office, or the SaaS with the fastest path to the edge MFN for the branch office through an edge MFN connected to the SaaS provider server).

[0104] To connect each computing node of an entity (e.g., a branch office, mobile user, etc.) to the nearest MFN via a VPN connection, in some embodiments, the virtual network provider deploys one or more authoritative domain name servers (DNS) in the public cloud for the computing nodes to contact. In some embodiments, each time a corporate computing node in some embodiments needs to establish a VPN connection to the virtual network provider's MFN (i.e., to initialize or reinitialize a VPN connection), the computing node first resolves the address associated with its virtual network (e.g., virtualnetworkX.net) with the authoritative DNS server to obtain the identity of the MFN that the server identifies as the MFN closest to the corporate computing node. To identify this MFN, in some embodiments, the authoritative DNS server provides an MFN identifier (e.g., the MFN's IP address). The corporate computing node then establishes a VPN connection to this managed forwarding node.

[0105] In other embodiments, the corporate computing node does not first perform a DNS resolution (i.e., does not first resolve the network address of a specific domain) each time it needs to establish a VPN connection to the MFN of the VPN. For example, in some embodiments, the corporate computing node persists using a DNS-resolved MFN for a specific duration (e.g., one day, one week, etc.) and then performs another DNS resolution to determine whether this MFN is still the optimal MFN to which it should connect.

[0106] When the source IP address in the DNS request is the source IP address of the local DNS server of the company's computing node rather than the source IP address of the node itself, in some embodiments, the authoritative DNS server identifies the MFN closest to the local DNS server rather than the MFN closest to the company's computing node. To address this issue, in some embodiments, the DNS request identifies the company's computing node by a domain name that includes one or more parts (labels) connected and delimited by dots, where one of the parts identifies the company and another part identifies the company's computing node.

[0107] In some embodiments, this domain name specifies a hierarchy of domains and subdomains in descending order from the right label to the left label in the domain name. The first label on the far right identifies the specific domain, the second label to the left of the first label identifies the corporate entity, and if the entity has more than one external machine location, the third label to the left of the second label identifies the external machine location of the entity. For example, in some embodiments, a DNS request identifies a corporate computing node as myNode of the company myCompany and requests resolution of the address myNode.myCompany.virtualnetwork.net. The DNS server then uses the myNode identifier to better select the entry MFN to which the corporate computing node should establish a VPN connection. In different embodiments, the myNode identifier is expressed in different ways. For example, it can be addressed as an IP address, a latitude / longitude description of a location, a GPS (Global Positioning System) location, a street address, etc.

[0108] Even if the IP address correctly reflects the location, there may be several potential entry routers, for example, belonging to different data centers in the same cloud or belonging to different clouds in the same region. In this case, in some embodiments, the virtual network authority server sends back a list of IP addresses of potential MFN CFEs (e.g., C5, C8, C12). Then, in some embodiments, the company computing node pings different CFEs in the list to generate measurements (e.g., distance or speed measurements) and selects the closest one by comparing the measurements between the set of CFE candidates.

[0109] Furthermore, the corporate compute node can establish this choice by identifying MFNs currently used by other compute nodes of the enterprise entity. For example, in some embodiments, the corporate compute node adds the connection cost to each MFN, so if many corporate branches are already connected to a given cloud, then the new compute node will have an incentive to connect to the same cloud, thereby minimizing inter-cloud costs in terms of processing, latency, and dollar cost.

[0110] Other embodiments use other DNS resolution technologies. For example, each time a corporate computing node (e.g., a branch office, a data center, a mobile user, etc.) needs to perform a DNS resolution, the corporate computing node (e.g., a mobile device or a local DNS resolver at a branch office or data center) communicates with a DNS service provider that acts as the authoritative DNS resolver for many entities. In some embodiments, this DNS service provider has a DNS resolver located in one or more private data centers, while in other embodiments, it is part of one or more public cloud data centers.

[0111] To identify which of the N managed forwarding nodes directly connected to the Internet should be used to reach the SaaS provider server, in some embodiments, the virtual network (e.g., the ingress MFN or the controller cluster that configures the MFN) identifies a set of one or more candidate edge MFNs from the N managed forwarding nodes. As further described below, in some embodiments, each candidate edge MFN is an edge MFN that is considered optimal based on a set of criteria (such as distance to the SaaS provider server, network connection speed, cost, latency and / or loss, network computation cost, etc.).

[0112] To help identify optimal edge points, the controller cluster of some embodiments maintains a list of the most popular SaaS provider and consumer web destinations and their IP address subnets for an entity. For each such destination, the controller cluster designates one or more of the optimal MFNs (again, based on physical distance, network connection speed, cost, loss and / or latency, computational cost, etc.) as candidate egress nodes. For each candidate egress MFN, the controller cluster then calculates the best route from each possible ingress MFN to that candidate MFN and sets the resulting next-hop table in the MFN accordingly, so that the Internet SaaS provider or web destination is associated with the correct virtual network next-hop node.

[0113] Given that a service destination can often be reached via multiple IP subnets located in multiple locations (e.g., provided by an authoritative DNS server), there are several potential egress nodes to minimize latency and provide load balancing. Therefore, in some embodiments, the controller cluster calculates the optimal location and egress node for each MFN and updates the next hop accordingly. Furthermore, the optimal egress node to a SaaS provider (e.g., office365.com) may be through one public cloud provider (e.g., Microsoft Azure), but the optimal ingress MFN, based solely on distance or connection speed, may be in another public cloud provider (e.g., AWS). In such cases, traversing another cloud (i.e., the public cloud with the best egress MFN) before leaving the virtual network may not be optimal in terms of latency, processing, and cost. In this case, providing multiple candidate edge nodes allows for the selection of the optimal edge MFN and the optimal path to that selected edge MFN.

[0114] To identify the optimal path through the virtual network to the egress MFN connected to the Internet or to the corporate compute node of the corporate entity, the controller cluster identifies the optimal routing path between the MFNs. As mentioned above, in some embodiments, the controller cluster identifies the optimal path between any two MFNs by first costing each link between a pair of directly connected MFNs, for example, based on a metric score reflecting a weighted sum of estimated latency and financial cost. In some embodiments, the latency and financial cost include (1) link delay measurements, (2) estimated message processing latency, (3) cloud fees for outgoing traffic from a particular data center or to another data center of the same public cloud provider or out of the public cloud provider's cloud (e.g., to another public cloud data center of another public cloud provider or to the Internet), and (4) estimated message processing costs associated with the MFNs executing on host computers in the public cloud.

[0115] Using the calculated costs for these paired links, the controller cluster can calculate the cost of each routing path using one or more of these paired links by aggregating the costs of the individual paired links used by the routing paths. As described above, the controller cluster then defines its routing graph based on the calculated routing path costs and generates a forwarding table for the MFN's cloud routers based on the defined routing graph. Furthermore, as mentioned above, the controller cluster repeats these cost calculation, graph construction, and forwarding table update and distribution operations periodically (e.g., every 12 hours, 24 hours, etc.) or upon receiving measurement updates from the MFN's measurement agent.

[0116] Whenever in MFN CFE C i The forwarding table at the address points to the next hop MFN CFE C j When CFE C i C j In some embodiments, CFE C i Established to CFE C j A secure, actively maintained VPN tunnel. In some embodiments, a secure tunnel is one that requires encryption of the payload of the encapsulated data message. Furthermore, in some embodiments, the tunnel is actively maintained by one or both endpoints of the tunnel sending a keep-alive signal to the other endpoint.

[0117] In other embodiments, the CFEs do not establish secure, actively maintained VPN tunnels. For example, in some embodiments, the tunnels between the CFEs are static tunnels that are not actively monitored by the transmission of keepalive signals. Furthermore, in some embodiments, these tunnels between the CFEs do not encrypt their payloads. In some embodiments, the tunnels between a pair of CFEs include two encapsulation headers, where an inner header identifies the tenant ID and the ingress and egress CFEs for data messages entering and leaving the virtual network (i.e., entering and leaving the public cloud), and an outer encapsulation header specifies the source and destination network addresses (e.g., IP addresses) of zero or more CFEs traversed from the ingress CFE to the egress CFE.

[0118] In addition to internal tunnels, in some embodiments, the virtual network also uses VPN tunnels to connect corporate computing nodes to its edge MFN, as mentioned above. Therefore, in embodiments where secure tunnels are used to connect CFEs, data messages are transported through the virtual network using a fully secure VPN path.

[0119] When using encapsulation within a virtual network to forward virtual network data messages, in some embodiments, the virtual network uses its own unique network address, which is different from the private addresses used by the tenant's different private networks. In other embodiments, the virtual network uses the private and public network address spaces of the public cloud on which it is defined. In still other embodiments, the virtual network uses its own unique network addresses for some components of the virtual network (e.g., its MFN, CFE, and / or some of its services), while using the private and public network address spaces of the public cloud for other components of the virtual network.

[0120] Furthermore, in some embodiments, the virtual network utilizes a clean-slate communication platform with its own proprietary protocol. In embodiments where data messages are forwarded entirely through software MFN routers (e.g., via software CFE), the virtual network can provide optimized rate control for long-distance end-to-end connections. In some embodiments, this is achieved by operating a TCP optimization proxy engine 220 at each MFN 150. In other embodiments that do not disrupt TCP itself (e.g., using HTTPS), this is achieved by the proxy engine 220 using intermediate per-flow buffering and TCP receiver window and ACK manipulation to partition rate control.

[0121] Due to its clean nature, in some embodiments, the virtual network optimizes many of its components to provide even better service. For example, in some embodiments, the virtual network uses multipath routing to support advanced bandwidth-guaranteed VPN setups for routing across virtual networks. In some embodiments, similar to ATM / MPLS routing, such VPNs include state data within each MFN, and their establishment and removal are centrally controlled. Some embodiments identify the available bandwidth of each outgoing link, either by direct measurement (via packet pairs or similar processing) or by having a given capacity for the link and reducing the traffic already sent over that link from this capacity.

[0122] Some embodiments use the remaining bandwidth of a link as a constraint. For example, when a link does not have at least 2 Mbps of available bandwidth, the controller cluster of some embodiments removes the link from the set of links used to calculate the least-cost path (e.g., shortest path) to any destination (e.g., removes the link from a routing map (such as map 400)). If an end-to-end route is still available after removing this link, then a new VPN will be routed across this new route. The removal of the VPN can restore the available capacity of a given link, which in turn can enable this link to be included in the least-cost path (e.g., shortest path) calculation. Some embodiments use other options for multipath routing, such as load balancing of traffic across multiple paths, for example, using MPTCP (Multipath TCP).

[0123] Some embodiments provide improved service for premium customers by leveraging path parallelism and inexpensive cloud links to replicate traffic from the ingress MFN to the egress MFN over two disjoint paths (e.g., maximally disjoint paths) within the virtual network. In this approach, the earliest arriving message is accepted, and later messages are discarded. This approach increases virtual network reliability and reduces latency at the expense of increased egress processing complexity. In some such embodiments, forward error correction (FEC) technology is used to increase reliability while reducing duplicate traffic. Due to its clean nature, the virtual network of some embodiments performs other upper-layer optimizations, such as application-layer optimizations (e.g., deduplication and caching) and security optimizations (e.g., the addition of encryption, DPI (deep packet inspection), and firewalls).

[0124] Some embodiments of the virtual network allow for collaboration with cloud providers to further improve virtual network setup by using anycast messaging. For example, in some embodiments, when all MFNs obtain the same external IP address, it is easier to connect any new corporate compute node to the optimal edge node (e.g., the closest edge node) using anycast connectivity. Similarly, any SaaS provider can obtain this IP address and connect to the optimal MFN (e.g., the closest MFN).

[0125] As mentioned above, different embodiments use different types of VPN connections to connect corporate computing nodes (e.g., branches and mobile devices) to the MFN that establishes a virtual network of corporate entities. Some embodiments use IPsec to set up these VPN connections. Figure 6 The IPsec data message formats of some embodiments are illustrated. Specifically, the figure illustrates the original format of a data message 605 generated by a machine at a corporate computing node, and an IPsec encapsulated data message 610 after the data message 605 has been encapsulated (e.g., at the corporate computing node or MFN) for transmission over an IPsec tunnel (e.g., to the MFN or to the corporate computing node).

[0126] In this example, an IPsec tunnel is set up using ESP tunnel mode (port 50). As shown in the figure, this mode is set up in this example by replacing the TCP protocol identifier in the IP header with the ESP protocol identifier. The ESP header identifies the beginning of message 615 (i.e., header 620 and payload 625). Message 615 must be authenticated by the recipient of the IPsec-encapsulated data message (e.g., by the MFN's IPsec gateway). The beginning of payload 625 is identified by the value of Next field 622 of message 615. Furthermore, payload 625 is encrypted. This payload includes the IP header, TCP header, and payload of the original data message 605, as well as padding field 630 including Next field 622.

[0127] In some embodiments, each MFN IPsec gateway can handle multiple IPsec connections for the same or different virtual network tenants (e.g., for the same company or for different companies). Thus, in some embodiments, the MFN IPsec gateway (e.g., gateway 230) identifies each IPsec connection by tunnel ID, tenant ID (TID), and company compute node subnet. In some embodiments, different company nodes (e.g., different branch offices) of a tenant do not have overlapping IP subnets (per RFC 1579). In some embodiments, the IPsec gateway maintains a table that maps each IPsec tunnel ID (contained in the IPsec tunnel header) to a tenant ID. For a given tenant that the IPsec gateway is configured to handle, the IPsec gateway also maintains a mapping of all of that tenant's subnets connected to the virtual network established by the MFN and its cloud forwarding elements.

[0128] When an ingress first MFN in a first public cloud data center receives a data message associated with a tenant ID and destined for a destination (e.g., a branch or data center subnet or SaaS provider) connected to an egress second MFN in a second public cloud data center via an IPsec tunnel, the first MFN's IPsec gateway removes the IPsec tunnel header. In some embodiments, the first MFN's CFE then encapsulates the message with two encapsulation headers that allow the message to traverse a path from the ingress first MFN to the egress second MFN, either directly or through one or more other intermediate MFNs. The first MFN's CFE identifies this path using a routing table configured by its controller.

[0129] As mentioned above, in some embodiments, the two encapsulation headers include (1) an outer header that specifies the next-hop MFN CFE to allow the encapsulated data message to traverse through the MFNs of the virtual network to reach the egress MFN CFE, and (2) an inner header that specifies the tenant ID and identifies the ingress and egress MFN CFEs of the MFNs through which the data message enters and leaves the virtual network.

[0130] Specifically, in some embodiments, the inner encapsulation header includes a valid IP header with the destination IP address of the CFE of the egress second MFN and the source IP address of the CFE of the ingress first MFN. This approach allows the use of standard IP router software in each CFE of the MFN. The encapsulation also includes a tenant ID (e.g., customer CID). When the message reaches the CFE of the egress second MFN, it is decapsulated and sent by the second MFN to its destination (e.g., by the second MFN's IPsec gateway via another IPsec tunnel associated with the message's tenant ID and destination subnet).

[0131] Some cloud providers prohibit machines from "spoofing" source IP addresses and / or impose other restrictions on TCP and UDP traffic. To address these potential restrictions, some embodiments use an outer header to connect adjacent pairs of MFNs used by one or more routers. In some embodiments, this header is a UDP header that specifies the source and destination IP addresses and UDP protocol parameters. In some embodiments, the ingress MFN CFE specifies its IP address as the source IP address in the outer header, while specifying the IP address of the next MFN CFE hop as the destination IP address in the outer header.

[0132] When the path to the egress MFN's CFE includes one or more intermediate MFN CFEs, the intermediate CFE replaces the source IP address in the outer header of the received double-encapsulated message with its own IP address. It also uses the destination IP address in the inner header to perform a route lookup in its routing table to identify the destination IP address of the next-hop MFN CFE on the path to the destination IP address in the inner header. The intermediate CFE then replaces the destination IP address in the outer header with the IP address it identified through its routing table lookup.

[0133] When a double-encapsulated data message arrives at the CFE of the egress MFN, the CFE determines that it is the egress node for the data message by retrieving the destination IP address in the inner header and determining that this destination IP address belongs to it. The CFE then removes both encapsulating headers from the data message and sends it to the destination (e.g., through the IPsec gateway of its MFN, via another IPsec tunnel associated with the tenant ID and destination IP address or subnet in the original header of the data message).

[0134] Figure 7 illustrates an example of two encapsulation headers of some embodiments, and Figure 8 An example is provided to illustrate how these two headers are used in some embodiments. In the following discussion, the inner header is referred to as the tenant header because it includes the tenant ID and the identity of the virtual network ingress / egress nodes connected to the tenant's corporate compute end nodes. The outer header is hereinafter referred to as the VN-hop tunnel header because it is used to identify the next hop through the virtual network as the data message traverses the path through the virtual network between the ingress and egress MFN CFEs.

[0135] Figure 7 A VN-hop tunnel header 705 and a tenant tunnel header 720 are shown, which encapsulate an original data message 750 having an original header 755 and a payload 760. As shown, in some embodiments, the VN-hop tunnel header 705 includes a UDP header 710 and an IP header 715. In some embodiments, the UDP header is defined according to the UDP protocol. In some embodiments, the VN-hop tunnel is a standard UDP tunnel, while in other embodiments, this tunnel is a proprietary UDP tunnel. In still other embodiments, this tunnel is a standard or proprietary TCP tunnel. In some embodiments, the tunnel header 705 is an encrypted tunnel header that encrypts its payload, while in other embodiments, it is an unencrypted tunnel.

[0136] As described further below, in some embodiments, a tunnel header 705 is used to define an overlay VPN network and is used by each MFN CFE to reach the next-hop MFN CFE on the underlying public cloud network. Accordingly, the IP header 715 of the tunnel header 705 identifies the source and destination IP addresses of the first and second CFEs of the first and second adjacent MFNs connected by the VPN tunnel. In some cases (e.g., when the next-hop destination MFN is in a different public cloud than the source MFN, which is a different public cloud provider), the source and destination IP addresses are public IP addresses used by the public cloud data center containing the MFNs. In other cases, when the source and destination MFN CFEs belong to the same public cloud, the source and destination IP addresses may be private IP addresses used only in the public cloud. Alternatively, in such cases, the source and destination IP addresses may still be public IP addresses of the public cloud provider.

[0137] like Figure 7 As shown in FIG, tenant tunnel header 720 includes an IP header 725, a tenant ID field 730, and a virtual circuit label (VCL) 735. Tenant tunnel header 720 is used by each CFE after the ingress CFE to identify the next hop for forwarding the data message to the egress CFE of the egress MFN. Accordingly, IP header 725 includes a source IP address, which is the IP address of the ingress CFE, and a destination IP address, which is the IP address of the egress CFE. Like the source IP address and destination IP address of VN-hop header 705, the source IP address and destination IP address of tenant header 720 can be either the private IP address of one public cloud provider (when the data message traverses a route through only one public cloud provider's data center) or the public IP addresses of one or more public cloud providers (for example, when the data message traverses a route through two or more public cloud providers' data centers).

[0138] In some embodiments, the IP header of the tenant header 720 can be routed using any standard software router and IP routing table. The tenant ID field 730 contains the tenant ID, which is a unique tenant identifier that can be used at the ingress and egress MFNs to uniquely identify the tenant. In some embodiments, the virtual network provider defines different tenant IDs for different corporate entities that serve as tenants of the provider. The VCL field 735 is an optional routing field that some embodiments use to provide alternative (non-IP-based) methods for forwarding messages across the network. In some embodiments, the tenant tunnel header 720 is a GUE (Generic UDP Encapsulation) header.

[0139] Figure 8An example is provided to illustrate how these two tunnel headers 705 and 710 may be used in some embodiments. In this example, a data message 800 is sent from a first machine 802 (e.g., a first VM) in a company's first branch office 805 to a second machine 804 (e.g., a second VM) in the company's second branch office 810. The two machines are located in two different subnets, 10.1.0.0 and 10.2.0.0, respectively. The first machine has an IP address of 10.1.0.17, while the second machine has an IP address of 10.2.0.22. In this example, the first branch office 805 is connected to an ingress MFN 850 in a first public cloud data center 830, while the second branch office 810 is connected to an egress MFN 855 in a second public cloud data center 838. Furthermore, in this example, the ingress and egress MFNs 850 and 855 of the first and second public cloud data centers are indirectly connected via an intermediate MFN 857 in a third public cloud data center 836.

[0140] As shown, data message 800 from machine 802 is sent to ingress MFN 850 along IPsec tunnel 870, which connects first branch office 805 to ingress MFN 850. This IPsec tunnel 870 is established between IPsec gateway 848 at the first branch office and IPsec gateway 852 at ingress MFN 850. This tunnel is established by encapsulating data message 800 with IPsec tunnel header 806.

[0141] The IPsec gateway 852 of the MFN 850 decapsulates the data message (i.e., removes the IPsec tunnel header 806) and forwards the decapsulated message directly or through one or more middlebox service machines (e.g., through a firewall machine such as Figure 2 The message is passed to the CFE 832 of the MFN (machine 210). When passing this message, in some embodiments, the IPsec gateway or some other module of the MFN 850 associates the message with the tunnel ID of the IPsec tunnel and the company's tenant ID. This tenant ID identifies the company in the virtual network provider's records.

[0142] Based on the associated tenant ID and / or IPsec tunnel ID, CFE 832 of ingress MFN 850 identifies a route for the message to reach its destination computer's subnet (i.e., to second branch office 810) through the virtual network, established by a MFN in a different public cloud data center. For example, CFE 832 uses the tenant ID and / or IPsec tunnel ID to identify the company's routing table. Within this routing table, CFE 832 then uses the received message's destination IP address, 10.2.022, to identify a record that identifies CFE 853 of egress MFN 855 in public cloud data center 838 as the destination egress forwarding node for data message 800. In some embodiments, the identified record maps the entire subnet 10.2.0.0 / 16 of second branch office 810 to CFE 853 of MFN 855.

[0143] After identifying the egress CFE 853, the CFE 832 of the ingress MFN 850 encapsulates the received data message with a tenant tunnel header 860. This tenant tunnel header 860 includes the source IP address of the ingress CFE 832 and the destination IP address of the egress CFE 853 in its IP header 725. In some embodiments, these IP addresses are defined in the public IP address space. Tunnel header 860 also includes the tenant ID associated with the data message at the ingress MFN 850. As mentioned above, in some embodiments, this tunnel header also includes the VCL header value.

[0144] In some embodiments, ingress CFE 832 also identifies the next-hop MFN on the desired CFE routing path to egress CFE 853. In some embodiments, ingress CFE 832 identifies this next-hop CFE in its routing table using the destination IP address of egress CFE 853. In this example, the next-hop MFN CFE is CFE 856 of third MFN 857 of third public cloud data center 836.

[0145] After identifying the next-hop MFN CFE, the ingress MFN CFE encapsulates the encapsulated data message 800 with a VN-hop, second tunnel header 862. This tunnel header allows the message to be routed to the next-hop CFE 856. In the IP header 715 of this outer header 862, the ingress MFN CFE 832 specifies the source and destination IP addresses as the source IP of the ingress CFE 832 and the destination IP of the intermediate CFE 856. In some embodiments, it also specifies its Layer 4 protocol as UDP.

[0146] When CFE 856 of third MFN 857 receives the double-encapsulated data message, it removes the VN-hop and second tunnel header 862 and extracts the destination IP address of CFE 853 of egress MFN 855 from tenant header 860. Since this IP address is not associated with CFE 856, the data message must still traverse another MFN to reach its destination. Therefore, CFE 856 uses the extracted destination IP address to identify the record in its routing table that identifies the next-hop MFN CFE 853. It then re-encapsulates the data message using the outer header 705 and specifies the source and destination IP addresses as its own IP address and the destination IP address of MFN CFE 853 in its IP header 715. CFE 856 then forwards the double-encapsulated data message 800 to egress CFE 853 via the intervening routing infrastructure of public cloud data centers 836 and 838.

[0147] After receiving the encapsulated data message, egress CFE 853 determines that the encapsulated message is intended for it when it retrieves the destination IP address in inner header 860 and determines that this destination IP address belongs to it. Egress CFE 853 removes both encapsulating headers 860 and 862 from data message 800 and extracts the destination IP address from the original header of the data message. This destination IP address identifies the IP address of second machine 804 in the subnet of the second branch office.

[0148] Using the tenant ID in the removed tenant tunnel header 860, egress CFE 853 identifies the correct routing table to search, then searches this routing table based on the destination IP address extracted from the original header value of the received data message. Based on this search, egress CFE 853 identifies the record that identifies the IPsec connection used to forward the data message to its destination. It then provides the data message and the IPsec connection identifier to the second MFN's IPsec gateway 858, which in turn encapsulates the message with an IPsec tunnel header 859 and forwards it to the IPsec gateway 854 at the second branch office 810. Gateway 854 then removes the IPsec tunnel header and forwards the data message to its destination machine 804.

[0149] Now refer to Figures 9-15 Several more detailed message processing examples are described. In these examples, it is assumed that each tenant IPsec interface is on the same local public IP address, just like a VPN tunnel. Accordingly, in some embodiments, the interfaces are attached to a single VRF (virtual routing and forwarding) namespace. This VRF namespace is hereinafter referred to as the VPN namespace.

[0150] Figures 9-11 Message handling processes 900-1100 are illustrated, which are performed by the ingress, intermediate, and egress MFNs, respectively, when the ingress, intermediate, and egress MFNs receive messages sent between two computing devices in two different external machine locations (e.g., branch offices, data centers, etc.) of a tenant. In some embodiments, the controller cluster 160 configures the CFE of each MFN to operate as an ingress, intermediate, or egress CFE when each such CFE is a candidate for use as an ingress, intermediate, or egress CFE for different data message flows of the tenant.

[0151] The following will refer to Figure 8 and Figure 12 The two examples in the following are used to explain the processing 900-1100. As mentioned above, Figure 8 The diagram illustrates an example when a data message passes through an intermediate MFN to reach an egress MFN. Figure 12 The diagram shows an example where no intermediate MFN is involved between the inlet and outlet MFNs. Specifically, Figure 12 The diagram illustrates a data message 1200 being sent from a first device 1202 in a first branch office 1205 to a second device 1210 in a second branch office 1220 when the two branch offices are connected to two public cloud data centers 1230 and 1238 using two directly connected MFNs 1250 and 1255. As shown, in these examples, the CFEs 1232 and 1253 of the MFNs perform routing operations associated with each MFN.

[0152] In some embodiments, the ingress CFE (e.g., ingress CFE 832 or 1232) of the ingress MFN 850 and 1250 performs process 900. Figure 9 As shown in FIG, ingress processing 900 begins by initially identifying (at 905) a tenant routing context based on an identifier of an IPsec tunnel (e.g., 806 or 1206) in a received data message. In some embodiments, an IPsec gateway or other MFN module stores a tenant ID for an IPsec tunnel ID in a mapping table. Whenever a data message is received along a particular IPsec tunnel, the IPsec gateway extracts the IPsec tunnel ID, which the gateway or other MFN module then uses to identify the associated tenant ID by referencing its mapping table. By identifying the tenant ID, the process identifies the tenant portion of the tenant routing table or VRF namespace for use.

[0153] At 910, the process increments the RX (receive) counter of the identified IPsec tunnel to account for the receipt of this data message. Next, at 915, the process performs a routing lookup (e.g., a longest prefix match (LPM) lookup) in the identified tenant routing context (e.g., in the tenant's portion of the VRF namespace) to identify the IP address of the egress interface for exiting the tenant's virtual network built on the public cloud data center. For the spoke-to-spoke example, the egress interface is the IP address of the egress CFE (e.g., CFE 853 or 1253) connected to the MFN of the destination spoke.

[0154] At 920, the process adds a tenant tunnel header (e.g., header 860 or 1260) to the received data message and embeds the source IP address of the ingress CFE (e.g., ingress CFE 832 or 1252) and the destination IP address of the egress CFE (e.g., egress CFE 853 or 1253) as the source and destination IP addresses in this tunnel header. In the tenant header, the process also stores the tenant ID (identified at 905) in the tenant header. At 920, the process adds a VN-hop tunnel header (e.g., header 862 or 1262) in addition to the tenant header and stores its IP address as the source IP address in this header. The process also specifies (at 920) UDP parameters (e.g., UDP port) for the VPN tunnel header.

[0155] Next, at 925, the process increments the tenant's VN transmission counter to account for the transmission of the data message. At 930, the process performs a routing lookup (e.g., an LPM lookup) in the identified VNP routing context (e.g., in the VNP's portion of the VRF namespace) to identify the next-hop interface for the data message. In some embodiments, the routing lookup is based at least in part on an LPM lookup of the destination IP of the egress CFE (e.g., in the VNP's portion of the VRF namespace).

[0156] At 935, the process determines whether the next-hop egress interface is a local interface (e.g., a physical or virtual port) of the ingress CFE. If so, the process defines (at 937) the destination IP address in the VN-hop outer tunnel header as the egress interface IP address identified at 915. Next, at 940, the process presents the double-encapsulated data message to its local interface so that it can be forwarded to the destination egress CFE. After 940, process 900 ends.

[0157] Figure 12The diagram illustrates an example of operations 905-940 for a data message 1200 received by the ingress CFE 1232 from the device 1202 of the first branch office 1205. As shown, the CFE's MFN 1250 receives this data message as an IPsec-encapsulated message at its IPsec gateway 1252 from the first branch office's IPsec gateway 1248. The ingress CFE 1232 encapsulates the received message 1200 (after its IPsec header has been removed by the IPsec gateway 1252) with a VN-hop tunnel header 1262 and a tenant tunnel header 1260, and forwards this dual-encapsulated message to the egress CFE 1253 of the MFN 1255 of the public cloud 1238. As shown, in this example, the source and destination IP addresses of both tunnel headers 1260 and 1262 are identical. Given that these two sets of IP addresses are identical, some embodiments forgo the use of outer IP header 1262 when the data message is not routed through any intermediate CFEs (such as CFE 856).

[0158] When the processing determines (at 935) that the next-hop egress interface is not the local interface of the ingress CFE but is the destination IP address of another router, the processing embeds (at 945) the destination IP address of the next-hop intermediate CFE (e.g., intermediate CFE 856) in the VN-hop tunnel header as the destination IP address of the VN-hop tunnel header.

[0159] Next, at 950, the process performs another routing lookup (e.g., an LPM lookup) within the identified VNP routing context (e.g., within the VNP's portion of the VRF namespace). This time, the lookup is based on the IP address of the intermediate CFE identified in the VNP tunnel header. Because the intermediate CFE (e.g., CFE 856) is the next-hop CFE in the virtual network for the ingress CFE (e.g., CFE 832), the routing table identifies the local interface (e.g., local port) for the data message sent to the intermediate CFE. Thus, this lookup within the VNP routing context identifies the local interface to which the ingress CFE provided (at 950) the double-encapsulated message. The process then increments (at 955) the VN intermediate counter to account for the transmission of this data message. After 955, the process ends.

[0160] Figure 10The diagram illustrates a process 1000 performed by a CFE (e.g., CFE 853 or 1253) of an egress MFN in some embodiments upon receiving a data message that should be forwarded to a corporate computing node (e.g., a branch office, a data center, a remote user location) connected to the MFN. As shown, the process initially receives (at 1005) a data message on an interface associated with a virtual network. This message is encapsulated with a VN-hop tunnel header (e.g., header 862 or 1262) and a tenant tunnel header (e.g., header 860 or 1260).

[0161] At 1010, the process determines that the destination IP address in the VN-hop tunnel header is the destination IP address of its CFE (e.g., the IP address of CFE 853 or 1253). Next, at 1015, the process removes both tunnel headers. The process then retrieves (at 1020) the tenant ID from the removed tenant tunnel header. To account for received data messages, the CFE then increments (at 1025) the RX (receive) counter it maintains for the tenant specified by the extracted tenant ID.

[0162] Next, at 1030, the process performs a routing lookup (e.g., an LPM lookup) in the identified tenant routing context (i.e., the routing context of the tenant identified by the tenant ID extracted at 1020) to identify the next-hop interface for the data message. In some embodiments, the process performs this lookup based on the destination IP address in the original header of the received data message (e.g., header 755). From the record identified by this lookup, process 1000 identifies the IPsec interface through which the data message must be sent to its destination. Process 1000 then sends the decapsulated, received data message to the IPsec gateway of its MFN (e.g., gateway 858 or 1258).

[0163] This gateway then encapsulates the data message with an IPsec tunnel header (e.g., tunnel header 859 or 1259) and forwards it to a gateway (e.g., gateway 854 or 1254) in the destination corporate computing node (e.g., the destination branch office), where it is decapsulated and forwarded to its destination. After 1030, the CFE or its MFN increments (at 1035) a counter it maintains for transmitting messages along the IPsec connection to the destination corporate computing node (e.g., the IPsec connection between gateways 854 and 858 or between gateways 1254 and 1258).

[0164] Figure 11The diagram illustrates a process 1100 performed by a CFE (e.g., CFE 856) of an intermediate MFN upon receiving a data message that should be forwarded to another CFE of another MFN in some embodiments. As shown, the process initially receives (at 1105) a data message on an interface associated with a virtual network. In some embodiments, this message is encapsulated with two tunnel headers: a VN tunnel header (e.g., header 862) and a tenant tunnel header (e.g., header 860).

[0165] At 1110, when the process determines that the destination IP address in this tunnel header is the destination IP address of its CFE (e.g., the destination IP address of CFE 856), the process terminates the VN-hop tunnel. Next, at 1115, the process determines whether the VN-hop tunnel header specifies the correct UDP port. If not, then the process ends. Otherwise, at 1120, the process removes the VN-hop tunnel header. To account for received data messages, the CFE then increments (at 1125) the RX (receive) counter it maintains to quantify the number of messages it has received as an intermediate hop CFE.

[0166] At 1130, the process performs a routing lookup (e.g., an LPM lookup) in the identified VNP routing context (e.g., in the VNP's portion of the VRF namespace) to identify the next-hop interface for the data message. In some embodiments, this routing lookup is an LPM lookup based at least in part on the destination IP of the egress CFE identified in the inner tenant tunnel header (e.g., in the VNP's portion of the VRF namespace).

[0167] The process then determines (at 1135) whether the next-hop egress interface is the local interface of the intermediate CFE. If so, the process adds (at 1140) a VN-hop tunnel header to the data message, which has been encapsulated with the tenant tunnel header. The process sets (at 1142) the destination IP address in the VN-hop tunnel header to the destination IP address of the egress CFE specified in the tenant tunnel header. It also sets (at 1142) the source IP address in the VN-hop tunnel header to the IP address of its CFE. In this tunnel header, the process also sets UDP attributes (e.g., UDP port, etc.).

[0168] Next, at 1144, the process provides the double encapsulated data message to its local interface (identified at 1130) so that it can be forwarded to the destination egress CFE. Figure 8The operation of CFE 856 in FIG1 illustrates an example of such VN-hop tunnel decapsulation and forwarding. To account for received data messages, the CFE then increments (at 1146) a TX (transmit) counter it maintains to quantify the number of messages it has transmitted as an intermediate hop CFE. After 1146, process 1100 ends.

[0169] On the other hand, when the process determines (at 1135) that the next-hop egress interface is not the local interface of its CFE but the destination IP address of another router, the process adds (at 1150) a VN-hop tunnel header to the data message from which it previously removed the VN-hop tunnel header. In the new VN-hop tunnel header, the process 1100 embeds (at 1150) the source IP address of its CFE and the destination IP address of the next-hop intermediate CFE (identified at 1130) as the source IP address and destination IP address of the VN-hop tunnel header. This VN-hop tunnel header also specifies the UDP layer 4 protocol with a UDP destination port.

[0170] Next, at 1155, the process performs another routing lookup (e.g., an LPM lookup) within the identified VNP routing context (e.g., within the VNP's portion of the VRF namespace). This time, the lookup is based on the IP address of the next-hop intermediate CFE identified in the new VN-hop tunnel header. Because this intermediate CFE is the next hop from the current intermediate CFE in the virtual network, the routing table identifies the local interface for data messages sent to the next-hop intermediate CFE. Thus, this lookup within the VNP routing context identifies the local interface to which the current intermediate CFE provides the dual-encapsulated message. The process then increments (at 1160) the VN intermediate TX (transmit) counter to account for the transmission of this data message. After 1160, the process ends.

[0171] Figure 13 Illustrated is a message handling process 1300 performed by the CFE of the ingress MFN when the ingress MFN receives a message intended for a tenant sent from a tenant's corporate computing device (e.g., in a branch office) to another tenant machine (e.g., in another branch office, a tenant data center, or a SaaS provider data center). As further described below, Figure 9 The process 900 is a subset of the process 1300. Figure 13 As shown in , process 1300 begins by initially identifying (at 905) a tenant routing context based on an identifier of an incoming IPsec tunnel.

[0172] At 1310, the process determines whether both the source IP address and the destination IP address in the header of the received data message are public IP addresses. If so, the process (at 1315) discards the data message and increments the discard counter for the IPsec tunnel it maintains for the received data message. At 1315, the process discards the counter because when it receives messages through the tenant's IPsec tunnel, it should not receive messages addressed to and from public IP addresses. In some embodiments, process 1300 also sends an ICMP error message back to the source company computer.

[0173] On the other hand, when the process determines (at 1310) that the data message is not from a public IP address and is destined for another public IP address, the process determines (at 1320) whether the destination IP address in the header of the received data message is a public IP address. If so, then the process transfers to 1325 to perform Figure 9 The process 1300 is as described above, except for operation 905, which was already performed at the beginning of the process 1300. After 1325, the process 1300 ends. On the other hand, when the process 1300 determines (at 1320) that the destination IP address in the header of the received data message is not a public IP address, the process increments (at 1330) the RX (receive) counter of the identified IPsec tunnel to account for the reception of this data message.

[0174] The process 1300 then performs (at 1335) a routing lookup (e.g., an LPM lookup) in the identified tenant routing context (e.g., in the tenant's portion of the VRF namespace). This lookup identifies the IP address of the egress interface for exiting the tenant's virtual network built on top of the public cloud data center. Figure 13 In the example shown, when the data message is intended for a machine in the SaaS provider data center, process 1300 reaches lookup operation 1335. Thus, this lookup identifies the IP address of the egress router used to exit the tenant's virtual network to reach the SaaS provider machine. In some embodiments, all SaaS provider routes are installed in one routing table or one portion of the VRF namespace, while in other embodiments, routes for different SaaS providers are stored in different routing tables or different portions of the VRF namespace.

[0175] At 1340, the process adds a tenant tunnel header to the received data message, embedding the ingress CFE's source IP address and the egress router's destination IP address as the source and destination IP addresses in the tunnel header. Next, at 1345, the process increments the tenant's VN-transmit counter to account for the transmission of the data message. At 1350, the process performs a routing lookup (e.g., an LPM lookup) within the VNP routing context (e.g., within the VNP's portion of the VRF namespace) to identify one of its local interfaces as the next-hop interface for the data message. When the next-hop is another CFE (e.g., in another public cloud data center), in some embodiments, the process also encapsulates the data message with a VN-hop header and embeds the IP address of its CFE and the IP address of the other CFE as the source and destination addresses in the VN-hop header. At 1355, the process presents the encapsulated data message to the identified local interface so that the data message can be forwarded to its egress router. After 1355, process 1300 ends.

[0176] In some cases, an ingress MFN may receive a data message for a tenant whose CFE can forward the message directly to the destination machine without passing through a CFE of another MFN. In some such cases, when the CFE does not need to relay any tenant-specific information to any subsequent VN processing module or the required information can be provided to the subsequent VN processing module through other mechanisms, it is not necessary to encapsulate the data message with a tenant header or a VN-hop header.

[0177] For example, to forward a tenant's data messages directly to an external SaaS provider data center, the ingress MFN's NAT engine 215 would have to perform NAT operations based on the tenant identifier, as described further below. The ingress CFE or another module in the ingress MFN would have to provide the tenant identifier to the associated NAT engine 215 of the ingress MFN. When the ingress CFE and NAT engine execute on the same computer, some embodiments share this information between the two modules by storing it in a shared memory location. On the other hand, when the CFE and NAT engine do not execute on the same computer, some embodiments use other mechanisms (e.g., out-of-band communication) to share the tenant ID between the ingress CFE and NAT engine. However, in such cases, other embodiments use encapsulation headers (i.e., use in-band communication) to store and share the tenant ID between different modules of the ingress MFN.

[0178] As described further below, some embodiments perform one or two source NAT operations on the source IP / port address of a data message before sending the data message outside the tenant's virtual network. Figure 14The diagram shows the NAT operation performed at the egress router. However, as further described below, even without reference to Figure 13 Describing this additional NAT operation, some embodiments also perform another NAT operation on the data message at the ingress router.

[0179] Figure 14 The diagram illustrates a process 1400 performed by an egress router in some embodiments when receiving a data message that should be forwarded to a SaaS provider data center over the Internet. As shown, the process initially receives (at 1405) a data message on an interface associated with a virtual network. This message is encapsulated with a tenant tunnel header.

[0180] At 1410, the process determines that the destination IP address in this tunnel header is the destination IP address of its router, so it removes the tenant tunnel header. The process then retrieves (at 1415) the tenant ID from the removed tunnel header. To account for received data messages, the process increments (at 1420) the RX (receive) counter it maintains for the tenant specified by the extracted tenant ID.

[0181] Next, at 1425, the process determines whether the destination IP in the original header of the data message is a public IP that is reachable through the local interface (e.g., local port) of the egress router. This local interface is an interface that is not associated with a VPN tunnel. If not, then the process ends. Otherwise, the process performs (at 1430) a source NAT operation to change the source IP / port address of the data message in the header of this message. The NAT operation and the reasons for performing it will be referenced below. Figure 16 and Figure 17 Further description.

[0182] After 1430, the process performs (at 1435) a routing lookup (e.g., an LPM lookup) in the Internet routing context (i.e., in the Internet routing portion of the routing data, such as the router's Internet VRF namespace) to identify the next-hop interface for this data message. In some embodiments, the process performs this lookup based on the destination network address (e.g., the destination IP address) of the original header of the received data message. From the record identified by this lookup, the process 1400 identifies the local interface through which the data message must be sent to its destination. Thus, at 1435, the process 1400 provides the source network address translated data message to its identified local interface for forwarding to its destination. After 1435, the process increments (at 1440) a counter it maintains for messages transmitted to the SaaS provider and then ends.

[0183] Figure 15The diagram illustrates a message handling process 1500 performed by an ingress router that receives a message sent from a SaaS provider machine to a tenant machine. As shown, the ingress process 1500 begins by initially receiving (at 1505) a data message on a dedicated input interface having a public IP address used for some or all SaaS provider communications. In some embodiments, this input interface is a different interface having an IP address different from the IP address used for communicating with the virtual network.

[0184] After receiving the message, the process performs (at 1510) a routing lookup in the public Internet routing context using the destination IP address contained in the header of the received data message. Based on this lookup, the process determines (at 1515) whether the destination IP address is a local IP address and is associated with an enabled NAT operation. If not, the process ends. Otherwise, the process increments (at 1520) the Internet RX (receive) counter to account for the receipt of the data message.

[0185] Next, at 1525, the process performs a reverse NAT operation that translates the destination IP / port address of the data message into a new destination IP / port address associated with the virtual network for a particular tenant. This NAT operation also generates a tenant ID (e.g., by retrieving the tenant ID from a mapping table that associates the tenant ID with the translated destination IP, or by retrieving the tenant ID from the same mapping table used to obtain the new destination IP / port address). In some embodiments, process 1500 performs (at 1525) its reverse NAT operation using the connection record that process 1400 created when it performed (at 1430) its SNAT operation. This connection record contains the mapping between internal and external IP / port addresses used by SNAT and DNAT operations.

[0186] Based on the translated destination network address, the process then performs (at 1530) a routing lookup (e.g., an LPM lookup) in the identified tenant routing context (i.e., the routing context specified by the tenant ID) to identify the IP address of the egress interface used to exit the tenant's virtual network and reach the tenant's machine in a corporate compute node (e.g., in a branch office). In some embodiments, this egress interface is the IP address of the egress CFE of the egress MFN. At 1530, the process adds a tenant tunnel header to the received data message, embedding the IP address of the ingress router and the IP address of the egress CFE as the source and destination IP addresses in this tunnel header. Next, at 1535, the process increments the tenant's VN-transmit counter to account for the transmission of this data message.

[0187] At 1540, the process performs a routing lookup (e.g., an LPM lookup) in the identified VNP routing context (e.g., in the portion of the VNP that routed the data, such as in the router's VRF namespace) to identify the local interface (e.g., its physical port or virtual port) to which the ingress router provided the encapsulated message. The process then adds (at 1540) a VN-hop header to the received data message, embedding the IP address of the ingress router and the IP address of the next-hop CFE as the source and destination IP addresses of this VN-hop header. After 1555, the process ends.

[0188] As mentioned above, in some embodiments, the MFN includes a NAT engine 215 that performs NAT operations on the ingress and / or egress paths of data messages entering and exiting the virtual network. Today, NAT operations are commonly performed in many contexts and by many devices (e.g., routers, firewalls, etc.). For example, NAT operations are often performed when traffic leaves a private network to isolate internal IP address space from the regulated public IP address space used on the Internet. NAT operations typically map one IP address to another.

[0189] With the rapid increase in computers connected to the Internet, the challenge is that the number of computers will exceed the number of available IP addresses. Unfortunately, even with 4,294,967,296 possible unique addresses, assigning a unique public IP address to each computer is no longer practical. One solution is to assign public IP addresses only to routers at the edge of the private network, while other devices inside the network obtain addresses that are unique only within their internal private network. When a device wants to communicate with devices outside its internal private network, its traffic typically passes through an Internet gateway, which performs a NAT operation to replace the source IP address of this traffic with the public source IP address of the Internet gateway.

[0190] Although the Internet gateway of a private network obtains a registered public address on the Internet, each device inside the private network connected to this gateway receives an unregistered private address. The private addresses of the internal private network can be in any IP address range. However, the Internet Engineering Task Force (IETF) has recommended several private address ranges for use in private networks. These ranges are generally not available on the public Internet so that routers can easily distinguish private addresses from public addresses. These private address ranges are known as RFC1918 and are: (1) Class A 10.0.0.0-10.255.255.255, (2) Class B 172.16.0.0-172.31.255.255, and (3) Class C 192.168.0.0-192.168.255.255.

[0191] It is important to perform source IP translation on data message streams leaving the private network so that external devices can distinguish between different devices within different private networks that use the same internal IP address. When the external device must send a reply message to a device within the private network, the external device must send its reply to a public address that is unique and routable on the Internet. It cannot use the original IP address of the internal device, which may be used by numerous devices in numerous private networks. The external device sends its reply to the public IP address, and the original NAT operation replaces the private source IP address of the internal device with this public IP address. After receiving this reply message, the private network (e.g., the network's gateway) performs another NAT operation to replace the public destination IP address in the reply with the IP address of the internal device.

[0192] Many devices within a private network, and many applications running on those devices, must share one or a limited number of public IP addresses associated with the private network. Consequently, NAT operations typically also translate layer 4 port addresses (e.g., UDP addresses, TCP addresses, RTP addresses, etc.) to uniquely associate external message flows with internal message flows that originate or terminate on different internal machines and / or different applications on those machines. NAT operations are also often stateful because, in many contexts, they require tracking connections and dynamically handling tables, message reassembly, timeouts, forced termination of expired tracked connections, and the like.

[0193] As mentioned above, the virtual network provider of some embodiments provides virtual networks as a service to different tenants through multiple public clouds. These tenants may use public IP addresses in their private networks, and they share a common set of network resources of the virtual network provider (e.g., public IP addresses). In some embodiments, data traffic for different tenants is carried between the CFEs of the overlay network through a tunnel, and the tunnel marks each message with a unique tenant ID. These tenant identifiers allow messages to be sent back to the source device even when the private tenant IP spaces overlap. For example, the tenant identifier allows a message sent from a branch of tenant 17 with a source address of 10.5.12.1 to Amazon.com to be distinguished from a message sent from a branch of tenant 235 with the same source address (and even with the same source port number 55331) to Amazon.com.

[0194] Standard NAT implemented according to RFC 1631 does not support the concept of leases and therefore cannot distinguish between two messages with the same private IP address. However, in many virtual network deployments of some embodiments, using a standard NAT engine is beneficial because many mature, open source, high-performance implementations exist today. In fact, many current Linux kernels have a running NAT engine as a standard feature.

[0195] In order to use a standard NAT engine for different tenants of a tenant virtual network, the virtual network provider of some embodiments uses a Tenant Mapping (TM) engine before using the standard NAT engine. Figure 16 Such a TM engine 1605 is illustrated, placed in each virtual network gateway 1602 on the virtual network's egress path to the Internet. As shown, each TM engine 1605 is placed before a NAT engine 1610 on the message egress path through the Internet 1625 to a SaaS provider data center 1620. In some embodiments, each NAT engine 215 of the MFN includes a TM engine (such as TM engine 1605) and a standard NAT engine (such as NAT engine 1610).

[0196] exist Figure 16 In the illustrated example, messages originate from two virtual network tenants, branch offices 1655 and 1660, and data center 1665, and enter virtual network 1600 through the same ingress gateway 1670, but this need not be the case. In some embodiments, virtual network 1600 is defined across multiple public cloud data centers across multiple public cloud providers. In some embodiments, the virtual network gateway is part of a managed forwarding node, and the TM engine is placed before the NAT engine 1610 in the egress MFN.

[0197] When data messages arrive at egress gateway 1602 to leave the virtual network on their way to SaaS provider data center 1620, each TM engine 1605 maps the source network address (e.g., source IP and / or port address) of the data messages to a new source network address (e.g., source IP and / or port address), and NAT engine 1610 maps the new source network address to yet another source network address (e.g., another source IP and / or port address). In some embodiments, the TM engine is a stateless element and performs mapping for each message using a static table without consulting any dynamic data structures. As a stateless element, the TM engine does not create a connection record when processing the first data message in a data message stream to use when performing its address mapping for subsequent messages in the data message stream.

[0198] On the other hand, in some embodiments, the NAT engine 1605 is a stateful element that performs its mappings by referencing a connection storage device that stores connection records reflecting its previous SNAT mappings. When the NAT engine receives a data message, in some embodiments, the engine first checks its connection storage device to determine whether it has previously created a connection record for the received message stream. If so, the NAT engine uses the mapping contained in this record to perform its SNAT operation. Otherwise, it performs the SNAT operation based on a set of criteria that it uses to derive a new address mapping for the new data message stream. To this end, in some embodiments, the NAT engine uses a common network address translation technology.

[0199] In some embodiments, when the NAT engine receives a reply data message from the SaaS provider machine, the NAT engine may also use the connection storage device in some embodiments to perform a DNAT operation to forward the reply data message to the tenant machine that sent the original message. In some embodiments, the connection record for each processed data message flow has a record identifier that includes an identifier of the flow (e.g., a five-tuple identifier with a translated source network address).

[0200] When performing mapping, the TM engine ensures that data message streams from different tenants using the same source IP and port addresses are mapped to unique, non-overlapping address spaces. For each message, the TM engine identifies the tenant ID and performs address mapping based on this identifier. In some embodiments, the TM engine maps source IP addresses for different tenants to different IP ranges, ensuring that no two messages from different tenants are mapped to the same IP address.

[0201] Therefore, each network type with a different tenant ID will be mapped to a full 2 ​​of the IP addresses 32Unique addresses within the range (0.0.0.0-255.255.255.255). Class A and Class B networks have 256 times and 16 times more possible IP addresses than Class C networks. Based on the ratio of the sizes of Class A, Class B, and Class C networks, the 256 Class A networks can be allocated as follows: (1) 240 to map 240 tenants with Class A networks, (2) 15 to map 240 tenants with Class B networks, and (3) a single Class A network to map 240 tenants with Class C networks. More specifically, in some embodiments, the lowest range of Class A networks (starting with 0.xxx / 24, 1.xxx / 24... up to 239.xxx / 24) will be used to map addresses from the 10.x Class A networks to 240 different target Class A networks. The next 15 Class A networks, 240.xxx / 24 through 254.xxx / 24, will each be used to include 16 Class B networks (e.g., 240 networks total (15*16)). The last Class A network, 255.xxx / 24, will be used to include up to 256 private Class C networks. Even though 256 tenants can be accommodated, only 240 are used, and the 16 Class C networks are unused.

[0202] In summary, some embodiments use the following mapping:

[0203] 10.xxx / 24 network → 1.xxx / 24-239.xxx / 24, resulting in 240 different mappings per tenant;

[0204] 172.16-31.xx / 12 network → 240x.xx / 24-254.xxx / 24, resulting in 240 different mappings per tenant.

[0205] 192.168.xx / 16 → 255.xxx / 24 network, resulting in 240 of the 256 possible mappings per tenant.

[0206] Assuming that it is not known in advance what type of network class the tenant will use, the above scheme can support up to 240 tenants. In some embodiments, the public cloud network uses private IP addresses. In this case, it is expected that it will not be mapped into the private address space again. Since some embodiments remove Class A and Class B networks, only 239 different tenants can be supported in these embodiments. In order to achieve unique mapping, some embodiments number all tenant IDs from 1 to 239, and then, modulo 240, add the least significant 8 bits of the unmasked part of the private domain to the tenant ID (expressed in 8 bits). In this case, for Class A addresses, the first tenant (number 1) will be mapped to 11.xx.xx.xx / 24, and the last one (239) will be mapped to 9.xx.xx.xx / 24.

[0207] exist Figure 16 In the illustrated implementation, some embodiments provide each TM engine 1605 with a list of potential tenant ID subnets, along with a method for routing messages back to any specific IP address within each such subnet. This information can change dynamically as tenants, branches, and mobile devices are added or removed. Therefore, this information must be dynamically distributed to the TM engines in the virtual network's Internet egress gateways. Because a virtual network provider's egress Internet gateways may be used by a large number of tenants, the volume of information distributed and regularly updated can be substantial. Furthermore, the 240 (or 239) limit on tenant IDs is a global limit and can only be addressed by adding multiple IP addresses to the egress point.

[0208] Figure 17 The double NAT method used in some embodiments is illustrated instead of Figure 16 The single NAT method shown in . Figure 17 The approach shown in requires that less tenant data be distributed to most (if not all) TM engines and allows more private tenant networks to be mapped to the virtual network provider's internal network. For a data message flow from a tenant machine traversing through the virtual network 1700 and then through the Internet 1625 to another machine (e.g., to a machine in the SaaS provider data center 1620), Figure 17 The approach shown in FIG places a NAT engine at the ingress gateway 1770 where a data message flow enters a virtual network, and at the egress gateway 1702 or 1704 where this flow leaves the virtual network and enters the Internet 1625. This approach also places a TM engine 1705 before the NAT engine 1712 of the ingress gateway 1770.

[0209] exist Figure 17In the illustrated example, messages originate from two branch offices 1755 and 1760 and data center 1765, belonging to two virtual network tenants, and enter virtual network 1700 through the same ingress gateway 1770, but this need not be the case. Like virtual network 1600, in some embodiments, virtual network 1700 is defined across multiple public cloud data centers across multiple public cloud providers. Furthermore, in some embodiments, virtual network gateways 1702, 1704, and 1770 are part of a managed forwarding node, and in these embodiments, the TM engine is placed before the NAT engine 215 in these MFNs.

[0210] TM Engines 1605 and 1705 Figure 16 and Figure 17 1605. Like TM engine 1605, when a data message is destined for SaaS provider data center 1620 (i.e., has a destination IP address for it), TM engine 1705 maps the source IP and port address of the data message entering the virtual network to a new source IP and port address. For each such data message, TM engine 1705 identifies the tenant ID and performs its address mapping based on this identifier.

[0211] Like TM engine 1605, in some embodiments, TM engine 1705 is a stateless element and performs mapping for each message using a static table without consulting any dynamic data structures. As a stateless element, the TM engine does not create a connection record when processing the first data message in a data message stream to use when performing its address mapping for processing subsequent messages in the data message stream.

[0212] When performing its mapping, the TM engine 1705 in the ingress gateway 1770 ensures that data message streams from different tenants using the same source IP and port addresses are mapped to unique, non-overlapping address spaces. In some embodiments, the TM engine maps the source IP addresses of different tenants to different IP ranges, so that no two messages from different tenants are mapped to the same IP address. In other embodiments, the TM engine 1705 may map the source IP addresses of two different tenants to the same source IP range but different source port ranges. In still other embodiments, the TM engine maps two tenants to different source IP ranges, while mapping two other tenants to the same source IP range but different source port ranges.

[0213] Unlike TM Engine 1605, TM Engine 1705 at the virtual network ingress gateway only needs to identify tenants of the branch office, corporate data center, and corporate compute node connected to the ingress gateway. This significantly reduces the amount of tenant data that needs to be initially provisioned and periodically updated for each TM Engine. Furthermore, as before, each TM Engine can only map 239 / 240 tenants to a unique address space. However, because the TM Engine is located at the virtual network provider's ingress gateway, each TM Engine can uniquely map 239 / 240 tenants.

[0214] In some embodiments, the NAT engine 1712 of the ingress gateway 1770 can use either an external public IP address or an internal IP address specific to the public cloud (e.g., AWS, GCP, or Azure) where the ingress gateway 1770 resides. In either case, the NAT engine 1712 maps the source network address of the incoming message (i.e., the message entering the virtual network 1700) to a unique IP address within the private cloud network of its ingress gateway. In some embodiments, the NAT engine 1712 translates the source IP address of each tenant's data message stream into a different unique IP address. However, in other embodiments, the NAT engine 1712 translates the source IP addresses of different tenants' data message streams into the same IP address, but uses the source port address to distinguish between different tenants' data message streams. In yet other embodiments, the NAT engine maps the source IP addresses of two tenants to different source IP ranges, while mapping the source IP addresses of the other two tenants to the same source IP range but different source port ranges.

[0215] In some embodiments, the NAT engine 1712 is a stateful element that performs its mappings by referencing a connection store that stores connection records reflecting its previous SNAT mappings. In some embodiments, when the NAT engine receives a reply data message from a SaaS provider machine, the NAT engine may also use the connection store in some embodiments to perform a DNAT operation to forward the reply data message to the tenant machine that sent the original message. In some embodiments, the TM and NAT engines 1705, 1710, and 1712 are configured by the controller cluster 160 (e.g., provided with tables describing mappings to be used for different tenants and different ranges of network address space).

[0216] Figure 18An example is provided illustrating source port translation by ingress NAT engine 1712. Specifically, the diagram shows source address mapping performed by tenant mapping engine 1705 and ingress NAT engine 1712 on data message 1800 as it enters virtual network 1700 through ingress gateway 1770 and as it leaves the virtual network at egress gateway 1702. As shown, tenant gateway 1810 sends data message 1800, which arrives at IPsec gateway 1805 with source IP address 10.1.1.13 and source port address 4432. In some embodiments, these source addresses are addresses used by tenant machines (not shown), while in other embodiments, one or both of these source addresses are source addresses generated by a source NAT operation performed by the tenant gateway or another network element in the tenant data center.

[0217] After this message has been processed by IPsec gateway 1805, this gateway or another module of the ingress MFN associates the message with tenant ID 15, which identifies the virtual network tenant to which message 1800 belongs. As shown, based on this tenant ID, tenant mapping engine 1705 then maps the source IP and port address to the source IP and port address pair of 15.1.1.13 and 253. This source IP and port address uniquely identifies the message flow of data message 1800. In some embodiments, TM engine 1705 performs this mapping in a stateless manner (i.e., without reference to connection tracking records). In other embodiments, the TM engine performs this mapping in a stateful manner.

[0218] Ingress NAT engine 1712 then (1) translates the source IP address of data message 1800 into a unique private or public (internal or external) IP address 198.15.4.33, and (2) translates the source port address of this message into port address 714. In some embodiments, the virtual network uses this IP address for other data message flows of the same or different tenants. Thus, in these embodiments, the source network address translation (SNAT) operation of NAT engine 1712 uses the source port address to distinguish different message flows of different tenants using the same IP address within the virtual network.

[0219] In some embodiments, the source port address assigned by the SNAT operation of the ingress NAT engine is also the source port address used to distinguish different message flows outside the virtual network 1700. Figure 18This is the case in the example shown. As shown, the egress NAT engine 1710 in this example does not change the source port address of the data message when performing the SNAT operation. Instead, it simply changes the source IP address to the external IP address 198.15.7.125. In some embodiments, this external IP address is the public IP address of the egress gateway(s) of the virtual network. In some embodiments, this public IP address is also the IP address of the public cloud data center in which the ingress and egress gateways 1770 and 1702 operate.

[0220] The data message is routed through the Internet to the gateway 1815 of the SaaS provider's data center using source IP and port addresses 198.15.7.125 and 714. In this data center, the SaaS provider machine performs actions based on this message and sends back a reply message 1900, which will be referenced below. Figure 19 Describe its processing. In some embodiments, the SaaS provider machine performs one or more service operations (e.g., middlebox service operations, such as firewall operations, IDS operations, IPS operations, etc.) on the data message based on one or more service rules defined by reference to the source IP and port addresses 198.15.7.125 and 714. In some of these embodiments, different service rules for different tenants may specify the same source IP address (e.g., 198.15.7.125) in the rule identifier while specifying different source port addresses in the rule identifiers. The rule identifier specifies a set of attributes for comparison with the data message flow attributes when performing a lookup operation to identify a rule that matches the data message.

[0221] Figure 19 The diagram illustrates the processing of a reply message 1900 sent by a SaaS machine (not shown) in response to the SaaS machine's processing of a data message 1800. In some embodiments, the reply message 1900 may be identical to the original data message 1800, may be a modified version of the original data message 1800, or may be an entirely new data message. As shown, the SaaS gateway 1815 sends the message 1900 based on the destination IP address and destination port address of 198.15.7.125 and 714, which are the source IP address and source port address of the data message 1800 when it arrives at the SaaS gateway 1815.

[0222] The message 1900 is received at the gateway of the virtual network (not shown), and this gateway provides the data message to the NAT engine 1710, which performs a final SNAT operation on the message 1800 before it is sent to the SaaS provider. Figure 19In the example shown, the data message 1900 is received at the same NAT engine 1710 that performed the last SNAT operation, but this need not be the case in every deployment.

[0223] NAT engine 1710 (now acting as an ingress NAT engine) performs a DNAT (destination NAT) operation on data message 1900. This operation changes the external destination IP address 198.15.7.125 to a destination IP address 198.15.4.33, which the virtual network uses to forward data message 1900 through the public cloud routing architecture and between virtual network components. Again, in some embodiments, IP address 198.15.4.33 can be a public IP address or a private IP address.

[0224] As shown, after NAT engine 1710 has translated the destination IP address of message 1900, NAT engine 1712 (now acting as the egress NAT engine) receives message 1900. NAT engine 1712 then performs a second DNAT operation on message 1900, replacing the destination IP address and destination port address of message 1900 with 15.1.1.13 and 253. These addresses are the addresses recognized by TM engine 1705. TM engine 1705 replaces these addresses with the destination IP address and destination port address 10.1.1.13 and 4432, associates data message 1900 with tenant ID 15, and provides message 1900, along with this tenant ID, to IPsec gateway 1805 for forwarding to tenant gateway 1810.

[0225] In some embodiments, a virtual network provider uses the above-described processes, systems, and components to provide multiple virtual WANs for multiple different tenants (eg, multiple different corporate WANs for multiple companies) on multiple public clouds of the same or different public cloud providers. Figure 20 An example is given showing M virtual corporate WANs 2015 for M tenants of a virtual network provider having network infrastructure and controller cluster(s) 2010 in N public clouds 2005 of one or more public cloud providers.

[0226] Each tenant's virtual WAN 2015 can span all N public clouds 2005 or a subset of these public clouds. Each tenant's virtual WAN 2015 connects one or more of the tenant's branch offices 2020, data centers 2025, SaaS provider data centers 2030, and remote devices. In some embodiments, each tenant's virtual WAN spans any public cloud 2005 that the VNP's controller cluster deems necessary to efficiently forward data messages between the tenant's various compute nodes 2020-2035. When selecting a public cloud, in some embodiments, the controller cluster also considers the tenant's selected public cloud and / or a public cloud in which the tenant or at least one of the tenant's SaaS providers has one or more machines.

[0227] Each tenant's virtual WAN 2015 allows a tenant's remote device 2035 (e.g., a mobile device or remote computer) to avoid interacting with the tenant's WAN gateway at any branch office or tenant data center in order to access SaaS provider services (i.e., access a SaaS provider machine or cluster of machines). In some embodiments, the tenant's virtual WAN allows remote devices to bypass WAN gateways at branch offices and tenant data centers by moving the functionality of these WAN gateways (e.g., WAN security gateways) to one or more machines in the public cloud spanned by the virtual WAN.

[0228] For example, to allow remote devices to access the computing resources of a tenant or its SaaS provider services, in some embodiments, the WAN gateway must enforce firewall rules that control how remote devices can access the tenant's computer resources or its SaaS provider services. To avoid the tenant's branch office or data center WAN gateway, the tenant's firewall engine 210 is placed in a virtual network MFN in one or more public clouds spanned by the tenant's virtual WAN.

[0229] The firewall engines 210 within these MFNs perform firewall service operations on data message flows to and from remote devices. By performing these operations within a virtual network deployed on one or more public clouds, data message traffic associated with a tenant's remote devices receives firewall rule processing without being unnecessarily routed through the tenant's data center(s) or branch offices. This alleviates traffic congestion in tenant data centers and branch offices and avoids consuming expensive ingress / egress network bandwidth at these locations to process traffic that is not destined for computing resources at these locations. This approach also helps expedite the forwarding of data message traffic to and from remote devices by allowing intrusive firewall rule processing within the virtual network as data message flows traverse to their destinations (e.g., at their ingress MFN, egress MFN, or intermediate MFN).

[0230] In some embodiments, the firewall enforcement engine 210 (e.g., a firewall service VM) of the MFN receives firewall rules from the VPN central controller 160. In some embodiments, the firewall rule includes a rule identifier and an action. In some embodiments, the rule identifier includes one or more match values ​​to be compared with data message attributes (such as layer 2 attributes (e.g., MAC address), layer 3 attributes (e.g., 5-tuple identifier, etc.), tenant ID, location ID (e.g., office location ID, data center ID, remote user ID, etc.)) to determine whether the firewall rule matches the data message.

[0231] In some embodiments, the action of a firewall rule specifies the action (e.g., allow, drop, redirect, etc.) that the firewall enforcement engine 210 must take on a data message when the firewall rule matches an attribute of the data message. To account for the possibility of multiple firewall rules matching a data message, the firewall enforcement engine 210 stores the firewall rules (received from the controller cluster 160) in a hierarchical manner in the firewall rule data storage device so that one firewall rule can have a higher priority than another firewall rule. In some embodiments, when a data message matches two firewall rules, the firewall enforcement engine applies the rule with the higher priority. In other embodiments, the firewall enforcement engine checks the firewall rules according to the hierarchy of firewall rules (i.e., checks higher priority rules before lower priority rules) to ensure that it matches the higher priority rule first in the event that another lower priority rule may also be a match for the data message.

[0232] Some embodiments allow a controller cluster to configure MFN components to cause a firewall services engine to inspect data messages at an ingress node (e.g., node 850) upon entry into a virtual network, at an intermediate node (e.g., node 857) on the virtual network, or at an egress node (e.g., node 855) upon exiting the virtual network. In some embodiments, at each of these nodes, a CFE (e.g., 832, 856, or 858) invokes its associated firewall services engine 210 to perform firewall services operations on the data message received by the CFE. In some embodiments, the firewall services engine returns the firewall services engine's decision to the module that invoked the firewall services engine (e.g., to the CFE) so that the module can perform firewall operations on the data message. In other embodiments, the firewall services engine performs its firewall actions on the data message.

[0233] In some embodiments, other MFN components direct the firewall services engine to perform its operations. For example, in some embodiments, at an ingress node, a VPN gateway (e.g., 225 or 230) instructs its associated firewall services engine to perform its operations to determine whether a data message should be passed to the ingress node's CFE. Furthermore, in some embodiments, at an egress node, the CFE passes the data message to its associated firewall services engine. If it determines to allow the data message to pass, the firewall services engine either passes the data message through an external network (e.g., the Internet) to its destination or passes the data message to its associated NAT engine 215 to perform its NAT operations before passing the data message through the external network to its destination.

[0234] Some embodiments of the virtual network provider allow tenants' WAN security gateways defined in the public cloud to implement other security services in addition to or in lieu of firewall services. For example, a tenant's distributed WAN security gateway (in some embodiments, distributed across every public cloud data center spanned by the tenant's virtual network) may include not only a firewall service engine, but also an intrusion detection engine and an intrusion prevention engine. In some embodiments, the intrusion detection engine and intrusion prevention engine are architecturally integrated into MFN 150, occupying a similar position to firewall service engine 210.

[0235] In some embodiments, each of these engines includes one or more storage devices that store intrusion detection / prevention policies distributed by the central controller cluster 160. In some embodiments, these policies configure the engines to detect / prevent unauthorized intrusions into tenants' virtual networks (deployed across multiple public cloud data centers) and take action (e.g., generating logs, sending notifications, shutting down services or machines, etc.) in response to detected intrusion events. Much like firewall rules, intrusion detection / prevention policies can be enforced at various managed forwarding nodes (e.g., ingress MFNs, intermediate MFNs, and / or egress MFNs of a data message flow) on which the virtual networks are defined.

[0236] As mentioned above, the virtual network provider deploys each tenant's virtual WAN by deploying at least one MFN in each public cloud spanned by the virtual WAN and configuring the deployed MFNs to define routes between the MFNs that allow tenants' message flows to enter and exit the virtual WAN. Furthermore, as mentioned above, in some embodiments, each MFN can be shared by different tenants, while in other embodiments, each MFN is deployed only for a specific tenant.

[0237] In some embodiments, each tenant's virtual WAN is a secure virtual WAN established by connecting the MFN used by that WAN via an overlay tunnel. In some embodiments, this overlay tunneling approach encapsulates each tenant's data message stream with a tunnel header unique to each tenant (e.g., containing a tenant identifier that uniquely identifies the tenant). For each tenant, in some embodiments, the virtual network provider's CFE uses one tunnel header to identify the ingress / egress forwarding element for entry / exit of the tenant's virtual WAN and another tunnel header to traverse intermediate forwarding elements of the virtual network. In other embodiments, the virtual WAN's CFE uses a different overlay encapsulation mechanism.

[0238] To deploy a virtual WAN for a tenant on one or more public clouds, the controller cluster of the VNP (1) identifies possible edge MFNs (which can serve as ingress or egress MFNs for different data message flows) for the tenant based on the location of the tenant's corporate computing nodes (e.g., branch offices, data centers, mobile users, and SaaS providers), and (2) identifies routes between all possible edge MFNs. Once these routes are identified, they are propagated to the forwarding table of the CFE (e.g., using OpenFlow to different OVS-based virtual network routers). Specifically, to identify the best routes through the tenant's virtual WAN, the MFNs associated with this WAN generate measurements that quantify the quality of the network connections between them and their neighboring MFNs and periodically provide these measurements to the controller cluster of the VNP.

[0239] As mentioned above, the controller cluster then aggregates measurements from the different MFNs, generates a routing map based on these measurements, defines routes through the tenant's virtual WAN, and then distributes these routes to the forwarding elements of the MFN's CFE. To dynamically update the defined routes for a tenant's virtual WAN, the MFN associated with this WAN periodically generates its measurements and provides these measurements to the controller cluster. The controller cluster then periodically repeats its measurement aggregation, routing map generation, route identification, and route distribution based on the updated measurements it receives.

[0240] When defining routes across a tenant's virtual WAN, the VPN's controller cluster optimizes the routes for desired end-to-end performance, reliability, and security, while attempting to minimize routing of tenants' message flows across the Internet. The controller cluster also configures MFN components to optimize Layer 4 processing of data message flows passing through the network (e.g., optimizing the end-to-end rate of a TCP connection by splitting rate control mechanisms across the connection path).

[0241] With the proliferation of public clouds, it is often easy to find a major public cloud data center near each of a company's branch offices. Similarly, SaaS providers are increasingly hosting their applications within the public cloud or similarly located near a public cloud data center. Thus, the virtual corporate WAN 2015 securely uses the public cloud 2005 as the corporate network infrastructure that exists near corporate computing nodes (e.g., branch offices, data centers, remote devices, and SaaS providers).

[0242] Corporate WANs require guaranteed bandwidth to consistently deliver business-critical applications with acceptable performance. These applications can include interactive data applications such as ERP, finance, or procurement, deadline-oriented applications (e.g., industrial or IoT control), and real-time applications (e.g., VoIP or video conferencing). Therefore, traditional WAN infrastructure (e.g., Frame Relay or MPLS) provides these guarantees.

[0243] The primary obstacle to providing bandwidth guarantees in multi-tenant networks is the need to reserve bandwidth for a particular customer on one or more paths. In some embodiments, VNPs provide QoS services and offer both ingress committed rate (ICR) and egress committed rate (ECR) guarantees. ICR refers to the rate of traffic entering the virtual network, while ECR refers to the rate of traffic leaving the virtual network to a tenant site.

[0244] In some embodiments, the virtual network can provide bandwidth and latency guarantees as long as traffic does not exceed ICR and ECR limits. For example, as long as HTTP ingress or egress traffic does not exceed 1Mbps, bandwidth and low latency are guaranteed. This is a point-to-cloud model because, for QoS purposes, the VPN does not need to track the destination of traffic as long as the destination is within the ICR / ECR boundaries. This model is sometimes referred to as the hose model.

[0245] For more stringent applications, where customers expect point-to-point guarantees, building virtual data pipes is necessary to deliver highly critical traffic. For example, an enterprise might want to connect two hub sites or data centers with a high service level agreement guarantee. To achieve this, VPN routing automatically selects a routing path that meets each customer's bandwidth constraints. This is known as the point-to-point model or pipe model.

[0246] A VPN's primary advantage in providing guaranteed bandwidth to end users is the ability to adapt VPN infrastructure to changing bandwidth demands. Most public clouds provide a minimum bandwidth guarantee between every two instances in different regions of the same cloud. If the current network lacks sufficient unused capacity to provide guaranteed bandwidth for new requests, the VPN can add new resources to its infrastructure. For example, a VPN could add a new CFE in a high-demand region.

[0247] One challenge is optimizing the performance and cost of this new dimension when planning routes and scaling infrastructure. To facilitate algorithms and bandwidth accounting, some embodiments assume that end-to-end bandwidth reservations are not split. Alternatively, if a certain bandwidth (e.g., 10 Mbps) is reserved between branch A and branch B of a certain tenant, the bandwidth is allocated along a single path starting from the ingress CFE to which branch A is connected, then traversing a set of zero or more intermediate CFEs to the egress CFE connected to branch B. Some embodiments also assume that the bandwidth-guaranteed path traverses only a single public cloud.

[0248] To account for various bandwidth reservations that intersect in the network topology, in some embodiments, the VNP statically defines routes along the reserved bandwidth path, so that a data message flow always traverses the same route reserved for that bandwidth requirement. In some embodiments, each route is identified by a single label, and each CFE traversed by that route matches that label to a single outgoing interface associated with that route. Specifically, each CFE matches a single outgoing interface to each data message that has that label in its header and arrives from a specific incoming interface.

[0249] In some embodiments, the controller cluster maintains a network graph formed by a number of interconnected nodes. Each node n in the graph has an allocated total guaranteed bandwidth (TBW) associated with the node. n ), and the amount of bandwidth (RBW) that this node has reserved (allocated to a reserved path) n ). In addition, for each node, the graph includes the cost in minutes per gigabyte (C ij ), and the delay in milliseconds associated with sending traffic between this node and all other nodes in the graph (D ij The weight associated with sending traffic between node i and node j is W ij =a*C ij +D ij , where a is a system parameter typically between 1 and 10.

[0250] When a bandwidth reservation request with a value of BW is accepted between branches A and B, the controller cluster first maps the request to the specific ingress and egress routers n and m, respectively, bound to branches A and B. The controller cluster then performs a routing process that calculates two minimum-cost (e.g., shortest-path) routes between n and m. The first is the minimum-cost (e.g., shortest-path) route between n and m, regardless of the available bandwidth along the calculated route. The total weight of this route is calculated as W1.

[0251] The second lowest cost (e.g., shortest path) is initially calculated by eliminating BW>TBW i -RBW i The modified graph is called a pruned graph. The controller cluster then performs a second minimum cost (e.g., shortest path) route calculation on the pruned graph. If the weight of the second route does not exceed K% of the first route (K is typically 10%-30%), then the second route is selected as the preferred path. On the other hand, when this requirement is not met, the controller cluster will have a TBW i -RBW i The node i with the minimum value of is added to the first path, and then the calculation of the minimum cost (e.g., shortest path) is repeated twice. The controller cluster will continue to add more routers until the condition is met. At that time, the reserved bandwidth BW is added to all RBW i , where i is the router on the selected route.

[0252] For the special case of requesting additional bandwidth for a route that already has reserved bandwidth, the controller cluster will first delete the current bandwidth reservation between nodes A and B and will calculate the path between these nodes for the total bandwidth requested. To this end, in some embodiments, the information saved for each node also includes the bandwidth reserved for each label or each source and destination branch, rather than just the total bandwidth reserved. After adding the bandwidth reservation to the network, some embodiments do not revisit the route as long as the measured network latency or cost through the virtual network has not changed significantly. However, when the measurements and / or costs change, these embodiments repeat the bandwidth reservation and route calculation process.

[0253] Figure 21 The process 2100 that is performed by the controller cluster 160 of a virtual network provider to deploy and manage a virtual WAN for a particular tenant is conceptually illustrated. In some embodiments, the process 2100 is performed by several different controller programs executing on the controller cluster 160. The operations of this process do not necessarily have to follow Figure 21 The order shown in FIG is not because these operations can be performed in parallel or in a different order by different programs. Therefore, these operations are shown in this figure only to describe an exemplary sequence of operations performed by the controller cluster.

[0254] As shown, the controller cluster initially deploys (at 2105) several MFNs in several public cloud data centers of several different public cloud providers (e.g., Amazon AWS, Google GCP, etc.) In some embodiments, the controller cluster configures (at 2105) these deployed MFNs for one or more other tenants that are different from the specific tenant for which process 2100 is shown.

[0255] At 2110, the controller cluster receives data from a particular tenant regarding external machine properties and the location of the particular tenant. In some embodiments, this data includes identifiers of a private subnet used by the particular tenant and one or more tenant offices and data centers where the particular tenant has external machines. In some embodiments, the controller cluster may receive tenant data via an API or through a user interface provided by the controller cluster.

[0256] Next, at 2115, the controller cluster generates a routing map for the specific tenant based on measurements collected by the measurement agent 205 of the MFN 150, which is a candidate MFN for establishing a virtual network for the specific tenant. As mentioned above, the routing map has nodes representing MFNs and links between the nodes representing network connections between the MFNs. Links have associated weights, which are cost values ​​that quantify the quality and / or cost of using the network connection represented by the link. As mentioned above, the controller cluster first generates a measurement map based on the collected measurements, and then generates the routing map by removing non-optimal links (e.g., links with high latency or packet loss) from the measurement map.

[0257] After constructing the routing graph, the controller cluster performs (at 2120) a path search to identify possible routes between different pairs of candidate ingress and egress nodes (i.e., MFNs) that the tenant's external machines can use to send data messages into and receive data messages from the virtual network (deployed by the MFN). In some embodiments, the controller cluster uses a known path search algorithm to identify different paths between each candidate ingress / egress pair of nodes. Each path for such a pair uses one or more links that, when cascaded, traverse from the ingress node through zero or more intermediate nodes to the egress node.

[0258] In some embodiments, the cost between any two MFNs comprises a weighted sum of an estimated latency and a financial cost of a connecting link between the two MFNs. In some embodiments, the latency and financial cost comprise one or more of: (1) a link delay measurement, (2) an estimated message processing latency, (3) cloud fees for outgoing traffic from a particular data center or to another data center of the same public cloud provider or out of the public cloud (PC) provider's cloud (e.g., to another public cloud data center of another public cloud provider or to the Internet), and (4) an estimated message processing cost associated with the MFN executing on a host computer in the public cloud.

[0259] Some embodiments evaluate the cost of a connection link between two MFNs that traverses the public Internet to minimize such traversal whenever possible. Some embodiments also incentivize the use of private network connections between two data centers (e.g., by reducing the cost of the connection link) to bias route generation toward using such connections. Using the calculated costs of these paired links, the controller cluster can calculate the cost of each routing path using one or more of these paired links by aggregating the costs of the individual paired links used by the routing path.

[0260] The controller cluster then selects (at 2120) one or up to N identified paths (where N is an integer greater than 1) based on the calculated cost (e.g., the lowest aggregate cost) of the identified candidate paths between each candidate ingress / egress node pair. In some embodiments, as mentioned above, the calculated cost of each path is based on the weighted cost of each link used by the path (e.g., the sum of the associated weight values ​​of each link). The controller cluster may select more than one path between a pair of ingress / egress nodes when more than one route is required between two MFNs to allow the ingress MFN or intermediate MFN to perform multipath operations.

[0261] After selecting (at 2120) one or N paths for each candidate ingress / egress node pair, the controller cluster defines one or N routes based on the selected paths and then generates a routing table or portion of a routing table for the MFN that implements the virtual network for the specific tenant. The generated routing records identify the edge MFNs that reach the different subnets of the specific tenant and identify the next-hop MFN for traversing the route from the ingress MFN to the egress MFN.

[0262] At 2125, the controller cluster distributes the routing record to the MFNs so as to configure the forwarding elements 235 of these MFNs to implement the virtual network for the particular tenant. In some embodiments, the controller cluster communicates with the forwarding elements to pass the routing record using a communication protocol currently used in software-defined multi-tenant data centers to configure software routers executing on host computers to implement logical networks across host computers.

[0263] Once the MFN has been configured and the virtual network is operational for a particular tenant, the edge MFN receives data messages from the tenant's external machines (i.e., machines outside the virtual network) and forwards these data messages to the edge MFN in the virtual network. The edge MFN, in turn, forwards the data messages to the tenant's other external machines. While performing these forwarding operations, the ingress, intermediate, and egress MFNs collect statistics about their forwarding operations. Furthermore, in some embodiments, one or more modules on each MFN collect other statistics about network or compute consumption in the public cloud data center. In some embodiments, the public cloud provider collects this consumption data and transmits it to the virtual network provider.

[0264] When a billing period approaches, the controller cluster collects (e.g., at 2130) statistics collected by the MFN and / or network / computing consumption data collected by the MFN or provided by the public cloud provider. Based on the collected statistics and / or provided network / computing consumption data, the controller cluster generates (at 2130) a billing report and sends the billing report to the specific tenant.

[0265] As mentioned above, the billed amount in the billing report takes into account the statistics and network / consumption data received by the controller cluster (e.g., at 2130). Furthermore, in some embodiments, the bill takes into account the costs incurred by the virtual network provider to operate the MFN (implementing the virtual network for the particular tenant) plus a rate of return (e.g., an additional 10%). This billing scheme is convenient for the particular tenant because the particular tenant does not have to deal with bills from multiple different public cloud providers on which the tenant's virtual network is deployed. In some embodiments, the costs incurred by the VNP include costs charged to the VNP by the public cloud provider. At 2130, the controller cluster also charges the credit card or electronically withdraws funds from the bank account for the charges reflected in the billing report.

[0266] At 2135, the controller cluster determines whether it has received new measurements from the measurement agent 205. If not, processing transitions to 2145, which will be described below. On the other hand, when the controller cluster determines that it has received new measurements from the measurement agent, it determines (at 2140) whether it needs to recheck its routing map for the particular tenant based on the new measurements. In the absence of an MFN failure, in some embodiments, the controller cluster updates its routing map for each tenant based on the received updated measurements at most once during a specific time period (e.g., once every 24 hours or once a week).

[0267] When the controller cluster determines (at 2140) that it needs to re-examine the routing map based on new measurements it has received, the process generates (at 2145) a new measurement map based on the newly received measurements. In some embodiments, the controller cluster blends each new measurement with the previous measurements using a weighted sum to ensure that the measurement values ​​associated with the links of the measurement map do not fluctuate wildly each time a new set of measurements is received.

[0268] At 2145, the controller cluster also determines whether the routing map needs to be adjusted based on the adjusted measurement map (e.g., whether weight values ​​for routing map links need to be adjusted, or links need to be added or removed from the routing map due to the adjusted measurement values ​​associated with the links). If so, the controller cluster (at 2145) adjusts the routing map, performs a path search operation (such as operation 2120) to identify routes between ingress / egress node pairs, generates routing records based on the identified routes, and distributes the routing records to the MFN. From 2145, processing transitions to 2150.

[0269] When the controller cluster determines (at 2140) that it does not need to recheck the routing map, processing also transitions to 2150. At 2150, the controller cluster determines whether it is approaching another billing cycle for which it must collect statistics about processed data messages and consumed network / computing resources. If not, processing returns to 2135 to determine whether it has received new measurements from the MFN measurement agent. Otherwise, processing returns to 2130 to collect statistics, network / computing consumption data, and generate and send billing reports. In some embodiments, the controller cluster repeats the operations of process 2100 until a particular tenant no longer requires the virtual network deployed across the public cloud data center.

[0270] In some embodiments, the controller cluster not only deploys virtual networks for tenants in public cloud data centers but also assists tenants in deploying and configuring compute node machines and service machines in the public cloud data centers. The deployed service machines can be separate from the service machines in the MFN. In some embodiments, the controller cluster billing reports issued to a specific tenant also account for the computing resources consumed by the deployed compute and service machines. Again, receiving a single bill from a single virtual network provider for network and computing resources consumed across multiple public cloud data centers across multiple public cloud providers is more advantageous for tenants than receiving multiple bills from multiple public cloud providers.

[0271] Many of the above features and applications are implemented as software processing designated as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more processing units (e.g., one or more processors, cores of a processor, or other processing units), they cause the (one or more) processing units to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard drives, EPROMs, and the like. Computer-readable media do not include carrier waves and electronic signals transmitted wirelessly or via wired connections.

[0272] In this specification, the term "software" is intended to include firmware residing in a read-only memory or an application stored in a magnetic storage device, which can be read into memory for processing by a processor. Moreover, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while retaining different software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement the software inventions described herein is within the scope of the present invention. In some embodiments, the software program, when installed to operate on one or more electronic systems, defines one or more specific machine implementations that execute and perform the operations of the software program.

[0273] Figure 22 A computer system 2200 is conceptually illustrated, and some embodiments of the present invention may be implemented using the computer system 2200. The computer system 2200 may be used to implement any of the hosts, controllers, and managers described above. As such, it may be used to perform any of the processes described above. This computer system includes various types of non-transitory machine-readable media and interfaces for various other types of machine-readable media. The computer system 2200 includes a bus 2205, processing unit(s) 2210, system memory 2225, read-only memory 2230, permanent storage device 2235, input device 2240, and output device 2245.

[0274] The bus 2205 collectively represents all system buses, peripheral, and chipset buses that communicatively connect the numerous internal devices of the computer system 2200. For example, the bus 2205 communicatively connects the processing unit(s) 2210 with the read-only memory 2230, the system memory 2225, and the permanent storage device 2235.

[0275] The processing unit(s) 2210 retrieves instructions to be executed and data to be processed from these various storage units in order to perform the processing of the present invention. In different embodiments, the processing unit(s) may be a single processor or a multi-core processor. The read-only memory (ROM) 2230 stores static data and instructions required by the processing unit(s) 2210 and other modules of the computer system. On the other hand, the permanent storage device 2235 is a read-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the computer system 2200 is turned off. Some embodiments of the present invention use a mass storage device (such as a magnetic disk or optical disk and its corresponding disk drive) as the permanent storage device 2235.

[0276] Other embodiments use removable storage devices (such as floppy disks, flash drives, etc.) as permanent storage devices. Like permanent storage device 2235, system memory 2225 is a read-write memory device. However, unlike storage device 2235, system memory is a volatile read-write memory, such as random access memory. System memory stores some instructions and data needed by the processor at runtime. In some embodiments, the processes of the present invention are stored in system memory 2225, permanent storage device 2235, and / or read-only memory 2230. The processing unit(s) 2210 retrieve instructions to execute and data to process from these various memory units in order to perform the processes of some embodiments.

[0277] The bus 2205 is also connected to input and output devices 2240 and 2245. Input devices enable a user to communicate information and select commands to the computer system. Input devices 2240 include an alphanumeric keyboard and a pointing device (also known as a "cursor control device"). Output devices 2245 display images generated by the computer system. Output devices include printers and display devices, such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs). Some embodiments include devices such as touch screens that function as both input and output devices.

[0278] Finally, if Figure 22 As shown in FIG, bus 2205 also couples computer system 2200 to a network 2265 via a network adapter (not shown). In this manner, the computer can be part of a computer network, such as a local area network (LAN), a wide area network (WAN), or an intranet, or a network of networks, such as the Internet. Any or all components of computer system 2200 may be used in conjunction with the present invention.

[0279] Some embodiments include electronic components such as a microprocessor, a storage device and memory that stores computer program instructions on a machine-readable or computer-readable medium (alternatively referred to as a computer-readable storage medium, a machine-readable medium, or a machine-readable storage medium). Some examples of such computer-readable media include RAM, ROM, compact disc read-only memory (CD-ROM), compact disc recordable memory (CD-R), compact disc rewritable memory (CD-RW), digital versatile disc read-only memory (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid-state hard drives, read-only and recordable A computer readable medium may store a computer program that is executable by at least one processing unit and includes a set of instructions for performing various operations. Examples of computer programs or computer code include machine code (such as produced by a compiler) and files including high-level code that is executed by a computer, electronic component, or microprocessor using an interpreter.

[0280] While the above discussion primarily refers to microprocessors or multi-core processors executing software, some embodiments are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuits themselves.

[0281] As used in this specification, the terms "computer," "server," "processor," and "memory" refer to electronic or other technical devices. These terms do not include a person or group of people. For the purposes of this specification, the term "display" refers to displaying on an electronic device. As used in this specification, the terms "computer-readable medium," "computer-readable medium," and "machine-readable medium" are entirely limited to tangible, physical objects that store information in a computer-readable form. These terms do not include any wireless signals, wired download signals, or any other short-lived or transient signals.

[0282] While the present invention has been described with reference to many specific details, it will be appreciated by those skilled in the art that the present invention may be implemented in other specific forms without departing from the spirit of the invention. For example, several of the examples described above illustrate virtual corporate WANs for corporate tenants of a virtual network provider. It will be appreciated by those skilled in the art that in some embodiments, a virtual network provider deploys virtual networks for non-corporate tenants (e.g., for schools, colleges, universities, non-profit entities, etc.) on several public cloud data centers of one or more public cloud providers. These virtual networks are virtual WANs that connect multiple computing endpoints of non-corporate entities (e.g., offices, data centers, computers and devices of remote users, etc.).

[0283] Several of the above-described embodiments include various data in the overlay encapsulation header. Those skilled in the art will appreciate that other embodiments may not use the encapsulation header to relay all of this data. For example, instead of including a tenant identifier in the overlay encapsulation header, other embodiments derive the tenant identifier from the address of the CFE forwarding the data message. For example, in some embodiments where different tenants deploy their own MFNs in the public cloud, the tenant identity is associated with the MFN that processes the tenant's message.

[0284] Moreover, several figures conceptually illustrate the processes of some embodiments of the present invention. In other embodiments, the specific operations of these processes may not be performed in the exact order shown and described in these figures. The specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different embodiments. In addition, the processes may be implemented using several sub-processes or as part of a larger macro-process. Therefore, it will be understood by those skilled in the art that the present invention is not limited to the foregoing illustrative details, but is defined by the appended claims.

Claims

1. A method for defining a route for a data message flow associated with an entity, the data message flow passing through a virtual network defined on a plurality of public cloud data centers, the method comprising: In the plurality of public cloud data centers, deploying a plurality of virtual machines (VMs) to operate as edge routers of the virtual network to connect machines outside the plurality of public cloud data centers; identifying, based on a set of one or more optimization criteria, different paths through different pairs of edge router VMs operating as ingress routers / egress routers in the plurality of public cloud data centers for a flow of data messages through the virtual network, wherein each path originates and terminates at a machine associated with the entity located outside the plurality of public cloud data centers and traverses at least (i) an ingress VM router deployed in a first public cloud data center and (ii) an egress VM router deployed in a second public cloud data center; defining routing data based on the identified path, the routing data for instructing a plurality of edge router VMs deployed in a set of two or more public cloud data centers to route the data message flow along the identified path through the virtual network, wherein the routing data includes a plurality of next-hop routing records identifying a next hop along each identified path between an ingress VM router at an ingress public cloud data center for the data message flow and an egress VM router at an egress public cloud data center for the data message flow; as well as The routing data is distributed to the deployed edge router VMs. 2 . The method of claim 1 , wherein the set of optimization criteria for a path includes the length of the path.

3. The method of claim 1 , wherein the set of optimization criteria for a path comprises a financial cost of the path, the financial cost comprising a set of costs charged to the edge router VM by a set of public cloud providers of a set of public cloud data centers through which the path is operated. 4 . The method of claim 1 , wherein the set of optimization criteria for a path includes minimizing a portion of the path that traverses the public Internet between different pairs of public cloud data centers.

5. The method of claim 1, wherein the set of optimization criteria for a path comprises maximizing the portion of the path that traverses a private public cloud connection between a pair of public cloud data centers traversed by the path. The method of claim 1 , wherein the edge router VM executes on a host computer in the public cloud data center.

7. The method of claim 6, wherein at least one set of the edge router VMs executes on a set of host computers in the public cloud data center along with other virtual machines that perform computing operations.

8. The method of claim 1, wherein the public cloud data center is a multi-tenant public cloud data center.

9. The method of claim 1 , wherein the external machine associated with the entity comprises: Remote user machines and machines at the entity's private data centers and office locations.

10. The method of claim 1, wherein the external machine associated with the entity further comprises: A machine at a private data center of the entity's Software as a Service (SaaS) provider.

11. The method of claim 1 , wherein the set of optimization criteria considers a set of costs for each hop along the identified path, the set of costs comprising at least one of: (a) the data message throughput rate of the network connection between the two forwarding elements corresponding to the hop, (b) Data message loss metrics for network connections, (c) financial costs associated with network connectivity, and (d) Financial charges associated with data computation nodes associated with one or both of the forwarding elements.

12. The method of claim 11, wherein the financial charges for the network connection or the data computing node are charges charged by one or more public cloud providers.

Citation Information

Patent Citations

  • Providing domain-joined remote applications in a cloud environment

    CN105378671A

  • Network system and method for improving routing capability

    US20140112171A1