Virtual layer 2 network
By integrating a virtual Layer 3 network with a virtual Layer 2 network and utilizing a VLAN Switching and Routing Service, the limitations of existing virtual networks are overcome, enhancing communication efficiency and scalability.
Patent Information
- Application Number
- JP2025157965
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-06
AI Technical Summary
Existing virtual networks in cloud computing environments have limitations that restrict their functionality and value, particularly in managing virtual Layer 2 networks.
Implementing a virtual Layer 3 network hosted by an underlying physical network and a virtual Layer 2 network, utilizing a VLAN with multiple endpoints, L2VNICs, and switches, along with a VLAN Switching and Routing Service (VSRS) to manage packet delivery and access control lists (ACLs) across virtual networks.
Enhances the functionality and scalability of virtual networks by enabling efficient packet delivery and access control across multiple virtual networks, ensuring reliable communication and isolation between tenants.
Smart Images

Figure 2026001078000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following applications: (1) U.S. Provisional Application No. 63 / 051,728, filed July 14, 2020, entitled "VLAN Switching and Routing Services and Layer 2 Virtual Networking in a Virtual Cloud Environment," and (2) U.S. Provisional Application No. 63 / 132,377, filed December 30, 2020, entitled "Layer 2 Virtual Networking in a Virtual Cloud Environment." The entire contents of the above-referenced provisional applications are incorporated by reference into this disclosure for all purposes.
[0002] This application is also related to U.S. Application No. 17 / 376002 (Attorney Docket No. 088325-1256547-276510US), filed on July 14, 2021, entitled "System and Method for VLAN Switching and Routing Services," and U.S. Application No. 17 / 376004 (Attorney Docket No. 088325-1256549-276520US), filed on July 14, 2021, entitled "Interface-Based ACLs in Layer 2 Networks," the entire contents of each of which are incorporated by reference into this disclosure for all purposes. [Background technology]
[0003] background Cloud computing provides on-demand availability of computing resources. Cloud computing can be based on data centers available to users over the Internet. Cloud computing can provide Infrastructure as a Service (IaaS). Virtual networks provide users with However, these virtual networks have limitations that limit their functionality and value. Therefore, further improvements are desirable. Summary of the Invention [Means for solving the problem]
[0004] overview One aspect of the present disclosure relates to a computer-implemented method that includes providing a virtual Layer 3 network in a virtualized cloud environment that is hosted by an underlying physical network, and providing a virtual Layer 2 network in the virtualized cloud environment that is hosted by the underlying physical network.
[0005] In some embodiments, the virtual Layer 2 network may be a virtual local area network (VLAN). In some embodiments, the VLAN includes multiple endpoints. In some embodiments, the multiple endpoints may be multiple compute instances. In some embodiments, the VLAN includes multiple L2 virtual network interface cards (L2VNICs) and multiple switches.
[0006] In some embodiments, each of the plurality of compute instances is communicatively connected to a pair including a unique L2 virtual network interface card (L2VNIC) and a unique switch. In some embodiments, the plurality of switches together can form a distributed switch. In some embodiments, each of the plurality of switches routes outbound traffic according to a mapping table received from the L2VNIC paired to the switch. In some embodiments, the mapping The mapping table specifies the interface-to-MAC address mapping of endpoints within a VLAN.
[0007] In some embodiments, the method further includes instantiating, on a network virtualization device (NVD), a pair including a unique L2VNIC and a unique switch. In some embodiments, the method includes receiving, on a unique L2VNIC of one of the plurality of compute instances, a packet addressed to one of the plurality of compute instances from another endpoint in the VLAN, and learning, with the unique L2VNIC of the one of the plurality of compute instances, a mapping of the other endpoint. In some embodiments, the mapping of the other endpoint includes an interface-to-MAC address mapping of the other endpoint.
[0008] In some embodiments, the method includes decapsulating the received packet using a unique L2VNIC of one of the plurality of compute instances and forwarding the decapsulated packet to one of the plurality of compute instances. In some embodiments, the method includes learning, using one of the plurality of compute instances, an IP address-to-MAC address mapping of the other endpoint.
[0009] In some embodiments, the method includes transmitting an IP packet from a first compute instance in a VLAN, the IP packet including a destination IP address of a second compute instance in the VLAN, receiving the IP packet on a first L2 VNIC associated with the first compute instance, encapsulating the IP packet on the first L2 VNIC, and forwarding the IP packet to the second compute instance via a first switch. In some embodiments, the first switch and the first L2 VNIC are communicatively coupled in a pair to the first compute instance. In some embodiments, the method further includes receiving the IP packet on a second VNIC associated with the second compute instance, decapsulating the IP packet on the second VNIC, and forwarding the IP packet from the second VNIC to the second compute instance.
[0010] In some embodiments, the virtual Layer 2 network includes a plurality of virtual local area networks (VLANs). In some embodiments, each of the plurality of VLANs includes a plurality of endpoints. In some embodiments, the plurality of VLANs includes a first VLAN and a second VLAN. In some embodiments, the first VLAN includes a plurality of first endpoints and the second VLAN includes a plurality of second endpoints. In some embodiments, each of the plurality of VLANs has a unique identifier. In some embodiments, one of the plurality of first endpoints in the first VLAN communicates with one of the plurality of second endpoints in the second VLAN.
[0011] One aspect of the present disclosure relates to a system including a physical network including at least one host machine and at least one network virtualization device, wherein the physical network can provide a virtual Layer 3 network hosted by the underlying physical network in a virtualized cloud environment, and can provide a virtual Layer 2 network hosted by the underlying physical network in the virtualized cloud environment.
[0012] One aspect of the present disclosure relates to a computer-readable storage medium storing instructions executable by one or more processors, the instructions, when executed by the one or more processors, causing the one or more processors to provide a virtual Layer 3 network hosted by an underlying physical network in a virtualized cloud environment, In this way, a virtual Layer 2 network is provided that is hosted by the underlying physical network.
[0013] One aspect of the present disclosure relates to a method that includes generating a table for an instance of a VLAN Switching and Routing Service (VSRS), the VSRS connecting a first virtual Layer 2 network to a second network. In some embodiments, the table includes information for identifying an IP address, a MAC address, and a virtual interface identifier of an instance in the first virtual Layer 2 network. The method includes receiving, using the VSRS, a packet to be delivered from the first instance to a second instance in the virtual Layer 2 network; identifying, using the VSRS, a second instance in the virtual Layer 2 network for delivering the packet based on the information received with the packet and the information included in the table; and delivering the packet to the identified second instance.
[0014] In some embodiments, the first virtual Layer 2 network includes a plurality of instances. In some embodiments, the first virtual Layer 2 network includes a plurality of L2 virtual network interface cards (L2VNICs) and a plurality of switches. In some embodiments, each of the plurality of instances is communicatively connected to a pair including a unique L2 virtual network interface card (L2VNIC) and a unique switch.
[0015] In some embodiments, using the VSRS to identify a second instance in the first virtual Layer 2 network for delivering the packet based on information received with the packet and information contained in the table includes: using the VSRS to determine that the table does not include mapping information for the second instance; suspending delivery of the packet by the VSRS; and using the VSRS to broadcast an ARP request to a VNIC in the first virtual Layer 2 network, the ARP request including an IP address of the second instance; and using the VSRS to receive an ARP response from an L2 VNIC of the second instance.
[0016] In some embodiments, the method further includes updating the table based on the received ARP response. In some embodiments, the first instance is located outside the first virtual Layer 2 network and inside the second network. In some embodiments, the second network may be an L3 network. In some embodiments, the second network may be a second virtual Layer 2 network. In some embodiments, the table is generated based on information received by the VSRS.
[0017] In some embodiments, the method includes instantiating the VSRS as a service on multiple hardware nodes. In some embodiments, the method includes distributing a table among the hardware nodes. In some embodiments, the table distributed among the hardware nodes is available to other VSRS instantiations. In some embodiments, the first instance is within a first virtual Layer 2 network.
[0018] In some embodiments, the method includes receiving a packet from a third instance in the first virtual Layer 2 network using the VSRS. In some embodiments, the packet is delivered to a fourth instance outside the first virtual Layer 2 network. In some embodiments, the method includes receiving a packet from a third instance in the first virtual Layer 2 network using the VSRS. In some embodiments, the packet is delivered to a service used by a third instance in the first virtual Layer 2 network. In some embodiments, the service may be at least one of DHCP, NTP, and DNS.
[0019] In some embodiments, the method includes receiving, using the VSRS, a packet from a third instance in the first virtual Layer 2 network. In some embodiments, the packet is delivered to a fourth instance in the second virtual Layer 2 network. In some embodiments, the method further includes distributing a table for the instance of the VSRS with Layer 2 and Layer 3 network information across multiple service nodes to provide highly reliable and highly scalable VSRS instantiation. In some embodiments, the method includes receiving, using the VSRS, a packet from a third instance in the first virtual Layer 2 network and learning, using the VSRS, a mapping for the third instance.
[0020] One aspect of the present disclosure relates to a system. The system includes a physical network. The physical network includes at least one processor and a network virtualization device. The at least one processor can instantiate an instance of a VLAN Switching and Routing Service (VSRS), the VSRS connecting a first virtual Layer 2 network to a second network, and the at least one processor can generate a table for the instance of the VSRS. In some embodiments, the table includes information for identifying an IP address, a MAC address, and a virtual interface identifier of an instance in the first virtual Layer 2 network. The at least one processor can receive, using the VSRS, a packet to be delivered from the first instance to a second instance in the virtual Layer 2 network, and, using the VSRS, identify the second instance in the first virtual Layer 2 network for delivery of the packet based on the information received with the packet and the information included in the table, and deliver the packet to the identified second instance.
[0021] One aspect of the present disclosure relates to a computer-readable storage medium storing a plurality of instructions executable by one or more processors. The plurality of instructions, when executed by the one or more processors, cause the one or more processors to instantiate an instance of a VLAN Switching and Routing Service (VSRS), the VSRS connecting a first virtual Layer 2 network to a second network, and generate a table for the instance of the VSRS. In some embodiments, the table includes information for identifying an IP address, a MAC address, and a virtual interface identifier of an instance in the first virtual Layer 2 network. The plurality of instructions, when executed by the one or more processors, cause the one or more processors to receive, using the VSRS, a packet to be delivered from the first instance to a second instance in the virtual Layer 2 network, identify a second instance in the virtual Layer 2 network for delivering the packet based on information received with the packet and information included in the table, and deliver the packet to the identified second instance.
[0022] One aspect of the present disclosure relates to a method, including: transmitting a packet from a source compute instance to a destination compute instance in a virtual network through a destination L2 virtual network interface card (destination L2 VNIC) in a first virtual Layer 2 network; evaluating an access control list (ACL) of the packet using the source virtual network interface card (source VNIC); embedding ACL information associated with the packet in the packet; forwarding the encapsulated packet to a virtual switching and routing service (VSRS) for connecting the first virtual Layer 2 network (VLAN) to a second network; and, using the VSRS, forwarding the encapsulated packet along with the packet. The method includes identifying a destination L2VNIC in the first virtual Layer 2 network for delivering the packet based on the received information and mapping information included in the mapping table, obtaining ACL information from the packet using the VSRS, and applying the obtained ACL information to the packet.
[0023] In some embodiments, the packets include IP packets. In some embodiments, the source computing instance is located in a virtual L3 network. In some embodiments, the source computing instance is located in a second virtual Layer 2 network.
[0024] In some embodiments, the method includes encapsulating the packet using a source VNIC. In some embodiments, the method includes receiving and decapsulating the packet using a VSRS. In some embodiments, identifying a destination L2 VNIC in a first virtual Layer 2 network for delivering the packet based on information received with the packet and mapping information included in a mapping table using the VSRS includes determining, using the VSRS, that the mapping table does not include mapping information for the destination compute instance, suspending forwarding of the packet using the VSRS, broadcasting, using the VSRS, an ARP request including an IP address of the destination compute instance to an L2 VNIC in the first virtual Layer 2 network, and receiving, using the VSRS, an ARP response from the L2 VNIC of the destination compute instance. In some embodiments, one of the L2 VNICs is the L2 VNIC of the destination compute instance.
[0025] In some embodiments, the method includes updating the table based on the received ARP response. In some embodiments, using the VSRS to identify a destination L2 VNIC in the first virtual Layer 2 network for delivering the packet based on information received with the packet and mapping information included in the mapping table includes determining that the mapping table includes mapping information for the destination compute instance, and identifying the destination L2 VNIC based on the mapping information included in the mapping table. In some embodiments, embedding ACL information associated with the packet in the packet includes storing the ACL information as metadata in the packet. In some embodiments, obtaining the ACL information for the packet using the VSRS includes extracting metadata including the ACL information in the packet.
[0026] In some embodiments, applying the obtained ACL information to the packet includes determining that the ACL information is not associated with the destination L2VNIC. In some embodiments, applying the obtained ACL information to the packet further includes forwarding the packet to the destination compute instance via the destination L2VNIC. In some embodiments, applying the obtained ACL information to the packet includes determining, using a VSRS, that the ACL information is associated with the destination L2VNIC. In some embodiments, applying the obtained ACL information to the packet further includes determining, using a VSRS, that the destination L2VNIC complies with the ACL information, and forwarding, using a VSRS, the packet to the destination compute instance via the destination L2VNIC.
[0027] In some embodiments, applying the obtained ACL information to the packet includes determining, with the VSRS, that the destination L2 VNIC does not comply with the ACL information, and the VSRS discarding the packet. In some embodiments, applying the obtained ACL information to the packet further includes, with the VSRS, sending a response to the source compute instance indicating the discarding of the packet.
[0028] One aspect of the present disclosure relates to a system including a physical network. The physical network includes at least one first processor, a network virtualization device, and at least one second processor. The at least one processor can transmit a packet from a source compute instance in a virtual network instantiated on the physical network to a destination compute instance in a virtual network instantiated on the physical network via a destination L2 virtual network interface card (destination L2 VNIC) in a first virtual Layer 2 network instantiated on the physical network. The network virtualization device can instantiate a source VNIC. The source VNIC can evaluate an access control list (ACL) of the packet, embed ACL information associated with the packet in the packet, and forward the packet to a virtual switching and routing service (VSRS), which connects the first virtual Layer 2 network (VLAN) to a second network. The at least one second processor can instantiate the VSRS. The VSRS can identify a destination L2 VNIC for delivering the packet based on information received with the packet and mapping information included in a mapping table, obtain ACL information from the packet, and apply the obtained ACL information to the packet.
[0029] In some embodiments, applying the obtained ACL information to the packet includes determining that the ACL information is associated with a destination L2VNIC, determining that the destination L2VNIC complies with the ACL information, and forwarding the packet to a destination compute instance via the destination L2VNIC using the VSRS.
[0030] One aspect of the present disclosure relates to a computer-readable storage medium storing a plurality of instructions executable by one or more processors, which, when executed by the one or more processors, cause the one or more processors to send a packet from a source compute instance to a destination compute instance in a virtual network through a destination L2 virtual network interface card (destination L2 VNIC) in a first virtual Layer 2 network, evaluate an access control list (ACL) of the packet using the source virtual network interface card (source VNIC), embed ACL information associated with the packet in the packet, forward the packet to a virtual switching and routing service (VSRS) for connecting the first virtual Layer 2 network (VLAN) to a second network, identify, using the VSRS, a destination L2 VNIC in the first virtual Layer 2 network for delivering the packet based on information received with the packet and mapping information included in a mapping table, obtain the ACL information from the packet, and apply the obtained ACL information to the packet. [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a high-level diagram of a distributed environment illustrating a virtual or overlay cloud network hosted by a cloud service provider infrastructure, according to certain embodiments. [Figure 2] 1 is an architectural schematic diagram illustrating physical elements of a physical network within a CSPI, according to certain embodiments. [Figure 3] FIG. 1 illustrates an exemplary arrangement of CSPI in which a host machine is connected to multiple network virtualization devices (NVDs), according to certain embodiments. [Figure 4] FIG. 1 illustrates a connection between a host machine and an NVD that provides I / O virtualization to support multi-tenancy, according to certain embodiments. [Figure 5]1 is a schematic block diagram illustrating a physical network provided by CSPI in accordance with certain embodiments. [Figure 6] FIG. 1 is a schematic diagram illustrating one embodiment of a computing network. [Figure 7] FIG. 1 is a logical and hardware schematic diagram of a virtual local area network (VLAN). [Figure 8] FIG. 1 is a logical schematic diagram illustrating multiple connected L2 VLANs. [Figure 9] FIG. 1 is a logical schematic diagram illustrating multiple connected L2 VLANs and subnets. [Figure 10] FIG. 1 is a schematic diagram illustrating one embodiment of intra-VLAN communication and learning within a VLAN. [Figure 11] FIG. 1 is a schematic diagram illustrating one embodiment of a VLAN implementation. [Figure 12] 10 is a flow chart illustrating one embodiment of a process for performing intra-VLAN communication. [Figure 13] FIG. 1 is a schematic diagram illustrating a process for performing intra-VLAN communication. [Figure 14] 1 is a flow diagram illustrating one embodiment of a process for inter-VLAN communication within a virtual L2 network. [Figure 15] FIG. 1 is a schematic diagram illustrating a process for performing inter-VLAN communication. [Figure 16] 10 is a flow diagram illustrating one embodiment of a process for performing a receive packet flow. [Figure 17] FIG. 1 is a schematic diagram illustrating a process for receiving communications. [Figure 18] 10 is a flow diagram illustrating one embodiment of a process for transmitting packet flow from a VLAN. [Figure 19] FIG. 1 is a schematic diagram illustrating a process for performing a transmit packet flow. [Figure 20]1 is a flow chart illustrating one embodiment of a process for performing deferred access control list (ACL) classification. [Figure 21] 1 is a flow chart illustrating an embodiment of a process for early classification of an ACL. [Figure 22] 1 is a flow diagram illustrating one embodiment of a process for sender-based next-hop routing. [Figure 23] 1 is a flow diagram illustrating one embodiment of a process for performing deferred next-hop routing. [Figure 24] FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 25] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 26] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 27] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 28] FIG. 1 is a block diagram illustrating an exemplary computer system in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0032] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of particular embodiments. It will be apparent, however, that various embodiments may be practiced without these specific details. The drawings and description are not intended to be limiting. The term "exemplary" is used in this disclosure to mean "serving as an example, instance, or illustration." Any embodiment or design described in this disclosure as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0033] Example Virtual Network Architecture The term cloud services generally refers to the services that a cloud service provider (CSP) provides to users or customers using systems and infrastructure (cloud infrastructure) that are available on demand (e.g., via a subscription model). Cloud services refer to services that can be accessed from a CSP's cloud infrastructure. Typically, the servers and systems that make up a CSP's infrastructure are separate from the customer's own on-premise servers and systems. Therefore, customers can use cloud services provided by CSPs without having to purchase separate hardware and software resources for the services. Cloud services are designed to provide subscribing customers with easy and scalable access to applications and computing resources without the customer having to invest in procuring the infrastructure used to deliver the services.
[0034] Several cloud service providers (CSPs) offer various types of cloud services, including a variety of different types or models such as Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS).
[0035] A customer can subscribe to one or more cloud services offered by a CSP. A customer can be any entity, such as an individual, an organization, or a business. When a customer subscribes or registers for a service offered by a CSP, a tenant or account is created for that customer. The customer can then access one or more subscribed cloud resources associated with the account through this account.
[0036] As mentioned above, IaaS (Infrastructure as a Service) is a service that provides services to a specific type of client. Cloud computing services. In the IaaS model, a CSP provides infrastructure (called cloud service provider infrastructure or CSPI) that customers can use to build their own customizable networks and deploy customer resources. Therefore, customer resources and networks are hosted in a distributed environment by the infrastructure provided by the CSP. This differs from traditional computing, where the customer's infrastructure hosts the customer's resources and networks.
[0037] CSPI may include interconnected high-performance computing resources, including various host machines, memory resources, and network resources, forming a physical network, also known as a substrate network or underlay network. CSPI resources may be distributed across one or more data centers, geographically dispersed across one or more geographic regions. Virtualization software can run on these physical resources to provide a virtualized distributed environment. Virtualization creates an overlay network (also known as a software-based network, software-defined network, or virtual network) on top of the physical network. The CSPI physical network provides the basis for creating one or more overlay or virtual networks on top of the physical network. The physical network (or substrate network or underlay network) includes physical network devices such as physical switches, routers, computers, and host machines. An overlay network is a logical (or virtual) network that operates on top of the physical substrate network. A given physical network can support one or more overlay networks. Overlay networks typically use encapsulation techniques to distinguish traffic belonging to different overlay networks. A virtual network or overlay network is also known as a virtual cloud network (VCN). Virtual networks are implemented using software virtualization technologies (e.g., hypervisors, virtualization functions implemented by network virtualization devices (NVDs) (e.g., smart NICs), top-of-rack (TOR) switches, smart TORs that implement one or more of the functions performed by NVDs, and other mechanisms) to create a network abstraction layer that can run on top of a physical network. Virtual networks can take various forms, such as peer-to-peer networks, IP networks, etc. A virtual network is typically either a Layer 3 IP network or a Layer 2 VLAN. Such a virtual network or overlay network is often called a virtual Layer 3 network or an overlay Layer 3 network. Examples of protocols developed for virtual networks are IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN - IETF RFC7348), virtual private networks (VPNs) (e.g., MPLS Layer 3 Virtual Private Networks (RFC4364)), VMware NSX, Generic Network Virtualization Encapsulation (GENEVE), etc.
[0038] In IaaS, the infrastructure provided by the CSP (CSPI) may be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing service provider may host infrastructure elements (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may offer various services (e.g., billing, monitoring, logging, security, load balancing, clustering, etc.) associated with those infrastructure elements. Because these services are policy-driven, IaaS users can maintain application availability and performance by implementing policies to drive load balancing. The CSPI provides infrastructure and a set of complementary cloud services, enabling customers to build and run a wide range of applications and services in a highly available, hosted, distributed environment. The CSPI provides high-performance computing resources and power, as well as storage capacity, over a flexible virtual network that can be securely accessed from various network locations, such as customer on-premises networks. When a customer subscribes or registers for an IaaS service offered by a CSP, the tenancy created for that customer is a secure, isolated partition from the CSP where the customer can create, organize, and manage their cloud resources.
[0039] Customers can build their own virtual networks using the compute, memory, and networking resources provided by CSPI. On these virtual networks, they can connect one or more customer resources or workloads, such as compute instances. For example, a customer can use resources provided by CSPI to build one or more customizable private virtual networks called virtual cloud networks (VCNs). The customer can then deploy one or more customer resources, such as compute instances, on the customer VCN. The compute instances may be virtual machines, bare metal instances, etc. Thus, CSPI provides infrastructure and a set of complementary cloud services that enable customers to build and run various applications and services in a highly available, virtualized host environment. While customers do not manage or control the underlying physical resources provided by CSPI, they do control the operating systems, storage, and deployed applications, and in some cases have limited control over some networking components (e.g., firewalls).
[0040] The CSP may provide a console that enables customers and network administrators to configure, access, and manage resources deployed in the cloud using CSPI resources. In certain embodiments, the console provides a web-based user interface that can be used to utilize and manage CSPI. In some embodiments, the console is a web-based application provided by the CSP.
[0041] CSPI can support single-tenancy or multi-tenancy architectures. In a single-tenancy architecture, software (e.g., applications, databases) or hardware elements (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenancy architecture, software or hardware elements serve multiple customers or tenants. Thus, in a multi-tenancy architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenancy environment, CSPI employs precautions and safeguards to ensure that each tenant's data is isolated and not visible to other tenants.
[0042] In a physical network, a network endpoint (endpoint) refers to a computing device or system that is connected to the physical network and communicates bidirectionally with the connected network. A network endpoint of a physical network may be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints of a physical network include modems, hubs, bridges, switches, routers, and other networking devices, physical computers (or host machines), etc. Each physical device of a physical network has a fixed network address that can be used to communicate with the device. This fixed network address may be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints may include various virtual endpoints, such as virtual machines hosted by elements of the physical network (e.g., hosted by a physical host machine). These endpoints of the virtual network are addressed by overlay addresses, such as overlay Layer 2 addresses (e.g., an overlay MAC address) and overlay Layer 3 addresses (e.g., an overlay IP address). Network overlays achieve flexibility by allowing network administrators to move overlay addresses associated with network endpoints using software management (e.g., via software implementing the virtual network's control plane). Thus, unlike physical networks, in virtual networks, overlay addresses (e.g., overlay IP addresses) can be moved from one endpoint to another using network management software. Because virtual networks are built on physical networks, both the virtual network and the underlying physical network are involved in communications between elements of the virtual network.To facilitate such communications, each element of the CSPI is configured to learn and store mappings that map overlay addresses in the virtual network to real physical addresses in the substrate network, or real physical addresses in the substrate network to overlay addresses in the virtual network. These mappings are used to facilitate communications. To facilitate virtual network routing, customer traffic is encapsulated.
[0043] Thus, physical addresses (e.g., physical IP addresses) are associated with elements of a physical network, and overlay addresses (e.g., overlay IP addresses) are associated with entities of a virtual network. A physical IP address is an IP address associated with a physical device (e.g., network device) in the substrate or physical network. For example, each NVD has an associated physical IP address. An overlay IP address is an overlay address associated with an entity in the overlay network, such as a compute instance in a customer's virtual cloud network (VCN). Two different customers or tenants, each with their own private VCN, can potentially use the same overlay IP address in their VCNs without knowing about each other. Both physical IP addresses and overlay IP addresses are real IP addresses. These are distinct from virtual IP addresses, which are typically a single IP address that represents or maps to multiple real IP addresses. A virtual IP address provides a one-to-many mapping between the virtual IP address and multiple real IP addresses. For example, a load balancer may use a VIP to map to or represent multiple servers, each with its own real IP address.
[0044] A cloud infrastructure or CSPI is physically hosted in one or more data centers in one or more regions around the world. A CSPI may include elements of a physical or substrate network and virtualized elements (e.g., virtual networks, compute instances, virtual machines) of a virtual network built on top of the physical network elements. In certain embodiments, a CSPI may be organized into realms, regions, and CSPI resources are organized and hosted in available domains. A region is typically a local geographic area containing one or more data centers. Regions are generally independent of one another and may be separated by vast distances, e.g., across countries or continents. For example, a first region may be in Australia, another region may be in Japan, and yet another region may be in India. CSPI resources are divided among these regions so that each region has an independent subset of CSPI resources. Each region can provide a set of core infrastructure services and resources, such as compute resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure), storage resources (e.g., block volume storage, file storage, object storage, archive storage), networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connectivity to on-premises networks), database resources, edge networking resources (e.g., DNS), access management, and monitoring resources. Each region typically has multiple routes connecting it to other regions within the realm.
[0045] Typically, applications are deployed in the region where they will be used the most (i.e., on infrastructure associated with that region) because using nearby resources is faster than using resources that are farther away. Applications may also be deployed in different regions for a variety of reasons, such as redundancy to mitigate the risk of region-wide events such as large weather systems or earthquakes, or to meet various requirements for legal jurisdictions, tax domains, and other business or societal criteria.
[0046] Data centers within a region may be further organized and subdivided into availability domains (ADs). An availability domain is one or more data centers located in a region. A region may correspond to a data center in the cloud. A region may be composed of one or more availability domains. In such a distributed environment, CSPI resources may be region-specific, such as a virtual cloud network (VCN), or availability domain-specific, such as a compute instance.
[0047] ADs within a region are isolated from each other to be fault tolerant, and are configured to have a very low probability of simultaneous failure. This is achieved by configuring ADs so that they do not share critical infrastructure resources such as networking, physical cables, cable routes, and cable entrances, so that a failure of one AD in a region has little impact on the availability of other ADs in the same region. Connecting ADs within the same region to each other with low-latency, high-bandwidth networks provides highly available connections to other networks (e.g., the Internet, customer on-premises networks), and multiple ADs can be used to build replicated systems for both high availability and disaster recovery. Cloud services utilize multiple ADs to provide high availability. This ensures availability and protects against resource failure. As the infrastructure provided by the IaaS provider grows, more regions and ADs may be added along with additional capacity. Traffic between available domains is typically encrypted.
[0048] In certain embodiments, regions are grouped into realms. A realm is a logical collection of regions. Realms are isolated from each other and do not share any data. Regions within the same realm can communicate with each other, but regions within different realms cannot. A CSP's customer tenancy or account exists in a single realm and can span one or more regions within that single realm. Typically, when a customer subscribes to an IaaS service, their tenancy or account is created in a customer-specified region (called their "home" region) within a realm. The customer can extend their tenancy to one or more other regions within the realm. The customer cannot access regions that do not exist within the realm in which the customer's tenancy resides.
[0049] An IaaS provider may offer multiple realms, each corresponding to a particular set of customers or users. For example, a commercial realm may be offered for commercial customers. As another example, a realm may be offered for a particular country or for customers in that country. As yet another example, a government realm may be offered, for example, for a government. For example, a government realm may be created for a particular government and may have a higher security level than a commercial realm. For example, Oracle® Cloud Infrastructure (OCI) currently offers a realm for the commercial domain and two realms for the government cloud domain (e.g., FedRAMP-authorized and IL5-authorized).
[0050] In certain embodiments, an AD can be subdivided into one or more fault domains. A fault domain is a grouping of infrastructure resources within an AD to provide anti-affinity. Fault domains can distribute compute instances so that compute instances are not located on the same physical hardware within an AD. This is known as anti-affinity. A fault domain refers to a collection of hardware elements (computers, switches, etc.) that share a single point of failure. A compute pool is logically divided into fault domains. Thus, a hardware failure or compute hardware maintenance event that affects one fault domain does not affect instances in other fault domains. Depending on the embodiment, the number of fault domains in each AD may vary. For example, in certain embodiments, each AD includes three fault domains. Fault domains function as logical data centers within an AD.
[0051] When a customer subscribes to an IaaS service, resources from CSPI are provisioned to the customer and associated with the customer's tenancy. Customers can use these provisioned resources to build private networks and deploy resources on these networks. A customer network hosted on the cloud by CSPI is called a Virtual Cloud Network (VCN). Customers can configure one or more Virtual Cloud Networks (VCNs) using the CSPI resources allocated for the customer. A VCN is a virtual or software-defined private network. Customer resources deployed in a customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances may represent various customer workloads, such as applications, load balancers, and databases. Compute instances deployed on a VCN can communicate with publicly accessible endpoints (public endpoints) over public networks such as the Internet, with other instances in the same VCN or other VCNs (e.g., other VCNs of the customer or VCNs not belonging to the customer), and with customer on-premises data centers or networks. It can communicate with service endpoints and can communicate with other types of endpoints.
[0052] CSPs can offer various services using CSPI. In some cases, customers of a CSPI themselves can act as service providers and provide services using CSPI resources. Service providers can expose service endpoints characterized by identifying information (e.g., IP addresses, DNS names, and ports). Customer resources (e.g., compute instances) can consume a particular service by accessing the service endpoint for that particular service exposed by the service. These service endpoints are generally publicly accessible over a public communications network, such as the Internet, with users using the public IP address associated with the endpoint. Publicly accessible network endpoints are sometimes referred to as public endpoints.
[0053] In certain embodiments, a service provider may expose a service through a service endpoint (sometimes referred to as a service endpoint). Customers of the service may access the service using this service endpoint. In certain embodiments, a service endpoint provided for a service may be accessed by multiple customers wishing to consume the service. In other implementations, a dedicated service endpoint may be provided to a customer. Thus, only that customer may access the service using that dedicated service endpoint.
[0054] In certain embodiments, when a VCN is created, it is assigned a private overlay IP address range (e.g., 10.0 / 16) called Private Overlay Classless Inter-Domain Routing (PIR). A VCN is associated with a CIDR (Communication Override) address space. It contains associated subnets, route tables, and gateways. A VCN exists within a single region but can extend to one or more or all available domains in the region. A gateway is a virtual interface configured for a VCN that enables traffic communication between the VCN and one or more endpoints outside the VCN. You can configure one or more different types of gateways for a VCN to enable communication between different types of endpoints.
[0055] A VCN may be subdivided into one or more subnetworks, such as one or more subnets. A subnet is thus a building block or division that can be created within a VCN. A VCN can have one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets within the VCN and represents a subset of the VCN's address space.
[0056] Each compute instance is associated with a virtual network interface card (VNIC), which allows it to participate in a subnet of a VCN. A VNIC is a logical representation of a physical network interface card (NIC). In general, a VNIC is an interface between an entity (e.g., compute instance, service) and a virtual network. A VNIC resides in a subnet and has one or more associated IP addresses and associated security rules or policies. A VNIC is equivalent to a Layer 2 port on a switch. A VNIC connects a compute instance to a subnet within a VCN. A VNIC allows a compute instance to be part of a subnet in a VCN and enables the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, with endpoints in a different subnet within the VCN, or with endpoints outside the VCN. Thus, the VNIC associated with a compute instance determines how the compute instance connects with endpoints inside and outside the VCN. A VNIC for a compute instance is created and associated with the compute instance when the compute instance is created and added to a subnet in a VCN. If a subnet consists of a set of compute instances, it contains VNICs corresponding to the set of compute instances, and each VNIC is connected to a compute instance in the set of compute instances.
[0057] Each compute instance is assigned a private overlay IP address via the VNIC associated with the compute instance. This private overlay IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic for the compute instance. All VNICs in a particular subnet use the same route table, security lists, and DHCP options. As described above, each subnet in a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that do not overlap with other subnets in the VCN and represent an address space subset of the VCN's address space. For a VNIC on a particular subnet of a VCN, the overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.
[0058] In certain embodiments, if desired, a compute instance can be assigned additional overlay IP addresses in addition to the private overlay IP address, for example, one or more public IP addresses in the case of a public subnet. These multiple addresses are assigned to the same VNIC or multiple VNICs associated with the compute instance. However, each instance has a primary VNIC associated with the overlay private IP address that is created and assigned to the instance at instance launch. This primary VNIC cannot be deleted. Additional VNICs, called secondary VNICs, can be added to an existing instance within the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. Secondary VNICs can be in the same subnetwork in the same VCN as the primary VNIC, or in different subnetworks in the same or a different VCN.
[0059] Compute instances can optionally be assigned public IP addresses if they are in a public subnet. When creating a subnet, you can specify that the subnet is either a public or private subnet. A private subnet means that resources (e.g., compute instances) and associated VNICs within the subnet cannot have public overlay IP addresses. A public subnet means that resources and associated VNICs within the subnet can have public IP addresses. Customers can specify subnets that exist across a single available domain or multiple available domains within a region or realm.
[0060] As mentioned above, a VCN may be subdivided into one or more subnets. In certain embodiments, a virtual router (referred to as a VCN VR or simply VR) configured for a VCN enables communication between the subnets of the VCN. For subnets within a VCN, the VR separates the subnet (i.e., the compute instances on that subnet) from the VCN-internal A VCN VR represents a logical gateway for a subnet, enabling communication with endpoints on other subnets in that subnet and with other endpoints outside the VCN. A VCN VR is a logical entity configured to route traffic between VNICs in a VCN and a virtual gateway (gateway) associated with the VCN. Gateways are further described below with respect to FIG. 1. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is one VCN VR per VCN. This VCN VR potentially has an unlimited number of ports addressed by IP addresses, one port for each subnet of the VCN. In this way, a VCN VR has a different IP address for each subnet of the VCN to which the VCN VR is connected. VRs are also connected to various gateways configured for the VCN. In certain embodiments, a specific overlay IP address from a subnet's overlay IP address range is reserved for a port in the VCN VR for that subnet. For example, consider a VCN with two subnets, each with associated address ranges 10.0 / 16 and 10.1 / 16. For the first subnet of a VCN with address range 10.0 / 16, addresses from this range are reserved for ports in the VCN VR for that subnet. In some cases, the first IP address from this range may be reserved for a VCN VR. For example, for a subnet with overlay IP address range 10.0 / 16, IP address 10.0.0.1 may be reserved for a port in the VCN VR for that subnet. For a second subnet in the same VCN with address range 10.1 / 16, the VCN VR may have a port in the second subnet with IP address 10.1.0.1. A VCN VR has a different IP address for each subnet in the VCN.
[0061] In some other embodiments, each subnet within a VCN may have its own associated VR that is addressable by the subnet using a reserved or default IP address associated with the VR. The reserved or default IP address may, for example, be the first IP address from a range of IP addresses associated with the subnet. VNICs within a subnet can use this default or reserved IP address to communicate (e.g., send and receive packets) with the VR associated with the subnet. In such embodiments, a VR is the ingress / egress point for that subnet. VRs associated with subnets within a VCN can communicate with other VRs associated with other subnets within the VCN. VRs can also communicate with gateways associated with the VCN. The VR functions for a subnet are performed on or by one or more NVDs that perform the VNIC functions for VNICs within the subnet.
[0062] Route tables, security rules, and DHCP options may be configured for a VCN. A route table is a virtual route table for a VCN and contains rules for routing traffic from subnets inside the VCN to destinations outside the VCN through gateways or specially configured instances. You can customize a VCN's route table to control the forwarding / routing of packets into and out of the VCN. DHCP options refer to configuration information that is automatically provided to an instance when it is launched.
[0063] The security rules configured for a VCN represent the overlay firewall rules for the VCN. Security rules can include inbound and outbound rules and can specify the type of traffic allowed in and out of instances in the VCN (e.g., based on protocol and port). Customers can choose whether a particular rule is stateful or stateless. For example, a customer can configure a stateful inbound rule with a source CIDR of 0.0.0.0 / 0 and a destination TCP port of 22. By configuring security rules, you can allow incoming SSH traffic from anywhere to a set of instances. Security rules may be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to resources in that group, while a security list contains rules that apply to all resources in a subnet that uses that security list. A VCN may also include default security rules and a default security list. DHCP options configured for a VCN provide configuration information that is automatically provided to instances in the VCN when they launch.
[0064] In particular embodiments, configuration information for a VCN is determined and stored by a VCN control plane. The configuration information for a VCN may include, for example, address ranges associated with the VCN, subnets and associated information within the VCN, one or more VRs associated with the VCN, compute instances and associated VNICs within the VCN, NVDs that perform various virtualized network functions (e.g., VNICs, VRs, gateways) associated with the VCN, VCN state information, and other VCN-related information. In particular embodiments, a VCN distribution service publishes the configuration information stored by the VCN control plane, or portions thereof, to the NVD. The distributed information can be used to forward packets to and from compute instances within the VCN by updating information (e.g., forwarding tables, routing tables, etc.) stored and used by the NVD.
[0065] In certain embodiments, VCN and subnet creation is handled by a VCN control plane (CP), and compute instance launch is handled by the compute control plane. The compute control plane is configured to allocate physical resources for the compute instance and then invoke the VCN control plane to create and attach VNICs to the compute instance. The VCN CP also sends VCN data mappings to a VCN data plane configured to perform packet forwarding and routing functions. In certain embodiments, the VCN CP provides a distribution service configured to provide updates to the VCN data plane. Examples of the VCN control plane are shown in Figures 24, 25, 26, and 27 (see reference numerals 2416, 2516, 2616, and 2716) and described below.
[0066] Customers can create one or more VCNs with resources hosted by CSPI. Compute instances deployed on a customer VCN can communicate with different endpoints. These endpoints can include endpoints hosted by CSPI and endpoints external to CSPL.
[0067] Various different architectures for implementing cloud-based services using CSPI are shown in Figures 1, 2, 3, 4, 5, 24, 25, 26, and 28 and described below. Figure 1 is a high-level diagram of a distributed environment 100 showing an overlay VCN or customer VCN hosted by CSPI, according to certain embodiments. The distributed environment shown in Figure 1 includes multiple elements in an overlay network. The distributed environment 100 shown in Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, alternatives, and modifications are possible. For example, in some implementations, the distributed environment shown in Figure 1 may have more or fewer systems or elements than those shown in Figure 1, may combine two or more systems, or may have a different system configuration or arrangement.
[0068] As shown in the example of FIG. 1, a distributed environment 100 provides services and resources that customers can subscribe to and use to build a virtual cloud network (VCN). The CSPI 101 includes a CSPI 101 that provides infrastructure-as-a-service (IaaS) services to subscribing customers. In a particular embodiment, the CSPI 101 provides infrastructure-as-a-service (IaaS) services to subscribing customers. Data centers within the CSPI 101 may be organized into one or more regions. FIG. 1 shows an example region, "US Region" 102. A customer configures a customer VCN 104 for the region 102. A customer can deploy various compute instances on the VCN 104, which may include virtual machines or bare metal instances. Example instances include applications, databases, load balancers, etc.
[0069] In the embodiment shown in Figure 1, customer VCN 104 includes two subnets, "Subnet-1" and "Subnet-2," each with its own CIDR IP address range. In Figure 1, the overlay IP address range of Subnet-1 is 10.0 / 16, and the address range of Subnet-2 is 10.1 / 16. VCN virtual router 105 represents the logical gateway of the VCN, enabling communication between subnets in VCN 104 and with other endpoints outside the VCN. VCN VR105 is configured to route traffic between VNICs in VCN104 and gateways associated with VCN104. VCN VR105 provides a port for each subnet of VCN104. For example, VR105 may provide a port with IP address 10.0.0.1 for subnet-1 and a port with IP address 10.1.0.1 for subnet-2.
[0070] Multiple compute instances can be deployed on each subnet. In this case, the compute instances may be virtual machine instances and / or bare metal instances. The compute instances within a subnet may be hosted by one or more host machines within CSPI 101. A compute instance joins a subnet via a VNIC associated with the compute instance. For example, as shown in FIG. 1, compute instance C1 is part of subnet-1 via a VNIC associated with the compute instance. Similarly, compute instance C2 is part of subnet-1 via a VNIC associated with C2. Similarly, multiple compute instances, which may be virtual machine instances or bare metal instances, may be part of subnet-1. Each compute instance is assigned a private overlay IP address and a media access control address (MAC address) via an associated VNIC. For example, in FIG. 1, compute instance C1 has an overlay IP address of 10.0.0.2 and a MAC address of M1, and compute instance C2 has a private overlay IP address of 10.0.0.3 and a MAC address of M2. Each compute instance in Subnet-1, including compute instances C1 and C2, has a default route to VCN VR105 using IP address 10.0.0.1, which is the IP address of a port in VCN VR105 in Subnet-1.
[0071] Subnet-2 may have multiple compute instances deployed, including virtual machine instances and / or bare metal instances. For example, as shown in FIG. 1, compute instances D1 and D2 are part of subnet-2 via VNICs associated with the respective compute instances. In the embodiment shown in FIG. 1, compute instance D1 has an overlay IP address of 10.1.0.2 and a MAC address of MM1, and compute instance D2 has a private overlay IP address of 10.1.0.3 and a MAC address of MM2. Each compute instance in subnet-2, including compute instances D1 and D2, has a default route to VCN VR105 using IP address 10.1.0.1, which is the IP address of a port in VCN VR105 in subnet-2.
[0072] VCN A 104 may also include one or more load balancers. A load balancer may be provided for a subnet and configured to load balance traffic among multiple compute instances on the subnet, and a load balancer may be provided to load balance traffic among subnets within a VCN.
[0073] A particular compute instance deployed on VCN 104 can communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 101 may include endpoints on the same subnet as the particular compute instance (e.g., communication between two compute instances in Subnet-1), endpoints in a different subnet but within the same VCN (e.g., communication between a compute instance in Subnet-1 and a compute instance in Subnet-2), endpoints in a different VCN in the same region (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN in the same region 106 or 110, or communication between a compute instance in Subnet-1 and an endpoint in the service network 110 in the same region), or endpoints in a VCN in a different region (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN in a different region 108). Additionally, compute instances in a subnet hosted by CSPI 101 can communicate with endpoints not hosted by CSPI 101 (i.e., external to CSPI 101). These external endpoints include endpoints within customer on-premise networks 116, endpoints within other remote cloud host networks 118, public endpoints 114 accessible via public networks such as the Internet, and other endpoints.
[0074] Communication between compute instances on the same subnet is facilitated using VNICs associated with the source and destination compute instances. For example, compute instance C1 in Subnet-1 may want to send a packet to compute instance C2 in Subnet-1. For a packet sent from a source compute instance whose destination is another compute instance in the same subnet, the packet is first processed by the VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance may include determining the packet's destination information from the packet header, identifying any policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the packet's next hop, performing any packet encapsulation / decapsulation functions as needed, and forwarding / routing the packet to the next hop to facilitate communication of the packet to its intended destination. If the destination compute instance is in the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing. The VNIC associated with the destination compute instance then executes and forwards the packet to the destination compute instance.
[0075] When communicating a packet from a compute instance in a subnet to an endpoint in a different subnet of the same VCN, the communication is facilitated by the VNICs associated with the source and destination compute instances and the VCN VRs. For example, if compute instance C1 in Subnet-1 in Figure 1 wants to send a packet to compute instance D1 in Subnet-2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR105 using the VCN VR's default route or port 10.0.0.1. VCN VR105 then routes the packet to VCN VR105 using port 10.1.0.1. The packet is then received and processed by the VNIC associated with D1, which forwards the packet to compute instance D1.
[0076] To communicate packets from a compute instance within VCN 104 to an endpoint outside VCN 104, the communication is facilitated by a VNIC associated with the source compute instance, VCN VR 105, and a gateway associated with VCN 104. One or more types of gateways can be associated with VCN 104. A gateway is an interface between a VCN and another endpoint, where the other endpoint is outside the VCN. A gateway is a Layer 3 / IP layer concept that enables a VCN to communicate with endpoints outside the VCN. Thus, a gateway facilitates traffic flow between a VCN and other VCNs or networks. A variety of different types of gateways can be configured in a VCN to facilitate different types of communications with different types of endpoints. Through gateways, communications may occur over a public network (e.g., the Internet) or a private network. These communications may use various communication protocols.
[0077] For example, compute instance C1 may wish to communicate with an endpoint outside VCN 104. The packet may first be processed by a VNIC associated with source compute instance C1. The VNIC processing determines that the packet's destination is outside of Cl's subnet-1. The VNIC associated with C1 may forward the packet to VCN VR105 of VCN 104. VCN VR105 then processes the packet and, as part of the processing, determines a particular gateway associated with VCN 104 as the packet's next hop based on the packet's destination. VCN VR105 may then forward the packet to the particular gateway. For example, if the destination is an endpoint within a customer's operating premises network, the packet may be forwarded by VCN VR105 to dynamic routing gateway (DRG) 122 configured for VCN 104. The packet may then be forwarded from the gateway to the next hop to facilitate communication of the packet to its intended final destination.
[0078] Various different types of gateways may be configured for a VCN. Examples of gateways that may be configured for a VCN are shown in FIG. 1 and described below. Examples of gateways associated with VCNs are also shown in FIGS. 24, 25, 26, and 27 (e.g., gateways indicated by reference numbers 2434, 2436, 2438, 2534, 2536, 2538, 2634, 2636, 2638, 2734, 2736, and 2738) and described below. As shown in the embodiment shown in FIG. 1, a dynamic routing gateway (DRG) 122 may be added to or associated with the customer VCN 104. The DRG 122 provides a path for private network traffic communication between the customer VCN 104 and another endpoint. The other endpoint may be a customer on-premises network 116, a VCN 108 in a different region of the CSPI 101, or another remote cloud network 118 not hosted by the CSPI 101. The customer on-premises network 116 may be a customer network or customer data center built using customer resources. Access to the customer on-premises network 116 is typically highly restricted. For a customer that has both a customer on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, the customer may want the on-premises network 116 and the cloud-based VCNs 104 to be able to communicate with each other. This allows the customer to build an extended hybrid environment that includes the on-premises network 116 and the customer VCNs 104 hosted by CSPI 101. The DRG 122 enables such communication. To enable such communication, a communication channel A communication channel 124 is established. In this case, one endpoint of the communication channel is located in the customer on-premises network 116, and the other endpoint is located in the CSPI 101 and connected to the customer VCN 104. The communication channel 124 can be over a public communication network, such as the Internet, or a private communication network. A variety of different communication protocols can be used, such as IPsec VPN technology over a public communication network, such as the Internet, or Oracle's FastConnect technology, which uses a private network instead of a public network. The device or equipment in the customer on-premises network 116 that forms one endpoint of the communication channel 124 is called customer premises equipment (CPE), such as the CPE 126 shown in Figure 1. The endpoint on the CSPI 101 side may be a host machine running the DRG 122.
[0079] In certain embodiments, remote peering connections (RPCs) can be added to a DRG, allowing customers to peer one VCN with another VCN in another region. Using such RPCs, a customer VCN 104 can connect to a VCN 108 in another region using a DRG 122. The DRG 122 can also connect to other remote cloud networks 118 not hosted by CSPI 101, such as Microsoft® Azure Cloud, Amazon® AWS Cloud, or other cloud services. It may also be used to communicate with the cloud.
[0080] As shown in Figure 1, an Internet Gateway (IGW) 120 can be configured in a customer VCN 104 to enable compute instances on the customer VCN 104 to communicate with public endpoints 114 accessible over a public network, such as the Internet. The IGW 120 is a gateway for connecting a VCN to a public network, such as the Internet. The IGW 120 enables public subnets in a VCN, such as VCN 104 (resources in the public subnet have public overlay IP addresses) to directly access public endpoints 112 on the public network 114, such as the Internet. The IGW 120 can be used to initiate connections from subnets in the VCN 104 or from the Internet.
[0081] Customer VCN 104 can be configured with a network address translation (NAT) gateway 128. NAT gateway 128 allows cloud resources in the customer VCN that do not have dedicated public overlay IP addresses to access the Internet without exposing them to direct incoming Internet connections (e.g., L4-L7 connections). This allows private subnets in a VCN, such as private subnet-1 in VCN 104, to privately access public endpoints on the Internet. With a NAT gateway, connections can be initiated from the private subnet to the public Internet, but connections cannot be initiated from the Internet to the private subnet.
[0082] In certain embodiments, customer VCN 104 can be configured with a service gateway (SGW) 126. SGW 126 provides a path for private network traffic between VCN 104 and service endpoints supported by service network 110. In certain embodiments, service network 110 may be provided by a CSP and may offer a variety of services. An example of such a service network is the Oracle® Service Network, which offers a variety of services available to customers. For example, a compute instance (e.g., a database system) in a private subnet of customer VCN 104 can connect to a service endpoint (e.g., a public IP address) without requiring access to the Internet. For example, a VCN can back up data to an object store. In some embodiments, a VCN can have only one SGW, and connections can be initiated only from subnets within the VCN, not from the service network 110. When a VCN is peered with another VCN, resources in the other VCN typically do not have access to the SGW. Resources in an on-premises network connected to a VCN with FastConnect or VPN Connect can also use a service gateway configured in that VCN.
[0083] In some implementations, the SGW 126 uses service classless inter-domain routing (CIDR) labels. A CIDR label is a string that represents all regional public IP address ranges for a service or group of services of interest. Customers use service CIDR labels to control traffic to services when configuring the SGW and associated routing rules. Customers can optionally use service CIDR labels when configuring security rules without having to adjust the security rules if the service's public IP addresses change in the future.
[0084] A local peering gateway (LPG) 132 is a gateway that can be added to a customer VCN 104 to enable the VCN 104 to peer with another VCN in the same region. Peering means that the VCNs communicate using private IP addresses without the traffic going over a public network such as the Internet or routing the traffic through the customer on-premises network 116. In a preferred embodiment, a VCN has a separate LPG for each peering it establishes. Local peering or VCN peering is a common practice used to establish network connectivity between different applications or infrastructure management functions.
[0085] A service provider, such as a provider of a service in service network 110, can provide access to a service using different access models. According to a public access model, the service may be exposed as a public endpoint publicly accessible by a compute instance in the customer VCN over a public network such as the Internet, or may be accessed privately through SGW 126. According to a specific private access model, the service may be accessed as a private IP endpoint in a private subnet in the customer VCN. This is called private endpoint (PE) access and allows a service provider to expose its service as an instance in the customer's private network. A private endpoint resource represents a service in a customer VCN. Each PE appears as a VNIC (called a PE-VNIC, which has one or more private IPs) that the customer selects from a subnet in the customer VCN. Thus, the PE provides a way to provide a service within the customer's private VCN subnet using a VNIC. Because the endpoint is exposed as a VNIC, the PE VNIC can utilize all the functionality associated with a VNIC, such as routing rules and security lists.
[0086] Service providers register services to make them accessible through PEs. Providers can associate policies with services that regulate the visibility of the service to customer tenants. Providers can register multiple services under a single virtual IP address (VIP), especially for multi-tenant services. There can also be multiple private endpoints (in multiple VCNs) that represent the same service.
[0087] Compute instances in the private subnet can then access the service using the private IP address or service DNS name of the PE VNIC. Compute instances in the customer VCN can access the service by sending traffic to the private IP address of the PE in the customer VCN. The Private Access Gateway (PAGW) 130 is a gateway resource that can connect to a service provider VCN (e.g., a VCN in the service network 110) and serves as the ingress / egress point for all traffic from / to the customer subnet private endpoints. The PAGW 130 allows providers to scale the number of PE connections without utilizing internal IP address resources. A provider only needs to configure one PAGW for any number of services registered in a single VCN. A provider can present services as private endpoints in multiple VCNs for one or more customers. From the customer's perspective, the PE VNIC appears not to be connected to the customer's instance but to the service the customer wants to interact with. Traffic destined for the private endpoint is routed to the service through the PAGW 130. These are called customer-to-service private connections (C2S connections).
[0088] Also, the PE concept is used to ensure that traffic is routed between the FastConnect / IPsec link and the customer VCN. Private access for services can also be extended to customer on-premises networks and data centers by allowing traffic to flow through private endpoints within LPG 132. Private access for services can also be extended to customer peering VCNs by allowing traffic to flow between PEs in LPG 132 and the customer VCN.
[0089] Customers can control VCN routing at the subnet level, allowing them to specify which subnets in a customer VCN, such as VCN 104, use each gateway. A VCN's route tables can be used to determine whether traffic can be routed outside the VCN through a particular gateway. For example, in a particular case, the route table for a public subnet in customer VCN 104 can send non-local traffic through IGW 120. The route table for a private subnet in the same customer VCN 104 can send traffic to CSP services through SGW 126. All remaining traffic may be sent through NAT gateway 128. Route tables only control traffic that leaves the VCN.
[0090] Security lists associated with a VCN are used to control traffic entering the VCN through inbound connections and gateways. All resources within a subnet use the same mute tables and security lists. Security lists may be used to control specific types of traffic entering and leaving instances within a VCN's subnets. Security list rules may include inbound (inbound) rules and outbound (outbound) rules. For example, inbound rules may specify allowed source address ranges, and outbound rules may specify allowed destination address ranges. Security rules may specify specific protocols (e.g., TCP, ICMP), specific ports (e.g., port 22 for SSH, port 3389 for Windows RDP), etc. In certain implementations, the instance's operating system may enforce its own firewall rules that match security list rules. Rules may be stateful (e.g., connections are tracked and responses are automatically allowed without explicit security list rules for the response traffic) or stateless.
[0091] Access from a customer VCN (i.e., resources or compute instances deployed on VCN 104) may be categorized as public access, private access, or dedicated access. Public access refers to an access model for accessing public endpoints using public IP addresses or NATs. Private access enables customer workloads in VCN 104 with private IP addresses (e.g., resources in a private subnet) to access a service without traversing a public network such as the Internet. In particular embodiments, CSPI 101 enables customer VCN workloads with private IP addresses to access the service's public service endpoint using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer VCN and the service's public endpoint, which resides outside the customer's private network.
[0092] Additionally, CSPI is working to develop dedicated public peering services using technologies such as FastConnect public peering. It can provide brick access, where a customer's on-premises instance can access the FastConnect connection without going through a public network such as the Internet. You can access one or more services in a customer VCN using FastConnect. CSPI also provides dedicated private access using FastConnect private peering. In this case, customer on-premises instances with private IP addresses can access workloads in the customer VCN using the FastConnect connection. FastConnect connects customers' on-premises networks using the public internet. FastConnect is a network connection used instead of connecting your network to CSPI and its services. FastConnect offers higher bandwidth options and It provides an easy, flexible and economical way to create dedicated, private connections with a reliable and consistent networking experience.
[0093] FIG. 1 and the accompanying description above illustrate various virtualized elements in an exemplary virtual network. As noted above, a virtual network is built on an underlying physical network or substrate network. FIG. 2 is a simplified architecture diagram illustrating physical elements within a physical network within CSPI 200 that provides the foundation for the virtual network, according to certain embodiments. As shown, CSPI 200 provides a distributed environment including elements and resources (e.g., compute, memory, and networking resources) provided by a cloud service provider (CSP). These elements and resources are used to provide cloud services (e.g., IaaS services) to subscribing customers, i.e., customers who subscribe to one or more services offered by the CSP. Based on the services to which the customer subscribes, CSPI 200 provides some resources (e.g., compute, memory, and networking resources) to the customer. The customer can then build their own cloud-based (i.e., CSPI-hosted), customizable private virtual network using the physical compute, memory, and networking resources provided by CSPI 200. As previously mentioned, these customer networks are referred to as virtual cloud networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, into these Customer VCNs. The compute instances may be virtual machines, bare metal instances, etc. CSPI200 provides infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available hosted environment.
[0094] In the exemplary embodiment shown in FIG. 2, the physical elements of CSPI 200 include one or more physical host machines or servers (e.g., 202, 206, 208), network virtual machines (VMs), and The VCN includes a virtual network device (NVD) (e.g., 210, 212), a top-of-rack (TOR) switch (e.g., 214, 216), a physical network (e.g., 218), and a switch within the physical network 218. The physical host machines or servers can host and execute various compute instances participating in one or more subnets of the VCN. The compute instances may include virtual machine instances and bare metal instances. For example, the various compute instances shown in FIG. 1 may be hosted by the physical host machines shown in FIG. 2. The virtual machine compute instances in the VCN may be executed by one host machine or by multiple different host machines. Also, the physical host machines can host virtual host machines, container-based hosts or functions, etc. The VIC and VCN VRs shown in FIG. 1 may be executed by the FTVD shown in FIG. 2. The gateways shown in FIG. 1 may be executed by the host machines and / or NVDs shown in FIG. 2.
[0095] A host machine or server may run a hypervisor (also called a virtual machine monitor or VMM) that creates and enables a virtualized environment on the host machine. Virtualization or a virtualized environment facilitates cloud-based computing. One or more computing instances may be created, executed, and managed on the host machine by the hypervisor on the host machine. The hypervisor on the host machine enables the host machine's physical computing resources (e.g., computing resources, memory resources, and networking resources) to be shared among various computing instances running on the host machine.
[0096] For example, as shown in FIG. 2, host machines 202 and 208 execute hypervisors 260 and 266, respectively. These hypervisors may be implemented using software, firmware, hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that resides in a host machine's operating system (OS), which executes on the host machine's hardware processor. The hypervisor provides a virtualized environment that allows the host machine's physical computing resources (e.g., processing resources such as processors / cores, memory resources, and networking resources) to be shared among various virtual machine computing instances executed by the host machine. For example, in FIG. 2, hypervisor 260 resides in the OS of host machine 202 and allows the host machine's computing resources (e.g., processing resources, memory resources, and networking resources) to be shared among computing instances (e.g., virtual machines) executed by host machine 202. A virtual machine can have its own OS (called a guest OS). This guest OS may be the same as or different from the host machine's OS. The OS of a virtual machine executed by a host machine may be the same as or different from the OS of other virtual machines executed by the same host machine. Thus, the hypervisor can run multiple OSs in parallel while sharing the same computing resources of the host machine. The host machines shown in Figure 2 may have the same type of hypervisor or different types of hypervisors.
[0097] A compute instance may be a virtual machine instance or a bare metal instance. In Figure 2, compute instance 268 on host machine 202 and compute instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.
[0098] In certain instances, an entire host machine may be provided to a single customer, and one or more compute instances (either virtual machines or bare metal instances) hosted by that host machine may all belong to the same customer. A host machine may be shared among multiple customers (i.e., multiple tenants). In such a multi-tenant scenario, a host machine can host virtual machine compute instances belonging to different customers. These compute instances may be members of different VCNs for different customers. In certain embodiments, bare metal compute instances are hosted by bare metal servers without a hypervisor. When bare metal compute instances are provided, a single customer or tenant maintains control of the physical CPU, memory, and network interfaces of the host machine that hosts the bare metal instance, and the host machine is not shared with other customers or tenants.
[0099] As previously described, each compute instance that is part of a VCN is associated with a VNIC that enables the compute instance to be a member of a subnet of the VCN. The VNIC associated with a compute instance facilitates communication of packets or frames to and from the compute instance. A VNIC is associated with the compute instance when the compute instance is created. In particular embodiments, for a compute instance executed by a host machine, the VNIC associated with the compute instance is executed by an NVD connected to the host machine. For example, in FIG. 2, host machine 202 executes virtual machine compute instance 268 associated with VNIC 276, which is executed by NVD 210 connected to host machine 202. As another example, bare metal instance 272 hosted by host machine 206 is associated with VNIC 280, which is executed by NVD 212 connected to host machine 206. As yet another example, VNIC 284 is associated with compute instance 274 executed by host machine 208, which is executed by NVD 212 connected to host machine 208.
[0100] For a compute instance hosted by a host machine, the NVD connected to that host machine executes a VCN VR corresponding to the VCN of which the compute instance is a member. For example, in the embodiment shown in Figure 2, NVD 210 executes VCN VR 277 corresponding to the VCN of which compute instance 268 is a member. NVD 212 may also execute one or more VCN VRs 283 corresponding to the VCNs corresponding to the compute instances hosted by host machines 206 and 208.
[0101] A host machine may include one or more network interface cards (NICs) for connecting the host machine to other devices. The NICs on a host machine may provide one or more ports (or interfaces) for communicatively connecting the host machine to another device. For example, one or more ports (or interfaces) on the host machine and the NVD may be used to connect the host machine to the NVD. The host machine may also be connected to other devices, such as other host machines.
[0102] 2, host machine 202 is connected to NVD 210 using link 220 extending between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD 210. Host machine 206 is connected to NVD 212 using link 224 extending between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD 212. Host machine 208 is connected to NVD 212 using link 226 extending between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD 212.
[0103] Similarly, the NVD connects the physical network to the network through a communication link (also known as a switch fabric). The NVDs are connected to top-of-rack (TOR) switches that are connected to network 218. In particular embodiments, the links between the host machines and the NVDs and the links between the NVDs and the TOR switches are Ethernet links. For example, in FIG. 2, NVDs 210 and 212 are connected to TOR switches 214 and 216, respectively, via links 228 and 230. In particular embodiments, links 220, 224, 226, 228, and 230 are Ethernet links. The collection of host machines and NVDs connected to a TOR may be referred to as a rack.
[0104] The physical network 218 provides a communications fabric that enables the TOR switches to communicate with each other. The physical network 218 may be a multi-tier network. In a particular implementation, the physical network 218 is a multi-tier Clos network of switches, with the TOR switches 214 and 216 representing leaf-level nodes of the multi-tier and multi-node physical switching network 218. Different Clos network configurations are possible, including, but not limited to, 2-tier networks, 3-tier networks, 4-tier networks, 5-tier networks, and generally "n"-tier networks. An example of a Clos network is shown in FIG. 5 and described below.
[0105] A variety of different connection configurations are possible between a host machine and the N virtual disks, including one-to-one, many-to-one, and one-to-many configurations. In a one-to-one implementation, each host machine is connected to its own separate virtual disk. For example, in FIG. 2, host machine 202 is connected to virtual disk 210 via NIC 232 of host machine 202. In a many-to-one configuration, multiple host machines are connected to a single virtual disk. For example, in FIG. 2, host machines 206 and 208 are connected to the same virtual disk 212 via NICs 244 and 250, respectively.
[0106] In a one-to-many configuration, one host machine is connected to multiple NVDs. Figure 3 shows an example of a CSPI 300 in which a host machine is connected to multiple NVDs. As shown in Figure 3, host machine 302 includes a network interface card (NIC) 304 including multiple ports 306 and 30S. Host machine 300 is connected to a first NVD 310 via port 306 and link 320, and to a second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between host machine 302 and NVDs 310 and 312 may be Ethernet links. NVD 310 is connected to a first TOR switch 314, and NVD 312 is connected to a second TOR switch 316. The links between NVDs 310 and 312 and TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent Tier-0 switching devices within a multi-tier physical network 318 .
[0107] 3 provides two separate physical network paths from the physical switch network 318 to the host machine 302: a first path from the TOR switch 314 to the host machine 302 via the NVD 310, and a second path from the TOR switch 316 to the host machine 302 via the NVD 312. The separate paths provide enhanced availability (referred to as high availability) for the host machine 302. If there is a problem with one of the paths (e.g., a link on one of the paths fails) or if there is a problem with a device (e.g., a particular NVD is not functioning), the other path can be used for communications to and from the host machine 302.
[0108] In the configuration shown in Figure 3, the host machine is connected to two different NVDs using two different ports provided by the host machine's NIC. In this case, the host machine may include multiple NICs, allowing the host machine to connect to multiple NVDs.
[0109] Referring again to Figure 2, an NVD is a physical device or element that performs one or more network virtualization functions and / or storage virtualization functions. An NVD may be any device that has one or more processing units (e.g., a CPU, a network processing unit (NPU), an FPGA, a packet processing pipeline), memory including cache, and ports. Various virtualization functions may be performed by software / firmware executed by one or more processing units of the NVD.
[0110] The NVD may be implemented in a variety of different ways. For example, in a particular embodiment, the NVD is implemented as an interface card with an embedded processor, called a smart NIC or intelligent NIC. The smart NIC is a separate device from the NIC on the host machine. In Figure 2, the NVD 210 may be implemented as a smart NIC connected to the host machine 202, and the NVD 212 may be implemented as a smart NIC connected to the host machines 206 and 208.
[0111] However, a smart NIC is only one example of an NVD implementation. Various other implementations are possible. For example, in some other implementations, the NVD or one or more functions performed by the NVD may be incorporated into or performed by one or more host machines, one or more TOR switches, and other elements of CSPI 200. For example, the NVD may be integrated into a host machine. In this case, the functions performed by the NVD are performed by the host machine. As another example, the NVD may be part of a TOR switch, or a TOR switch may be configured to perform the functions performed by the NVD that enable the TOR switch to perform various complex packet transformations used in public clouds. A TOR that performs the functions of an NVD may be referred to as a smart TOR. In yet other implementations that provide customers with virtual machine (VM) instances rather than bare metal (BM) instances, the functions provided by the NVD may be implemented inside the hypervisor of the host machine. In some other implementations, some of the NVD's functions may be offloaded to a centralized service running on a set of host machines.
[0112] In certain embodiments, such as when implemented as a smart NIC, as shown in FIG. 2, an NVD may include multiple physical ports that allow the NVD to connect to one or more host machines and one or more TOR switches. Ports on an NVD can be categorized as host-facing ports (also called "south ports") or network-facing or TOR-facing ports (also called "north ports"). A host-facing port of an NVD is a port used to connect the NVD to a host machine. Examples of host-facing ports in FIG. 2 include port 236 of NVD 210 and ports 248 and 254 of NVD 212. A network-facing port of an NVD is a port used to connect the NVD to a TOR switch. Examples of network-facing ports in FIG. 2 include port 256 of NVD 210 and port 258 of NVD 212. As shown in FIG. 2, the NVD 210 is connected to the TOR switch 214 via link 228 extending from port 256 of NVD 210 to the TOR switch 214. Similarly, the NVD 212 is connected to the TOR switch 216 via a link 230 that extends from a port 258 of the NVD 212 to the TOR switch 216 .
[0113] The NVD receives packets and frames (e.g., packets and frames generated by compute instances hosted by the host machine) from the host machine via its host-facing ports, performs any necessary packet processing, and then transmits the packets to the NVD's network. The NVD can receive packets and frames from the TOR switch through its network-facing port, perform necessary packet processing, and then forward the packets and frames to the host machine through its host-facing port.
[0114] In certain embodiments, multiple ports and associated links may be provided between the NVD and the TOR switch. These ports and links can be aggregated to form a link aggregator group (LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and the TOR switch) to be treated as a single logical link. All physical links within a given LAG can operate at the same speed and in full-duplex mode. LAGs help increase the bandwidth and reliability of the connection between two endpoints. If one of the physical links in a LAG fails, traffic is dynamically and transparently reassigned to another physical link within the LAG. The aggregated physical link provides higher bandwidth than individual links. Multiple ports associated with a LAG are treated as a single logical port. Traffic can be load-balanced across the multiple physical links in the LAG. One or more LAGs can be configured between two endpoints. The two endpoints may be, for example, between the NVD and the TOR switch, or between a host machine and the NVD.
[0115] The NVD implements or performs network virtualization functions. These functions are performed by software / firmware executed by the NVD. Examples of network virtualization functions include, but are not limited to, packet encapsulation and decapsulation functions, functions for creating VCN networks, functions for implementing network policies such as VCN security list (firewall) functions, functions for facilitating routing and forwarding of packets to and from compute instances within a VCN, etc. In particular embodiments, upon receiving a packet, the NVD is configured to execute a packet processing pipeline that processes the packet and determines how to forward or route the packet. As part of this packet processing pipeline, the NVD provides execution of one or more virtual functions associated with the overlay network, such as execution of VNICs associated with compute instances in the VCN, execution of virtual routers (VRs) associated with the VCN, encapsulation and decapsulation of packets to facilitate forwarding or routing within the virtual network, execution of specific gateways (e.g., local peering gateways), implementation of security lists, network security groups, network address translation (NAT) functions (e.g., public IP to private IP translation per host), throttling functions, and other functions.
[0116] In some embodiments, the packet processing data path within the NVD may include multiple packet pipelines. Each packet pipeline consists of a series of packet transformation stages. In some implementations, upon receiving a packet, the packet is parsed and sorted into a single pipeline. The packet is then processed stage by stage in a linear fashion until it is discarded or sent out through an interface of the NVD. These stages provide packet processing building blocks of basic functionality (e.g., validating headers, performing throttling, inserting new Layer 2 headers, performing L4 firewalling, VCN encapsulation / decapsulation), such that new pipelines can be constructed by assembling existing stages, and new functionality can be added by creating and inserting new stages into existing pipelines.
[0117] The NVD can perform both control plane and data plane functions corresponding to the control and data planes of the VCN. An example of a VCN control plane is shown in Figure 1. 24, 25, 26, and 27 (see reference numbers 2416, 2516, 2616, and 2716) and described below. Examples of the VCN data plane are shown in Figures 24, 25, 26, and 27 (see reference numbers 2418, 2518, 2618, and 2718) and described below. The control plane functions include functions used to configure the network (e.g., setting routes and route tables, configuring VNICs) to control how data is forwarded. In certain embodiments, a VCN control plane is provided that centrally computes and exposes all overlay-to-substrate mappings to the NVD and virtual network edge devices (e.g., various gateways such as DRGs, SGWs, and IGWs). Firewall rules can also be exposed using the same mechanism. In certain embodiments, the NVD retrieves only the mappings that are relevant to that NVD. The data plane functions include functions that perform the actual routing / forwarding of packets based on the configuration set using the control plane. The VCN data plane is implemented by encapsulating customer network packets before they traverse the backbone network. The encapsulation / decapsulation functionality is implemented in the NVD. In certain embodiments, the NVD is configured to intercept all network packets entering and leaving the host machine and perform network virtualization functions.
[0118] As described above, the NVD performs various virtualization functions, including VNICs and VCN VRs. The NVD can execute VNICs associated with compute instances hosted by one or more host machines connected to the VNICs. For example, as shown in FIG. 2, NVD 210 executes the functions of VNIC 276 associated with compute instance 268 hosted by host machine 202 connected to NVD 210. As another example, NVD 212 executes VNIC 280 associated with bare metal compute instance 272 hosted by host machine 206 and VNIC 284 associated with compute instance 274 hosted by host machine 208. The host machines can host compute instances that belong to different VCNs that belong to different customers. The NVDs connected to the host machines can execute VNICs (i.e., execute functions associated with the VNICs) corresponding to the compute instances.
[0119] NVDs also execute VCN virtual routers corresponding to the VCNs of the compute instances. For example, in the embodiment shown in FIG. 2, NVD 210 executes VCN VR 277 corresponding to the VCN to which compute instance 268 belongs. NVD 212 executes one or more VCN VRs 283 corresponding to one or more VCNs to which compute instances hosted on host machines 206 and 208 belong. In particular embodiments, a VCN VR corresponding to a VCN is executed by all NVDs connected to a host machine that hosts at least one compute instance belonging to that VCN. If a host machine hosts compute instances that belong to different VCNs, the NVDs connected to that host machine may execute VCN VRs corresponding to the different VCNs.
[0120] In addition to VNICs and VCN VRs, the NVD may include one or more hardware elements that run various software (e.g., daemons) and facilitate various network virtualization functions performed by the NVD. For simplicity, these various elements are grouped as "packet processing elements" shown in FIG. 2. For example, NVD 210 includes packet processing element 286, and NVD 212 includes packet processing element 288. For example, the packet processing element of the NVD may include a packet processor configured to monitor all packets received and communicated using the NVD by interacting with the NVD's ports and hardware working interfaces and to store network information. The network information may include, for example, network flow information for identifying different network flows processed by the NVD and information about each flow (e.g., statistics for each flow). In certain embodiments, the network flow information may include: As another example, the packet processing element may include a replication agent configured to replicate information stored by the NVD to one or more different replication target stores. The packet processing element may include a logging agent configured to perform logging functions for the NVD. The packet processing element may also include a logging agent configured to monitor the performance and It may also include software for monitoring the health and possibly the status and health of other elements connected to the NVD.
[0121] FIG. 1 illustrates elements of an exemplary virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, VRs for the VCN, and a set of gateways configured for the VCN. The overlay elements illustrated in FIG. 1 may be executed or hosted by one or more of the physical elements illustrated in FIG. 2. For example, compute instances within a VCN may be executed or hosted by one or more host machines illustrated in FIG. 2. For compute instances hosted by a host machine, the VNICs associated with the compute instance are typically executed by an NVD connected to the host machine (i.e., the VNIC functionality is provided by an NVD connected to the host machine). The VCN VR functionality is performed by all NVDs connected to the host machines that host or execute compute instances that are part of the VCN. Gateways associated with a VCN may be executed by one or more different types of NVDs. For example, some gateways may be executed by smart NICs, and other gateways may be executed by one or more host machines or other implementations of NVDs.
[0122] As described above, compute instances within a customer VCN can communicate with a variety of different endpoints. These endpoints may be in the same subnet as the source compute instance, in a different subnet but in the same VCN as the source compute instance, or may include endpoints outside the VCN of the source compute instance. These communications are facilitated using the VNICs associated with the compute instances, the VCN VRs, and the gateways associated with the VCN.
[0123] Communication between two compute instances on the same subnet within a VCN is facilitated using VNICs associated with the source and destination compute instances. The source and destination compute instances may be hosted by the same host machine or different host machines. A packet originating from a source compute instance may be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. In the NVD, the packet is processed using a packet processing pipeline, which may include the execution of a VNIC associated with the source compute instance. Because the packet's destination endpoint is in the same subnet, the execution of a VNIC associated with the source compute instance forwards the packet to an NVD running a VNIC associated with the destination compute instance, which processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances may run on the same NVD (e.g., if both the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., if the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNIC can use the routing / forwarding tables stored by the NVD to determine the next hop for a packet.
[0124] When communicating a packet from a compute instance in a subnet to an endpoint in a different subnet within the same VCN, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. In the NVD, the packet is processed using a packet processing pipeline that may include running one or more VNICs and VRs associated with the VCN. For example, the NVD executes or invokes a function corresponding to a VNIC associated with the source compute instance (also referred to as executing a VNIC) as part of the packet processing pipeline. The function executed by the VNIC may include looking up a VLAN identifier on the packet. Because the packet's destination is outside the subnet, a VCN VR function is invoked and executed by the NVD. The VCN VR then routes the packet to the NVD executing the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards the packet to the destination compute instance. The VNICs associated with the source compute instance and the destination compute instance may run on the same NVD (e.g., if both the source compute instance and the destination compute instance are hosted by the same host machine) or may run on different NVDs (e.g., if the source compute instance and the destination compute instance are hosted by different host machines connected to different NVDs).
[0125] If the packet's destination is outside the VCN of the source compute instance, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD runs the VNIC associated with the source compute instance. Because the packet's destination endpoint is outside the VCN, the packet is processed by the VCN VR for that VCN. The NVD invokes a VCN VR function, which may result in the packet being forwarded to an NVD running the appropriate gateway associated with the VCN. For example, if the destination is an endpoint in a customer's on-premises network, the packet may be forwarded by the VCN VR to an NVD running a DRG gateway configured for the VCN. The VCN VR may run on the same NVD as the NVD running the VNIC associated with the source compute instance, or it may be run by a different NVD. The gateway may be run by a smart NIC, a host machine, or another NVD implementation. The packet is then processed by the gateway and forwarded to the next hop that facilitates communication of the packet to the intended destination endpoint. 2, a packet originating from compute instance 268 may be communicated from host machine 202 to NVD 210 via link 220 (using NIC 232). VNIC 276 on NVD 210 is called out because it is the VNIC associated with source compute instance 268. VNIC 276 is configured to inspect encapsulation information in the packet, determine a next hop for forwarding the packet to facilitate communication of the packet to its intended destination endpoint, and forward the packet to the determined next hop.
[0126] Compute instances deployed on a VCN can communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 200 may include instances in the same VCN or other VCNs (which may be customer VCNs or VCNs not belonging to the customer). Communication between endpoints hosted by CSPI 200 may be performed over physical network 218. Compute instances can also communicate with endpoints not hosted by or external to CSPI 200. Examples of these endpoints are instances in a customer on-premises network or data center. 2. The CSPI 200 includes endpoints accessible via a public network, such as the Internet, or public endpoints accessible via a public network, such as the Internet. Communications with endpoints external to CSPI 200 may be performed via a public network (e.g., the Internet) (not shown in FIG. 2) or a private network (not shown in FIG. 2) using various communication protocols.
[0127] The architecture of CSPI 200 shown in FIG. 2 is merely an example and is not intended to be limiting. Variations, substitutions, and modifications are possible in alternative embodiments. For example, in some implementations, CSPI 200 may have more or fewer systems or elements than those shown in FIG. 2, may combine two or more systems, or may have a different system configuration or arrangement. The systems, subsystems, and other elements shown in FIG. 2 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or a combination thereof. Software may be stored in a non-transitory storage medium (e.g., a memory device).
[0128] FIG. 4 illustrates connections between host machines and an NVD to provide I / O virtualization to support multi-tenancy, according to certain embodiments. As shown in FIG. 4, a host machine 402 runs a hypervisor 404 that provides a virtualized environment. The host machine 402 runs two virtual machine instances: VM1 406, which belongs to customer / tenant #1, and VM2 408, which belongs to customer / tenant #2. The host machine 402 includes a physical NIC 410 connected to an NVD 412 via link 414. Each of the compute instances is connected to a VNIC run by the NVD 412. In the embodiment of FIG. 4, VM1 406 is connected to VNIC-VM1 420, and VM2 408 is connected to VNIC-VM2 422.
[0129] 4, NIC 410 includes two logical NICs: logical NIC A 416 and logical NIC B 418. Each virtual machine is connected to and configured to operate with its own logical NIC. For example, VM1 406 is connected to logical NIC A 416, and VM2 408 is connected to logical NIC B 418. The logical NICs allow each tenant's virtual machine to believe it owns its own host machine and NIC, even though host machine 402 consists of only one physical NIC 410 shared by multiple tenants.
[0130] In particular embodiments, each logical NIC is assigned its own VLAN ID. Thus, logical NIC A 416 for tenant #1 is assigned a particular VLAN ID, and logical NIC B 418 for tenant #2 is assigned a different VLAN ID. When a packet is communicated from VM1 406, the hypervisor attaches a tag assigned to tenant #1 to the packet before communicating the packet from host machine 402 to NVD 412 over link 414. Similarly, when a packet is communicated from VM2 408, the hypervisor attaches a tag assigned to tenant #2 to the packet before communicating the packet from host machine 402 to NVD 412 over link 414. Thus, a packet 424 communicated from host machine 402 to NVD 412 has an associated tag 426 that identifies the particular tenant and associated VM. When a packet 424 is received on the NVD from host machine 402, the tag 426 associated with the packet is used to determine whether the packet should be processed by VNIC-VM1 420 or VNIC-VM2 422. The packet is then processed by the corresponding VNIC. The configuration shown in Figure 4 allows each tenant's compute instance to believe it owns its own host machine and NIC. The configuration shown in Figure 4 provides I / O virtualization to support multi-tenancy. do.
[0131] FIG. 5 is a schematic block diagram illustrating a physical network 500 according to a particular embodiment. The embodiment illustrated in FIG. 5 is constructed as a Clos network. A Clos network is a particular type of network topology designed to provide connection redundancy while maintaining high bisection bandwidth and maximum resource utilization. A Clos network is a type of non-blocking, multi-stage or multi-layer switching network, and the number of stages or layers may be 2, 3, 4, 5, etc. The embodiment illustrated in FIG. 5 is a three-layer network, including layers 1, 2, and 3. TOR switch 504 represents a layer-0 switch in the Clos network. One or more NVDs are connected to the TOR switch. The layer-0 switch is also referred to as an edge device of the physical network. The layer-0 switch is connected to a layer-1 switch, also referred to as a leaf switch. In the embodiment illustrated in FIG. 5, “n” layer-0 TOR switches are connected to “n” layer-1 switches to form a pod. Each layer-0 switch in a pod is interconnected to all layer-1 switches in the pod, but switches between pods are not connected. In a specific implementation, the two pods are referred to as blocks. Each block is served by or connected to "n" layer-2 switches (also referred to as spine switches). A physical network topology may include multiple blocks. Similarly, the layer-2 switches are connected to "n" layer-3 switches (also referred to as super-spine switches). Communication of packets through the physical network 500 is typically performed using one or more layer-3 communication protocols. Typically, all layers of the physical network, except for the TOR layer, are n-way redundant, thus achieving high availability. The physical network can be scaled by specifying policies on pods and blocks to control the mutual visibility of switches in the physical network.
[0132] A characteristic of Clos networks is that the maximum hop count from one tier-0 switch to another tier-0 switch (or from an NVD connected to a tier-0 switch to another NVD connected to a tier-0 switch) is constant. For example, in a three-tier Clos network, a packet requires a maximum of seven hops to travel from one NVD to another. In this case, the source and target NVDs are connected to the leaf layers of the Clos network. Similarly, in a four-tier Clos network, a packet requires a maximum of nine hops to travel from one NVD to another. In this case, the source and target NVDs are connected to the leaf layers of the Clos network. Therefore, the Clos network architecture maintains constant overall network latency, which is important for intra- and inter-datacenter communications. Clos topologies are horizontally scalable and cost-effective. The network bandwidth / throughput capacity can be easily increased by adding more switches (e.g., more leaf switches and spine switches) at each tier and by increasing the number of links between switches in adjacent tiers.
[0133] In certain embodiments, each resource in CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource's information. This identifier can be used to manage the resource, for example, through a console or API. An exemplary syntax for a CID is as follows:
[0134] ocid1.<RESOURCE TYPE> . <realm>.[REGION].[FUTURE USE].<UNIQUE ID> is. During the ceremony, "ocid1" is a string that indicates the version of the CID.
[0135] "RESOURCE TYPE" is the type of resource (e.g., instance, volume, VCN) , subnet, user, group).
[0136] "REALM" represents the region where the resource resides. An example value is "c1" "c1" represents the government cloud region, "c2" represents the government cloud region, or "c3" represents the federal government cloud region. Each region can have its own domain name.
[0137] "REGION" represents the region that the resource belongs to. If a region does not apply to the resource, this part may be blank.
[0138] "FUTURE USE" indicates that the item is reserved for future use. "UNIQUE ID" is the unique ID part. This format is used for resources or servers. This may vary depending on the type of screw.
[0139] Layer 2 Virtual Network The number of enterprise customers migrating their on-premises applications to cloud environments offered by cloud service providers (CSPs) continues to grow rapidly. However, many of these customers quickly realize that migrating to a cloud environment is quite challenging and that their existing applications must be redesigned and re-architected to run in a cloud environment. This is because applications developed in an on-premises environment often rely on the capabilities of the physical network for monitoring, availability, and scalability. These on-premises applications must be redesigned and re-architected before they can run in a cloud environment.
[0140] There are several reasons why on-premises applications cannot be easily migrated to cloud environments. One of the main reasons is that current cloud virtual networks operate at Layer 3 of the OSI model, e.g., the IP layer, and therefore cannot provide the Layer 2 functionality required by applications. Layer 3-based routing or forwarding involves determining where to send a packet (e.g., to which customer instance) based on information contained in the packet's Layer 3 header, e.g., the destination IP address contained in the packet's Layer 3 header. To facilitate this, the location of IP addresses within the virtualized cloud network is determined via a centralized control and orchestration system or controller. This location may include, for example, IP addresses associated with customer entities or resources within the virtualized cloud environment.
[0141] Many customers run applications in on-premise environments that have stringent requirements for Layer 2 network functions that are not addressed by current cloud and IaaS service providers. For example, current cloud service traffic is routed using Layer 3 protocols that use Layer 3 headers, and the Layer 2 functions required by the applications are not supported. These Layer 2 functions include Address Resolution Protocol (ARP) processing, media access control, and network management. The network may include features such as Layer 2 broadcast functionality, Layer 2 (MAC-based) forwarding, Layer 2 networking configuration, and Layer 2 MAC address learning. By providing virtualized Layer 2 networking functionality in a virtualized cloud network as described in this disclosure, customers can smoothly migrate legacy applications to a cloud environment without requiring substantial redesign or re-architecting. For example, the virtualized Layer 2 networking functionality described in this disclosure can be used to seamlessly ... These applications can run in the same versions and configurations on the public cloud, allowing customers to easily deploy legacy on-premise applications, including existing knowledge, tools, and processes associated with those applications. Customers can also access native cloud services from their applications (for example, using VMware Software-Defined Data Center (SDDC)).
[0142] As another example, some legacy on-premises applications (e.g., enterprise clustering software applications, network virtual appliances) require Layer 2 broadcast support to accommodate failover. Exemplary applications include Fortinet FortiGate, IBM® QRadar, Palo Alto Firewall, Cisco ASA, Juniper SRX, and Oracle® Real Application Clustering (RAC). By providing virtualized Layer 2 networking in a virtualized public cloud as described herein, these applications can operate unchanged in the virtualized public cloud environment. As described herein, virtualized Layer 2 networking functionality comparable to on-premises functionality is provided. The virtualized Layer 2 networking functionality described herein supports traditional Layer 2 networking. This includes support for unicast Layer 2 traffic functionality, broadcast Layer 2 traffic functionality, and multicast Layer 2 traffic functionality, along with support for customer-defined VLANs. Layer 2-based packet routing and forwarding includes, for example, routing or forwarding packets based on the destination MAC address contained in the Layer 2 header using the Layer 2 protocol and information contained in the packet's Layer 2 header. Protocols used by enterprise applications (e.g., clustering software applications), such as ARP (Address Resolution Protocol), GARP (Gratuitous ARP), and RARP (Reverse Address Resolution Protocol), The Reverse Address Resolution Protocol (ARP) can now also operate in cloud environments.
[0143] There are several reasons why traditional virtualized cloud infrastructures support virtualized Layer 3 networks but not Layer 2 networks. Layer 2 networks typically do not scale as well as Layer 3 networks. Layer 2 network control protocols do not have the level of sophistication required for scaling. For example, Layer 3 networks do not need to worry about packet looping, which Layer 2 networks must address. IP packets (i.e., Layer 3 packets) have a time to live (TTL), but Layer 2 packets do not. IP addresses contained within Layer 3 packets have topology, such as subnets and CIDR ranges, but Layer 2 addresses (e.g., MAC addresses) do not. Layer 3 IP networks have built-in tools to facilitate troubleshooting, such as ping and traceroute to find routing information. Such tools are not available for Layer 2. Layer 3 networks support multipathing, which is not available for Layer 2. Layer 2 networks require advanced control protocols, particularly for exchanging information between entities within the network, such as Border Gateway Protocol (BGP). Layer 2 does not have Layer 3 (OSPF) and Layer 2 Routing (Layer 3 Routing) protocols, so it must rely on broadcast and multicast to learn the network, which can negatively impact network performance. Also, as the network changes, Layer 2 requires a learning process to be repeated, whereas Layer 3 does not. For these and other reasons, it is more desirable for cloud IaaS service providers to offer infrastructure that operates at Layer 3 rather than Layer 2.
[0144] However, despite its multiple drawbacks, many on-premise applications require Layer 2 functionality. For example, whether the instance is a compute instance (e.g. bare metal, virtual machine or container) or a service instance (e.g. load balancing) Assume a virtualized cloud configuration in which a customer (Customer 1) has two instances, Instance A with IP 1 and Instance B with IP 2, in virtual network "V," which could be a DNS server, an NFS mount point, or another service instance. Virtual network V is a separate address space isolated from other virtual networks and the underlying physical network. This isolation can be achieved using various techniques, including, for example, packet encapsulation or NAT. Therefore, the IP addresses of instances in the customer's virtual network are different from the addresses in the physical network hosting those instances. A centralized SDN (Software-Defined Networking) control plane is provided that keeps track of the physical IPs and virtual interfaces for all virtual IP addresses. When sending a packet from Instance A to a destination IP 2 in virtual network V, the virtual network SDN stack needs to know the location of IP 2 in advance so that it can send the packet to the IP in the physical network hosting virtual IP address IP 2 in virtual network V. Because the location of virtual IP addresses can change on the cloud, the relationship between physical IPs and virtual IP addresses can also change. Every time a virtual IP address is moved (e.g., the IP address associated with a virtual machine is moved to another virtual machine, or the virtual machine is migrated to a new physical host), an API call must be made to the SDN control plane to inform the controller that the IP has been moved so that all participants in the SDN stack, including the packet processors (data plane), can be updated. However, there are some types of applications that do not make such API calls. Examples include various on-premise applications and applications provided by various virtualization software vendors, such as VMware.The utility of facilitating virtual Layer 2 networks in a virtualized cloud environment is to enable support for applications that are not programmed to make API calls or that rely on other Layer 2 networking capabilities, such as support for non-IP Layer 3 and MAC learning.
[0145] A virtual Layer 2 network creates a broadcast domain, and members of the broadcast domain are learned. In a virtual Layer 2 domain, any IP can exist on any MAC on any host within this Layer 2 domain. The system learns using standard Layer 2 networking protocols. The system virtualizes these networking primitives without requiring a centralized controller to explicitly tell it where MACs and IPs exist on the virtual Layer 2 network. This enables applications that require low-latency failover, applications that need to support broadcast or multicast protocols to multiple nodes, and legacy applications that do not know how to make API calls to an SDN control plane or API endpoint to determine the location of IP and MAC addresses. Therefore, providing Layer 2 networking capabilities in a virtualized cloud environment is necessary to support functionality not available at the IP Layer 3 level.
[0146] Another technical advantage of providing a virtual Layer 2 in a virtualized cloud environment is the ability to support a variety of different Layer 3 protocols (e.g., IPV4, IPV6, etc.), including non-IP protocols. For example, various non-IP protocols such as IPX, AppleTalk, etc. Existing cloud IaaS providers are unable to support these non-IP protocols because they do not provide Layer 2 functionality in their virtualized cloud networks. By providing Layer 2 networking functionality as described in this disclosure, support can be provided for Layer 3 protocols and for applications that require and depend on the availability of Layer 2 level functionality.
[0147] The techniques described in this disclosure are used to provide both Layer 3 and Layer 2 functionality in a virtualized cloud infrastructure. As previously mentioned, Layer 3-based networking offers certain capabilities not offered by Layer 2 networking, particularly capabilities suited to scaling. By providing Layer 2 functionality in addition to Layer 3 functionality, it is possible to provide Layer 2 functionality in a more scalable manner while also leveraging the capabilities offered by Layer 3 (e.g., to provide a more scalable solution). For example, virtualizing Layer 3 avoids having to use broadcast for learning. By providing Layer 3 capabilities while also enabling applications that require Layer 2 functionality and applications that cannot function without it, and providing a virtualized Layer 2 to support non-IP protocols, etc., customers can be provided with the full flexibility of their virtualized cloud environment.
[0148] The customer themselves has a hybrid environment where a Layer 2 environment exists alongside a Layer 3 environment, and the virtualized cloud environment can support both of these environments. The customer can have Layer 3 networks, such as subnets, and / or Layer 2 networks, such as VLANs, and these two environments can talk to each other within the virtualized cloud environment.
[0149] Additionally, virtualized cloud environments need to support multi-tenancy, which makes it technically difficult and complex to provision both Layer 3 and Layer 2 functionality in the same virtualized cloud environment. For example, Layer 2 broadcast domains must be managed across many different customers within a cloud provider's infrastructure. The embodiments described in this disclosure overcome these technical challenges.
[0150] For virtualization providers (e.g., VMware), a virtualized Layer 2 network that emulates a physical Layer 2 network allows workloads to run unchanged. Applications provided by such virtualization providers can run on the virtualized Layer 2 network provided by the cloud infrastructure. For example, such applications may include a set of instances that must run on a Layer 2 network. If a customer wants to lift and shift such applications from their own on-premises environment to a virtualized cloud environment, they cannot run these applications in the cloud environment because they depend on underlying Layer 2 networks (e.g., Layer 2 network functions used to migrate or move virtual machines to where their MAC and IP addresses reside) that are not provided by the current virtualization cloud provider. For these reasons, such applications cannot run natively in the virtualized cloud environment. Using the technology described in this disclosure, cloud providers not only provide virtualized Layer 3 networks but also virtualized Layer 2 networks. This allows such application stacks to run unchanged in the cloud environment, enabling nested virtualization within the cloud environment. Customers can run and manage their own Layer 2 applications on the cloud. Application providers do not need to make changes to their software to facilitate this. This allows such legacy applications or workloads (e.g., legacy load balancers, legacy applications, KVM, Openstack, clustering software) to be deployed unchanged on the virtualized cluster. It may be implemented in a loud environment.
[0151] By providing virtualized Layer 2 functionality as described in this disclosure, a virtualized cloud environment can support a variety of Layer 3 protocols, including non-IP protocols. Taking Ethernet as an example, it is possible to support a variety of different EtherTypes (a field in the Layer 2 header that contains the type of Layer 3 packet being sent or the Layer 3 protocol expected), including various non-IP protocols. The EtherType is a two-octet field within the Ethernet frame. The EtherType is used to indicate which protocol is encapsulated in the payload of the frame and is used at the receiving end by the Data Link Layer to determine how to process the payload. The EtherType is used as the basis for 802.1Q VLANs, which tag and encapsulate packets from a VLAN for multiplexing with other VLAN traffic over an Ethernet trunk. Examples of EtherTypes are IPV4, IPV6, Address Resolution Protocol (ARP), AppleTalk, and IP X, etc. A cloud network that supports Layer 2 protocols can support any Layer 3 protocol. Similarly, when a cloud infrastructure supports Layer 3 protocols, it can support a variety of Layer 4 protocols, such as TCP, UDP, and ICMP. When a network is virtualized at Layer 3, it is independent of the Layer 4 protocol. Similarly, when a network is virtualized at Layer 2, it is independent of the Layer 3 protocol. This technology can be extended to support any Layer 2 network, including FDDI, InfiniBand, etc.
[0152] Consequently, many applications written for physical networks, especially those that run on clusters of computer nodes that share a broadcast domain, use Layer 2 features that are not supported by L3 virtual networks.The following six examples highlight the complications that can arise from not providing Layer 2 networking capabilities:
[0153] (1) MAC and IP assignment without a preceding API call. Network appliances and hypervisors (e.g., VMware) were not built for virtual cloud networks. They assume that you can use any MAC as long as it's unique, get a dynamic address from a DHCP server, or use any IP assigned to the cluster. There is often no mechanism that can be configured to notify the control plane of these Layer 2 and Layer 3 address assignments. If the Layer 3 virtual network doesn't know where the MAC and IP are, it doesn't know where to send traffic.
[0154] (2) Low-latency reassignment of MAC and IP for high availability and live migration. Many on-premise applications use ARP to reassign IP and MAC for high availability. When an instance in a cluster or HA pair becomes unresponsive, the newly active instance reassigns the service IP to its MAC by sending a Gratuitous ARP (GARP) or the service MAC to its interface by sending a Reverse ARP (RARP). This is also important when live migrating instances on a hypervisor: when a guest is migrated, the new host must send a RARP so that guest traffic is sent to the new host. Not only does this allocation have to happen without an API call, but it also has to happen with very low latency (sub-milliseconds), which cannot be achieved with an HTTPS call to a REST endpoint.
[0155] (3) Interface multiplexing by MAC address. When a hypervisor hosts multiple virtual machines on a single host and all virtual machines are on the same network, guest interfaces are differentiated by MAC addresses. This allows multiple virtual machines to share the same virtual interface. It is necessary to support multiple MACs on the interface.
[0156] (4) VLAN support. A single physical virtual machine host may need to reside on multiple broadcast domains, as indicated by the use of VLAN tags. For example, VMware ESX uses VLANs to separate traffic (e.g., guest virtual machines can communicate on one VLAN, be stored on another, and host virtual machines on yet another).
[0157] (5) Use of Broadcast and Multicast Traffic: ARP requires L2 broadcast, and exemplary on-premise applications use broadcast and multicast traffic in cluster and HA use cases.
[0158] (6) Support for non-IP traffic. L3 networks require an IPv4 or IPv6 header for communication, so they do not work if an L3 protocol other than IP is used. L2 virtualization makes networks within a VLAN independent of the L3 protocol. In this case, the L3 header can be IPv4, IPv6, IPX, or something else, or it can be absent altogether.
[0159] As described in this disclosure, a Layer 2 (L2) network can be created within a cloud network. This virtual L2 network includes one or more virtual L2 VLANs (referred to as VLANs in this disclosure). Each VLAN can include multiple compute instances, and each compute instance can be associated with at least one L2 virtual interface (e.g., L2 VNIC) and a local switch. In some embodiments, each pair of L2 VNIC and switch is hosted on an NVD. The NVD can host multiple such pairs, each pair associated with a different compute instance. A collection of local switches represents a single switch for an emulated VLAN. An L2 VNIC represents a collection of ports on the single emulated switch. A VLAN can be connected to other VLANs, Layer 3 (L3) networks, on-premises networks, and / or other networks via a VLAN Switching and Routing Service (VSRS), also referred to as a Real Virtual Router (RVR) or L2 VSRS in this disclosure.
[0160] Referring to FIG. 6, FIG. 6 illustrates a schematic diagram of a computing network of one embodiment. VCN 602 resides in CSPI 601. VCN 602 includes multiple gateways for connecting VCN 602 to other networks. These gateways include, for example, DRG 604, which can connect VCN 602 to an on-premises network, such as on-premises data center 606. The gateways can further include gateway 600, which can include, for example, an LPG for connecting VCN 602 to another VCN, and / or an IGW and / or NAT gateway for connecting VCN 602 to the Internet. The gateways of VCN 602 can further include service gateway 610 for connecting VCN 602 to a service network 612. Service network 612 can include one or more databases and / or stores, including, for example, autonomous database 614 and / or object store 616. The service network can include a conceptual network including a collection of IP ranges, which can be, for example, public IP ranges. In some embodiments, these IP ranges can cover some or all of the public services offered by the CSPI 601 provider. These services can be accessed, for example, through an Internet gateway or a NAT gateway. In some embodiments, the service network provides a dedicated gateway (service gateway) for accessing the services from a local area within the service network. In some embodiments, the backend of these services may be on its own private network, for example. In some embodiments, the service network 612 may include further additional databases.
[0161] VCN 602 can include multiple virtual networks. Each of these networks can include one or more compute instances that can communicate within the network, between networks, or outside of VCN 602. One of VCN 602's virtual networks is L3 subnet 620. L3 subnet 620 is a unit of organization or partition created within VCN 602. Subnet 620 can include a virtual Layer 3 network within VCN 602's virtualized cloud environment, which is hosted on the underlying physical network of CPSI 601. While FIG. 6 shows only one subnet 620, VCN 602 can include one or more subnets. Each subnet within VCN 602 can be associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets within the VCN and represents an address space subset of the VCN's address space. In some embodiments, this IP address space may be separate from the address space associated with CPSI 601.
[0162] Subnet 620 includes one or more compute instances, specifically, first compute instance 622-A and second compute instance 622-B. Compute instances 622-A and 622-B can communicate with each other within subnet 620 or with other instances, devices, and / or networks outside subnet 620. Virtual router (VR) 624 enables communication outside subnet 620. VR 624 enables communication between subnet 620 and other networks in VCN 602. In the case of subnet 620, VR 624 represents a logical gateway that enables subnet 620 (i.e., compute instances 622-A and 622-B) to communicate with endpoints on other networks within VCN 602 and with other endpoints outside VCN 602.
[0163] VCN 602 may further include additional networks, specifically, one or more L2 VLANs (referred to as VLANs in this disclosure), which are an example of a virtual L2 network. Each of these one or more VLANs may include a virtual Layer 2 network located in VCN 602's cloud environment and / or hosted by the underlying physical network of CPSI 601. In the embodiment of FIG. 6, VCN 602 includes VLAN A 630 and VLAN B 640. Each VLAN 630, 640 in VCN 602 may be associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other networks, such as other subnets or VLANs, within the VCN and represents an address space subset of the VCN's address space. In some embodiments, the IP address space of this VLAN may be isolated from the address space associated with CPSI 601.
[0164] Each of the VLANs 630, 640 can include one or more compute instances. Specifically, VLAN A 630 can include, for example, a first compute instance 632-A and a second compute instance 632-B. In some embodiments, VLAN A 630 can include additional compute instances. VLAN B 640 can include, for example, a first compute instance 642-A and a second compute instance 642-B. Each of the compute instances 632-A, 632-B, 642-A, and 642-B can have an IP address and a MAC address. These The addresses may be assigned or generated in any desired manner. In some embodiments, these addresses may be within the CIDR of the compute instance's VLAN. In some embodiments, these addresses may be arbitrary addresses. In embodiments in which the compute instances of a VLAN communicate with endpoints outside the VLAN, one or both of these addresses will be from the VLAN CIDR, but if all communications are intra-VLAN communications, these addresses are not limited to addresses within the VLAN CIDR. In contrast to networks in which addresses are assigned by the control plane, the IP addresses and / or MAC addresses of compute instances in a VLAN may be assigned by users / customers of that VLAN, and these IP addresses and / or MAC addresses may be discovered and / or learned by the compute instances in the VLAN according to the learning process described below.
[0165] Each VLAN may include a VLAN Switching and Routing Service (VSRS). Specifically, VLAN A 630 includes VSRS A 634, and VLAN B 640 includes VSRS B 644. Each VSRS 634, 644 participates in Layer 2 switching and local learning within the VLAN and performs all necessary Layer 3 network functions, including ARP, NDP, and routing. The VSRS performs ARP (a Layer 2 protocol) because it is required to map IP to MAC.
[0166] In these cloud-based VLANs, each virtual interface or virtual gateway may be associated with one or more media access control (MAC) addresses, and the one or more MAC addresses may be virtual MAC addresses. One or more compute instances 632-A, 632-B, 642-A, 642-B and / or one or more service instances within a VLAN, which may be bare metal, VMs, or containers, can communicate directly with each other through a virtual switch. Communication outside the VLAN with other VLANs or L3 networks is enabled through a VSRS 634, 644. The VSRS 634, 644 is a distributed service that provides Layer 3 functions, such as IP routing, to the VLAN network. In some embodiments, the VSRS 634, 644 is a horizontally scalable, highly available routing service located at the intersection of IP and L2 networks and capable of participating in IP routing and L2 learning within the cloud-based L2 domain.
[0167] The VSRSs 634, 644 may be distributed across multiple nodes within the infrastructure. The VSRS 634, 644 functionality is scalable, specifically horizontally scalable. In some embodiments, each of the nodes implementing the VSRS 634, 644 functionality shares and replicates router and / or switch functionality with each other. These nodes may also present themselves as a single VSRS 634, 644 to all instances within the VLANs 630, 640. The VSRSs 634, 644 may be implemented on any virtualization device within the CSPI 601, specifically a virtual network. Thus, in some embodiments, the VSRSs 634, 644 may be implemented on any virtual network virtualization device, including a NIC, a smart NIC, a switch, a smart switch, or a general-purpose compute host.
[0168] The VSRS 634, 644 may be a service that resides on one or more hardware nodes, e.g., one or more servers, such as one or more x86 servers, or one or more networking devices, e.g., one or more NICs, particularly one or more smart NICs, and supports a cloud network. In some embodiments, the VSRS 634, 644 may be implemented on a cluster of servers. Thus, the VSRS 634, 644 may be a service that is distributed across a cluster of nodes, which may provide routing and security services. The VSRS evaluates security policies and participates in and shares L2 and L3 learning, and may be a centrally managed virtual networking enforcer cluster or may be distributed at the edge of the virtual networking enforcer cluster. In some embodiments, each VSRS instance can update other VSRS instances with new mapping information learned by one VSRS instance. For example, when one VSRS instance learns the IP, interface, and / or MAC mappings of one or more CIs in a VLAN, that VSRS instance can provide that update information to other VSRS instances in the VCN. Through this mutual update, a VSRS instance associated with a first VLAN can learn mappings including the IP, interface, and / or MAC mappings of CIs in other VLANs, in some embodiments, in VCN 602. These updates can be significantly accelerated if the VSRS resides on a cluster of servers and / or is distributed across a cluster of nodes.
[0169] In some embodiments, the VSRS 634, 644 can host one or more higher-level services required for networking, including, but not limited to, DHCP relay, DHCP (hosting), DHCPv6, neighbor discovery protocols such as IPv6 Neighbor Discovery Protocol, DNS, hosting DNSv6, SLAAC for IPv6, NTP, metadata services, and block store mount points. In some embodiments, the VSRS can support one or more network address translation (NAT) functions for translating network address spaces. In some embodiments, the VSRS can incorporate anti-spoofing, MAC spoofing anti-spoofing, ARP cache poisoning protection for IPv4, IPv6 route advertisement (RA) defense, DHCP defense, packet filtering using access control lists (ACLs), and / or reverse path forwarding checks. The VSRS can implement functions including, for example, ARP, GARP, packet filters (ACLs), DHCP relay, and / or IP routing protocols. The VSRS 634, 644 may, for example, learn MAC addresses, invalidate expired MAC addresses, handle MAC address movement, obtain MAC address information, handle MAC information flooding, handle storm control, prevent loops, perform Layer 2 multicast via protocols such as IGMP within the cloud, collect statistics including logs and statistics using SNMP, and / or monitor, collect, and use statistics such as broadcasts, total traffic, bits, spanning tree packets, etc.
[0170] In the virtual network, the VSRS 634, 644 can appear as different instantiations. In some embodiments, each of these VSRS instantiations can be associated with a VLAN 630, 640, and in some embodiments, each VLAN 630, 640 can have an instantiation of the VSRS 634, 644. In some embodiments, each instantiation of the VSRS 634, 644 can have one or more unique tables corresponding to the VLAN 630, 640 with which the VSRS 634, 644 instantiation is associated. Each instantiation of the VSRS 634, 644 can generate and / or curate unique tables associated with the VSRS 634, 644 instantiation. Thus, while a single service can provide VSRS 634, 644 functionality for one or more cloud networks, individual instantiations of the VSRS 634, 644 within a cloud network can have unique Layer 2 and Layer 3 forwarding tables, and such multiple customer networks can have overlapping Layer 2 and Layer 3 forwarding tables.
[0171] In some embodiments, the VSRS 634, 644 may be configured to communicate across multiple tenants. Conflicting VLANs and IP spaces can be supported. This can include having multiple tenants on the same VSRS 634, 644. In some embodiments, some or all of these tenants can choose to use some or all of the same IP address space, the same MAC space, and the same VLAN space. This can provide great flexibility to users in selecting addresses. In some embodiments, this multi-tenancy is supported by providing each tenant with a separate virtual network, which is a private network within the cloud network. Each virtual network is given a unique identifier. Similarly, in some embodiments, each host can have a unique identifier, and / or each virtual interface or virtual gateway can have a unique identifier. In some embodiments, these unique identifiers, specifically the unique identifier of the tenant's virtual network, can be encoded in each communication. By giving each virtual network a unique identifier and including it in the communication, a single instantiation of the VSRS 634, 644 can serve multiple tenants with overlapping addresses and / or namespaces.
[0172] The VSRSs 634, 644 may perform these switching and / or routing functions to facilitate and / or enable the creation of and / or communication with L2 networks within the VLANs 630, 640. The VLANs 630, 640 may be deployed in a cloud computing environment, and more specifically, in virtual networks within the cloud computing environment.
[0173] For example, each of VLANs 630, 640 includes multiple compute instances 632-A, 632-B, 642-A, and 642-B. VSRS 634, 644 enable communication between compute instances in one VLAN 630 or 640 and compute instances in another VLAN 630, VLAN 640, or subnet 620. In some embodiments, VSRS 634, 644 enable communication between compute instances in one VLAN 630 or 640 and another network outside the VCN, including another VCN, the Internet, an on-premises data center, etc. In such embodiments, a compute instance, such as compute instance 632-A, can send information to an endpoint outside the VLAN, in this example, L2VLAN A 630. The computation instance (632-A) can send the information to VSRS A 634, which can send the information to a router 624, 644 or a gateway 604, 608, 610 communicatively connected to the desired endpoint. The router 624, 644 or a gateway 604, 608, 610 communicatively connected to the desired endpoint can receive the information from the computation instance (632-A) and send it to the desired endpoint.
[0174] Referring to FIG. 7, FIG. 7 illustrates a logical and hardware schematic of a VLAN 700. As illustrated, the VLAN 700 includes multiple endpoints, specifically multiple compute instances and a VSRS. Multiple compute instances (CIs) are instantiated on one or more host machines. In some embodiments, this may be a one-to-one relationship, where each CI is instantiated on a unique host machine, and / or in some embodiments, this may be a many-to-one relationship, where multiple CIs are instantiated on a single common host machine. In various embodiments, the CIs may be Layer 2 CIs configured to communicate with each other using an L2 protocol. FIG. 7 illustrates a scenario in which some CIs are instantiated on unique host machines and some CIs share a common host machine. As shown in FIG. 7, Instance 1 (CI1) 704-A is instantiated on Host Machine 1 702-A, and Instance 2 (CI2) 704-B is instantiated on Host Machine 2 702-B. B, and instance 3 (CI3) 704-C and instance 4 (CI4) 704-D are instantiated on common host machine 702-C.
[0175] Each of CIs 704-A, 704-B, 704-C, and 704-D is communicatively connected to other CIs 704-A, 704-B, 704-C, and 704-D in the VLAN 700 and to the VSRS 714. Specifically, each of CIs 704-A, 704-B, 704-C, and 704-D is connected to other CIs 704-A, 704-B, 704-C, and 704-D in the VLAN 700 via an L2 VNIC and a switch and is connected to the VSRS 714. Each CI 704-A, 704-B, 704-C, and 704-D is associated with a unique L2 VNIC and switch. The switch may be a local L2 virtual switch uniquely associated with the L2 VNIC and deployed for the L2 VNIC. Specifically, CI1 704-A is associated with L2VNIC 1 708-A and switch 1 710-A, CI2 704-B is associated with L2VNIC 2 708-B and switch 710-B, CI3 704-C is associated with L2VNIC 3 708-C and switch 3 710-C, and CI4 704-D is associated with L2VNIC4 708-D and switch 4 710-D.
[0176] In some embodiments, each L2VNIC 708 and associated switch 710 may be instantiated on the NVD 706. This instantiation may be a one-to-one relationship, where a single L2VNIC 708 and associated switch 710 is instantiated on a unique NVD 706, or this instantiation may be a many-to-one relationship, where multiple L2VNICs 708 and associated switches 710 are instantiated on a single, common NVD 706. Specifically, L2VNIC1 708-A and switch 1 710-A are instantiated on NVD1 706-A, L2VNIC2 708-B and switch 2 710-B are instantiated on NVD2, and both L2VNIC3 708-C and switch 3 710-C and L2VNIC4 708-D and switch 710-D are instantiated on a common NVD, i.e., NVD 706-C.
[0177] In some embodiments, the VSRS 714 can support competing VLANs and IP spaces across multiple tenants. This can include having multiple tenants on the same VSRS 714. In some embodiments, some or all of these tenants can select and use some or all of the same IP address space, the same MAC space, and the same VLAN space. This can provide users with great flexibility in selecting addresses. In some embodiments, this multi-tenancy is supported by providing each tenant with a separate virtual network, which is a private network within the cloud network. Each virtual network (e.g., each VLAN or VCN) is given a unique identifier, such as a VCN identifier, which may be a VLAN identifier. This unique identifier may be selected, for example, by the control plane, specifically the CSPI control plane. In some embodiments, this unique identifier can include one or more bits that can be included and / or used in packet encapsulation.
[0178] Similarly, in some embodiments, each host may have a unique identifier, and / or each virtual interface or virtual gateway may have a unique identifier. In some embodiments, these unique identifiers, specifically the unique identifier of a tenant's virtual network, may be encoded into each communication. By giving each virtual network a unique identifier and including it in the communication, a single instantiation of the VSRS can communicate with multiple tenants with overlapping addresses and / or namespaces. It can provide services to clients.
[0179] In some embodiments, the VSRS 714 can determine to which tenant a packet belongs based on a VCN identifier and / or a VLAN identifier associated with the communication, specifically based on the VCN identifier and / or the VLAN identifier in the VCN header of the communication. In embodiments disclosed in this disclosure, communications entering or leaving a VLAN can have a VCN header that may include a VLAN identifier. Based on the VCN header that includes the VLAN identifier, the VSRS can determine the tenant. In other words, the receiving VSRS can determine to which VLAN and / or which tenant to send the communication.
[0180] Additionally, each compute instance (e.g., L2 compute instance) belonging to a VLAN is given a unique interface identifier that identifies the L2 VNIC associated with that compute instance. The interface identifier may be included in traffic from and / or to the computer instance (e.g., in a frame header) and may be used by the NVD to identify the L2 VNIC associated with the compute instance. In other words, the interface identifier may uniquely indicate the compute instance and / or its associated L2 VNIC. As shown in FIG. 7 , switches 710-A, 710-B, 710-C, and 710-D together may form an L2 distributed switch 712, also referred to as a distributed switch 712 in this disclosure. From a customer's perspective, each switch 710-A, 710-B, 710-C, and 710-D in the L2 distributed switch 712 is a single switch connected to all CIs in the VLAN. However, a distributed switch that emulates the user experience of a single switch is infinitely scalable and includes a collection of local switches (e.g., in the example of FIG. 7, switches 710-A, 710-B, 710-C, and 710-D). As shown in FIG. 7, each CI runs on a host machine connected to the NVD. For each CI on a host connected to the NVD, the NVD hosts a Layer 2 VNIC associated with the compute instance and a local switch (e.g., an L2 virtual switch located locally on the NVD, associated with the Layer 2 VNIC, and a member or component of the L2 distributed switch 712). The Layer 2 VNIC represents a port of the compute instance within a Layer 2 VLAN. The local switch connects the L2 VNIC to other L2 VNICs (e.g., other ports) associated with other compute instances within the Layer 2 VLAN.
[0181] Each of CIs 704-A, 704-B, 704-C, and 704-D can communicate with others of CIs 704-A, 704-B, 704-C, and 704-D in VLAN 700 or can communicate with VSRS 714. One of CIs 704-A, 704-B, 704-C, and 704-D transmits a packet to another of CIs 704-A, 704-B, 704-C, and 704-D or VSRS 714 by sending the packet to the MAC address and interface of the receiving CI of one of CIs 704-A, 704-B, 704-C, and 704-D or VSRS 714. The MAC address and interface identifier may be included in the packet's header. As explained above, the interface identifier may indicate the L2VNIC of the receiving CI of one of CIs 704-A, 704-B, 704-C, and 704-D, or the L2VNIC of VSRS 714.
[0182] In one embodiment, CI1 704-A may be the source CI, L2VNIC 708-A may be the source VNIC, and switch 710-A may be the source switch. In this embodiment, CI3 704-C may be the destination CI, and L2VNIC3 708-C may be the destination VNIC. The source CI may send a packet that includes a source MAC address and a destination MAC address. This packet may be intercepted by the NVD 706-A, which instantiates the source VNIC and source switch.
[0183] Each of the L2 VNICs 708-A, 708-B, 708-C, and 708-D of the VLAN 700 can learn the MAC address-to-interface identifier mapping for the L2 VNIC. This mapping can be learned based on packets and / or information received from the VLAN 700. The source VNIC can determine the interface identifier of the destination interface associated with the destination CI within the VLAN based on this predetermined mapping and can encapsulate the packet. In some embodiments, this encapsulation can include GENEVE encapsulation, specifically, L2GENEVE encapsulation, which includes encapsulating the packet's L2 (Ethernet) header. The encapsulated packet can identify the destination MAC, the destination interface identifier, the source MAC, and the source interface identifier.
[0184] The source VNIC can send the encapsulated packet to the source switch, and the source switch can send the packet to the destination VNIC. Upon receiving the packet, the destination VNIC can decapsulate the packet and then provide the packet to the destination CI.
[0185] Referring to Figure 8, a logical schematic diagram of multiple connected L2 VLANs 800 is shown. In the particular embodiment shown in Figure 8, both VLANs are located in the same VCN. As shown, the multiple connected L2 VLANs 800 can include a first VLAN, VLAN A 802-A, and a second VLAN, VLAN B 802-B. Each of these VLANs 802-A, 802-B can include one or more CIs, and each CI can have an associated L2 VNIC and an associated L2 virtual switch. Additionally, each of these VLANs 802-A, 802-B can include a VSRS.
[0186] Specifically, VLAN A 802-A can include instance 1 804-A connected to L2VNIC1 806-A and switch 1 808-A, instance 2 804-B connected to L2VNIC2 806-B and switch 808-B, and instance 3 804-C connected to L2VNIC3 806-C and switch 3 808-C. VLAN B 802-B can include instance 4 804-D connected to L2VNIC4 806-D and switch 4 808-D, instance 5 804-E connected to L2VNIC5 806-E and switch 808-E, and instance 6 804-F connected to L2VNIC6 806-F and switch 3 808-F. 804-F and 804-F. Also, VLAN A 802-A can include VSRS A VLAN A 802-B may include VSRS B 810-A, and VLAN B 802-B may include VSRS B 810-B. Each of CIs 804-A, 804-B, and 804-C in VLAN A 802-A may be communicatively connected to VSRS A 810-A, and each of CISs 804-D, 804-E, and 804-F in VLAN B 802-B may be communicatively connected to VSRS B 810-B.
[0187] VLAN A 802-A may be communicatively connected to VLAN B 802-B through corresponding VSRSs 810-A, 810-B. Similarly, each VSRS may be connected to a gateway 812, which can provide access from CIs 804-A, 804-B, 804-C, 804-D, 804-E, and 804-F in each VLAN 802-A, 802-B to other networks outside the VCN in which the VLANs 802-A, 802-B are located. In some embodiments, these The network can include, for example, one or more on-premises networks, another VCN, a service network, or a public network such as the Internet.
[0188] CIs 804-A, 804-B, and 804-C in VLAN A 802-A communicate with each other via VSRSs 810-A and 810-B in each VLAN 802-A and 802-B. CIs 804-D, 804-E, and 804-F in VLAN B 802-B can communicate with CIs 804-D, 804-E, and 804-F in VLAN B 802-B. For example, one of CIs 804-A, 804-B, 804-C, 804-D, 804-E, and 804-F located in one of VLANs 802-A and 802-B can transmit a packet to CIs 804-A, 804-B, 804-C, 804-D, 804-E, and 804-F located in the other of VLANs 802-A and 802-B. This packet may be transmitted from the source VLAN via the VSRS of the source VLAN, received by the destination VLAN, and forwarded to the destination CI via the destination VSRS.
[0189] In one embodiment, C1 1 804-A may be the source C1, L2VNIC 806-A may be the source VNIC, and switch 808-A may be the source switch. In this embodiment, C1 5 804-E may be the destination CI, and L2VNIC5 806-E may be the destination VNIC. VSRS A 810-A may be the source VSRS identified as an SVSRS, and VSRS B 810-B may be the destination VSRS identified as a DVSRS.
[0190] The source CI can send a packet including the MAC address. This packet may be intercepted by the NVD and the source switch that instantiates the source VNIC. The source VNIC encapsulates the packet. In some embodiments, this encapsulation can include Geneve encapsulation, specifically L2 Geneve encapsulation. The encapsulated packet can identify a destination address of the destination CI. In some embodiments, this destination address can also include a destination address of the destination VSRS. The destination address of the destination CI can include a destination IP address, a destination MAC of the destination CI, and / or a destination interface identifier of a destination VNIC of the destination CI. The destination address of the destination VSRS can include an IP address of the destination VSRS, an interface identifier of a destination VNIC associated with the destination VSRS, and / or a MAC address of the destination VSRS.
[0191] The source VSRS can receive the packet from the source switch, look up the VNIC mapping from the packet's destination address, which may be the destination IP address, and forward the packet to the destination VSRS. The destination VSRS can receive the packet. Based on the destination address included in the packet, the destination VSRS can forward the packet to the destination VNIC. After receiving and decapsulating the packet, the destination VNIC can provide the packet to the destination CI.
[0192] Referring to Figure 9, Figure 9 illustrates a logical schematic diagram of multiple connected L2 VLANs and subnets 900. In the particular embodiment illustrated in Figure 9, both the VLANs and subnets are located in the same VCN, which means that the virtual routers and VSRSs for both the VLANs and subnets are directly connected rather than connected through a gateway.
[0193] As shown, this can include a first VLAN, VLAN A 902-A, a second VLAN, VLAN B 902-B, and a subnet 930. Each of these VLANs 902-A, 902-B can include one or more CIs, and each CI can include an associated L2 VNIC and an associated L2 switch. Also, each of these VLANs 902-A, 902-B can include a VSRS. Additionally, the subnet 930, which may be an L3 subnet, may include one or more CIs, each CI may include an associated L3 VNIC, and the L3 subnet 930 may include a virtual router 916.
[0194] Specifically, VLAN A 902-A can include instance 1 904-A connected to L2VNIC1 906-A and switch 1 908-A, instance 2 904-B connected to L2VNIC2 906-B and switch 908-B, and instance 3 904-C connected to L2VNIC3 906-C and switch 3 908-C. VLAN B 902-B can include instance 4 904-D connected to L2VNIC4 906-D and switch 4 908-D, instance 5 904-E connected to L2VNIC5 906-E and switch 908-E, and instance 6 904-F connected to L2VNIC6 906-F and switch 3 908-F. 904-F. VLAN A 902-A may further include VSRS A 910-A, and VLAN B 902-B may include VSRS B 910-B. Each of the CIs 904-A, 904-B, and 904-C of VLAN A 902-A may be communicatively connected to VSRS A 910-A, and each of the CISs 904-D, 904-E, and 904-F of VLAN B 902-B may be communicatively connected to VSRS B 910-B. L3 subnet 930 may include one or more CIs, and specifically may include instance 7 904-G communicatively connected to L3 VNIC7 906-G. L3 subnet 930 may include a virtual router 916.
[0195] VLAN A 902-A may be communicatively connected to VLAN B 902-B through corresponding VSRSs 910-A, 910-B. L3 subnet 930 may be communicatively connected to VLAN A 902-A and VLAN B 902-B through a virtual router 916. Similarly, virtual router 916 and each of VSRS instances 910-A, 910-B may be connected to a gateway 912, which can provide access from CIs 904-A, 904-B, 904-C, 904-D, 904-E, 904-F, and 904-G in each VLAN 902-A, 902-B and subnet 930 to other networks outside the VCN in which VLAN 902-A, 902-B and subnet 930 are located. In some embodiments, these networks may include, for example, one or more on-premises networks, another VCN, a service network, or a public network such as the Internet.
[0196] Each VSRS instance 910-A, 910-B can provide a transmit path for packets exiting the associated VLAN 902-A, 902-B and a receive path for packets entering the associated VLAN 902-A, 902-B. Packets may be sent from the VSRS instance 910-A, 910-B of a VLAN 902-A, 902-B to any desired endpoint, including an L2 endpoint, e.g., an L2 CI, in another VLAN on the same VCN, a different VCN, or network, and / or an L3 endpoint, e.g., an L3 CI, in a subnet on the same VCN, a different VCN, or network.
[0197] In one embodiment, CI1 904-A may be the source CI, L2VNIC 906-A may be the source VNIC, and switch 908-A may be the source switch. In this embodiment, CI7 904-G may be the destination CI, and VNIC7 906-G may be the destination VNIC. VSRS A 910-A may be the source VSRS identified as an SVSRS, and virtual router (VR) 916 may be the destination VR.
[0198] The source CI can send a packet containing the MAC address. The packet may be intercepted by the NVD and source switch instantiating the source VNIC. The source VNIC encapsulates the packet. In some embodiments, this encapsulation may include Geneve encapsulation, specifically L2 Geneve encapsulation. The encapsulated packet may identify a destination address of the destination CI. In some embodiments, this destination address may also include a destination address of a VSRS of the VLAN of the source CI. The destination address of the destination CI may include a destination IP address, a destination MAC of the destination CI, and / or a destination interface identifier of the destination VNIC of the destination CI.
[0199] The source VSRS can receive the packet from the source switch, look up the VNIC mapping from the packet's destination address, which may be the destination IP address, and forward the packet to the destination VR. The destination VR can receive the packet. Based on the destination address contained in the packet, the destination VR can forward the packet to the destination VNIC. After receiving and decapsulating the packet, the destination VNIC can provide the packet to the destination CI.
[0200] Learning in a virtual L2 network 10, a schematic diagram of one embodiment of intra-VLAN communication and learning within a VLAN 1000 is shown. This learning is specific to how L2 VNICs, VSRS VNICs, and / or L2 virtual switches learn associations between MAC addresses and L2 VNICs / VSRS VNICs (more specifically, between MAC addresses associated with L2 compute instances or VSRSs and interface identifiers associated with L2 VNICS of those L2 compute instances or associated with VSRS VNICs). Generally, learning is based on ingress traffic. In the case of interface-to-MAC address learning aspects, In this case, this learning is distinct from the learning process (e.g., ARP process) that the L2 compute instance performs to learn the destination MAC address. The two learning processes (e.g., of the L2 VNIC / L2 virtual switch and the L2 compute instance) are shown as being performed jointly in FIG. 12.
[0201] As shown, VLAN 1000 includes compute instance 1 1000-A, which is communicatively connected to NVD1 1001-A, which instantiates L2VNIC1 1002-A and L2 switch 1 1004-A. VLAN 1000 also includes compute instance 2 1000-B, which is communicatively connected to NVD2 1001-B, which instantiates L2VNIC2 1002-B and L2 switch 2 1004-A. VLAN 1000 further includes VSRS 1010, which runs on a group of servers and includes VSRS VNIC 1002-C and VSRS switch 1004-C. All of switches 1004-A, 1004-B, and 1004-C together form a distributed switch. VSRS 1010 is communicatively connected to endpoint 1008. The endpoint 1008 may include a gateway, and in particular may include an L2 / L3 router, for example in the form of another VSRS, or an L3 router, for example in the form of a virtual router.
[0202] The control plane 1001 of the VCN hosting VLAN 1000 maintains information to identify each L2 VNIC on VLAN 1000 and the network placement of the L2 VNIC. For example, this information may include an interface identifier associated with the L2 VNIC and / or the physical IP address of the NVD hosting the L2 VNIC. The control plane 1001 uses this information to update (e.g., periodically or on-demand) the interfaces in VLAN 1000. Thus, each L2 VNIC 1002-A, 1002-B, 1002-C in the VLAN receives information from the control plane 1001 to identify the interfaces in the VLAN and stores this information in Populate the table. The table populated by the L2VNIC may be stored locally on the NVD hosting the L2VNIC. If the L2VNIC 1002-A, 1002-B, 1002-C already contains a current table, the L2VNIC 1002-A, 1002-B, 1002-C can determine a discrepancy between the current table of the L2VNIC 1002-A, 1002-B, 1002-C and the information / table received from the control plane 1001. In some embodiments, the L2VNIC 1002-A, 1002-B, 1002-C can update the table to match the information received from the control plane 1001.
[0203] 10, a packet is transmitted through L2 switches 1004-A, 1004-B, and 1004-C and received by receiving L2 VNICs 1002-A, 1002-B, and 1002-C. When a packet is received by an L2 VNIC 1002-A, 1002-B, or 1002-C, that VNIC learns the packet's source interface (source VNIC) and source MAC address mapping. Based on a table of information received from control plane 1010, the VNIC can map the source MAC address (from the received packet, also referred to as a frame in this disclosure) to the source VNIC's interface identifier, the VNIC's IP address, and / or the IP address of the NVD hosting the VNIC (the interface identifier and IP address are available from the table). Thus, the L2VNICs 1002-A, 1002-B, and 1002-C can learn interface identifier-to-MAC address mappings based on received information and / or packets, and the L2VNICs 1002-A, 1002-B, and 1002-C can use the learned mapping information to update their tables, i.e., L2 forwarding tables 1006-A, 1006-B, and 1006-C. In some embodiments, the L2 forwarding tables include MAC addresses and associate the MAC addresses with at least one of interface identifiers or physical IP addresses. In such embodiments, the MAC addresses are addresses assigned to the L2 compute instances and correspond to ports emulated by the L2VNICs associated with the L2 compute instances. The interface identifiers can uniquely identify the L2VNICs and / or the L2 compute instances. The virtual IP addresses may be the virtual IP addresses of the L2VNICs. The physical IP addresses may be the IP addresses of the NVDs hosting the L2VNICs. The L2 forwarding table updated by the L2VNIC is stored locally on the NVD hosting the L2VNIC and can be used by the L2 virtual switch associated with the L2VNIC to transmit frames.In some embodiments, L2 VNICs in a common VLAN can share all or part of a mapping table.
[0204] Traffic flow will be described below with reference to the network architecture described above. For clarity, traffic flow will be described with respect to compute instance 2 1000-B, L2VNIC2 10002-B, L2 switch 2 1004-B, and NVD2 1001-B. This description applies equally to traffic flow to and / or from other compute instances.
[0205] As described above, a VLAN is implemented in a VCN as an overlay L2 network on an L3 physical network. An L2 compute instance of a VLAN can send or receive an L2 frame that includes an overlay MAC address (also called a virtual MAC address) as the source MAC address and destination MAC address. The L2 frame can also encapsulate a packet that includes an overlay IP address (also called a virtual IP address) as the source IP address and destination IP address. In some embodiments, the overlay IP address of a compute instance is configured to be overlaid on the VLAN. Other overlay IP addresses can be located within the CIDR range (in this case, the L2 frame flows within the VLAN) or outside the CIDR range (in this case, the L2 frame is sent to or received from another network). L2 frames can include a VLAN tag to uniquely identify the VLAN. This VLAN tag can be used to distinguish between multiple L2 VNICs on the same NVD. L2 frames can be received via a tunnel as packets encapsulated by the NVD from the host machine of a compute instance, from another NVD, or from a group of servers hosting a VSRS. In these different cases, the encapsulated packets can be L3 packets sent over the physical network, with the source and destination IP addresses being physical IP addresses. Different types of encapsulation are possible, including Geneve encapsulation. The NVD can decapsulate the received packets to extract the L2 frames. Similarly, to transmit an L2 frame, the NVD can encapsulate the L2 frame in an L3 packet and send it over the physical board.
[0206] For intra-VLAN outbound traffic from compute instance 2 1000-B, NVD2 1001-B receives a frame from the host machine of instance 2 1000-B over an Ethernet link. The frame includes an interface identifier to identify L2VNIC2 1000-B. The frame includes the overlay MAC address of compute instance 2 1000-B (e.g., M.2) as the source MAC address and the overlay MAC address of compute instance 1 1000-A (e.g., M.1) as the destination MAC address. Given the interface identifier, NVD2 1001-B routes the frame to L2VNIC2 for further processing. L2VNIC2 1002-B passes the frame to L2 switch 2 1004-B. L2 switch 2 1004-B determines, based on L2 forwarding table 1006-B, whether the destination MAC address is known (e.g., whether it matches an entry in L2 forwarding table 1006-B).
[0207] If so, L2 Switch 2 1004-B determines that L2VNIC1 1002-A is the associated tunnel endpoint and forwards the frame to L2VNIC1 1002-A. This forwarding may include encapsulating the frame into a packet and decapsulating the packet (e.g., Geneve encapsulation and decapsulation), where the packet includes the frame, the physical IP address of NVD1 1001-A (e.g., IP.1) as the destination address, and the physical IP address of NVD2 1001-B (e.g., IP.2) as the source address.
[0208] If unknown, L2 switch 2 1004-B broadcasts the frame to the various L2 VNICs of the VLAN (e.g., including L2VNIC1 1002-A and any other L2 VNICs of the VLAN). The broadcasted frame is processed (e.g., encapsulated, transmitted, decapsulated) between the associated NVDs. In some embodiments, this broadcast is performed, or more specifically emulated, in a physical network. This physical network can encapsulate the frame to each L2 VNIC that includes the VSRS of the VLAN. Thus, the broadcast is emulated in the physical network via a series of replicated unicast packets. Each L2 VNIC then receives the frame and learns the association between L2VNIC2 1002-B's interface identifier and the source MAC address (e.g., M.2) and source physical IP address (e.g., IP.2).
[0209] For intra-VLAN inbound traffic from compute instance 1 1000-A to compute instance 2 1000-B, NVD2 1001-B receives the packet from NVD1. The packet has IP.1 as the source address and the frame includes M.2 as the destination MAC address and M.1 as the source MAC address. The frame also includes the network identifier of L2VNIC1 1002-A. Upon decapsulation, L2VNIC2 receives the frame, learns that this interface identifier is associated with M.1 and / or IP.1, and stores this learned information in switch 2's L2 forwarding table 1006-B for later outgoing traffic if it was previously unknown. Alternatively, upon decapsulation, L2VNIC2 receives the frame, learns that this interface identifier is associated with M.1 and / or IP.1, and refreshes the expiration date if this information is known.
[0210] For outbound traffic sent from instance 2 1000-B in VLAN 1000 to an instance in another VLAN, the flow is similar to the outbound traffic flow described above, except that a VSRS VNIC and a VSRS switch are used. Specifically, the destination MAC address is not within the L2 broadcast range of VLAN 1000 (it is within another L2 VLAN). Therefore, the overlay destination IP address of the destination instance (e.g., IP.A) is used to send this outbound traffic. For example, L2VNIC2 1002-B determines that IP.A is outside the CIDR range of VLAN 1000. Therefore, L2VNIC2 1002-B sets the destination MAC address to the default gateway MAC address (e.g., M.DG). L2 Switch 2 1004-B sends the outbound traffic to the VSRS VNIC (e.g., via a tunnel, with appropriate end-to-end encapsulation) based on M.DG. The VSRS VNIC forwards the outbound traffic to the VSRS switch, which performs the routing function. The VSRS switch in VLAN 1000 sends the outgoing traffic to the VSRS switch in the other VLAN (e.g., via a virtual router between these two VLANs, with appropriate end-to-end encapsulation) based on the overlay destination IP address (e.g., IP.A). The VSRS switch in the other VLAN then determines that IP.A falls within the CIDR range of this VLAN, performs its switching function, and determines the destination MAC address associated with IP.A by looking up its ARP cache based on IP.A. If no match exists in the ARP cache, it determines the destination MAC address by sending an ARP request to a different L2 VNIC in the other VLAN. Otherwise, the VSRS switch sends the outgoing traffic to the associated VNIC (e.g., via a tunnel, with appropriate encapsulation).
[0211] For inbound traffic from an instance in another VLAN to an instance in VLAN1000, the traffic flow is similar to the above, except that it occurs in the reverse direction. For outbound traffic from an instance in VLAN1000 to the L3 network, the traffic flow is similar to the above, except that the VSRS switch for VLAN1000 routes the packet directly to the destination VNIC in the virtual L3 network via the virtual router (e.g., without routing the packet through another VSRS switch). For inbound traffic from the virtual L3 network to an instance in VLAN1000, the traffic flow is similar to the above, except that the VSRS switch for VLAN1000A, which sent the packet as a frame within the VLAN, receives the packet. For traffic between VLAN1000 and other networks (outbound or inbound), the VSRS switch similarly uses its routing function to send the packet through the appropriate gateway (e.g., IGW, NGW, DRG, SGW, LPG) for outbound traffic, and its switching function to send the frame within VLAN1000 for inbound traffic.
[0212] Referring to FIG. 11, a schematic diagram of one embodiment of a VLAN 1100 (e.g., a cloud-based virtual L2 network) is shown, specifically illustrating an implementation of a VLAN. .
[0213] As described above in this disclosure, a VLAN can include “n” compute instances 1102-A, 1102-B, 1102-N, each running on a host machine. As described above, there may be a one-to-one association between a compute instance and a host machine, or a many-to-one association between multiple compute instances and a single host machine. Each compute instance 1102-A, 1102-B, 1102-N may be an L2 compute instance and is associated with at least one virtual interface (e.g., L2 VNIC) 1104-A, 1104-B, 1104-N and a switch 1106-A, 1106-B, 1106-N. The switches 1106-A, 1106-B, 1106-N are L2 virtual switches and together form an L2 distributed switch 1107.
[0214] A pair of L2VNICs 1104-A, 1104-B, 1104-N and switches 1106-A, 1106-B, 1106-N associated with compute instances 1102-A, 1102-B, 1102-N on a host machine is a pair of software modules on NVDs 1108-A, 1108-B, 1108-N connected to the host machine. Each L2VNIC 1104-A, 1104-B, 1104-N represents an L2 port of a single customer-perceived switch (referred to as a vSwitch in this disclosure). Generally, host machine "i" runs compute instance "i" and is connected to NVD "i". Similarly, NVD "i" runs L2VNIC "i" and switch "i". L2VNIC "i" represents L2 port "i" of the vSwitch, where i is a positive integer from 1 to n. Again, while a one-to-one association has been described, other types of associations are possible. For example, a single NVD can be connected to multiple hosts, each running one or more compute instances belonging to a VLAN. In this case, the NVD hosts multiple pairs of L2VNICs and switches, each corresponding to one of the compute instances.
[0215] A VLAN may include an instance of VSRS 1110. VSRS 1110 performs switching and routing functions and includes instances of VSRS VNIC 1112 and VSRS Switch 1114. VSRS VNIC 1112 represents a port on the vSwitch, which connects the vSwitch to other networks through a virtual router. As shown, VSRS 1110 may be instantiated on a cluster of servers 1116.
[0216] The control plane 1118 can track information to identify the L2 VNICs 1104-A, 1104-B, 1104-N and their placement within a VLAN, and the control plane 1110 can provide this information to the interfaces 1104-A, 1104-B, 1104-N within the VLAN.
[0217] 11, a VLAN may be a cloud-based virtual L2 network that may be built on a physical network 1120. In some embodiments, this physical network 1120 may include NVDs 1108-A, 1108-B, and 1108-N.
[0218] In general, a first L2 compute instance of a VLAN (e.g., compute instance 1 1102-A) can communicate with a second compute instance of the VLAN (e.g., compute instance 2 1102-B) using an L2 protocol. For example, a frame may be transmitted between two L2 compute instances via a VLAN. Nevertheless, the frame may be encapsulated, tunneled, routed, and / or otherwise processed as it is transmitted over the underlying physical network 1120.
[0219] For example, compute instance 1 1102-A sends a frame to compute instance 2 1102-B. Different types of processing can be applied to the frame depending on the network connection between host machine 1 and NVD1, the network connection between NVD1 and physical network 1120, the network connection between physical network 1120 and NVD2, and the network connection between NVD2 and host machine 2 (e.g., TCP / IP connection, Ethernet connection, tunneling connection). For example, the frame is received by NVD1, encapsulated, and this processing is repeated until it reaches compute instance 2. This processing is assumed to allow the frame to be transmitted between underlying physical resources, and for brevity and clarity, its description is omitted from the description of VLANs and related L2 operations.
[0220] Virtual L2 network communication Multiple types of communications can occur within or between virtual L2 networks. These communications can include intra-VLAN communications. In such embodiments, a source compute instance can send packets to a destination compute instance that is in the same VLAN as the source compute instance (CI). The communications can further include sending packets to an endpoint outside the source CI's VLAN. The communications can include, for example, communications from a source CI in a first VLAN to a destination CI in a second VLAN, communications from a source CI in the first VLAN to a destination CI in an L3 subnetwork, and / or communications from a source CI in the first VLAN to a destination CI outside the VCN that includes the source CI's VLAN. The communications can further include, for example, a destination CI receiving information from a source CI outside the destination CI's VLAN. This source CI may be located in another VLAN, in an L3 subnetwork, or outside the VCN that includes the source CI's VLAN.
[0221] Each CI in a VLAN can play an active role in traffic flow, including learning interface identifiers to MAC addresses, also referred to in this disclosure as interface-to-MAC addresses, mapping instances within a VLAN to maintain an L2 forwarding table within the VLAN, and transmitting and / or receiving communication packets. A VSRS can play an active role in communication within a VLAN and with source or destination CIs outside the VLAN. A VSRS can exist within an L2 network and an L3 network to enable transmission and reception.
[0222] Intra-VLAN communication 12, which is a flowchart illustrating one embodiment of a process 1200 for performing intra-VLAN communication. In some embodiments, process 1200 may be performed by a compute instance in a common VLAN. This process is specifically performed when a source CI sends a packet to a destination CI in a VLAN but does not know the IP-to-MAC address mapping of the destination CI. This may occur, for example, when the source CI sends a packet to a destination CI with an IP address in the VLAN but does not know the MAC address of the IP address. In this case, the destination MAC address and the IP-to-MAC address mapping can be learned by performing an ARP process.
[0223] If the source CI knows the IP-to-MAC address mapping, it can send the packet directly to the destination CI without having to perform the ARP process. In some embodiments, this packet may be intercepted by the source VNIC. This source VNIC is an L2 VNIC during intra-VLAN communication. If the source VNIC knows the interface-to-MAC address mapping of the destination MAC address, For example, the packet can be encapsulated using L2 encapsulation, and the encapsulated packet can be forwarded to a destination VNIC, which is an L2 VNIC for the destination MAC address in intra-VLAN communication.
[0224] If the source VNIC does not know the interface-to-MAC address mapping of the MAC address, it can perform an interface-to-MAC address learning process. This learning process can include the source VNIC sending a packet to all interfaces in the VLAN. In some embodiments, this packet can be sent to all interfaces in the VLAN via broadcast. In some embodiments, this broadcast can be implemented in the physical network in the form of a serial unicast. This packet can include a destination MAC address, a destination IP address, and the interface, MAC address, and IP address of the source VNIC. Each of the VNICs in the VLAN can receive this packet and can learn the interface-to-MAC address mapping of the source VNIC.
[0225] Furthermore, each receiving VNIC can decapsulate the packet and forward the decapsulated packet to an associated CI. Each CI can include a network interface that can evaluate the forwarded packet. If the network interface determines that the CI that received the forwarded packet does not match the destination MAC and / or IP address, the packet is discarded. If the network interface determines that the CI that received the forwarded packet matches the destination MAC and / or IP address, the packet is received by the CI. In some embodiments, a CI having a MAC and / or IP address that matches the destination MAC and / or IP address of the packet can send a response to the source CI, thereby allowing the source VNIC to learn the interface-to-MAC address mapping of the destination CI, and therefore the source CI to learn the IP-to-MAC address mapping of the destination CI.
[0226] Process 1200 may be performed if the source CI does not know the IP-to-MAC address mapping or if the source CI's IP-to-MAC address mapping for the destination CI is out of date.
[0227] Thus, if the IP to MAC address mapping is known, the source CI can send the packet. If the IP to MAC address mapping is unknown, process 1200 can be performed. If the interface to MAC address mapping is unknown, the interface to MAC address learning process outlined above can be performed. If the interface to MAC address mapping is known, the VNIC can send the packet to the destination CI.
[0228] Process 1200 begins at block 1202. In block 1202, the source CI determines that the IP-to-MAC address mapping of the destination CI is unknown to the source CI. In some embodiments, this may include the source CI determining the destination IP address of the packet and determining that the destination IP address is not associated with a MAC address stored in the source CI's mapping table. Alternatively, the source CI may determine that the IP-to-MAC address mapping of the destination CI is stale. In some embodiments, the mapping may be determined to be stale if the mapping has not been updated and / or verified within a certain time limit. Upon determining that the IP-to-MAC address mapping of the destination CI is unknown and / or stale to the source CI, the source CI initiates an ARP request for the destination IP and forwards the ARP request to the destination CI. Send to Ethernet broadcast.
[0229] At block 1204, the source VNIC, also referred to in this disclosure as the source interface, receives an ARP request from the source CI. The source interface identifies all interfaces on the VLAN and sends the ARP request to all interfaces in the VLAN broadcast domain. As previously described, because the control plane knows all of the interfaces on the VLAN and provides that information to the interfaces on the VLAN, the source interface also knows all of the interfaces in the VLAN and can send an ARP request to each interface in the VLAN. To this end, the source interface replicates the ARP request and encapsulates it for each interface on the VLAN. Each encapsulated ARP request includes the interface identifier of the source CI, the MAC address of the source CI, the IP address of the source CI, the target IP address, and the interface identifier of the destination CI. The source CI interface replicates the Ethernet broadcast by sending the replicated and encapsulated ARP requests as serial unicasts to each interface in the VLAN, one at a time.
[0230] In block 1206, all interfaces in the VLAN broadcast domain receive and decapsulate the packet. Because the packet identifies the MAC address, IP address, and interface identifier of the source CI, each interface in the VLAN broadcast domain that received the packet learns the interface-to-MAC address mapping of the source VNIC of the source CI (e.g., the mapping of the interface identifier of the source interface to the MAC address of the source CI). As part of learning the interface-to-MAC address mapping of the source CI, each interface can update a mapping table (e.g., an L2 forwarding table) and provide the updated mapping to an associated switch and / or CI. Each receiving interface, except the VSRS, can forward the decapsulated packet to an associated CI. The CI that receives the forwarded and decapsulated packet, specifically the network interface of the CI, can determine whether the target IP address of the packet matches the IP address of the CI. In some embodiments, if the IP address of the CI associated with an interface does not match the IP address of the destination CI specified in the received packet, the CI discards the packet and takes no further action. In the case of a VSRS, the VSRS can determine whether the target IP address of the packet matches the IP address of the VSRS. In some embodiments, if the IP address of the VSRS does not match the target IP address specified in the received packet, the VSRS discards the packet and takes no further action.
[0231] If it is determined that the IP address of the destination CI specified in the received packet matches the IP address of the CI (destination CI) associated with the receiving interface, the destination CI sends a response, which may be a unicast ARP response, to the source interface, as shown in block 1208. The response includes the MAC address and IP address of the destination CI and the IP address and MAC address of the source CI. As described below, the VSRS may send an ARP response if it determines that the target IP address matches its own IP address.
[0232] This response is received by the destination interface for encapsulating a unicast ARP response, as shown in block 1210. In some embodiments, this encapsulation may include Geneve encapsulation. The destination interface forwards the encapsulated packet to the source interface through the destination switch. The encapsulated packet includes the MAC address, IP address and interface identifier of the destination CI, and the MAC address, IP address and interface identifier of the source CI.
[0233] In block 1212, the source interface receives and decapsulates the ARP response. The source interface may also learn the interface-to-MAC address mapping of the destination CI based on information contained in the encapsulation and / or the encapsulated packet. In some embodiments, the source interface may forward the ARP response to the source CI.
[0234] At block 1214, the source CI receives the ARP response. In some embodiments, the source CI can update its mapping table based on the information included in the ARP response, specifically, to reflect an IP-to-MAC address mapping based on the MAC and IP addresses of the destination CI. The source CI can then send a packet, which may be any packet including an IP packet, specifically, an IPv4 or IPv6 packet, to the destination CI. The packet may include the MAC address and IP address of the source CI as the packet's source MAC address and source IP address, and the MAC address and IP address of the destination CI as the packet's destination MAC address and destination IP address.
[0235] At block 1216, the source interface may receive a packet from a source CI. The source interface may encapsulate the packet, and in some embodiments, may encapsulate the packet using Geneve encapsulation. The source interface may forward the encapsulated packet to a destination CI, specifically to the destination interface. The encapsulated packet may include the IP address, MAC address, and interface identifier of the source CI as a source MAC address, a source IP address, and a source interface identifier, and may include the MAC address, IP address, and interface identifier of the destination CI as a destination MAC address, a destination IP address, and a destination interface identifier.
[0236] In block 1218, the destination interface receives the packet from the source interface. The destination interface can decapsulate the packet and then forward the packet to the destination CI. In block 1220, the destination CI receives the packet from the destination interface.
[0237] Referring to FIG. 13, FIG. 13 is a schematic diagram 1300 illustrating the process 1200 for performing intra-VLAN communication. As shown, VLAN A 1302 has a VLAN CIDR with 10.0.3.0 / 24. VLAN A 1302 includes a VSRS VNIC (VRVI) 1304 that may be instantiated on one or more hardware devices, specifically on a server group 1306. VRVI 1304 may have an IP address of 10.0.3.1. The VLAN may include a compute instance 1 (CI1) 1310 having an IP address of 10.0.3.2 and communicatively connected to an NVD1 (SN1) 1312 that may instantiate an L2 VNIC1 (VI1) 1314 and an L2 switch 1. The VLAN can include compute instance 2 (CI2) 1320 having an IP address of 10.0.3.3 and communicatively connected to NVD2 (SN2) 1322, which can instantiate L2VNIC2 (VI2) 1324 and L2 switch 2. The VLAN can include compute instance 3 (CI3) 1320 having an IP address of 10.0.3.4 and communicatively connected to NVD3 (SN3) 1332, which can instantiate L2VNIC3 (VI3) 1334 and L2 switch 3. May contain CI3 (CI3) 1330.
[0238] Applying method 1200 of FIG. 12 to the example of FIG. 13, CI3 1330 is the source CI and VI3 1334 is the source interface. CI2 1320 is the destination CI and VI2 1324 is the destination interface. When CI3 determines that it does not have an IP-to-MAC mapping for the destination IP address (10.0.3.3), CI3 1330 sends an ARP request. The ARP request may be for a known address, and specifically, for the known IP address of CI2 1320. Thus, in some embodiments, the ARP request may be for 10.0.3.3.
[0239] This ARP request is received by SN3 1332 and VI3 1334. VI3 1334 creates an ARP request for each CI in VLAN 1302 by duplicating the ARP request. VI3 1334 encapsulates each ARP request and sends the ARP request to each interface in the VLAN. These encapsulated ARP requests may include information to identify the MAC address of the source CI (CI3), the interface identifier of the source interface (VI3), and the destination MAC address. These requests may be sent to each interface in the VLAN, as indicated by arrow 1350. In a VLAN, these requests may be broadcast ARP requests.
[0240] The interfaces in the VLAN receive the encapsulated ARP request and decapsulate the ARP request. Based on the information contained in and / or related to the ARP request, the interfaces in the VLAN update their mappings. Specifically, for example, VI1 1314, VI2 1324, and VRVI 1304 each receive the ARP request from CI3 1330, decapsulate the ARP request, and learn the interface-to-MAC address mapping of the source CI. Both the interface identifier and the MAC address of the source CI are included in the encapsulated packet. Furthermore, VRVI 1304 can update the IP-to-MAC address mapping of the source CI based on the information contained in the encapsulated packet.
[0241] As indicated by arrow 1352, the interface in the VLAN having the CI with the requested IP address can transmit an ARP response to CI3 1330. Specifically, as shown in FIG. 13, VI2 1324 is an interface in CI2 1320 that can transmit the ARP response. The ARP response from CI2 is received by VI2, encapsulated, and transmitted as an ARP unicast to the requesting interface, specifically VI3. As previously described herein, transmitting this ARP response includes providing the encapsulated ARP response to an associated switch, which can transmit the encapsulated ARP response to VI3.
[0242] The ARP response may be received by VI3 and decapsulated. VI3 may learn the interface-to-MAC mapping for CI2 based on the received ARP response and provide the learned mapping to the switch associated with VI3. The decapsulated packet may be provided to CI3. CI3 may learn the IP-to-MAC address mapping for CI2 1320 based on the decapsulated packet. CI3 may send a packet to CI2. This packet may be an IP packet, such as an IPv4 or IPv6 packet. This packet may include the MAC and IP addresses of CI3 as a source address and the IP and MAC addresses of CI2 as a destination address.
[0243] A packet sent by CI3 may be received by VI3, which can encapsulate the packet and forward it to interface VI2, which can receive the packet, decapsulate the packet, and forward it to CI2.
[0244] Inter-VLAN communication Referring now to FIG. 14, FIG. 14 is a flow chart illustrating one embodiment of a process 1400 for inter-VLAN communication in a virtual L2 network. Process 1400 may be performed by all or a portion of two connected VLANs, such as multiple connected L2 VLANs 800 as shown in FIG. 8. In some embodiments, process 1400 may be performed when a compute instance (source CI) in a first VLAN (source VLAN) sends a packet to a destination compute instance (destination CI) in a second VLAN (destination VLAN). In some embodiments, the source CI may determine that the destination CI is outside the source VLAN based on the IP address of the destination CI. For example, the source CI may determine that the destination IP is outside the source VLAN CIDR. In such a case, the source CI may decide to send the IP packet to the destination CI via the VSRS of the source VLAN. If the source CI already knows the mapping of the VSRS of the first VLAN (source VSRS), the source CI may send the packet directly to the source VSRS. If the source CI does not know the source VSRS mapping, the source CI and associated VNIC (source interface or source VNIC) first learn the source VSRS mapping. In an embodiment of inter-VLAN communication, both the source VNIC and the destination VNIC are L2 VNICs. The first steps 1402-1410 of process 1400 involve learning the source VSRS mapping by the source VNIC and source CI.
[0245] Process 1400 begins at block 1402. In block 1402, a source CI having a destination IP address initiates an ARP request. The ARP request is used to determine the IP-to-MAC address mapping of the source VSRS. The source CI sends the ARP request to an Ethernet broadcast. The ARP request includes the IP address of the source CI as the source IP address and MAC address.
[0246] In block 1404, the source VNIC receives the ARP request and replicates the ARP request. Specifically, the source VNIC receives the ARP request from the source CI, identifies all interfaces on the VLAN, and sends the ARP request to all interfaces on the VLAN broadcast domain. As described above, because the control plane knows all interfaces on the VLAN and provides this information to interfaces in the VLAN, the source interface also knows all interfaces in the VLAN and can send an ARP request to each interface in the VLAN. Therefore, the source interface replicates the ARP request and encapsulates the ARP request for each interface on the VLAN. Each encapsulated ARP request includes the interface identifier of the source CI and the IP address and MAC address of the source CI as the source address, and the interface identifier and IP address of the destination CI as the destination address. The source CI interface realizes an Ethernet broadcast by sending the replicated and encapsulated ARP request to each interface in the VLAN via serial unicast. In some embodiments, the source VNIC can encapsulate the ARP request using Geneve encapsulation.
[0247] In block 1406, all interfaces in the VLAN broadcast domain receive and decapsulate the packet. To identify the source IP address, IP address, and source interface identifier, each interface in the VLAN broadcast domain that received the packet learns the interface-to-MAC address mapping of the source VNIC of the source CI. As part of learning the interface-to-MAC address mapping of the source CI, each interface can update its corresponding mapping table and provide the updated mapping table to its associated switch and / or CI. Each receiving interface, except for the VSRS, forwards the decapsulated packet to its associated CI. The CI that received the forwarded decapsulated packet, specifically the network interface of that CI, can determine whether the packet's target IP address matches the IP address of that CI. If the IP address of the CI associated with that interface does not match the IP address of the destination CI specified in the received packet, no further action is taken.
[0248] If the source VSRS determines that the target IP address matches the IP address of the source VSRS, it encapsulates a response, which may be a unicast ARP response, and sends it to the source interface, as shown in block 1408. The response includes the MAC address, IP address, and interface identifier of the source CI as a destination address. Additionally, the response includes the MAC address, IP address, and interface identifier of the VSRS as a source address. In some embodiments, the encapsulation of the ARP response may include Geneve encapsulation.
[0249] At block 1410, the source interface receives and decapsulates the ARP response. The source interface may also learn the interface-to-MAC address mapping of the source VSRS based on information contained in the encapsulation and / or the encapsulated packet. In some embodiments, the source interface may forward the ARP response to the source CI.
[0250] At block 1412, the source CI receives the ARP response. In some embodiments, the source CI may update its mapping table based on information included in the ARP response, specifically, based on the MAC address and IP address of the source VSRS. In some embodiments, for example, the source CI may update its mapping table to reflect the IP-to-MAC address mapping of the source VSRS. The source CI may then send a packet to the source VSRS, which may be any packet including an IP packet, specifically, an IPv4 or IPv6 packet. In some embodiments, this may include sending an IP packet that includes the IP address of the destination CI as a destination address. In some embodiments, the IP address of the destination CI may be included in a header of the packet, e.g., the L3 header of the packet. This header may further include the MAC address of the source VSRS in another header of the packet, e.g., the L2 header of the packet. The packet may further include the MAC address of the source CI and the IP address of the source CI as a source MAC address and a source IP address.
[0251] In block 1414, the source interface may receive a packet from the source CI. The source interface may encapsulate the packet. The source interface may forward the encapsulated packet to the source VSRS, specifically to the VNIC of the source VSRS. In addition to the packet's address, the encapsulated packet may include the MAC address and interface identifier of the source CI as the source MAC address and source interface identifier, and the MAC address and interface identifier of the source VSRS as the destination MAC address and destination interface identifier.
[0252] At block 1416, the source VSRS receives the encapsulated packet. The source VSRS decapsulates the packet and removes any address information associated with the source VSRS (e.g., including the IP address, MAC address, and / or interface identifier of the source VSRS) from the packet. The source VSRS identifies the destination CI of the packet. In some embodiments, the source VSRS identifies the destination CI of the packet based on the IP address of the destination CI included in the packet. The source VSRS searches for a mapping for the destination IP address in the packet. If the IP address is within the IP address space of the VCN, the source VSRS searches for a mapping for the packet's destination IP address in the IP address space of the VCN. The source VSRS then re-encapsulates the packet using L3 encapsulation. In some embodiments, this L3 encapsulation may include, for example, MPLSoUDP L3 encapsulation. The source VSRS then forwards the packet to the VSRS of the destination VLAN (the destination VSRS). In some embodiments, the source VSRS may forward the packet to the destination VSRS such that the destination VSRS becomes the tunnel endpoint (TEP) of the packet. The L3 encapsulated packet contains the IP address and MAC address of the source CI and the IP address of the destination CI.
[0253] At block 1418, the destination VSRS receives the packet and decapsulates the packet. In some embodiments, this decapsulation may include removing the L3 encapsulation. If the destination VSRS knows the IP-to-MAC and MAC-to-interface mappings for the destination CI, the destination VSRS identifies the interface and MAC address of the destination CI within the destination VLAN. Alternatively, if the destination VSRS does not know the mappings for the destination CI, it may perform steps 1612 through 1622 of process 1600.
[0254] At block 1420, the destination VSRS forwards the packet to the destination interface of the destination CI. In some embodiments, this may include encapsulating the packet using L2 encapsulation. In some embodiments, the packet may include the IP address and MAC address of the source CI and / or the MAC address and interface identifier of the destination VSRS. In some embodiments, the packet may further include the MAC address and interface identifier of the destination CI.
[0255] At block 1422, the destination interface receives and decapsulates the packet. Specifically, the destination VNIC removes the L2 encapsulation. In some embodiments, the destination VNIC forwards the packet to the destination CI. At block 1424, the destination CI receives the packet.
[0256] Referring to FIG. 15, FIG. 15 is a schematic diagram 1500 illustrating the process 1400 for performing inter-VLAN communication. As shown, VLAN A 1502-A has a VLAN CIDR with 10.0.3.0 / 24. VLAN A 1502-A includes VSRS VNIC A (VRVI A) 1504-A, which may be instantiated on one or more hardware devices, specifically on a server group 1506. VRVI A 1504-A may have an IP address of 10.0.3.1. The VLAN may include a compute instance 1 (CI1) 1510 having an IP address of 10.0.3.2 and communicatively connected to NVD1 (SN1) 1512, which may instantiate an L2 VNIC1 (VI1) 1514 and an L2 switch 1. The VLAN is connected to NVD2 (SN2) 1522, which has IP address 10.0.3.3 and can instantiate L2VNIC2 (VI2) 1524 and L2 switch 2. The VLAN may include a compute instance 2 (CI2) 1520 that is communicatively connected to the VLAN. The VLAN may include a compute instance 3 (CI3) 1530 that has an IP address 10.0.3.4 and is communicatively connected to an NVD3 (SN3) 1532 that may instantiate an L2VNIC3 (VI3) 1534 and an L2 switch 3.
[0257] VLAN B 1502-B has a VLAN CIDR with 10.0.34.0 / 24. VLAN B 1502-B may be instantiated on one or more hardware devices, and specifically, includes VSRS VNIC B (VRVI B) 1504-B, which may be instantiated on server group 1506. VRVI B 1504-B may have an IP address of 10.0.4.1. The VLAN may include compute instance 4 (CI4) 1540 having an IP address of 10.0.4.2 and communicatively connected to NVD4 (SN4) 1542, which may instantiate L2 VNIC4 (VI3) 1544 and L2 switch 4.
[0258] Applying the method 1400 of FIG. 14 to the example of FIG. 15, CI3 1530 is the source CI and VI3 1534 is the source interface. CI4 1540 is the destination CI and VI4 1544 is the destination interface. Upon determining that CI3 does not have an IP-MAC mapping for VRVI A 1504-A, CI3 1530 sends an ARP request. The ARP request may be for an address, specifically, a known IP address for VRVI A 1504-A. Thus, in some embodiments, the ARP request may be for 10.0.3.1.
[0259] This ARP request is received by SN3 1532 and VI3 1534. VI3 1534 creates an ARP request for each CI in VLAN 1502-A by duplicating the ARP request. VI3 1534 encapsulates each ARP request using L2 encapsulation and sends the ARP request to each interface in the VLAN. These encapsulated ARP requests may include information identifying the MAC address and IP address of the source CI (CI3), the interface identifier of the source interface (VI3), and the target IP address. These requests may be sent to each interface in the VLAN, as indicated by arrow 1550, or specifically, may be sent as serial unicast so that each interface in the VLAN receives the ARP request.
[0260] The interfaces in the VLAN receive the encapsulated ARP request and decapsulate the ARP request. Based on the information contained in and / or related to the ARP request, the interfaces in the VLAN update their mappings. Specifically, for example, VI1 1514, VI2 1524, and VRVI A 1504-A each receive a unicast ARP request from CI3 1530, decapsulate the unicast ARP request, and learn the interface-to-MAC address mapping for the source CI. Both the interface identifier and the MAC address are included in the encapsulated packet.
[0261] As indicated by arrow 1552, if VRVI A 1504-A determines that the IP address matches the requested IP address, it transmits an ARP response to CI3 1530. For example, VRVI A 1504-A can encapsulate the ARP response from VRVI A 1504-A using L2 encapsulation and transmit it as an ARP unicast to the requesting interface, specifically VI3. As previously described herein, transmitting this ARP response can include providing the encapsulated ARP response to a switch associated with VRVI A 1504-A, which can then unicast the ARP response to CI3. The received ARP response can be sent to VI3.
[0262] The ARP response may be received by VI3 and decapsulated. VI3 may learn the interface-to-MAC mapping for VRVI A 1504-A based on the received ARP response and provide the learned mapping to the switch associated with VI3. The decapsulated packet may be provided to CI3. CI3 may learn the IP-to-MAC address mapping for VRVI A 1504-A and transmit the packet. The packet may be an IP packet, such as an IPv4 or IPv6 packet. The packet may include the MAC address and IP address of CI3 as a source address. The packet may further include the IP address of the destination CI (CI4 1540) as a destination address, which may include the MAC address and IP address of VRVI A 1504-A.
[0263] VI3 can receive packets sent by CI3, encapsulate the packets, and forward the packets to VRVI A 1504-A. VRVI A 1504-A can receive the packets, decapsulate the packets, and look up a mapping for the packet's destination IP address (the IP address of CI4). In some embodiments, decapsulating the packets can include removing the L2 header from the packets, i.e., removing information related to VLAN A 1502-A from the header. The removed information can include, for example, the MAC address of CI3, the interface identifier of interface VI3, and / or the MAC address and interface identifier of VRVI A 1504-A. In some embodiments, the VSRS, and therefore VRVI A 1504-A, can reside in both an L2 network and an L3 network. In the L2 network, and therefore the VLAN, the VSRS can utilize an L2 communication protocol, but can utilize an L3 communication protocol when communicating with the L3 network. In contrast to learning performed in a VLAN, the VSRS can learn the mapping of endpoints in the L3 network from the control plane. In some embodiments, for example, the control plane can provide mapping information of IP addresses and / or MAC addresses of instances in the L3 network.
[0264] VRVI A 1504-A can look up a mapping for the IP destination address contained in the packet, in other words, look up a mapping for the IP address of the destination CI. In some embodiments, looking up this mapping is performed by VRVI A 1504-A. The forwarding of the encapsulated packet may include identifying VRVI A 1504-B as the VSRS for VLAN B 1502-B. VRVI A 1504-A may encapsulate the packet and forward the encapsulated packet to VRVI B 1504-B. VRVI B 1504-B may be a tunnel endpoint (TEP) for the destination VLAN and / or destination interface. The forwarding of the encapsulated packet is indicated by block 1556. The forwarded packet may include the IP address of the source CI (CI3) as the source address and the IP address of the destination CI (CI4) as the destination address.
[0265] VRVI B 1504-B can receive and decapsulate the packet. VRVI B 1504-B can further identify the interface in VLAN B 1502-B that corresponds to the destination IP address. VRVI B 1504-B can encapsulate the packet and add an L2 header to tunnel it within VLAN B 1504-B. VRVI B 1504-B can then forward the packet to the destination CI. Destination interface VI4 1544 receives and decapsulates the packet, then forwards the Ethernet frame or packet to the destination CI. It can be transferred to destination CI1540.
[0266] Incoming Packet Flow Referring to Figure 16, Figure 16 is a flow chart illustrating one embodiment of a process 1600 for performing receive packet flow. Specifically, Figure 16 illustrates one embodiment of a process 1600 for receiving packets from a subnet. This process may be performed by all or a portion of system 600, and specifically, by entities in a VLAN and external source CIs (with respect to the VLAN). The external source CIs may reside in an L3 subnet.
[0267] Process 1600 begins at block 1602. In block 1602, the source L3 CI determines to send a packet to the destination CI, specifically to a destination IP address within a VLAN. Because the source L3 CI does not know the mapping, it sends an ARP request for the MAC address of the virtual router of the subnet that contains the source L3 CI. In some embodiments, sending this ARP request can include sending an ARP request to an Ethernet broadcast. In block 1604, the source L3 CI The source interface of the CI responds to the ARP request with the MAC address of the VR. In some embodiments, the source interface can determine the MAC address of the VR based on mapping information accessed by the source interface.
[0268] Upon receiving the ARP response from the source interface, the source L3 C1 transmits the IP packet to the VR, as shown in block 1606. In some embodiments, this may include the source L3 CI transmitting the packet and the source L3 interface receiving, encapsulating, and forwarding the packet. In some embodiments, the source L3 interface may encapsulate the packet using L3 encapsulation. The L3 encapsulation includes the original packet beginning with an L3 header. The source L3 interface may forward the packet to the VR. This IP packet may be sent to the VR MAC address and the VR interface and may include the MAC address of the source L3 CI as the MAC address and the interface identifier of the source L3 CI as the source interface identifier.
[0269] The VR receives the packet and decapsulates it. The VR can then look up the VNIC mapping for the destination IP address in the packet. The VR can then determine whether the packet is in a VLAN. The VR may determine that the destination CI is within a CIDR, then encapsulate the packet and forward the encapsulated packet to the VSRS of the VLAN that includes the destination CI. In some embodiments, the VR may encapsulate the packet using L3 encapsulation. The encapsulated packet may include the destination IP address, MAC address, and interface identifier of the VSRS, and the source IP address.
[0270] As shown in block 1610, the VSRS receives and decapsulates the packet. If the VSRS knows the mapping of the destination CI, specifically the mapping of the destination IP address in the received packet, it forwards the IP packet to the TEP of the CI corresponding to the destination IP address. This may include generating and encapsulating an L2 packet with a destination MAC corresponding to the MAC address of the destination CI and a destination interface identifier corresponding to the destination interface. In an embodiment of a received packet flow, the destination interface is an L2 VNIC.
[0271] If the VSRS does not know the destination CI mapping, process 1600 proceeds to block 1612, where the VSRS receives and decapsulates the IP packet.
[0272] If the VSRS does not know the destination MAC address-to-VNIC mapping, it can perform an interface-to-MAC address learning process. This can include the VSRS sending a packet to all interfaces in the VLAN. In some embodiments, this packet can be sent to all interfaces in the VLAN via broadcast. This packet can include the destination MAC and IP addresses, the interface identifier and MAC address of the VSRS, and the IP address of the source CI. Each VNIC in the VLAN can receive this packet and learn the interface-to-MAC address mapping of the VSRS.
[0273] Each receiving VNIC can further decapsulate the packet and forward the decapsulated packet to an associated CI. Each CI can include a network interface that can evaluate the forwarded packet. If the network interface determines that the CI that received the forwarded packet does not match the destination MAC and / or IP address, the packet is discarded. If the network interface determines that the CI that received the forwarded packet matches the destination MAC and / or IP address, the packet is received by the CI. In some embodiments, a CI with a MAC and / or IP address that matches the destination MAC and / or IP address of the packet can send a response to the VSRS. This allows the VSRS to learn the interface-to-MAC address mapping and IP-to-MAC address mapping of the destination CI.
[0274] Alternatively, if the VSRS does not know the destination CI mapping, specifically the destination IP address to MAC address mapping, it withholds the IP packet. The VSRS then generates an ARP request for the destination IP address to all interfaces of the VSRS within the VLAN broadcast domain. This ARP request contains the VSRS MAC as the source MAC, the interface identifier of the VSRS interface as the source interface, and the destination IP address as the target IP address. In some embodiments, this ARP request can be broadcast to all interfaces within the VLAN.
[0275] In block 1614, all interfaces in the VLAN broadcast domain receive and decapsulate the packet. Because the packet identifies the MAC address and interface identifier of the VSRS, each interface in the VLAN broadcast domain that received the packet learns the interface-to-MAC address mapping of the VSRS. As part of learning the interface-to-MAC address mapping of the VSRS, each interface can update its mapping table and provide the updated mapping table to its associated switch and / or CI.
[0276] Each receiving interface can forward the decapsulated packet to the associated CI. The CI that receives the forwarded decapsulated packet, specifically the network interface of the CI, can determine whether the packet's destination IP address matches the IP address of the CI. If the IP address of the CI associated with the interface does not match the IP address of the destination CI specified in the received packet, no further action is taken.
[0277] If it is determined that the IP address of the destination CI specified in the received packet matches the IP address of the CI associated with the receiving interface (destination CI), then the process proceeds to block 1616. As shown, the destination CI sends a response, which may be a unicast ARP response, to the source interface. This response includes the MAC address and IP address of the destination CI and the MAC address and IP address of the VSRS. As shown in block 1618, this response is received by the destination interface, which encapsulates the unicast ARP response. The destination interface can forward the encapsulated packet to the VSRS through the destination switch. The encapsulated packet includes the MAC address, IP address, and interface identifier of the destination CI and the MAC address, IP address, and interface identifier of the VSRS.
[0278] The VSRS, and in particular the VSRS interface, receives and decapsulates the ARP response in block 1620. The VSRS, and in particular the VSRS interface, may further learn the interface-to-MAC address mapping of the destination CI based on information contained in the encapsulation and / or the encapsulated packet.
[0279] In block 1622, the VSRS may encapsulate and add an L2 header to the previously deferred IP packet, and then forward the previously deferred IP packet to a destination CI, specifically, a destination interface. The destination interface may decapsulate the packet and provide the decapsulated packet to the destination CI. This packet forwarded by the VSRS may include the MAC address and interface identifier of the VSRS as the source MAC address and source interface identifier, and the MAC address and interface identifier of the destination CI as the destination MAC address and destination interface identifier. This packet may further include the IP address of the destination CI and the IP address of the source CI.
[0280] The destination interface receives the packet from the VSRS and decapsulates the packet. This decapsulation may include removing a header added by the VSRS. The header may be a VCN header. The destination interface can then forward the packet to the destination CI, and the destination CI can receive the packet from the destination interface.
[0281] Referring to FIG. 17, FIG. 17 is a schematic diagram 1700 illustrating a process 1600 for performing inbound communications. As shown, VLAN A 1502-A has a VLAN CIDR with 10.0.3.0 / 24. VLAN A 1502-A includes VSRS VNIC A (VRVI A) 1504-A, which may be instantiated on one or more hardware devices, specifically on a server group 1506. VRVI A 1504-A may have an IP address of 10.0.3.1. The VLAN may include a compute instance 1 (CI1) 1510 having an IP address of 10.0.3.2 and communicatively connected to NVD1 (SN1) 1512, which may instantiate an L2 VNIC1 (VI1) 1514 and an L2 switch 1. The VLAN may include compute instance 2 (CI2) 1520 having an IP address of 10.0.3.3 and communicatively connected to NVD2 (SN2) 1522, which may instantiate L2VNIC2 (VI2) 1524 and L2 switch 2. The VLAN may include compute instance 3 (CI3) 1530 having an IP address of 10.0.3.4 and communicatively connected to NVD3 (SN3) 1532, which may instantiate L2VNIC3 (VI3) 1534 and L2 switch 3.
[0282] L3 Compute Instance 4 (CI4) 1744 is a compute instance that can be CI4 may reside on a subnet 1739 external to A 1702-A. CI4 may have an IP address 10.0.4.4 and may communicatively connect to NVD4 (SN4) 1742, which may instantiate L3 VNIC4 (VI4) 1744. It is possible.
[0283] Applying method 1600 of FIG. 16 to the example of FIG. 17, CI4 1740 is the source CI and VI4 1744 is the source interface. CI3 1730 is the destination CI and VI3 1734 is the destination interface. CI4 determines that it does not know the mapping to send the packet to CI3. Therefore, CI4 sends an ARP request for the IP address of the subnet virtual router. In response to this request, VI4, which in some embodiments may reside on an NVD containing an instance of the subnet virtual router, responds directly to CI4 with the IP address of the VR. CI4 learns the IP address of the VR and sends the packet to the VR. The VR receives the packet, encapsulates the packet based on the mapping information, and forwards the packet to the VSRS, as indicated by arrow 1748.
[0284] When the VSRS receives a packet and does not know the mapping to the destination CI, specifically the IP-to-MAC address mapping, the VSRS withholds the packet and sends an ARP request to all interfaces in VLAN A 1704-A. This ARP request includes the MAC address and interface identifier of the VSRS. Each receiving interface in VLAN A 1704-A learns the mapping of the VSRS based on the ARP request. Similarly, each interface decapsulates the packet and sends it to its CI. Upon receiving the decapsulated packet, CI3 determines that it is the CI specified in the packet and generates an ARP response responding with its MAC address. VI3 receives and encapsulates the ARP response and forwards the encapsulated ARP response to the VSRS. The encapsulated ARP response includes the MAC address and interface identifier of the VSRS and the MAC address and interface identifier of the interface in CI3.
[0285] The VSRS receives the ARP response and learns the IP address-to-MAC address mapping and MAC address-to-interface mapping of CI3. The VSRS then forwards the previously deferred packet to CI3. This packet is received by VI3, decapsulated, and forwarded to CI3.
[0286] Outbound Packet Flow Referring now to FIG. 18, FIG. 18 is a flow chart illustrating one embodiment of a process 1800 for transmitting packet flow from a VLAN. In some embodiments, packets can be transmitted from a VLAN to another VLAN, to a subnet, or to another network. Process 1800 may be performed by all or part of system 600, and specifically, by an entity in a VLAN. In some embodiments, parts of the process may be performed by a source CI external (to the VLAN). The external source CI may reside in an L3 subnet.
[0287] In some embodiments, process 1800 may be performed when a compute instance (source CI) within a VLAN (source VLAN) sends a packet to a destination compute instance (destination CI) outside the VLAN. If the source CI already knows the mapping of the VSRS of the first VLAN (source VSRS), it can send the packet directly to the source VSRS. If the source CI does not know the mapping of the source VSRS, the source CI and associated VNIC (source interface or source VNIC) first learn the mapping of the source VSRS, specifically the IP-to-MAC address mapping of the source VSRS. In an embodiment of a transmit packet flow from an L2 VLAN, the source VNIC is the L2 VNIC. The first step 1802 of process 1800 1810 relates to learning source VSRS mapping with source VNIC and source CI.
[0288] Process 1800 begins at block 1802. In block 1802, the source CI initiates an ARP request. The ARP request is used to determine the IP-to-MAC address mapping of the source VSRS. The ARP request is sent by the source CI to an Ethernet broadcast. The ARP request includes the MAC address, IP address, and interface identifier of the source CI as the source address and source interface identifier. The ARP request further includes the IP address of the source VSRS.
[0289] In block 1804, the source VNIC receives the ARP request and replicates the ARP request. Specifically, the source VNIC receives the ARP request from the source CI, identifies all interfaces on the VLAN, and sends the ARP request to all interfaces on the VLAN broadcast domain. As previously described, because the control plane knows all of the interfaces on the VLAN and provides that information to the interfaces in the VLAN, the source interface similarly knows all of the interfaces in the VLAN and can send an ARP request to each interface in the VLAN. To do this, the source interface replicates the ARP request and encapsulates one ARP request for each interface on the VLAN. Each encapsulated ARP request includes the source CI's interface identifier and the source CI's MAC and / or IP address as the source address and the destination CI's interface identifier as the destination address. The source CI's interface replicates the Ethernet broadcast by sending the replicated and encapsulated ARP request via serial unicast. In some embodiments, the source VNIC can encapsulate the ARP request using Geneve encapsulation.
[0290] In block 1806, all interfaces in the VLAN broadcast domain receive and decapsulate the packet. Because the packet identifies the MAC address and interface identifier of the source CI, each interface in the VLAN broadcast domain that received the packet learns the interface-to-MAC address mapping of the source CI. In addition, the VSRS learns the IP-to-MAC address mapping of the source CI. As part of learning the interface-to-MAC address mapping of the source CI, each interface can update its mapping table and provide the updated mapping table to its associated switch and / or CI. Each receiving interface, except the VSRS, can forward the decapsulated packet to its associated CI. The CI that received the forwarded decapsulated packet, specifically the network interface of that CI, can determine whether the packet's destination IP address matches the CI's IP address. If the IP address of the CI associated with the interface does not match the IP address of the destination CI specified in the received packet, no further action is taken.
[0291] If the source VSRS determines that the destination IP address matches the IP address of the source VSRS, it encapsulates a response, which may be a unicast ARP response, and sends it to the source interface, as shown in block 1808. The response includes the MAC address and IP address of the source CI as a destination address and the interface identifier of the source CI as a destination interface identifier. The response also includes the MAC address and IP address of the source VSRS as a source address and the interface identifier of the source VSRS as a source interface identifier.
[0292] In block 1810, the source interface receives the ARP response and The source interface may further learn the interface-to-MAC address mapping of the source VSRS based on information contained in the encapsulation and / or the encapsulated packet. In some embodiments, the source interface may forward the ARP response to the source CI.
[0293] At block 1812, the source CI receives the ARP response. In some embodiments, the source CI may update its mapping table based on information included in the ARP response, specifically, based on the MAC address of the source VSRS and the IP address of the source VSRS. The source CI may then send a packet to the source VSRS, which may be any packet including an IP packet, specifically, an IPv4 or IPv6 packet. In some embodiments, this may include sending an IP packet with a destination address of the MAC address of the source VSRS and the IP address of the source VSRS. The packet may further include the MAC address and IP address of the source CI as source addresses.
[0294] In block 1814, the source interface receives a packet from the source CI. The source interface encapsulates the packet. The source interface can forward the encapsulated packet to the source VSRS, specifically to the VNIC of the source VSRS. The encapsulated packet can include the MAC address and interface identifier of the source CI as the source MAC address and source interface identifier, and the MAC address and interface identifier of the source VSRS as the destination MAC address and destination interface identifier. The encapsulated packet can further include the IP address of the source CI and the IP address of the source VSRS.
[0295] At block 1816, the source VSRS receives the encapsulated packet. The source VSRS determines the destination CI of the packet. In some embodiments, the source VSRS determines the destination CI of the packet based on the IP address of the destination CI included in the packet. The source VSRS looks up a mapping for the packet's destination IP address. If the IP address is within the IP address space of the VCN, the source VSRS looks up a mapping for the packet's destination IP address in the IP address space of the VCN. The source VSRS then re-encapsulates the packet using L3 encapsulation.
[0296] The source VSRS then forwards the packet to the destination CI. In some embodiments, this may include forwarding the encapsulated packet to a VR associated with the subnet containing the destination CI and / or forwarding the packet to a gateway to allow the packet to exit the VCN. In some embodiments, forwarding the packet to the destination CI may include forwarding the packet to a TEP associated with the destination CI. The L3 encapsulated packet includes the IP address of the source CI as a source address and the IP address of the destination CI as a destination address.
[0297] At block 1820, the destination interface receives and decapsulates the packet. In some embodiments, the destination VNIC forwards the packet to the destination CI. At block 1822, the destination CI receives the packet.
[0298] 19, which is a schematic diagram 1900 illustrating a process 1800 for performing outbound packet flow. As shown, VLAN A 1502-A has a VLAN CIDR of 10.0.3.0 / 24. VLAN A 1502-A is connected to VSRS VNIC A (VRVI A) 1504-A, which may be instantiated on one or more hardware devices, specifically on server cluster 1506. VRVI A 1504-A may have an IP address of 10.0.3.1. The VLAN may include compute instance 1 (CI1) 1510 having an IP address of 10.0.3.2 and communicatively connected to NVD1 (SN1) 1512, which may instantiate L2VNIC1 (VI1) 1514 and L2 switch 1. The VLAN may include compute instance 2 (CI2) 1520 having an IP address of 10.0.3.3 and communicatively connected to NVD2 (SN2) 1522, which may instantiate L2VNIC2 (VI2) 1524 and L2 switch 2. The VLAN may include compute instance 3 (CI3) 1530 having an IP address of 10.0.3.4 and communicatively connected to NVD3 (SN3) 1532, which may instantiate L2VNIC3 (VI3) 1534 and L2 switch 3.
[0299] L3 Compute Instance 4 (CI4) 1944 is a compute instance that can be CI4 may reside on a subnet 1939 external to A1902-A. CI4 may have an IP address of 10.0.4.4 and may be communicatively connected to NVD4 (SN4) 1942, which may instantiate L3 VNIC4 (VI4) 1944. Subnet 1939 may include a virtual router (VR) 1948. VR 1948 may have an IP address of 10.0.4.1. VR 1948 may be instantiated on, for example, a smart NIC, a server, or a cluster of servers.
[0300] Applying the method 1800 of FIG. 18 to the example of FIG. 19, CI3 1930 is the source CI and VI3 1934 is the source interface. CI4 1940 is the destination CI and VI4 1944 is the destination interface. Upon determining that CI3 does not have an IP-to-MAC mapping for VRVI A 1904-A, CI3 1930 sends an ARP request. The ARP request may be for a known address, and specifically, may be for a known IP address of VRVI A 1904-A. Thus, in some embodiments, the ARP request may be for 10.0.3.1.
[0301] This ARP request is received by SN3 1932 and VI3 1934. VI3 1934 creates an ARP request for each CI in VLAN 1902-A by duplicating the ARP request. VI3 1934 encapsulates each ARP request using L2 encapsulation and sends the ARP request to each interface in the VLAN. These encapsulated ARP requests may include information identifying the MAC address of the source CI (CI3), the source interface identifier VI3, and the target IP address. These requests may be sent to each interface in the VLAN, as indicated by arrow 1950, or may be broadcast so that each interface in the VLAN receives the ARP request.
[0302] The interfaces in the VLAN receive the encapsulated ARP request and decapsulate the ARP request. Based on the information contained in and / or related to the ARP request, the interfaces in the VLAN update their mappings. Specifically, for example, VI1 1914, VI2 1924, and VRVI A 1904-A each receive a unicast ARP request from CI3 1930, decapsulate the unicast ARP request, and learn the interface-to-MAC address mapping for the source CI. Both the interface identifier and the MAC address are included in the encapsulated packet.
[0303] As shown by arrow 1952, VRVI A 1904-A determines that the IP address matches the requested IP address and sends an ARP response to CI3 1930. The ARP response from VRVI A 1904-A may be, for example, , and may be sent as an ARP unicast to the requesting interface, specifically VI3. As previously described herein, sending this ARP response may include providing the encapsulated ARP response to a switch associated with VRVI A 1904-A, which may then send the encapsulated ARP response to VI3.
[0304] The ARP response may be received by VI3 and decapsulated. VI3 may learn the interface-to-MAC mapping for VRVI A 1904-A based on the received ARP response and provide the learned mapping to a switch associated with VI3. The decapsulated packet may be provided to CI3. CI3 may learn the IP-to-MAC address mapping based on the received packet. CI3 may transmit a packet, which may be an IP packet such as an IPv4 or IPv6 packet. The packet may include the MAC address and IP address of CI3 as a source address and the IP address of the destination CI (CI4 1940) as a destination address. In some embodiments, the destination address may further include the MAC address of VRVI A 1904-A.
[0305] Packets sent by CI3 can be received by VI3, which can encapsulate the packets and forward them to VRVI A 1904-A. VRVI A 1904-A can receive the packets, decapsulate them, and look up the mapping of the packet's destination IP address (the IP address of CI4). In some embodiments, the VSRS, and therefore VRVI A 1904-A, can reside in both an L2 network and an L3 network. In the L2 network, and therefore a VLAN, the VSRS can utilize L2 communication protocols, but can utilize L3 communication protocols when communicating with the L3 network. In contrast to learning done in a VLAN, the VSRS can learn the mapping of endpoints in the L3 network from the L3 control plane. In some embodiments, for example, the L3 control plane can provide information mapping IP addresses, MAC addresses, and / or interface identifiers of instances in the L3 network.
[0306] VRVI A 1904-A can look up a mapping for the IP destination address contained in the packet, in other words, look up a mapping for the IP address of the destination CI. In some embodiments, looking up this mapping can include identifying the subnet 1939 on which CI4 1940 resides and / or identifying the VR associated with the subnet 1939 on which CI4 resides. VRVI A 1904-A can encapsulate the packet and forward the encapsulated packet to VR 1948, which may be a tunnel endpoint (TEP) for subnet 1939 and / or destination CI 1940. This forwarding of the encapsulated packet is indicated by block 1956.
[0307] The VR 1948 can receive and decapsulate the packet. The VR 1948 can further identify the interface within the subnet 1939 that corresponds to the destination IP address. In some embodiments, the VR 1948 can be located on the same NVD as the destination interface VI4 1944. Thus, the VR 1948 can forward the packet directly to the destination CI, specifically, CI4 1940, as indicated by block 1958.
[0308] Interface-based Access Control List filtering The VSRS can provide interface-based access control list (ACL) filtering. This can include evaluating the inbound security policy of a VLAN. This can also include evaluating the outbound security policy of the VSRS's sender based on the learned VLAN interface-to-MAC address and IP address mappings when the VSRS determines where to send a received packet. This can result in delays in ACL classification.
[0309] In some embodiments, an ACL may include a permission list associated with objects in a system. These objects may include hardware in a physical network, which may include, for example, one or more servers, smart NICs, host machines, etc. These objects may include one or more virtual objects in a virtual network. These virtual objects may include, for example, one or more interfaces, compute instances, addresses such as IP addresses and / or MAC addresses. In some embodiments, an ACL may specify users allowed to access the objects, the objects and / or system processes, and / or the actions allowed for a given object.
[0310] An ACL may be specific to one or more CIs. Thus, in some embodiments, some or all of the CIs may have and / or maintain their own ACLs. In some embodiments, a CI's ACL may specify interfaces and / or addresses (either MAC addresses or IP addresses) through which the CI can send packets, interfaces and / or addresses (either MAC addresses or IP addresses) through which the CI cannot send packets, one or more interfaces and / or addresses through which the CI can send one or more types of packets, and / or one or more interfaces through which the CI cannot send one or more types of packets. In some embodiments, a CI's ACL may be stored in a location accessible by other entities in the network, including other entities such as one or more VRs or VSRSs that can enforce the ACL.
[0311] For example, when receiving a communication from an IP network for one or more intended recipients within a VLAN, the VSRS may determine and apply filtering and / or delivery restrictions to the communication based on the ACL of the sender of that communication. In some embodiments, this may be accomplished, for example, by (1) the VSRS making a communication decision based on accessing a copy of the sender's ACL, or (2) the VSRS making a communication decision based on information encoded in the packet metadata of the communication received.
[0312] 20, which is a flow chart illustrating one embodiment of a process 2000 for performing deferred access control list (ACL) classification. Process 2000 may be performed by all or a portion of system 600, and in particular, by VSRS 624, 634.
[0313] Process 2000 begins at block 2002. In block 2002, a source CI sends a packet to a destination MAC or IP address. In some embodiments, the source CI may be outside the VLAN that includes the destination MAC or IP address to which the packet is being sent.
[0314] In block 2006, the VSRS for the VLAN containing the destination MAC or IP address receives the packet. In some embodiments, the packet has L2 encapsulation. The packet may be encapsulated using L3 encapsulation, and in some embodiments, the packet may be encapsulated using L4 encapsulation. As shown in block 2008, the VSRS may decapsulate the packet and identify the source CI. Upon identifying the source CI, the VSRS may obtain the ACL of the source CI. In some embodiments, for example, the ACL of the source CI may be stored in a location accessible by the VSRS. In some embodiments, accessing the ACL of the source CI may include retrieving information within the ACL of the source CI. This information may, for example, identify one or more restrictions on delivery of the packet. In some embodiments, this information may include one or more rules based on one or more IP addresses, MAC addresses, source and destination TCP and / or UDP ports, Ethertype, etc.
[0315] In block 2010, the VSRS determines the destination interface for the packet. In embodiments where the VSRS does not have a mapping for the destination address, the VSRS may determine the mapping information as described above with reference to steps 1612 through 1620 of process 1600 of FIG. 16. In embodiments where the mapping was previously learned, or where the mapping was learned by performing some or all of steps 1612 through 1620 of process 1600, the VSRS may determine the destination interface based on the mapping learned by the VSRS through communication with the VSRS's interfaces in the VLAN. In some embodiments, determining the destination interface may include looking up the destination interface based on the destination address, specifically the destination IP address and / or MAC address of the packet.
[0316] In block 2012, the VSRS applies the source CI's ACL to the destination interface. This may include determining whether any portion of the source CI's ACL is relevant to the destination interface and, if so, applying the source CI's ACL to that portion. In block 2014, if the destination interface complies with the source CI's ACL—in other words, if the source CI's ACL allows transmission of the packet from the source CI to the destination interface—the VSRS forwards the packet to the destination interface. In some embodiments, this forwarding of the packet may include encapsulating the packet. In some embodiments, the packet may be encapsulated according to an L2 encapsulation, such as L2 Geneve encapsulation. The VSRS may forward the packet to the destination interface, or more specifically, to a destination CI that has the destination interface as a TEP. The destination interface may receive the packet, decapsulate the packet, and forward the packet to the destination CI.
[0317] Alternatively, in block 2016, if the destination interface does not comply with the source CI's ACL, in other words, if the source CI's ACL does not allow the packet to be transmitted from the source CI to the destination interface, the VSRS discards the packet. In some embodiments, the VSRS can respond to the source CI indicating the discard of the packet and / or the reason for the discard of the packet. In some embodiments, the VSRS can update operational metrics / statistics associated with the transmitting interface to reflect the ACL decision.
[0318] 21, which is a flow chart illustrating one embodiment of a process 2100 for early classification of ACLs and incorporating the classification into metadata. Process 2100 may be performed by all or part of system 600, and in particular, by source C1 and VSRS 624, 634.
[0319] In block 2102, the source CI determines to send a packet to the destination CI. In some embodiments, this may include the source CI determining to send the packet to a MAC address and / or IP address of the destination CI. The source VNIC may send the packet, as shown in block 2104. The source VNIC may receive the packet, evaluate the packet's ACL, and embed ACL information associated with the packet into a portion of the packet. In some embodiments, this information may include one or more rules based on one or more IP addresses, MAC addresses, source and destination TCP and / or UDP ports, Ethertype, etc. The source VNIC may further encapsulate the packet. In some embodiments, the ACL may be stored in the encapsulated packet as metadata, specifically in the packet's header. In some embodiments, the source CI may be outside the VLAN that includes the destination CI to which the packet is being sent.
[0320] At block 2106, the VSRS for the VLAN containing the destination MAC or IP address receives the packet. In some embodiments, the packet may be encapsulated using L2 encapsulation, and in some embodiments, the packet may be encapsulated using L3 encapsulation. The VSRS can decapsulate the packet. In some embodiments, the VSRS can extract information from the packet that identifies the packet's destination, specifically, information that identifies the destination IP address.
[0321] At block 2108, the VSRS determines a destination interface for the packet. In embodiments where the VSRS does not have a mapping for the destination address, the VSRS may determine the mapping information as described above with reference to steps 1612 through 1620 of process 1600 of FIG. 16. In embodiments where the mapping was previously learned, or where the mapping was learned by performing some or all of steps 1612 through 1620 of process 1600, the VSRS may determine the destination interface based on the mapping learned by the VSRS through communication with an interface of the VSRS within the VLAN. Thus, in some embodiments, the VSRS may determine that it has mapping information for the destination address and may determine the destination interface based on this mapping information. In some embodiments, determining the destination interface may include looking up the destination interface based on the destination address, specifically the destination IP address and / or MAC address of the packet.
[0322] At block 2110, the VSRS obtains the ACL information contained in a portion of the packet. In some embodiments, this may include extracting metadata from the packet, specifically extracting metadata from the packet header. In some embodiments, this may include decoding information encoded in the packet metadata in the packet header.
[0323] In block 2112, the VSRS applies the ACL information extracted from the portion of the packet. Specifically, this may include applying the encoded security information to the destination interface. This may include determining any portion of the ACL information that is relevant to the destination interface and, if so, applying the ACL information for that portion. In block 2114, if the destination interface complies with the ACL information and / or one or more rules of the ACL information, in other words, if the ACL information allows the packet to be transmitted from the source CI to the destination interface, the VSRS forwards the packet to the destination interface. In some embodiments, Thus, forwarding the packet may include encapsulating the packet. In some embodiments, the packet may be encapsulated according to an L2 encapsulation, such as L2 Geneve encapsulation. The VSRS may forward the packet to a destination interface, more specifically, to a destination CI having the destination interface as a TEP. The destination interface may receive the packet, decapsulate the packet, and forward the packet to the destination CI.
[0324] Alternatively, if the destination interface does not comply with the ACL information, in other words, if the ACL information does not allow the packet to be transmitted from the source CI to the destination interface, the VSRS discards the packet at block 2116. In some embodiments, the VSRS may respond to the source CI indicating the discard of the packet and / or the reason for the discard of the packet.
[0325] Next-hop routing Some embodiments of the VSRS can facilitate next-hop routing, specifically delaying evaluation of the next hop until the VSRS receives the communication. Because the sender of the communication outside the VLAN may not know the updated virtual IP address of the instance within the VLAN, the sender outside the VLAN cannot accurately specify next-hop routing.
[0326] In some embodiments, the VSRS can facilitate next-hop routing specification. In some embodiments, this can be accomplished, for example, by (1) the sender making an initial specification of next-hop routing for a communication and re-evaluating the next-hop specification once the VSRS receives the communication, or (2) delaying the next-hop specification until the VSRS receives the communication.
[0327] Referring to FIG. 22, FIG. 22 is a flow chart illustrating one embodiment of a process 2200 for performing sender-based next-hop routing. In some embodiments, this may involve separating the evaluation of the next-hop route from the specification of the next-hop route. This may result in the source CI evaluating the route policy and encoding the next-hop route into packet metadata in the virtual packet header of the communication. In such an embodiment, the packet may include the intended virtual IP address of the next-hop destination, and the next-hop route may be encoded into the packet metadata.
[0328] The communication is received by the VSRS, and the VSRS uses the next hop route encoded in the packet metadata to determine the instance within the VLAN that will receive the communication and the instance's destination virtual IP address. This instance may be determined using tables generated and / or curated by the VSRS, specifically tables that link virtual IP addresses, MAC addresses, and / or virtual interface IDs. Using these tables, the VSRS can identify the virtual IP address that corresponds to the MAC address and / or virtual interface ID of the intended next hop destination. In some embodiments, the VSRS's identification of the recipient instance based on the encoded next hop route included in the packet metadata can result in the VSRS sending the communication to a destination virtual IP address indicated by the virtual IP address in the packet that differs from the next hop destination specified by the sender of the packet.
[0329] The process 2200 may be performed by all or part of the system 600, and in particular by the source CI and VSRS 624, 634.
[0330] The process begins at block 2202. In block 2202, a source CI may transmit a packet. In some embodiments, the source CI may be outside the VLAN that includes the destination MAC address or IP address to which the packet is being transmitted.
[0331] At block 2204, the source VNIC may make a routing decision based on one or more routing rules. In some embodiments, this may include the source VNIC looking up and / or obtaining one or more routing rules and making a routing decision based on these routing rules. In some embodiments, this routing decision may be a next-hop routing decision for the packet. In some embodiments, these routing rules may be stored in a route table in the network of the source CI, for example. In some embodiments, this may include a subnet route table, for example.
[0332] At block 2206, the source VNIC may embed the routing decision into a portion of the packet. In some embodiments, this may include encapsulating the packet and embedding the routing decision into metadata of the packet. The metadata may be encoded, for example, in a header of the packet. After encapsulating the packet and embedding the routing decision into the encapsulated packet, the source VNIC may forward the encapsulated packet to the VSRS.
[0333] At block 2208, a VSRS in a VLAN containing the destination MAC address or IP address receives the packet. In some embodiments, the packet may be encapsulated using L2 encapsulation, and in some embodiments, the packet may be encapsulated using L3 encapsulation. In some embodiments, receiving the packet by the VSRS may include decapsulating the packet. In some embodiments, receiving the packet by the VSRS may include extracting a routing decision embedded in a portion of the packet. In some embodiments, this may include decoding encoded metadata containing the routing decision.
[0334] At block 2210, the VSRS determines a destination interface for the packet by applying the routing decision to the VSRS routing information. In some embodiments, this may include determining a destination CI within the VLAN that corresponds to the routing information by applying the decoded routing decision to the VSRS routing information. At block 2212, the VSRS transmits the packet to a CI within the determined VLAN. In some embodiments, this may include the VSRS forwarding the packet to the destination interface. In some embodiments, this forwarding of the packet may include encapsulating the packet, which may be performed according to L2 encapsulation, such as L2 Geneve encapsulation. The VSRS may forward the packet to the destination interface, and more specifically, to a destination CI that has the destination interface as a TEP. The destination interface may receive the packet, decapsulate the packet, and forward the packet to the destination CI.
[0335] 23, a flow chart illustrating one embodiment of a process 2300 for performing deferred next-hop routing is shown. The process 2300 may be performed by all or a portion of the system 600, and in particular, by the source VNIC and VSRS 624, 634.
[0336] In some embodiments, for example, a source VNIC can make a routing decision for a communication traversing a VLAN based on a routing rule that may be included in a routing table. The communication is received by a VSRS, which can re-evaluate the next-hop specification and identify the routing rule by consulting a copy of the same routing table. Based on the routing rule, the VSRS determines an instance within the VLAN that corresponds to the routing rule and a virtual IP address for that instance, and sends the communication to that instance. Determining the instance that corresponds to the routing rule may include the VSRS retrieving information from a table that links virtual IP addresses, MAC addresses, and / or virtual interface IDs, and determining the virtual IP address associated with the virtual interface and / or instance that is the intended next-hop destination.
[0337] Process 2300 begins at block 2302. At block 2302, a source CI may transmit a packet. In some embodiments, this may include the source CI transmitting the packet to a MAC address and / or IP address of a destination CI. In some embodiments, the source CI may be outside of a VLAN that includes the destination MAC address or IP address to which the packet is being transmitted.
[0338] At block 2204, the source VNIC may receive the packet. The source VNIC may then make a routing decision based on one or more routing rules. In some embodiments, this may include the source VNIC looking up and / or obtaining one or more routing rules and making a routing decision based on those routing rules. In some embodiments, this routing decision may be a next-hop routing decision for the packet. In some embodiments, these routing rules may be stored, for example, in a route table in the network of the source CI. In some embodiments, this may include, for example, a subnet route table. The source VNIC may encapsulate the packet and send the packet to the VSRS.
[0339] At block 2206, a VSRS in a VLAN containing the destination MAC or IP address receives the packet. In some embodiments, the packet may be encapsulated using L2 encapsulation, and in some embodiments, the packet may be encapsulated using L3 encapsulation. In some embodiments, receiving the packet by the VSRS may include decapsulating the packet.
[0340] At block 2308, the VSRS obtains routing information associated with the packet. In some embodiments, this may include obtaining a routing table and / or a portion of a routing table associated with the received packet. In some embodiments, this routing table may be received, for example, from a control plane. Upon receiving the routing information, the VSRS determines a routing rule associated with the received packet.
[0341] In block 2310, the VSRS determines a destination interface of a VLAN for the packet by applying routing rules to the VSRS routing information. In some embodiments, this may include determining a destination CI of a VLAN corresponding to the routing information by applying the routing rules to the VSRS routing information. In block 2312, the VSRS transmits the packet to the determined CI of the VLAN. In some embodiments, this may include the VSRS forwarding the packet to the destination interface. In some embodiments, Thus, forwarding the packet may include encapsulating the packet, which may be performed according to an L2 encapsulation such as L2 Geneve encapsulation. The VSRS may forward the packet to a destination interface, more specifically, to a destination CI having the destination interface as a TEP. The destination interface may receive the packet, decapsulate the packet, and forward the packet to the destination CI.
[0342] Exemplary Implementation As mentioned above, IaaS (Infrastructure as a Service) is a specific type of Cloud computing. IaaS may be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider may host infrastructure elements (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may offer various services (e.g., billing, monitoring, logging, load balancing, clustering, etc.) associated with the infrastructure elements. Therefore, because these services can be policy-driven, IaaS users can implement policies to drive load balancing to maintain application availability and performance.
[0343] In some examples, IaaS customers can access resources and services over a wide area network (WAN), such as the Internet, and use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and install enterprise software on the VMs. Customers can use the provider's services to perform a variety of functions, including balancing network traffic, troubleshooting applications, monitoring performance, managing disaster recovery, and more.
[0344] In most cases, the cloud computing model requires the participation of a cloud provider, which can be, but does not have to be, a third-party service provider specializing in providing (e.g., offering, renting, or selling) IaaS. Alternatively, an enterprise can deploy a private cloud and become a provider of infrastructure services.
[0345] In some examples, IaaS deployment is the process of deploying a new application or a new version of an application onto a provisioned application server or the like. IaaS deployment may include the process of provisioning the server (e.g., installing libraries, daemons, etc.). IaaS deployment is often managed by the cloud provider below the hypervisor layer (e.g., server, storage, network hardware, and virtualization). Thus, customers can deploy the OS, middleware, and / or applications (e.g., self-service virtual machines (e.g., that can be spun up on demand)).
[0346] In some instances, IaaS provisioning may include obtaining the computer or virtual host to be used and installing the necessary libraries or services on the computer or virtual host. In most cases, deployment does not include provisioning, which must be performed first.
[0347] In some cases, IaaS provisioning presents two distinct challenges. First, there is the challenge of provisioning an initial set of infrastructure before anything can be done. Second, there is the challenge of evolving the existing infrastructure (e.g., adding new services, modifying services, removing services) after everything has been provisioned. In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., which elements are needed and how these elements interact) may be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which and how they work together) can be described declaratively. In some instances, once the topology is defined, workflows can be generated to create and / or manage the different elements described in the configuration files.
[0348] In some examples, the infrastructure can include many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potentially on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there may be one or more inbound and / or outbound traffic group rules and one or more virtual machines (VMs) that are provisioned to define how to configure inbound and / or outbound traffic for the network. Other infrastructure elements, such as load balancers, databases, etc., may also be provisioned. The infrastructure can evolve incrementally as more infrastructure elements are desired and / or added.
[0349] In some examples, continuous deployment techniques may be employed to enable deployment of infrastructure code across various virtual computing environments. The described techniques may also enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, typically many, different production environments (e.g., across a variety of different geographic locations, sometimes spanning the entire world). However, in some examples, the infrastructure for deploying the code must first be set up. In some examples, provisioning may be done manually, with provisioning tools used to provision resources and / or deployment tools used to deploy the code after the infrastructure has been provisioned.
[0350] FIG. 24 is a block diagram 2400 illustrating an example pattern of an IaaS architecture, according to at least one embodiment. A service operator 2402 may be communicatively connected to a secure host tenancy 2404, which may include a virtual cloud network (VCN) 2406 and a secure host subnet 2408. In some examples, the service operator 2402 may use one or more client computing devices. The one or more client computing devices may run, for example, software such as Microsoft Windows Mobile®, and / or various operating systems such as iOS, Windows Phone, Android®, BlackBerry 8, and Palm OS. The device may be a handheld mobile device (e.g., iPhone®, mobile phone, iPad®, tablet, personal digital assistant (PDA) or wearable device (e.g., Google® Glass® head-mounted display) capable of running any mobile operating system and enabled for Internet, email, short message service (SMS), Blackberry® or other communication protocols. The client computing device illustratively includes a Microsoft Windows Operating System, Apple Macintosh® Operating System and and / or running various versions of the Linux operating system. The client computing devices may be general-purpose personal computers, including personal computers and / or laptop computers running Windows 2000 and Windows Server 2008 R2. Alternatively, the client computing devices may be workstation computers running various commercially available UNIX or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems, e.g., Google Chrome OS. Alternatively or additionally, the client computing devices may be other electronic devices capable of communicating over the VCN 2406 and / or a network with access to the Internet, such as thin-client computers, Internet-enabled gaming systems (e.g., Microsoft Xbox game consoles with or without Kinect gesture input devices), and / or personal messaging devices.
[0351] VCN 2406 may include a local peering gateway (LPG) 2410 that may be communicatively connected to a secure shell (SSH) VCN 2412 via an LPG 2410 included in SSH VCN 2412. SSH VCN 2412 may include an SSH subnet 2414, which may be communicatively connected to a control plane VCN 2416 via an LPG 2410 included in control plane VCN 2416. SSH VCN 2412 may also be communicatively connected to a data plane VCN 2418 via LPG 2410. Control plane VCN 2416 and data plane VCN 2418 may be included in a service tenancy 2419, which may be owned and / or operated by the IaaS provider.
[0352] The control plane VCN 2416 may include a control plane demilitarized zone (DMZ) tier 2420 that functions as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers may have a particular level of trust and contain security breaches. Additionally, the DMZ tier 2420 may include one or more load balancer (LB) subnets 2422, a control plane app tier 2424 that may include an app subnet 2426, and a control plane data tier 2428 that may include a database (DB) subnet 2430 (e.g., a front-end DB subnet and / or a back-end DB subnet). LB subnet 2422 included in control plane DMZ layer 2420 may be communicatively connected to app subnet 2426 included in control plane app layer 2424 and to Internet gateway 2434, which may be included in control plane VCN 2416, and applisub 2426 may be communicatively connected to DB subnet 2430, service gateway 2436, and network address translation (NAT) gateway 2438 included in control plane data layer 2428. Control plane VCN 2416 may include service gateway 2436 and NAT gateway 2438.
[0353] The control plane VCN 2416 can include a data plane mirror app layer 2440, which can include an app subnet 2426. The app subnet 2426 included in the data plane mirror app layer 2440 can include a virtual network interface controller (VNIC) 2442 on which a compute instance 2444 can run. The compute instance 2444 can communicatively connect the app subnet 2426 of the data plane mirror app layer 2440 to the app subnet 2426, which can be included in the data plane app layer 2446.
[0354] Data plane VCN 2418 may include a data plane app layer 2446, a data plane DMZ layer 2448, and a data plane data layer 2450. Data plane DMZ layer 2448 may include LB subnet 2422, which may be communicatively connected to app subnet 2426 of data plane app layer 2446 and internet gateway 2434 of data plane VCN 2418. App subnet 2426 may be communicatively connected to service gateway 2436 of data plane VCN 2418 and NAT gateway 2438 of data plane VCN 2418. Data plane data layer 2450 may also include DB subnet 2430, which may be communicatively connected to app subnet 2426 of data plane app layer 2446.
[0355] The internet gateway 2434 of the control plane VCN 2416 and the internet gateway 2434 of the data plane VCN 2418 may be communicatively connected to a metadata management service 2452, which may be communicatively connected to the public internet 2454. The public internet 2454 may be communicatively connected to a NAT gateway 2438 of the control plane VCN 2416 and the NAT gateway 2438 of the data plane VCN 2418. The service gateway 2436 of the control plane VCN 2416 and the service gateway 2436 of the data plane VCN 2418 may be communicatively connected to cloud services 2456.
[0356] In some examples, the service gateway 2436 of the control plane VCN 2416 or the data plane VCN 2418 can make application programming interface (API) calls to the cloud services 2456 without traversing the public internet 2454. The API calls from the service gateway 2436 to the cloud services 2456 can be one-way. The service gateway 2436 can make API calls to the cloud services 2456, and the cloud services 2456 can send request data to the service gateway 2436. However, the cloud services 2456 may not initiate the API calls to the service gateway 2436.
[0357] In some examples, secure host tenancy 2404 may be directly connected to service tenancy 2419, which may be an orphan. Secure host subnet 2408 can communicate with SSH subnet 2414 through LPG 2410, which allows bidirectional communication with the orphan system. By connecting secure host subnet 2408 to SSH subnet 2414, secure host subnet 2408 can access other entities in service tenancy 2419.
[0358] Control plane VCN 2416 allows users of service tenancy 2419 to configure or provision desired resources. The desired resources provisioned in control plane VCN 2416 may be deployed or used in data plane VCN 2418. In some examples, control plane VCN 2416 may be isolated from data plane VCN 2418, and data plane mirror app layer 2440 of control plane VCN 2416 can communicate with data plane app layer 2446 of data plane VCN 2418 via VNIC 2442, which may be included in data plane mirror app layer 2440 and data plane app layer 2446.
[0359] In some examples, a user or customer of the system can make a request, such as, for example, a create, read, update, or delete (CRUD) operation, via the public internet 2454, which can communicate the request to a metadata management service 2452. The metadata management service 2452 can communicate the request to the control plane VCN 2416 via an internet gateway 2434. The request is routed to the control plane DMZ The request may be received by LB subnet 2422 included in layer 2420. LB subnet 2422 may determine that the request is valid, and in response to this determination, LB subnet 2422 may send the request to app subnet 2426 included in control plane app layer 2424. If the request is validated and requires a call to the public Internet 2454, the call to the public Internet 2454 may be sent to NAT gateway 2438, which can make the call to the public Internet 2454. A memory for storing the request may be stored in DB subnet 2430.
[0360] In some examples, the data plane mirror app layer 2440 can facilitate direct communication between the control plane VCN 2416 and the data plane VCN 2418. For example, it may be desirable for changes, updates, or other suitable modifications to a configuration to be applied to resources included in the data plane VCN 2418. The control plane VCN 2416 can communicate directly with the resources included in the data plane VCN 2418 via VNIC 2442, allowing the changes, updates, or other suitable modifications to the configuration to be implemented.
[0361] In some embodiments, the control plane VCN 2416 and the data plane VCN 2418 may be included in the service tenancy 2419. In this case, a user or customer of the system may not own or operate either the control plane VCN 2416 or the data plane VCN 2418. Instead, an IaaS provider may own or operate the control plane VCN 2416 and the data plane VCN 2418, both of which may be included in the service tenancy 2419. This embodiment can prevent users or customers from interacting with other users' or customers' resources by enabling network isolation. This embodiment can also allow users or customers of the system to store databases privately without having to rely on the public internet 2454, which may not have the desired level of threat protection for storage.
[0362] In another embodiment, LB subnet 2422 included in control plane VCN 2416 may be configured to receive signals from service gateway 2436. In this embodiment, control plane VCN 2416 and data plane VCN 2418 may be configured to be called by customers of the IaaS provider without calling the public internet 2454. Customers of the IaaS provider may desire this embodiment because databases used by the customers may be stored in service tenancy 2419, which is controlled by the IaaS provider and may be isolated from the public internet 2454.
[0363] Figure 25 is a block diagram 2500 illustrating another example parameter of an IaaS architecture, according to at least one embodiment. A service operator 2502 (e.g., service operator 2402 of Figure 24) may be communicatively connected to a secure host tenancy 2504 (e.g., secure host tenancy 2404 of Figure 24), which may include a virtual cloud network (VCN) 2506 (e.g., VCN 2406 of Figure 24) and a secure host subnet 2508 (e.g., secure host subnet 2408 of Figure 24). VCN 2506 may include a local peering gateway (LPG) 2510 (e.g., LPG 2410 of Figure 24), which may be communicatively connected to a secure shell (SSH) VCN 2512 (e.g., SSH VCN 2412 of Figure 24) via an LPG 2410 included in SSH VCN 2512. The SSH VCN 2512 can include an SSH subnet 2514 (e.g., SSH subnet 2414 in FIG. 24 ), and the SSH VCN 2512 can communicate with the control plane VCN 2516 via an LPG 2510 included in the control plane VCN 2516. It can be communicatively connected to control plane VCN 2524 (e.g., control plane VCN 2416 in FIG. 24 ). Control plane VCN 2524 may be included in service tenancy 2519 (e.g., service tenancy 2419 in FIG. 24 ), and data plane VCN 2518 (e.g., data plane VCN 2418 in FIG. 24 ) may be included in customer tenancy 2521, which may be owned or operated by a user or customer of the system.
[0364] The control plane VCN 2516 may include a control plane DMZ layer 2520 (e.g., control plane DMZ layer 2420 in FIG. 24 ) that may include a LB subnet 2522 (e.g., LB subnet 2422 in FIG. 24 ), a control plane app layer 2516 (e.g., control plane app layer 2424 in FIG. 24 ) that may include an app subnet 2526 (e.g., app subnet 2426 in FIG. 24 ), and a control plane data layer 2528 (e.g., control plane data layer 2428 in FIG. 24 ) that may include a database (DB) subnet 2530 (e.g., similar to DB subnet 2430 in FIG. 24 ). The LB subnet 2522 included in the control plane DMZ layer 2520 may be communicatively connected to the app subnet 2526 included in the control plane app layer 2516 and to an Internet gateway 2534 (e.g., Internet gateway 2434 in FIG. 24 ), which may be included in the control plane VCN 2516. The app subnet 2526 may be communicatively connected to a DB subnet 2530, a service gateway 2536 (e.g., the service gateway in FIG. 24 ), and a network address translation (NAT) gateway 2538 (e.g., the NAT gateway 2438 in FIG. 24 ) included in the control plane data layer 2528. The control plane VCN 2516 may include the service gateway 2536 and the NAT gateway 2538.
[0365] The control plane VCN 2516 can include a data plane mirror app layer 2540 (e.g., data plane mirror app layer 2440 of FIG. 24 ), which can include an app subnet 2526. The app subnet 2526 included in the data plane mirror app layer 2540 can include a virtual network interface controller (VNIC) 2542 (e.g., VNIC 2442) on which a compute instance 2544 (e.g., similar to compute instance 2444 of FIG. 24 ) can run. The compute instance 2544 can facilitate communication between the app subnet 2526 of the data plane mirror app layer 2540 and the app subnet 2526, which can be included in the data plane app layer 2546 (e.g., data plane app layer 2446 of FIG. 24 ), via the VNIC 2542 included in the data plane mirror app layer 2540 and the VNIC 2542 included in the data plane app layer 2546.
[0366] The internet gateway 2534 included in the control plane VCN 2516 may be communicatively connected to a metadata management service 2552 (e.g., metadata management service 2452 of FIG. 24), which may be communicatively connected to the public internet 2554 (e.g., public internet 2454 of FIG. 24). The public internet 2554 may be communicatively connected to a NAT gateway 2538 included in the control plane VCN 2516. The service gateway 2536 included in the control plane VCN 2516 may be communicatively connected to cloud services 2556 (e.g., cloud services 2456 of FIG. 24).
[0367] In some examples, the data plane VCN 2518 may be included in the customer tenancy 2521. In this case, the IaaS provider may provide a control plane VCN 2516 for each customer, and the IaaS provider may configure a unique compute instance 2544 for each customer, which is included in the service tenancy 2519. Each compute instance 2544 allows communication between the control plane VCN 2516, which is included in the service tenancy 2519, and the data plane VCN 2518, which is included in the customer tenancy 2521. Compute instance 2544 can allow resources provisioned in control plane VCN 2516 contained in service tenancy 2519 to be deployed or used in data plane VCN 2518 contained in customer tenancy 2521.
[0368] In another example, a customer of the IaaS provider may have a database that resides in customer tenancy 2521. In this example, control plane VCN 2516 may include data plane minor app tier 2540, which may include app subnet 2526. Data plane mirror app tier 2540 may reside in data plane VCN 2518, but data plane mirror app tier 2540 may not reside in data plane VCN 2518. That is, data plane mirror app tier 2540 has access to customer tenancy 2521, but data plane mirror app tier 2540 may not reside in data plane VCN 2518 and may not be owned or operated by the IaaS provider's customer. Data plane mirror app tier 2540 may be configured to make calls to data plane VCN 2518, but may not be configured to make calls to any entities included in control plane VCN 2516. A customer may desire to deploy or use resources in the data plane VCN 2518 provisioned to the control plane VCN 2516, and the data plane mirror application tier 2540 may facilitate the desired deployment or other use of the customer's resources.
[0369] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 2518. In this embodiment, the customer can determine what the data plane VCN 2518 can access, and the customer can restrict access from the data plane VCN 2518 to the public internet 2554. The IaaS provider may not be able to apply filters or control access from the data plane VCN 2518 to any external network or database. Applying filters and controls to the data plane VCN 2518 included in the customer tenancy 2521 by the customer can help isolate the data plane VCN 2518 from other customers and the public internet 2554.
[0370] In some embodiments, cloud services 2556 can be called by service gateway 2536 to access services that may not reside on the public internet 2554, on the control plane VCN 2516, or on the data plane VCN 2518. The connection between cloud services 2556 and the control plane VCN 2516 or the data plane VCN 2518 may not be live or continuous. Cloud services 2556 may reside on a separate network owned or operated by the IaaS provider. Cloud services 2556 may be configured to receive calls from service gateway 2536 and may not be configured to receive calls from the public internet 2554. Some cloud services 2556 may be isolated from other cloud services 2556, and control plane VCN 2516 may be isolated from cloud services 2556 that may not be located in the same region as control plane VCN 2516. For example, control plane VCN 2516 may be located in “Region 1,” and cloud service “Deployment 24” may be located in “Region 1” and “Region 2.” If a call to a deployment 24 is made by a service gateway 2536 included in a control plane VCN 2516 located in Region 1, the call may be sent to the deployment 24 in Region 1. In this example, the control plane VCN 2516 or the deployment 24 in Region 1 may not be communicatively connected to the deployment 24 in Region 2.
[0371] 26 is a block diagram 2600 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 2602 (e.g., service operator 2402 of FIG. 24) may be communicatively connected to a secure host tenancy 2604 (e.g., secure host tenancy 2404 of FIG. 24), which may include a virtual cloud network (VCN) 2606 (e.g., VCN 2406 of FIG. 24) and a secure host subnet 2608 (e.g., secure host subnet 2408 of FIG. 24). VCN 2606 may include an LPG 2610 (e.g., LPG 2410 of FIG. 24), which may be communicatively connected to an SSH VCN 2612 (e.g., SSH VCN 2412 of FIG. 24) via an LPG 2610 included in SSH VCN 2612. SSH VCN 2612 can include SSH subnet 2614 (e.g., SSH subnet 2414 in FIG. 24 ), which may be communicatively connected to control plane VCN 2616 (e.g., control plane VCN 2416 in FIG. 24 ) via LPG 2610 included in control plane VCN 2616, and may be communicatively connected to data plane VCN 2618 (e.g., data plane 2418 in FIG. 24 ) via LPG 2610 included in data plane VCN 2618. Control plane VCN 2616 and data plane VCN 2618 may be included in service tenancy 2619 (e.g., service tenant 2419 in FIG. 24 ).
[0372] Control plane VCN 2616 may include a control plane DMZ layer 2620 (e.g., control plane DMZ layer 2420 in FIG. 24 ) that may include a load balancer (LB) subnet 2622 (e.g., LB subnet 2422 in FIG. 24 ), a control plane app layer 2624 (e.g., control plane app layer 2424 in FIG. 24 ) that may include an app subnet 2626 (e.g., similar to app subnet 2426 in FIG. 24 ), and a control plane data layer 2628 (e.g., control plane data layer 2428 in FIG. 24 ) that may include a DB subnet 2630. LB subnet 2622 included in control plane DMZ layer 2620 may be communicatively connected to app subnet 2626 included in control plane app layer 2624 and to an Internet gateway 2634 (e.g., Internet gateway 2434 in FIG. 24 ), which may be included in control plane VCN 2616. App subnet 2626 may be communicatively connected to DB subnet 2630 included in control plane data layer 2628, and to service gateway 2636 (e.g., service gateway in FIG. 24) and network address translation (NAT) gateway 2638 (e.g., NAT gateway 2438 in FIG. 24). Control plane VCN 2616 may include service gateway 2636 and NAT gateway 2638.
[0373] Data plane VCN 2618 may include a data plane app layer 2646 (e.g., data plane app layer 2446 in FIG. 24 ), a data plane DMZ layer 2648 (e.g., data plane DMZ layer 2448 in FIG. 24 ), and a data plane data layer 2650 (e.g., data plane data layer 2450 in FIG. 24 ). Data plane DMZ layer 2648 may include LB subnet 2622, which may be communicatively connected to data plane app layer 2646 and trusted app subnet 2660 and untrusted app subnet 2662 of internet gateway 2634 included in data plane VCN 2618. Trusted app subnet 2660 may be communicatively connected to service gateway 2636 included in data plane VCN 2618, NAT gateway 2638 included in data plane VCN 2618, and DB subnet 2630 included in data plane data layer 2650. The untrusted app subnet 2662 may be communicatively connected to a service gateway 2636 included in the data plane VCN 2618 and to a DB subnet 2630 included in the data plane data layer 2650. The data plane data layer 2650 includes a DB subnet 2630 that may be communicatively connected to a service gateway 2636 included in the data plane VCN 2618. It may include a net 2630 .
[0374] The untrusted app subnet 2662 may include one or more primary VNICs 2664(1)-(N), which may be communicatively connected to tenant virtual machines (VMs) 2666(1)-(N). Each tenant VM 2666(1)-(N) may be communicatively connected to a respective app subnet 2667(1)-(N), which may be included in a respective container egress VCN 2668(1)-(N), which may be included in a respective customer tenancy 2670(1)-(N). Each secondary VNIC 2672(1)-(N) may facilitate communication between the untrusted app subnet 2662 included in the data plane VCN 2618 and the app subnet included in the container egress VCN 2668(1)-(N). Each container egress VCN 2668(1)-(N) may include a NAT gateway 2638, which may be communicatively connected to the public internet 2654 (e.g., public internet 2454 in FIG. 24 ).
[0375] The internet gateway 2634 included in the control plane VCN 2616 and the internet gateway 2634 included in the data plane VCN 2618 may be communicatively connected to a metadata management service 2652 (e.g., metadata management system 2452 of FIG. 24 ), which may be communicatively connected to the public internet 2654. The public internet 2654 may be communicatively connected to a NAT gateway 2638 included in the control plane VCN 2616 and the NAT gateway 2638 included in the data plane VCN 2618. The service gateway 2636 included in the control plane VCN 2616 and the service gateway 2636 included in the data plane VCN 2618 may be communicatively connected to cloud services 2656.
[0376] In some embodiments, data plane VCN 2618 may be integrated into customer tenancy 2670. This integration may be useful or desirable for an IaaS provider's customer in some cases, such as when they may want support when running their code. Customers may provide code that, when executed, may be disruptive, communicate with other customer resources, or cause undesirable effects. Thus, the IaaS provider can determine whether or not to run code that a customer has provided to the IaaS provider.
[0377] In some examples, an IaaS provider's customer can grant temporary network access to the IaaS provider and request a feature to be added to data plane app layer 2646. The code to perform the feature may run in VMs 2666(1)-(N) but cannot be configured to run elsewhere on data plane VCN 2618. Each VM 2666(1)-(N) may be connected to one customer tenancy 2670. Each container 2671(1)-(N) included in a VM 2666(1)-(N) may be configured to run code. In this case, double isolation (e.g., containers 2671(1)-(N) may run code, and containers 2671(1)-(N) may be included in at least VMs 2666(1)-(N) included in untrusted app subnet 2662) may exist, which can help prevent erroneous or unwanted code from damaging the IaaS provider's network or from damaging a different customer's network. Containers 2671(1)-(N) may be communicatively connected to customer tenancy 2670 and may be configured to send or receive data from customer tenancy 2670. Containers 2671(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 2618. Once code execution is complete, the IaaS provider can kill or discard containers 2671(I)-(N).
[0378] In some embodiments, trusted app subnet 2660 can execute code that may be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 2660 may be communicatively connected to DB subnet 2630 and configured to perform CRUD operations on DB subnet 2630. Untrusted app subnet 2662 may be communicatively connected to DB subnet 2630, but in this embodiment, the untrusted app subnet may be configured to perform read operations within DB subnet 2630. Containers 2671(1)-(N) included in each customer's VMs 2666(1)-(N) and capable of executing code from the customer may not be communicatively connected to DB subnet 2630.
[0379] In other embodiments, the control plane VCN 2616 and the data plane VCN 2618 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 2616 and the data plane VCN 2618. However, there may be indirect communication by at least one method. An LPG 2610 may be established by an IaaS provider that can facilitate communication between the control plane VCN 2616 and the data plane VCN 2618. In another example, the control plane VCN 2616 or the data plane VCN 2618 can make a call to a cloud service 2656 through a service gateway 2636. For example, a call from the control plane VCN 2616 to the cloud service 2656 can include a request for a service that can communicate with the data plane VCN 2618.
[0380] 27 is a block diagram 2700 illustrating another example parameter of an IaaS architecture, according to at least one embodiment. A service operator 2702 (e.g., service operator 2402 of FIG. 24) may be communicatively connected to a secure host tenancy 2704 (e.g., secure host tenancy 2404 of FIG. 24), which may include a virtual cloud network (VCN) 2706 (e.g., VCN 2406 of FIG. 24) and a secure host subnet 2708 (e.g., secure host subnet 2408 of FIG. 24). VCN 2706 may include an LPG 2710 (e.g., LPG 2410 of FIG. 24), which may be communicatively connected to an SSH VCN 2712 (e.g., SSH VCN 2412 of FIG. 24) via an LPG 2710 included in the SSH VCN 2712. VCN 2712 can include SSH subnet 2714 (e.g., SSH subnet 2414 in FIG. 24), which may be communicatively connected to control plane VCN 2716 (e.g., control plane VCN 2416 in FIG. 24) via LPG 2710 included in control plane VCN 2716, and may be communicatively connected to data plane VCN 2718 (e.g., data plane 2418 in FIG. 24) via LPG 2710 included in data plane VCN 2718. Control plane VCN 2716 and data plane VCN 2718 may be included in service tenancy 2719 (e.g., service tenancy 2419 in FIG. 24).
[0381] The control plane VCN 2716 may include a control plane DMZ layer 2720 (e.g., control plane DMZ layer 2420 of FIG. 24 ) that may include a LB subnet 2722 (e.g., LB subnet 2422 of FIG. 24 ), a control plane app layer 2724 (e.g., control plane app layer 2424 of FIG. 24 ) that may include an app subnet 2726 (e.g., app subnet 2426 of FIG. 24 ), and a control plane data layer 2728 (e.g., control plane data layer 2428 of FIG. 24 ) that may include a DB subnet 2730 (e.g., DB subnet 2630 of FIG. 26 ). The LB subnet 2722 included in the control plane DMZ layer 2720 is connected to the app subnet 2726 included in the control plane app layer 2724 and an Internet gateway 2734 (e.g., , internet gateway 2434 in FIG. 24 ). App subnet 2726 may be communicatively connected to DB subnet 2730, which is included in control plane data layer 2728, and to service gateway 2736 (e.g., service gateway in FIG. 24 ) and network address translation (NAT) gateway 2738 (e.g., NAT gateway 2438 in FIG. 24 ). Control plane VCN 2716 may include service gateway 2736 and NAT gateway 2738.
[0382] Data plane VCN 2718 can include a data plane app layer 2746 (e.g., data plane app layer 2446 in FIG. 24 ), a data plane DMZ layer 2748 (e.g., data plane DMZ layer 2448 in FIG. 24 ), and a data plane data layer 2750 (e.g., data plane data layer 2450 in FIG. 24 ). Data plane DMZ layer 2748 can include trusted app subnet 2760 (e.g., trusted app subnet 2660 in FIG. 26 ) and untrusted app subnet 2762 (e.g., untrusted app subnet 2662 in FIG. 26 ) of data plane app layer 2746 and LB subnet 2722, which can be communicatively connected to an Internet gateway 2734 included in data plane VCN 2718. Trusted app subnet 2760 may be communicatively connected to service gateway 2736 included in data plane VCN 2718, NAT gateway 2738 included in data plane VCN 2718, and DB subnet 2730 included in data plane data layer 2750. Untrusted app subnet 2762 may be communicatively connected to service gateway 2736 included in data plane VCN 2718 and DB subnet 2730 included in data plane data layer 2750. Data plane data layer 2750 may include DB subnet 2730, which may be communicatively connected to service gateway 2736 included in data plane VCN 2718.
[0383] The untrusted app subnet 2762 may include primary VNICs 2764(1)-(N) that may be communicatively connected to tenant virtual machines (VMs) 2766(1)-(N) that reside in the untrusted app subnet 2762. Each tenant VM 2766(1)-(N) may execute code in a respective container 2767(1)-(N) and may be communicatively connected to an app subnet 2726 that may be included in a data plane app layer 2746 that may be included in a container egress VCN 2768. Each secondary VNIC 2772(1)-(N) may facilitate communication between the untrusted app subnet 2762 included in the data plane VCN 2718 and the app subnet included in the container egress VCN 2768. The container egress VCN may include a NAT gateway 2738 that may be communicatively connected to the public internet 2754 (e.g., public internet 2454 in FIG. 24 ).
[0384] The internet gateway 2734 included in the control plane VCN 2716 and the internet gateway 2734 included in the data plane VCN 2718 may be communicatively connected to a metadata management service 2752 (e.g., metadata management system 2452 of FIG. 24 ), which may be communicatively connected to the public internet 2754. The public internet 2754 may be communicatively connected to the internet gateway 2734 included in the control plane VCN 2716 and the NAT gateway 2738 included in the data plane VCN 2718. The internet gateway 2734 included in the control plane VCN 2716 and the service gateway 2736 included in the data plane VCN 2718 may be communicatively connected to cloud services 2756.
[0385] In some examples, the architecture shown in block diagram 2700 of FIG. This pattern may be considered an exception to the pattern illustrated by the architecture of block diagram 2600 in FIG. 26 and may be desirable for an IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (e.g., in an unconnected region). The customer may have real-time access to each of the containers 2767(1)-(N) contained in each customer's VMs 2766(1)-(N). The containers 2767(1)-(N) may be configured to call each of the secondary VNICs 2772(1)-(N) contained in the app subnet 2726 of the data plane app tier 2746, which may be included in the container egress VCN 2768. The secondary VNICs 2772(1)-(N) may send the call to a NAT gateway 2738, which may send the call to the public Internet 2754. In this example, containers 2767(1)-(N) that a customer can access in real time may be isolated from control plane VCN 2716 and may be isolated from other entities included in data plane VCN 2718. Containers 2767(1)-(N) may also be isolated from resources of other customers.
[0386] In another example, a customer can invoke cloud service 2756 using containers 2767(1)-(N). In this example, the customer can execute code in containers 2767(1)-(N) that requests a service from cloud service 2756. Containers 2767(1)-(N) can send the request to secondary VNICs 2772(1)-(N), which can send the request to a NAT gateway that can send the request to public internet 2754. Public internet 2754 can send the request to LB subnet 2722, which is included in control plane VCN 2716, via internet gateway 2734. In response to determining that the request is valid, LB subnet 2726 can send the request to app subnet 2726, which can send the request to cloud service 2756 via service gateway 2736.
[0387] It should be noted that the illustrated IaaS architectures 2400, 2500, 2600, and 2700 may include elements other than those shown. Furthermore, the illustrated embodiments are only examples of some of the cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS system may have more or fewer elements than those shown, may combine two or more elements, or may have a different configuration or arrangement of elements.
[0388] In certain embodiments, the IaaS system described in this disclosure may include a suite of application, middleware, and database services that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example of such an IaaS system is the Oracle® Cloud Infrastructure (OCI) offered by the present applicant.
[0389] 28 illustrates an exemplary computer system 2800 upon which various embodiments may be implemented. System 2800 may be used to implement any of the computer systems described above. As shown, computer system 2800 includes a processing unit 2804 that communicates with a number of peripheral subsystems via a bus subsystem 2802. These peripheral subsystems may include a processing acceleration unit 2...
Claims
1. 1. A computer-implemented method comprising: providing a virtual Layer 3 network in a virtualized cloud environment, the virtual Layer 3 network being hosted by an underlying physical network; 10. A method comprising: providing a virtual Layer 2 network in the virtualized cloud environment, the virtual Layer 2 network being hosted by the underlying physical network.
2. The method of claim 1 , wherein the virtual Layer 2 network comprises a virtual local area network (VLAN).
3. The method of claim 2 , wherein the VLAN includes multiple endpoints.
4. the plurality of endpoints comprises a plurality of compute instances; the VLAN includes a plurality of L2 virtual network interface cards (VNICs) and a plurality of switches; The method of claim 3 , wherein each of the plurality of compute instances is communicatively connected to a pair comprising a unique L2 virtual network interface card (L2VNIC) and a unique switch.
5. The method of claim 4 , wherein the plurality of switches together form a distributed switch.
6. The method of claim 4 or 5, wherein each of the plurality of switches routes outbound traffic according to a mapping table received from the VNIC paired with the switch.
7. The method of claim 6 , wherein the mapping table identifies interface-to-MAC address mappings for the endpoints within the VLAN.
8. The method of claim 4 , further comprising instantiating the pair including the unique L2V NIC and the unique switch on a network virtualization device (NVD).
9. receiving, on the unique L2V NIC of one of the plurality of compute instances, a packet addressed to one of the plurality of compute instances from another endpoint in the VLAN; and learning a mapping of the other endpoint using the unique L2VNIC of the one of the plurality of compute instances.
10. The method of claim 9 , wherein the mapping of the other endpoint comprises an interface-to-MAC address mapping of the other endpoint.
11. decapsulating the received packet using the unique L2V NIC of the one of the plurality of compute instances; and forwarding the decapsulated packet to the one of the plurality of computing instances.
12. Learning with the one of the plurality of computation instances includes 12. The method of claim 11, further comprising IP address to MAC address mapping of the endpoint.
13. sending an IP packet from a first compute instance in the VLAN to a second compute instance in the VLAN, the second compute instance having a destination IP address; receiving the IP packet on a first VNIC associated with the first compute instance; encapsulating the IP packet on the first VNIC; forwarding the IP packet to the second compute instance via a first switch; The method of claim 4 , wherein the first switch and the first VNIC are communicatively connected to the first compute instance in a pairwise fashion.
14. receiving the IP packet on a second VNIC associated with the second compute instance; decapsulating the IP packet on the second VNIC; and forwarding the IP packet from the second VNIC to the second compute instance.
15. the virtual Layer 2 network includes a plurality of virtual local area networks (VLANs); The method of claim 1 , wherein each of the plurality of VLANs includes a plurality of endpoints.
16. the plurality of VLANs include a first VLAN and a second VLAN; the first VLAN includes a plurality of first endpoints; The method of claim 15 , wherein the second VLAN includes a plurality of second endpoints.
17. The method of claim 16 , wherein each of the plurality of VLANs has a unique identifier.
18. 18. The method of claim 16 or 17, wherein one of the plurality of first endpoints in the first VLAN communicates with one of the plurality of second endpoints in the second VLAN.
19. 1. A system comprising: With a physical network, the physical network includes at least one host machine and at least one network virtualization appliance; The physical network includes: In a virtualized cloud environment, a virtual Layer 3 network is provided that is hosted by an underlying physical network; A system configured to provide a virtual Layer 2 network in the virtualized cloud environment, the virtualized Layer 2 network being hosted by the underlying physical network.
20. 1. A non-transitory computer-readable storage medium storing a plurality of instructions executable by one or more processors, the plurality of instructions, when executed by the one or more processors, causing the one or more processors to: In a virtualized cloud environment, a virtual Layer 3 network is provided that is hosted by an underlying physical network; In the virtualized cloud environment, the underlying physical network A non-transitory computer-readable storage medium for providing a virtual Layer 2 network.