Virtual l2 network loop prevention
By enforcing specific rules and using lightweight STP on NICs, virtual L2 networks can prevent loops and maintain multipathing, addressing the complexity and inefficiencies of traditional STP.
Patent Information
- Application Number
- JP2025169417
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-03-04
- Filing Date
- 2025-10-07
- Publication Date
- 2026-01-27
AI Technical Summary
Virtual Layer 2 (L2) networks are susceptible to loops, which can lead to broadcast storms and lack of multipathing due to the inherent lack of loop prevention mechanisms in L2 switches, while traditional Spanning Tree Protocol (STP) is complex and prevents multipathing.
Implementing specific rules on network interface cards (NICs) and using lightweight single-port STP to prevent loops, enabling multipathing by managing ports connected to hosts without complex root bridge selection logic.
Prevents loops effectively while maintaining multipathing, avoiding broadcast storms and reducing complexity compared to traditional STP implementations.
Smart Images

Figure 2026012719000001_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 031,325, filed May 28, 2020, entitled "Loop Prevention for Layer 2 Virtual Networks Without Global Spanning Tree Protocol," and to U.S. Patent Application No. 17 / 192,681, filed March 4, 2021, entitled "Loop Prevention for Virtual L2 Networks," the entire contents of which are incorporated by reference into this disclosure for all purposes.
[0002] background Cloud infrastructures such as Oracle® Cloud Infrastructure (OCI) offer a suite of cloud services. Entities (e.g., businesses) subscribing to these cloud services can build and run a wide range of applications and services in a highly available, cloud-hosted environment. The subscribing entities are called customers of the cloud service provider. The cloud infrastructure can provide high-performance compute, storage, and network capabilities in a flexible overlay virtual network that runs on a physical underlay network and is securely accessible from a business's on-premises network. Cloud infrastructures such as OCI enable customers to manage their cloud-based workloads in the same way that they typically manage on-premises workloads. Thus, businesses can get all the benefits of the cloud with the same control, isolation, security, and predictable performance as their on-premises networks.
[0003] Virtual networking is the foundation of cloud infrastructure and applications because it enables access, connection, security, and modification of cloud resources. Virtual networking enables communication between multiple computers, virtual machines (VMs), virtual servers, or other devices in different physical locations. While physical networking connects computer systems through cables and other hardware, virtual networking uses software management to connect computers and servers in different physical locations over the Internet. Virtual networks use virtualized versions of traditional network elements, such as network switches, routers, and adapters, to enable more efficient routing and easier network configuration / reconfiguration. Summary of the Invention
[0004] overview
[0001] The present disclosure relates generally to virtual networking. More specifically, techniques for preventing loops while supporting multipathing in virtual Layer 2 (L2) networks are described. According to certain embodiments, loops associated with virtual L2 networks can be prevented by implementing specific rules on network interface cards (NICs) and / or by using lightweight single-port STP rather than using a global spanning tree protocol (STP). Various inventive embodiments are described herein, including methods, systems, programs, code, non-transitory computer-readable storage media storing instructions executable by one or more processors, and the like.
[0005] According to a specific embodiment, a method for preventing loops while supporting multipathing in a virtual L2 network includes: a network virtualization device (NVD) configured to: The NVD may include receiving an L2 frame including a source media access control (MAC) address and a destination MAC address via a port of the NVD, associating the source MAC address with a first port of the NVD, and transmitting the L2 frame via at least a second port of the NVD rather than the first port of the NVD based on a rule that prevents the NVD from transmitting using the port on which the L2 frame was received. The NVD may provide an instance of a virtual NIC (VNIC). The virtual L2 network may include an L2 virtual LAN (L2 VLAN). The first port may be connected to a host or to a switch network, such as a Clos switch network including multiple switches.
[0006] According to certain embodiments, a further method for preventing loops while supporting multiple paths in a virtual L2 network may include using a lightweight single-port STP to send an L2 frame executing a compute instance to a host via a first port of an NVD; receiving the L2 frame from the host via the first port of the NVD; the NVD determining that the L2 frame is looped back; and the NVD disabling the first port of the NVD to stop sending and receiving frames using the first port.
[0007] According to particular embodiments, a non-transitory computer-readable memory may store a plurality of instructions executable by one or more processors, the plurality of instructions including instructions that, when executed by the one or more processors, cause the one or more processors to perform any one of the methods described above.
[0008] According to particular embodiments, a system may include one or more processors and a memory coupled to the one or more processors, wherein the memory may store a plurality of instructions executable by the one or more processors, the plurality of instructions including instructions that, when executed by the one or more processors, cause the one or more processors to perform any one of the methods described above.
[0009] According to certain embodiments, the NVD may be configured to receive an L2 frame including a source MAC address and a destination MAC address via a first port of the NVD, associate the source MAC address with the first port of the NVD, and transmit the L2 frame via at least a second port of the NVD rather than the first port of the NVD based on a rule that prevents transmission using the port on which the L2 frame was received. In certain embodiments, the NVD may be further configured to transmit a bridge protocol data unit (BPDU) to a host of a compute instance via the first port of the NVD, receive a BPDU from the host via the first port of the NVD, determine that the BPDU is looped back, and disable the first port of the NVD to stop sending and receiving frames using the first port.
[0010] The terms and expressions employed are used for purposes of description and not limitation. In using these terms and expressions, there is no intention to exclude equivalents of the illustrated and described features or portions thereof. However, it is recognized that various modifications are possible within the scope of the claimed system and method. Thus, while the present system and method have been specifically disclosed by way of example and optional features, it should be understood that variations and modifications of the concepts disclosed herein may be adopted by those skilled in the art, and that such variations and modifications are considered to be within the scope of the system and method as defined by the appended claims.
[0011] This Summary is intended to identify key or essential features of the claimed subject matter. It is not intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to appropriate portions of the entire specification of this patent application, any and all drawings, and each claim.
[0012] These and other features and embodiments will become more apparent with reference to the following specification, claims and accompanying drawings.
[0013] Exemplary embodiments will now be described in detail with reference to the drawings. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a high-level diagram of a distributed environment illustrating a virtual or overlay cloud network hosted by a cloud service provider infrastructure, according to certain embodiments. [Figure 2] 1 is an architectural schematic diagram illustrating physical elements of a physical network within a CSPI, according to certain embodiments. [Figure 3] FIG. 1 illustrates an exemplary arrangement of CSPI in which a host machine is connected to multiple network virtualization devices (NVDs), according to certain embodiments. [Figure 4] FIG. 1 illustrates a connection between a host machine and an NVD that provides I / O virtualization to support multi-tenancy, according to certain embodiments. [Figure 5] 1 is a schematic block diagram illustrating a physical network provided by CSPI in accordance with certain embodiments. [Figure 6] FIG. 1 illustrates an example of a virtual cloud network (VCN) including a virtual L2 network, according to certain embodiments. [Figure 7] FIG. 1 illustrates an example of a VCN including a virtual L2 network and a virtual L3 network, according to certain embodiments. [Figure 8] FIG. 1 illustrates an example of an infrastructure supporting a virtual L2 network of a VCN, according to certain embodiments. [Figure 9] FIG. 1 illustrates an example of a loop between network virtualization devices (NVDs), according to certain embodiments. [Figure 10]A diagram illustrating an example of preventing loops between NVDs associated with a virtual L2 network, according to certain embodiments. [Figure 11] FIG. 1 illustrates an example of a loop between an NVD and a compute instance running on a host, according to certain embodiments. [Figure 12] FIG. 1 illustrates an example of preventing loops between an NVD and a compute instance running on a host, according to certain embodiments. [Figure 13] FIG. 10 illustrates an example of a flow for preventing a loop related to a virtual L2 network. [Figure 14] A diagram showing an example of a flow for preventing loops between NVDs associated with a virtual L2 network. [Figure 15] FIG. 1 illustrates an example of a flow for preventing loops between NVDs and compute instances associated with a virtual L2 network. [Figure 16] FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 17] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 18] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 19] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 20] FIG. 1 is a block diagram illustrating an exemplary computer system in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Detailed Description The present disclosure relates generally to virtual networking, and more particularly to techniques for preventing loops while supporting multipathing in virtual Layer 2 (L2) networks. According to certain embodiments, loops in virtual L2 networks can be prevented by enforcing certain rules regarding the use of ports for transmitting frames and / or the implementation of Lightweight Spanning Tree Protocol (STP) on each port. Various embodiments are described herein, including methods, systems, programs, code, non-transitory computer-readable storage media storing instructions executable by one or more processors, and the like.
[0016] When loops are introduced, the network may become susceptible to broadcast storms. Networks also need to include multiple paths (referred to herein as "multipaths") to provide redundant paths in the event of a link failure. L2 switches inherently do not allow multipaths and may suffer from loop problems because L2 control protocols may not inherently have loop prevention mechanisms. STP is generally effective for preventing loops in L2 networks. However, while STP prevents multipaths, it is generally complex to implement. For example, STP requires each switch or network virtualization device (NVD) to implement logic for root bridge selection or root bridge priority, and logic for root port (RP) selection or RP priority for various ports, and must select a single path upon loop detection.
[0017] According to some embodiments, a switch or NVD may implement specific rules to avoid loops at least between the switch and / or NVD while enabling multipathing. In some embodiments, lightweight STP may be implemented to avoid loops caused by software bugs, or more generally, by software code of compute instances on hosts connected via ports. In this case, the lightweight STP may only need to manage ports connected to hosts and may not need to implement root bridge selection or root bridge priority logic and RP selection or RP priority logic.
[0018] Example Virtual Networking Architecture The term cloud services generally refers to services that a cloud service provider (CSP) makes available to users or customers on demand (e.g., via a subscription model) using systems and infrastructure (cloud infrastructure). Typically, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premise servers and systems. Therefore, customers can use cloud services provided by CSPs without separately purchasing hardware and software resources for the services. Cloud services are designed to provide subscribing customers with easy and scalable access to application and computing resources without requiring the customer to invest in procuring the infrastructure used to deliver the services.
[0019] There are several cloud service providers that offer different types of cloud services. Cloud services come in a variety of different types or types, such as Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS). Includes models.
[0020] A customer can subscribe to one or more cloud services offered by a CSP. A customer can be any entity, such as an individual, an organization, or a business. When a customer subscribes or registers for a service offered by a CSP, a tenant or account is created for the customer. The customer can then access one or more subscribed cloud resources associated with the account through this account.
[0021] As mentioned above, IaaS (Infrastructure as a Service) is a service that provides services to a specific type of client. Cloud computing services. In the IaaS model, a CSP provides infrastructure (called cloud service provider infrastructure or CSPI) that customers can use to build their own customizable networks and deploy customer resources. Therefore, customer resources and networks are hosted in a distributed environment by the infrastructure provided by the CSP. This differs from traditional computing, where the customer's infrastructure hosts the customer's resources and networks.
[0022] CSPI may include interconnected high-performance computing resources, including various host machines, memory resources, and network resources, forming a physical network, also known as a substrate network or underlay network. CSPI resources may be distributed across one or more data centers, geographically dispersed across one or more geographic regions. Virtualization software can run on these physical resources to provide a virtualized distributed environment. Virtualization creates an overlay network (also called a software-based network, software-defined network, or virtual network) on the physical network. The CSPI physical network provides the basis for creating one or more overlay or virtual networks on top of the physical network. A virtual or overlay network can include one or more virtual cloud networks (VCNs). A virtual network is implemented using software virtualization technologies (e.g., hypervisors, functions performed by network virtualization devices (NVDs) (e.g., smart NICs), top-of-rack (TOR) switches, smart TORs that implement one or more functions performed by the NVDs, and other mechanisms) to create a network abstraction layer that can run on top of the physical network. Virtual networks can take various forms, such as peer-to-peer networks, IP networks, etc. Virtual networks are typically either Layer 3 IP networks or Layer 2 VLANs. Such virtual or overlay networks are often called virtual Layer 3 networks or overlay Layer 3 networks. Examples of protocols developed for virtual networks are IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN), and so on. -IETF RFC7348), virtual private networks (VPNs) (e.g., MPLS Layer 3 Virtual Private Networks (RFC4364)), VMware NSX, Generic Network Virtualization Encapsulation (GENEVE), etc.
[0023] In the case of IaaS, the infrastructure provided by the CSP (CSPI) may be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing service provider may host infrastructure elements (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may offer various services (e.g., billing, monitoring, logging, security, load balancing, and clustering) that accompany those infrastructure elements. Because these services are policy-driven, IaaS users can easily manage the load balancing and clustering. By implementing policies to drive the distribution of resources, CSPs can maintain application availability and performance. CSPIs provide infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available, hosted, distributed environment. CSPIs provide high-performance compute resources and power, as well as storage capacity, on flexible virtual networks that can be securely accessed from various network locations, including customer on-premises networks. When a customer subscribes to or registers for an IaaS service offered by a CSP, the tenancy created for that customer is a secure, isolated partition from CSPI, allowing the customer to create, organize, and manage their cloud resources.
[0024] Customers can build their own virtual networks using the compute, memory, and networking resources provided by CSPI. On these virtual networks, they can connect one or more customer resources or workloads, such as compute instances. For example, a customer can use resources provided by CSPI to build one or more customizable private virtual networks called virtual cloud networks (VCNs). The customer can then deploy one or more customer resources, such as compute instances, on the customer VCN. The compute instances may be virtual machines, bare metal instances, etc. Thus, CSPI provides infrastructure and a set of complementary cloud services that enable customers to build and run various applications and services in a highly available, virtualized host environment. While customers do not manage or control the underlying physical resources provided by CSPI, they do control the operating systems, storage, and deployed applications, and in some cases have limited control over some networking components (e.g., firewalls).
[0025] The CSP may provide a console that enables customers and network administrators to configure, access, and manage resources deployed in the cloud using CSPI resources. In certain embodiments, the console provides a web-based user interface that can be used to utilize and manage CSPI. In some embodiments, the console is a web-based application provided by the CSP.
[0026] CSPI can support single-tenancy or multi-tenancy architectures. In a single-tenancy architecture, software (e.g., applications, databases) or hardware elements (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenancy architecture, software or hardware elements serve multiple customers or tenants. Thus, in a multi-tenancy architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenancy environment, CSPI employs precautions and safeguards to ensure that each tenant's data is isolated and not visible to other tenants.
[0027] In a physical network, a network endpoint (endpoint) refers to a computing device or system that is connected to the physical network and communicates bidirectionally with the connected network. A network endpoint of a physical network may be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints of a physical network include modems, hubs, bridges, switches, routers, and other networking devices, physical computers (or host machines), etc. Each physical device of a physical network has a fixed network address that can be used to communicate with that device. This The fixed network address may be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints may include various virtual endpoints, such as virtual machines hosted by elements of a physical network (e.g., hosted by a physical host machine). These endpoints in the virtual network are addressed by overlay addresses, such as overlay Layer 2 addresses (e.g., an overlay MAC address) and overlay Layer 3 addresses (e.g., an overlay IP address). Network overlays achieve flexibility by allowing network administrators to move overlay addresses associated with network endpoints using software management (e.g., via software implementing the virtual network's control plane). Thus, unlike physical networks, in virtual networks, overlay addresses (e.g., overlay IP addresses) can be moved from one endpoint to another using network management software. Because virtual networks are built on physical networks, both the virtual network and the underlying physical network are involved in communications between elements of the virtual network. To facilitate such communications, each element of the CSPI is configured to learn and store mappings that map overlay addresses in the virtual network to real physical addresses in the substrate network, or real physical addresses in the substrate network to overlay addresses in the virtual network. These mappings are used to facilitate communications. To facilitate virtual network routing, customer traffic is encapsulated.
[0028] Thus, physical addresses (e.g., physical IP addresses) are associated with elements of the physical network, and overlay addresses (e.g., overlay IP addresses) are associated with entities of the virtual network. Both physical and overlay IP addresses are real IP addresses. They are distinct from virtual IP addresses, which are mapped to multiple real IP addresses. Virtual IP addresses provide a one-to-many mapping between virtual IP addresses and multiple real IP addresses.
[0029] A cloud infrastructure or CSPI is physically hosted in one or more data centers in one or more regions around the world. A CSPI may include elements of a physical or substrate network and virtualized elements (e.g., virtual networks, compute instances, virtual machines) of a virtual network built on the physical network elements. In certain embodiments, a CSPI may be organized into realms, regions, and available services. CSPI resources are organized and hosted in distinct domains. A region is a local geographic area that typically contains one or more data centers. Regions are generally independent of one another and may be separated by vast distances, e.g., across countries or continents. For example, a first region may be in Australia, another region may be in Japan, and yet another region may be in India. CSPI resources are divided among these regions so that each region has an independent subset of CSPI resources. Each region may provide a set of core infrastructure services and resources, such as compute resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure), storage resources (e.g., block volume storage, file storage, object storage, archive storage), networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connectivity to on-premises networks), database resources, edge networking resources (e.g., DNS), access management and monitoring resources, etc. Each region typically has multiple routes connecting it to other regions within the realm.
[0030] In general, applications should use resources that are close by rather than far away. Applications are deployed in the region where they will be used most (i.e., on infrastructure relevant to that region) because it is faster than using the underlying infrastructure. Applications may also be deployed in different regions for a variety of reasons, such as redundancy to mitigate the risk of region-wide events such as large weather systems or earthquakes, or to meet different requirements for legal jurisdictions, tax domains, and other business or societal criteria.
[0031] Data centers within a region may be further organized and subdivided into availability domains (ADs). An availability domain is one or more data centers located in a region. A region may correspond to a data center on the cloud. A region may consist of one or more availability domains. In such a distributed environment, CSPI resources may be region-specific, such as a virtual cloud network (VCN), or availability domain-specific, such as a compute instance.
[0032] ADs within a region are isolated from each other to be fault-tolerant, with the likelihood of simultaneous failures extremely low. This is achieved by configuring ADs so that they do not share critical infrastructure resources, such as networking, physical cables, cable routes, and cable entrances, so that a failure of one AD in a region rarely impacts the availability of other ADs in the same region. Connecting ADs within the same region to each other via low-latency, high-bandwidth networks provides highly available connections to other networks (e.g., the Internet, customer on-premises networks), allowing multiple ADs to be replicated for both high availability and disaster recovery. Crowdsense utilizes multiple ADs to ensure high availability and protect against resource failures. As the infrastructure provided by an IaaS provider grows, more regions and ADs may be added along with additional capacity. Traffic between available domains is typically encrypted.
[0033] In certain embodiments, regions are grouped into realms. A realm is a logical collection of regions. Realms are isolated from each other and do not share any data. Regions within the same realm can communicate with each other, but regions within different realms cannot. A CSP's customer tenancy or account exists in a single realm and can span one or more regions within that single realm. Typically, when a customer subscribes to an IaaS service, their tenancy or account is created in a customer-specified region (called their "home" region) within a realm. The customer can extend their tenancy to one or more other regions within the realm. The customer cannot access regions that do not exist within the realm in which the customer's tenancy resides.
[0034] An IaaS provider may offer multiple realms, each corresponding to a particular set of customers or users. For example, a commercial realm may be offered for commercial customers. As another example, a realm may be offered for a particular country or for customers in that country. As yet another example, a government realm may be offered, for example, for a government. For example, a government realm may be created for a particular government and may have a higher security level than a commercial realm. For example, Oracle® Cloud Infrastructure (OCI) currently offers a realm for the commercial domain and two realms for the government cloud domain (e.g., FedRAMP-authorized and IL5-authorized).
[0035] In certain embodiments, an AD can be subdivided into one or more fault domains. A fault domain can be A failure domain is a grouping of infrastructure resources within a D. Fault domains allow for the distribution of compute instances so that they are not placed on the same physical hardware within one AD. This is known as anti-affinity. A failure domain refers to a collection of hardware elements (computers, switches, etc.) that share a single point of failure. A compute pool is logically divided into failure domains. Thus, a hardware failure or compute hardware maintenance event that affects one failure domain does not affect instances in other failure domains. In some embodiments, the number of failure domains in each AD may vary. For example, in a particular embodiment, each AD includes three failure domains. Failure domains function as logical data centers within an AD.
[0036] When a customer subscribes to an IaaS service, resources from CSPI are provisioned to the customer and associated with the customer's tenancy. Customers can use these provisioned resources to build private networks and deploy resources on these networks. A customer network hosted on the cloud by CSPI is called a virtual cloud network (VCN). Customers can configure one or more virtual cloud networks (VCNs) using the CSPI resources allocated for the customer. A VCN is a virtual or software-defined private network. Customer resources deployed in a customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances may represent various customer workloads, such as applications, load balancers, and databases. Compute instances deployed on a VCN can communicate with publicly accessible endpoints (public endpoints) over a public network, such as the Internet; with other instances in the same VCN or other VCNs (e.g., other VCNs of the customer or VCNs not belonging to the customer); with customer on-premises data centers or networks; with sendee endpoints; and with other types of endpoints.
[0037] CSPs can offer a variety of services using CSPI. In some cases, customers of a CSPI themselves can act as service providers and provide services using CSPI resources. Service providers can expose service endpoints characterized by identifying information (e.g., IP addresses, DNS names, and ports). Customer resources (e.g., compute instances) can consume a particular service by accessing the service endpoint for that particular service exposed by the service. These service endpoints are generally publicly accessible over a public communications network, such as the Internet, with users using the public IP address associated with the endpoint. Publicly accessible network endpoints are sometimes referred to as public endpoints.
[0038] In certain embodiments, a service provider may expose a service through a service endpoint (sometimes referred to as a service endpoint). Customers of the service may access the service using this service endpoint. In certain embodiments, a service endpoint provided for a service may be accessed by multiple customers wishing to consume the service. In other implementations, a dedicated service endpoint may be provided to a customer. Thus, only that customer may access the service using that dedicated sender endpoint.
[0039] In certain embodiments, when a VCN is created, it is associated with a private overlay Classless Inter-Domain Routing (CIDR) address space, which is a private overlay IP address range (e.g., 10.0 / 16) assigned to the VCN. A VCN also contains associated subnets, route tables, and VCNs contain gateways, rules, and gateways. A VCN exists within a single region but can extend to one or more or all available domains in the region. A gateway is a virtual interface configured for a VCN that enables traffic communication between the VCN and one or more endpoints outside the VCN. You can configure one or more different types of gateways for a VCN to enable communication between different types of endpoints.
[0040] A VCN may be subdivided into one or more sub-networks, such as one or more subnetworks. A subnet is thus a unit or division that can be created within a VCN. A VCN can have one or more subnets. Each subnet in a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets in the VCN and represents a subset of the VCN's address space.
[0041] Each compute instance is associated with a virtual network interface card (VNIC), which allows it to participate in a subnet of a VCN. A VNIC is a logical representation of a physical network interface card (NIC). Generally, a VNIC is an interface between an entity (e.g., a compute instance, a service) and a virtual network. A VNIC resides in a subnet and has one or more associated IP addresses and associated security rules or policies. A VNIC corresponds to a Layer 2 port on a switch. A VNIC connects a compute instance to a subnet within a VCN. A VNIC associated with a compute instance allows the compute instance to be part of a subnet of a VCN and enables the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, endpoints in a different subnet within the VCN, or endpoints outside the VCN. Thus, the VNIC associated with a compute instance determines how the compute instance connects with endpoints inside and outside the VCN. A VNIC for a compute instance is created and associated with the compute instance when the compute instance is created and added to a subnet within the VCN. If a subnet consists of a set of compute instances, it includes VNICs corresponding to the set of compute instances, each VNIC being connected to a compute instance in the set of compute instances.
[0042] Each compute instance is assigned a private overlay IP address via the VNIC associated with the compute instance. This private overlay IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic for the compute instance. All VNICs in a particular subnet use the same route table, security lists, and DHCP options. As described above, each subnet in a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that do not overlap with other subnets in the VCN and represent an address space subset of the VCN's address space. For a VNIC on a particular subnet of a VCN, the overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.
[0043] In certain embodiments, if desired, a compute instance can be assigned additional overlay IP addresses in addition to its private overlay IP address, e.g., one or more public IP addresses in the case of a public subnet. These multiple addresses can be associated with the same VNIC or multiple VNICs associated with the compute instance. However, each instance has a primary VNIC associated with an overlay private IP address that is created and assigned to the instance when the instance is launched. This primary VNIC cannot be deleted. Additional VNICs, called secondary VNICs, can be added to an existing instance within the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. Secondary VNICs can be in the same subnetwork in the same VCN as the primary VNIC, or in different subnetworks in the same or a different VCN.
[0044] Compute instances can optionally be assigned public IP addresses if they are in a public subnet. When creating a subnet, you can specify that the subnet is either a public or private subnet. A private subnet means that resources (e.g., compute instances) and associated VNICs within the subnet cannot have public overlay IP addresses. A public subnet means that resources and associated VNICs within the subnet can have public IP addresses. Customers can specify subnets that exist across a single available domain or multiple available domains within a region or realm.
[0045] As described above, a VCN may be subdivided into one or more subnets. In certain embodiments, a virtual router (referred to as a VCN VR or simply VR) configured for a VCN enables communication between subnets of the VCN. For a subnet within a VCN, the VR represents the logical gateway for that subnet, enabling communication between the subnet (i.e., the compute instances on that subnet) and endpoints on other subnets within the VCN and other endpoints outside the VCN. A VCN VR is a logical entity configured to route traffic between VNICs in a VCN and a virtual gateway (gateway) associated with the VCN. Gateways are further described below with respect to FIG. 1. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is one VCN VR per VCN. This VCN VR potentially has an unlimited number of ports addressed by IP addresses, one port for each subnet of the VCN. In this way, a VCN VR has a different IP address for each subnet of the VCN to which the VCN VR is connected. VRs are also connected to various gateways configured for the VCN. In certain embodiments, a particular overlay IP address from a subnet's overlay IP address range is reserved in a port of the subnet's VCN VR. For example, consider a VCN with two subnets, each with associated address ranges 10.0 / 16 and 10.1 / 16. For a first subnet in the VCN with address range 10.0 / 16, addresses from this range are reserved in a port of the subnet's VCN VR. In some cases, a first IP address from this range may be reserved in a VCN VR. For example, for a subnet with overlay IP address range 10.0 / 16, IP address 10.0.0.1 may be reserved in a port of the subnet's VCN VR. For a second subnet in the same VCN with address range 10.1 / 16, the VCN VR may have a port of the second subnet with IP address 10.1.0.1.A VCN VR has a different IP address for each subnet in the VCN.
[0046] In some other embodiments, each subnet within a VCN may have its own associated VR that is addressable by the subnet using a reserved or default IP address associated with the VR. The reserved or default IP address may be, for example, the first IP address from a range of IP addresses associated with that subnet. VNICs in a subnet may use this default or reserved IP address to communicate (e.g., send and receive packets) with the VR associated with the subnet. In such an embodiment, the VR is the ingress / egress point for that subnet. VRs associated with subnets in a VCN may communicate with other VRs associated with other subnets in the VCN. VRs may also communicate with gateways associated with the VCN. The VR functions for a subnet are performed on or by one or more NVDs that perform the VNIC functions for VNICs in the subnet.
[0047] Route tables, security rules, and DHCP options may be configured for a VCN. A route table is a virtual route table for a VCN and contains rules for routing traffic from subnets inside the VCN to destinations outside the VCN through gateways or specially configured instances. You can customize a VCN's route table to control the forwarding / routing of packets into and out of the VCN. DHCP options refer to configuration information that is automatically provided to an instance when it is launched.
[0048] Security rules configured for a VCN represent the VCN's overlay firewall rules. Security rules can include inbound and outbound rules and can specify the type of traffic allowed in and out of instances in the VCN (e.g., based on protocol and port). Customers can choose whether certain rules are stateful or stateless. For example, a customer can allow incoming SSH traffic from anywhere to a set of instances by configuring a stateful inbound rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22. Security rules may be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to resources within that group. On the other hand, a security list contains rules that apply to all resources in subnets that use that security list. A VCN may also include default security rules and a default security list. DHCP options configured for a VCN provide configuration information that is automatically provided when instances in the VCN are launched.
[0049] In particular embodiments, configuration information for a VCN is determined and stored by a VCN control plane. The configuration information for a VCN may include, for example, address ranges associated with the VCN, subnets and associated information within the VCN, one or more VRs associated with the VCN, compute instances and associated VNICs within the VCN, NVDs that perform various virtualized network functions (e.g., VNICs, VRs, gateways) associated with the VCN, VCN state information, and other VCN-related information. In particular embodiments, a VCN distribution service publishes the configuration information stored by the VCN control plane, or portions thereof, to the NVD. The distributed information can be used to forward packets to and from compute instances within the VCN by updating information (e.g., forwarding tables, routing tables, etc.) stored and used by the NVD.
[0050] In certain embodiments, VCN and subnet creation is handled by a VCN control plane (CP), and compute instance launch is handled by a compute control plane. The compute control plane is configured to allocate physical resources for the compute instance and then invoke the VCN control plane to create and attach VNICs to the compute instance. The VCN CP also sends VCN data mappings to a VCN data plane, which is configured to perform packet forwarding and routing functions. In, the VCN CP provides a distribution service configured to provide updates to the VCN data plane. Examples of the VCN control plane are shown in Figures 16, 17, 18, and 19 (see reference numbers 1616, 1716, 1816, and 1916) and described below.
[0051] Customers can create one or more VCNs with resources hosted by CSPI. Compute instances deployed on a customer VCN can communicate with different endpoints. These endpoints can include endpoints hosted by CSPI and endpoints external to CSPL.
[0052] Various different architectures for implementing cloud-based services using CSPI are shown in Figures 1, 2, 3, 4, 5, 16, 17, 18, and 20 and described below. Figure 1 is a high-level diagram of a distributed environment 100 illustrating an overlay VCN or customer VCN hosted by CSPI, according to certain embodiments. The distributed environment shown in Figure 1 includes multiple elements in an overlay network. The distributed environment 100 shown in Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, alternatives, and modifications are possible. For example, in some implementations, the distributed environment shown in Figure 1 may have more or fewer systems or elements than those shown in Figure 1, may combine two or more systems, or may have a different system configuration or arrangement.
[0053] As shown in the example of FIG. 1 , distributed environment 100 includes CSPI 101, which provides services and resources that customers can subscribe to and use to build a virtual cloud network (VCN). In a particular embodiment, CSPI 101 provides IaaS services to subscribing customers. Data centers within CSPI 101 may be organized into one or more regions. FIG. 1 shows an example region, "US Region" 102. A customer configures a customer VCN 104 for region 102. A customer can deploy various compute instances on VCN 104, which may include virtual machines or bare metal instances. Example instances include applications, databases, load balancers, etc.
[0054] In the embodiment shown in Figure 1, customer VCN 104 includes two subnets, "Subnet-1" and "Subnet-2," each with its own CIDR IP address range. In Figure 1, the overlay IP address range of Subnet-1 is 10.0 / 16, and the address range of Subnet-2 is 10.1 / 16. VCN virtual router 105 represents the logical gateway of the VCN, enabling communication between subnets in VCN 104 and with other endpoints outside the VCN. VCN VR105 is configured to route traffic between VNICs in VCN104 and gateways associated with VCN104. VCN VR105 provides a port for each subnet of VCN104. For example, VR105 may provide a port with IP address 10.0.0.1 for subnet-1 and a port with IP address 10.1.0.1 for subnet-2.
[0055] Multiple compute instances can be deployed on each subnet. In this case, the compute instances may be virtual machine instances and / or bare metal instances. The compute instances within a subnet may be hosted by one or more host machines within CSPI 101. A compute instance joins a subnet via a VNIC associated with the compute instance. For example, as shown in Figure 1, compute instance C1 is part of subnet-1 via a VNIC associated with the compute instance. Similarly, compute instance C2 is part of subnet-1 via a VNIC associated with C2. Subnet-1 is part of the VCN VR105. Similarly, multiple compute instances, which may be virtual machine instances or bare metal instances, may be part of Subnet-1. Each compute instance is assigned a private overlay IP address and MAC address via an associated VNIC. For example, in FIG. 1, compute instance C1 has overlay IP address 10.0.0.2 and MAC address M1, and compute instance C2 has private overlay IP address 10.0.0.3 and MAC address M2. Each compute instance in Subnet-1, including compute instances C1 and C2, has a default route to VCN VR105 using IP address 10.0.0.1, which is the IP address of a port in VCN VR105 in Subnet-1.
[0056] Subnet-2 may have multiple compute instances deployed, including virtual machine instances and / or bare metal instances. For example, as shown in FIG. 1, compute instances D1 and D2 are part of subnet-2 via VNICs associated with the respective compute instances. In the embodiment shown in FIG. 1, compute instance D1 has an overlay IP address of 10.1.0.2 and a MAC address of MM1, and compute instance D2 has a private overlay IP address of 10.1.0.3 and a MAC address of MM2. Each compute instance in subnet-2, including compute instances D1 and D2, has a default route to VCN VR105 using IP address 10.1.0.1, which is the IP address of a port in VCN VR105 in subnet-2.
[0057] VCN A 104 may also include one or more load balancers. For example, a load balancer may be provided for a subnet and configured to load balance traffic among multiple compute instances on the subnet. A load balancer may also be provided to load balance traffic among subnets within the VCN.
[0058] A particular compute instance deployed on VCN 104 can communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 101 may include endpoints on the same subnet as the particular compute instance (e.g., communication between two compute instances in Subnet-1), endpoints in a different subnet but within the same VCN (e.g., communication between a compute instance in Subnet-1 and a compute instance in Subnet-2), endpoints in a different VCN in the same region (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN in the same region 106 or 110, or communication between a compute instance in Subnet-1 and an endpoint in the service network 110 in the same region), or endpoints in a VCN in a different region (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN in a different region 108). Additionally, compute instances in a subnet hosted by CSPI 101 can communicate with endpoints not hosted by CSPI 101 (i.e., external to CSPI 101). These external endpoints include endpoints within customer on-premise networks 116, endpoints within other remote cloud host networks 118, public endpoints 114 accessible via public networks such as the Internet, and other endpoints.
[0059] Communication between compute instances on the same subnet is facilitated using the VNICs associated with the source and destination compute instances. For example, compute instance C1 in subnet-1 sends a packet to compute instance C2 in subnet-1. For a packet sent from a source compute instance whose destination is another compute instance in the same subnet, the packet is first processed by a VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance may include determining the packet's destination information from the packet header, identifying any policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the packet's next hop, performing any packet encapsulation / decapsulation functions as needed, and forwarding / routing the packet to the next hop to facilitate communication to the packet's intended destination. If the destination compute instance is in the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing. The VNIC associated with the destination compute instance then executes and forwards the packet to the destination compute instance.
[0060] When communicating a packet from a compute instance in a subnet to an endpoint in a different subnet of the same VCN, the communication is facilitated by the VNICs and VCN VRs associated with the source and destination compute instances. For example, if compute instance C1 in Subnet-1 in Figure 1 wants to send a packet to compute instance D1 in Subnet-2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR105 using the VCN VR's default route or port 10.0.0.1. VCN VR105 is configured to route the packet to Subnet-2 using port 10.1.0.1. The packet is then received and processed by the VNIC associated with D1, which forwards the packet to compute instance D1.
[0061] To communicate packets from a compute instance within VCN 104 to an endpoint outside VCN 104, the communication is facilitated by a VNIC associated with the source compute instance, VCN VR 105, and a gateway associated with VCN 104. One or more types of gateways can be associated with VCN 104. A gateway is an interface between a VCN and another endpoint, where the other endpoint is outside the VCN. A gateway is a Layer 3 / IP layer concept that enables a VCN to communicate with endpoints outside the VCN. Thus, a gateway facilitates traffic flow between a VCN and other VCNs or networks. A variety of different types of gateways can be configured in a VCN to facilitate different types of communications with different types of endpoints. Through gateways, communications may occur over a public network (e.g., the Internet) or a private network. These communications may use various communication protocols.
[0062] For example, compute instance C1 may wish to communicate with an endpoint outside of VCN 104. The packet may first be processed by the VNIC associated with the source compute instance C1. The VNIC processing determines that the packet's destination is outside of Cl's subnet-1. The VNIC associated with C1 may forward the packet to VCN VR105 of VCN 104. VCN VR105 then processes the packet and, as part of the processing, determines a particular gateway associated with VCN 104 as the packet's next hop based on the packet's destination. VCN VR105 may then forward the packet to the particular gateway. For example, if the destination is an endpoint within a customer's operating premises network, the packet may be forwarded by VCN VR105 to dynamic routing gateway (DRG) 122 configured for VCN 104. The packet is then forwarded from the gateway to the next hop and forwarded to the intended destination. This can facilitate communication of the packet to its final destination.
[0063] Various different types of gateways may be configured for a VCN. An example of a gateway that may be configured for a VCN is shown in FIG. 1 and described below. Examples of gateways associated with a VCN are also shown in FIGS. 16, 17, 18, and 19 (e.g., gateways indicated by reference numbers 1634, 1636, 1638, 1734, 1736, 1738, 1834, 1836, 1838, 1934, 1936, and 1938) and described below. As shown in the embodiment shown in FIG. 1, a dynamic routing gateway (DRG) 122 may be added to or associated with the customer VCN 104. The DRG 122 provides a path for private network traffic communication between the customer VCN 104 and another endpoint. The other endpoint may be a customer on-premises network 116, a VCN 108 in a different region of the CSPI 101, or another remote cloud network 118 not hosted by the CSPI 101. The customer on-premises network 116 may be a customer network or customer data center built using customer resources. Access to the customer on-premises network 116 is typically highly restricted. For a customer that has both the customer on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, the customer may want the on-premises network 116 and the cloud-based VCNs 104 to be able to communicate with each other. This allows the customer to build an extended hybrid environment that includes the on-premises network 116 and the customer's VCNs 104 hosted by CSPI 101. The DRG 122 enables such communication. To enable such communication, a communication channel 124 is established. In this case, one endpoint of the communication channel is located in the customer on-premises network 116, and the other endpoint is located in CSPI 101 and connected to the customer VCN 104. The communication channel 124 can traverse a public communication network, such as the Internet, or a private communication network.A variety of different communication protocols can be used, such as IPsec VPN technology over a public communication network such as the Internet, or Oracle's FastConnect technology, which uses a private network instead of a public network. The device or equipment in the customer on-premises network 116 that forms one endpoint of the communication channel 124 is called customer premises equipment (CPE), such as the CPE 126 shown in Figure 1. The endpoint on the CSPI 101 side may be a host machine running the DRG 122.
[0064] In certain embodiments, remote peering connections (RPCs) can be added to a DRG, allowing customers to peer one VCN with another VCN in another region. Using such RPCs, a customer VCN 104 can connect to a VCN 108 in another region using a DRG 122. The DRG 122 can also connect to other remote cloud networks 118 not hosted by CSPI 101, such as Microsoft® Azure Cloud, Amazon® AWS Cloud, or other cloud services. It may also be used to communicate with the cloud.
[0065] As shown in FIG. 1, an Internet Gateway (IGW) 120 can be configured in a customer VCN 104 to allow compute instances on the customer VCN 104 to communicate with public endpoints 114 accessible over a public network, such as the Internet. The IGW 120 is a gateway for connecting a VCN to a public network, such as the Internet. The IGW 120 allows public subnets in a VCN, such as VCN 104 (resources in the public subnet have public overlay IP addresses) to directly access public endpoints 112 on the public network 114, such as the Internet. GW 120 can be used to initiate connections from subnets within VCN 104 or from the Internet.
[0066] Customer VCN 104 can be configured with a network address translation (NAT) gateway 128. NAT gateway 128 allows cloud resources in the customer VCN that do not have dedicated public overlay IP addresses to access the Internet without exposing them to direct incoming Internet connections (e.g., L4-L7 connections). This allows private subnets in a VCN, such as private subnet-1 in VCN 104, to privately access public endpoints on the Internet. With a NAT gateway, connections can be initiated from the private subnet to the public Internet, but connections cannot be initiated from the Internet to the private subnet.
[0067] In certain embodiments, a service gateway (SGW) 126 can be configured in customer VCN 104. SGW 126 provides a pathway for private network traffic between VCN 104 and service endpoints supported by service network 110. In certain embodiments, service network 110 may be provided by a CSP and may offer a variety of services. An example of such a service network is the Oracle® Service Network, which offers a variety of services available to customers. For example, a compute instance (e.g., a database system) in a private subnet of customer VCN 104 can back up data to a service endpoint (e.g., an object store) without requiring a public IP address or access to the Internet. In some embodiments, a VCN can have only one SGW, and connections can be initiated only from subnets within the VCN, not from service network 110. When a VCN is peered with another VCN, resources in the other VCN typically cannot access the SGW. Resources in an on-premises network connected to a VCN with FastConnect or VPN Connect can also use a service gateway configured in that VCN.
[0068] In some implementations, the SGW 126 uses service classless inter-domain routing (CIDR) labels. A CIDR label is a string that represents all regional public IP address ranges for a service or group of services of interest. Customers use service CIDR labels to control traffic to services when configuring the SGW and associated routing rules. Customers can optionally use service CIDR labels when configuring security rules without having to adjust the security rules if the service's public IP addresses change in the future.
[0069] A local peering gateway (LPG) 132 is a gateway that can be added to a customer VCN 104 to enable the VCN 104 to peer with another VCN in the same region. Peering means that the VCNs communicate using private IP addresses without the traffic going over a public network such as the Internet or routing the traffic through the customer on-premises network 116. In a preferred embodiment, a VCN has a separate LPG for each peering it establishes. Local peering or VCN peering is a common practice used to establish network connectivity between different applications or infrastructure management functions.
[0070] Service providers, such as providers of services in service network 110, may provide access to their services using different access models. According to a specific access model, a service may be exposed as a public endpoint publicly accessible by a compute instance in the customer VCN over a public network such as the Internet, or may be accessed privately through the SGW 126. According to a specific private access model, a service may be accessed as a private IP endpoint in a private subnet in the customer VCN. This is called private endpoint (PE) access and allows a service provider to expose its services as instances in the customer's private network. A private endpoint resource represents a service in a customer VCN. Each PE appears as a VNIC (called a PE-VNIC, with one or more private IPs) that the customer selects from a subnet in the customer VCN. Thus, the PE provides a way to provide services within the customer's private VCN subnet using VNICs. Because the endpoints are exposed as VNICs, the PE VNIC can utilize all the functionality associated with a VNIC, such as routing rules and security lists.
[0071] Service providers register services to make them accessible through PEs. Providers can associate policies with services that regulate the visibility of the service to customer tenants. Providers can register multiple services under a single virtual IP address (VIP), especially for multi-tenant services. There can also be multiple private endpoints (in multiple VCNs) that represent the same service.
[0072] Compute instances in the private subnet can then access the service using the private IP address or service DNS name of the PE VNIC. Compute instances in the customer VCN can access the service by sending traffic to the private IP address of the PE in the customer VCN. The Private Access Gateway (PAGW) 130 is a gateway resource that can connect to a service provider VCN (e.g., a VCN in the service network 110) and serves as the ingress / egress point for all traffic from / to the customer subnet private endpoints. The PAGW 130 allows providers to scale the number of PE connections without utilizing internal IP address resources. A provider only needs to configure one PAGW for any number of services registered in a single VCN. A provider can present services as private endpoints in multiple VCNs for one or more customers. From the customer's perspective, the PE VNIC appears not to be connected to the customer's instance but to the service the customer wants to interact with. Traffic destined for the private endpoint is routed to the service through the PAGW 130. These are called customer-to-service private connections (C2S connections).
[0073] Also, the PE concept is used to ensure that traffic is routed between the FastConnect / IPsec link and the customer VCN. It can also extend private access of services to customer on-premises networks and data centers by allowing traffic to flow through private endpoints within LPG132 and the customer VCN, and it can extend private access of services to customer peering VCNs by allowing traffic to flow between PEs in LPG132 and the customer VCN.
[0074] Customers can control VCN routing at the subnet level, allowing them to specify which subnets in a customer VCN, such as VCN 104, use each gateway. The VCN's route table can be used to determine whether traffic can be routed outside the VCN through a particular gateway. For example, In the example, a route table for a public subnet in customer VCN 104 can send non-local traffic through IGW 120. A route table for a private subnet in the same customer VCN 104 can send traffic to a CSP service through SGW 126. All remaining traffic may be sent through NAT gateway 128. Route tables only control traffic that leaves the VCN.
[0075] Security lists associated with a VCN are used to control traffic entering the VCN through inbound connections and gateways. All resources within a subnet use the same mute tables and security lists. Security lists may be used to control specific types of traffic entering and leaving instances within a VCN's subnets. Security list rules may include inbound (inbound) rules and outbound (outbound) rules. For example, inbound rules may specify allowed source address ranges, and outbound rules may specify allowed destination address ranges. Security rules may specify specific protocols (e.g., TCP, ICMP), specific ports (e.g., port 22 for SSH, port 3389 for Windows RDP), etc. In certain implementations, the instance's operating system may enforce its own firewall rules that match security list rules. Rules may be stateful (e.g., connections are tracked and responses are automatically allowed without explicit security list rules for the response traffic) or stateless.
[0076] Access from a customer VCN (i.e., resources or compute instances deployed on VCN 104) can be categorized as public access, private access, or dedicated access. Public access refers to an access model for accessing public endpoints using public IP addresses or NAT. Private access enables customer workloads in VCN 104 with private IP addresses (e.g., resources in a private subnet) to access a service without traversing a public network such as the Internet. In particular embodiments, CSPI 101 enables customer VCN workloads with private IP addresses to access the service's public service endpoint using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer VCN and the service's public endpoint, which resides outside the customer's private network.
[0077] Additionally, CSPI is working to develop dedicated public peering services using technologies such as FastConnect public peering. It can provide brick access, where a customer's on-premises instance can access the FastConnect connection without going through a public network such as the Internet. You can access one or more services in a customer VCN using FastConnect. CSPI also provides dedicated private access using FastConnect private peering. In this case, customer on-premises instances with private IP addresses can access workloads in the customer VCN using the FastConnect connection. FastConnect connects customers' on-premises networks using the public internet. FastConnect is a network connection used instead of connecting your network to CSPI and its services. FastConnect offers higher bandwidth options and It provides an easy, flexible and economical way to create dedicated, private connections with a reliable and consistent networking experience.
[0078] FIG. 1 and the accompanying description above illustrate various virtualized elements in an exemplary virtual network. As described above, virtual networks are built on an underlying physical or substrate network. FIG. 2 is a simplified architecture diagram illustrating physical elements within a physical network within CSPI 200 that provides the foundation for the virtual network, according to a particular embodiment. As shown, CSPI 200 provides a distributed environment including elements and resources (e.g., compute resources, memory resources, and networking resources) provided by a cloud service provider (CSP). These elements and resources are used to provide cloud services (e.g., IaaS services) to subscribing customers, i.e., customers who subscribe to one or more services offered by the CSP. Based on the services to which the customer subscribes, CSPI 200 provides some resources (e.g., compute resources, memory resources, and networking resources) to the customer. The customer can then build their own cloud-based (i.e., CSPI-hosted), customizable private virtual network using the physical compute resources, memory resources, and networking resources provided by CSPI 200. As previously described, these customer networks are referred to as virtual cloud networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, into these Customer VCNs. The compute instances may be virtual machines, bare metal instances, etc. CSPI200 provides infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available hosted environment.
[0079] In the exemplary embodiment shown in FIG. 2, the physical elements of CSPI 200 include one or more physical host machines or physical servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), a physical network (e.g., 218), and switches within physical network 218. The physical host machines or servers can host and execute various compute instances participating in one or more subnets of a VCN. The compute instances may include virtual machine instances and bare metal instances. For example, the various compute instances shown in FIG. 1 may be hosted by the physical host machines shown in FIG. 2. The virtual machine compute instances in a VCN may be executed by one host machine or by multiple different host machines. Additionally, the physical host machines may host virtual host machines, container-based hosts or functions, etc. The VICs and VCN VRs shown in FIG. 1 may be executed by the FTVDs shown in FIG. 2. The gateway shown in FIG. 1 may be implemented by the host machine and / or NVD shown in FIG.
[0080] A host machine or server may run a hypervisor (also called a virtual machine monitor or VMM) that creates and enables a virtualized environment on the host machine. Virtualization or a virtualized environment facilitates cloud-based computing. One or more computing instances may be created, executed, and managed on the host machine by the hypervisor on the host machine. The hypervisor on the host machine enables the host machine's physical computing resources (e.g., computing resources, memory resources, and networking resources) to be shared among various computing instances running on the host machine.
[0081] For example, as shown in Figure 2, host machines 202 and 208 execute hypervisors 260 and 266, respectively. These hypervisors may be implemented using software, firmware, hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that resides in the host machine's operating system (OS), which executes on the host machine's hardware processor. The hypervisor manages the physical computing resources (e.g., A hypervisor 260 provides a virtualization environment that allows processing resources (e.g., processors / cores, memory resources, networking resources, etc.) to be shared among various virtual machine computing instances executed by a host machine. For example, in FIG. 2, hypervisor 260 resides in the OS of host machine 202 and allows the computing resources (e.g., processing resources, memory resources, and networking resources) of host machine 202 to be shared among computing instances (e.g., virtual machines) executed by host machine 202. A virtual machine can have its own OS (called a guest OS). This guest OS may be the same as or different from the OS of the host machine. The OS of a virtual machine executed by a host machine may be the same as or different from the OS of other virtual machines executed by the same host machine. Thus, a hypervisor can run multiple OSs in parallel while sharing the same computing resources of the host machine. The host machines shown in FIG. 2 may have the same type of hypervisor or different types of hypervisors.
[0082] A compute instance may be a virtual machine instance or a bare metal instance. In Figure 2, compute instance 268 on host machine 202 and compute instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.
[0083] In certain examples, an entire host machine may be provided to a single customer, and one or more compute instances (either virtual machines or bare metal instances) hosted by that host machine may all belong to the same customer. In other examples, a host machine may be shared among multiple customers (i.e., multiple tenants). In such a multi-tenant scenario, a host machine may host virtual machine compute instances belonging to different customers. These compute instances may be members of different VCNs for different customers. In certain embodiments, bare metal compute instances are hosted by bare metal servers without a hypervisor. When bare metal compute instances are provided, a single customer or tenant maintains control of the physical CPU, memory, and network interfaces of the host machine hosting the bare metal instance, and the host machine is not shared with other customers or tenants.
[0084] As previously described, each compute instance that is part of a VCN is associated with a VNIC that enables the compute instance to be a member of a subnet of the VCN. The VNIC associated with a compute instance facilitates communication of packets or frames to and from the compute instance. A VNIC is associated with the compute instance when the compute instance is created. In particular embodiments, for a compute instance executed by a host machine, the VNIC associated with the compute instance is executed by an NVD connected to the host machine. For example, in FIG. 2, host machine 202 executes virtual machine compute instance 268 associated with VNIC 276, which is executed by NVD 210 connected to host machine 202. As another example, bare metal instance 272 hosted by host machine 206 is associated with VNIC 280, which is executed by NVD 212 connected to host machine 206. As yet another example, VNIC 284 is associated with compute instance 274 executed by host machine 208, which is executed by NVD 212 connected to host machine 208.
[0085] For a compute instance hosted by a host machine, the NVD connected to that host machine runs the VCN VR corresponding to the VCN of which the compute instance is a member. For example, in the embodiment shown in FIG. 2, the NVD 210 NVD 212 runs VCN VR 277 corresponding to the VCN of which instance 268 is a member. NVD 212 may also run one or more VCN VRs 283 corresponding to the VCNs corresponding to the compute instances hosted by host machines 206 and 208.
[0086] A host machine may include one or more network interface cards (NICs) for connecting the host machine to other devices. The NICs on a host machine may provide one or more ports (or interfaces) for communicatively connecting the host machine to another device. For example, one or more ports (or interfaces) on the host machine and the NVD may be used to connect the host machine to the NVD. The host machine may also be connected to other devices, such as other host machines.
[0087] 2, host machine 202 is connected to NVD 210 using link 220 extending between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD 210. Host machine 206 is connected to NVD 212 using link 224 extending between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD 212. Host machine 208 is connected to NVD 212 using link 226 extending between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD 212.
[0088] Similarly, the NVDs are connected via communication links to top-of-rack (TOR) switches, which are connected to a physical network 218 (also called a switch fabric). In particular embodiments, the links between the host machines and the NVDs and the links between the NVDs and the TOR switches are Ethernet links. For example, in FIG. 2, NVDs 210 and 212 are connected to TOR switches 214 and 216, respectively, via links 228 and 230. In particular embodiments, links 220, 224, 226, 228, and 230 are Ethernet links. A collection of host machines and NVDs connected to a TOR may be referred to as a rack.
[0089] The physical network 218 provides a communications fabric that enables the TOR switches to communicate with each other. The physical network 218 may be a multi-tier network. In a particular implementation, the physical network 218 is a multi-tier Clos network of switches, with the TOR switches 214 and 216 representing leaf-level nodes of the multi-tier and multi-node physical switching network 218. Different Clos network configurations are possible, including, but not limited to, 2-tier networks, 3-tier networks, 4-tier networks, 5-tier networks, and generally "n"-tier networks. An example of a Clos network is shown in FIG. 5 and described below.
[0090] A variety of different connection configurations are possible between a host machine and the N virtual disks, including one-to-one, many-to-one, and one-to-many configurations. In a one-to-one implementation, each host machine is connected to its own separate virtual disk. For example, in FIG. 2, host machine 202 is connected to virtual disk 210 via NIC 232 of host machine 202. In a many-to-one configuration, multiple host machines are connected to a single virtual disk. For example, in FIG. 2, host machines 206 and 208 are connected to the same virtual disk 212 via NICs 244 and 250, respectively.
[0091] In a one-to-many configuration, one host machine is connected to multiple NVDs. Figure 3 shows an example of a CSPI 300 in which a host machine is connected to multiple NVDs. As shown, host machine 302 includes a network interface card (NIC) 304 including multiple ports 306 and 30S. Host machine 300 is connected to a first NVD 310 via port 306 and link 320, and to a second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between host machine 302 and NVDs 310 and 312 may be Ethernet links. NVD 310 is connected to a first TOR switch 314, and NVD 312 is connected to a second TOR switch 316. The links between NVDs 310 and 312 and TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent layer-0 switching devices within a multi-tier physical network 318.
[0092] 3 provides two separate physical network paths from the physical switch network 318 to the host machine 302: a first path from the TOR switch 314 to the host machine 302 via the NVD 310, and a second path from the TOR switch 316 to the host machine 302 via the NVD 312. The separate paths provide enhanced availability (referred to as high availability) for the host machine 302. If there is a problem with one of the paths (e.g., a link on one of the paths fails) or if there is a problem with a device (e.g., a particular NVD is not functioning), the other path can be used for communications to and from the host machine 302.
[0093] In the configuration shown in Figure 3, the host machine is connected to two different NVDs using two different ports provided by the host machine's NIC. In other embodiments, the host machine may include multiple NICs allowing the host machine to connect to multiple NVDs.
[0094] Referring again to Figure 2, an NVD is a physical device or element that performs one or more network virtualization and / or storage virtualization functions. An NVD may be any device that has one or more processing units (e.g., a CPU, a network processing unit (NPU), an FPGA, a packet processing pipeline), memory including cache, and ports. Various virtualization functions may be performed by software / firmware executed by one or more processing units of the NVD.
[0095] The NVD may be implemented in a variety of different ways. For example, in a particular embodiment, the NVD is implemented as an interface card with an embedded processor, called a smart NIC or intelligent NIC. The smart NIC is a separate device from the NIC on the host machine. In Figure 2, the NVD 210 may be implemented as a smart NIC connected to the host machine 202, and the NVD 212 may be implemented as a smart NIC connected to the host machines 206 and 208.
[0096] However, a smart NIC is only one example of an NVD implementation. Various other implementations are possible. For example, in some other implementations, the NVD or one or more functions performed by the NVD may be incorporated into or performed by one or more host machines, one or more TOR switches, and other elements of CSPI200. For example, the NVD may be integrated into a host machine. In this case, the functions performed by the NVD are performed by the host machine. As another example, the NVD may be part of a TOR switch, or a TOR switch may be configured to perform the functions performed by the NVD, allowing the TOR switch to perform various complex packet transformations used in public clouds. A TOR that performs the functions of an NVD may be referred to as a smart TOR. A virtual machine (VM) rather than a bare metal (BM) instance may be used. In yet other implementations that provide thin (VM) instances to customers, the functionality provided by the NVD may be implemented inside the hypervisor of a host machine, hi some other implementations, some of the functionality of the NVD may be offloaded to a centralized service running on a set of host machines.
[0097] In certain embodiments, such as when implemented as a smart NIC, as shown in FIG. 2, an NVD may include multiple physical ports that allow the NVD to connect to one or more host machines and one or more TOR switches. Ports on an NVD can be categorized as host-facing ports (also called "south ports") or network-facing or TOR-facing ports (also called "north ports"). A host-facing port of an NVD is a port used to connect the NVD to a host machine. Examples of host-facing ports in FIG. 2 include port 236 of NVD 210 and ports 248 and 254 of NVD 212. A network-facing port of an NVD is a port used to connect the NVD to a TOR switch. Examples of network-facing ports in FIG. 2 include port 256 of NVD 210 and port 258 of NVD 212. As shown in FIG. 2, the NVD 210 is connected to the TOR switch 214 via link 228 extending from port 256 of NVD 210 to the TOR switch 214. Similarly, the NVD 212 is connected to the TOR switch 216 via a link 230 that extends from a port 258 of the NVD 212 to the TOR switch 216 .
[0098] The NVD can receive packets and frames (e.g., packets and frames generated by compute instances hosted by the host machine) from the host machine via its host-facing port, perform any necessary packet processing, and then forward the packets and frames to the TOR switch via the NVD's network-facing port. The NVD can receive packets and frames from the TOR switch via the NVD's network-facing port, perform any necessary packet processing, and then forward the packets and frames to the host machine via the NVD's host-facing port.
[0099] In certain embodiments, multiple ports and associated links may be provided between the NVD and the TOR switch. These ports and links can be aggregated to form a link aggregator group (LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and the TOR switch) to be treated as a single logical link. All physical links within a given LAG can operate at the same speed and in full-duplex mode. LAGs help increase the bandwidth and reliability of the connection between two endpoints. If one of the physical links in a LAG fails, traffic is dynamically and transparently reassigned to another physical link within the LAG. The aggregated physical link provides higher bandwidth than individual links. Multiple ports associated with a LAG are treated as a single logical port. Traffic can be load-balanced across the multiple physical links in the LAG. One or more LAGs can be configured between two endpoints. The two endpoints may be, for example, between the NVD and the TOR switch, or between a host machine and the NVD.
[0100] The NVD implements or performs network virtualization functions. These functions are performed by software / firmware executed by the NVD. Examples of network virtualization functions include, but are not limited to, packet encapsulation and decapsulation functions, functions for creating VCN networks, functions for implementing network policies such as VCN security list (firewall) functions, functions for facilitating the routing and forwarding of packets to and from compute instances within a VCN, etc. In certain embodiments, upon receiving a packet, the NVD processes the packet and determines how the packet will be routed. As part of this packet processing pipeline, the NVD is responsible for the execution of the VNICs associated with the cis in the VCN, the execution of the virtual routers (VRs) associated with the VCN, , encapsulating and decapsulating packets to facilitate forwarding or routing within the virtual network, enforcing specific gateways (e.g., local peering gateways), implementing security lists, network security groups, network address translation (NAT) functions (e.g., public IP to private IP translation per host), throttling functions, and other functions.
[0101] In some embodiments, the packet processing data path in the NVD may include multiple packet pipelines. Each packet pipeline consists of a series of packet transformation stages. In some implementations, upon receiving a packet, the packet is parsed and sorted into a single pipeline. The packet is then processed stage by stage in a linear fashion until it is discarded or sent out through an interface of the NVD. These stages provide packet processing building blocks of basic functions (e.g., validating headers, performing throttling, inserting new Layer 2 headers, performing L4 firewalling, VCN encapsulation / decapsulation), such that new pipelines can be constructed by assembling existing stages, and new functionality can be added by creating and inserting new stages into existing pipelines.
[0102] The NVD can perform both control plane and data plane functions corresponding to the VCN's control plane and data plane. Examples of the VCN control plane are shown in Figures 16, 17, 18, and 19 (see reference numbers 1616, 1716, 1816, and 1916) and described below. Examples of the VCN data plane are shown in Figures 16, 17, 18, and 19 (see reference numbers 1618, 1718, 1818, and 1918) and described below. Control plane functions include functions used to configure the network to control how data is forwarded (e.g., setting routes and route tables, configuring VNICs). In certain embodiments, a VCN control plane is provided that centrally computes and exposes all overlay-to-substrate mappings to the NVD and virtual network edge devices (e.g., various gateways such as DRGs, SGWs, and IGWs). Firewall rules can also be exposed using the same mechanism. In certain embodiments, the NVD retrieves only mappings relevant to the NVD. The data plane functions include those that perform the actual routing / forwarding of packets based on the configurations established using the control plane. The VCN data plane is implemented by encapsulating customer network packets before they traverse the backbone network. The encapsulation / decapsulation functions are implemented in the NVD. In certain embodiments, the NVD is configured to intercept all network packets entering and leaving the host machine and perform network virtualization functions.
[0103] As described above, the NVD performs various virtualization functions, including VNICs and VCN VRs. The NVD can execute VNICs associated with compute instances hosted by one or more host machines connected to the VNICs. For example, as shown in FIG. 2, NVD 210 executes the functions of VNIC 276 associated with compute instance 268 hosted by host machine 202 connected to NVD 210. As another example, NVD 212 executes VNIC 280 associated with bare metal compute instance 272 hosted by host machine 206 and VNIC 284 associated with compute instance 274 hosted by host machine 208. The host machines can host compute instances that belong to different VCNs that belong to different customers. The NVDs connected to the host machines can execute VNICs (i.e., execute functions associated with the VNICs) corresponding to the compute instances.
[0104] NVDs also execute VCN virtual routers corresponding to the VCNs of the compute instances. For example, in the embodiment shown in FIG. 2, NVD 210 executes VCN VR 277 corresponding to the VCN to which compute instance 268 belongs. NVD 212 executes one or more VCN VRs 283 corresponding to one or more VCNs to which compute instances hosted on host machines 206 and 208 belong. In particular embodiments, a VCN VR corresponding to a VCN is executed by all NVDs connected to a host machine that hosts at least one compute instance belonging to that VCN. If a host machine hosts compute instances that belong to different VCNs, the NVDs connected to that host machine may execute VCN VRs corresponding to the different VCNs.
[0105] In addition to VNICs and VCN VRs, an NVD may include one or more hardware elements that run various software (e.g., daemons) and facilitate various network virtualization functions performed by the NVD. For simplicity, these various elements are grouped as "packet processing elements" shown in FIG. 2. For example, NVD 210 includes packet processing element 286, and NVD 212 includes packet processing element 288. For example, the packet processing element of an NVD may include a packet processor configured to monitor all packets received and communicated using the NVD by interacting with the NVD's ports and hardware working interfaces and to store network information. The network information may include, for example, network flow information for identifying different network flows processed by the NVD and information about each flow (e.g., statistics for each flow). In certain embodiments, the network flow information may be stored on a per-VNIC basis. As another example, the packet processing element may include a replication agent configured to replicate information stored by the NVD to one or more different replication target stores. As yet another example, the packet processing element may include a packet processor configured to monitor all packets received and communicated using the NVD by interacting with the NVD's ports and hardware working interfaces and to store network information. The packet processing element may include a logging agent configured to perform logging functions for the NVD. The packet processing element may also include a logging agent configured to monitor the performance and It may also include software for monitoring the health and possibly the status and health of other elements connected to the NVD.
[0106] FIG. 1 illustrates elements of an exemplary virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, VRs for the VCN, and a set of gateways configured for the VCN. The overlay elements illustrated in FIG. 1 may be executed or hosted by one or more of the physical elements illustrated in FIG. 2. For example, compute instances within a VCN may be executed or hosted by one or more host machines illustrated in FIG. 2. For compute instances hosted by a host machine, the VNICs associated with the compute instance are typically executed by an NVD connected to the host machine (i.e., the VNIC functionality is provided by an NVD connected to the host machine). The VCN VR functionality is performed by all NVDs connected to the host machines that host or execute compute instances that are part of the VCN. Gateways associated with a VCN may be executed by one or more different types of NVDs. For example, some gateways may be executed by smart NICs, and other gateways may be executed by one or more host machines or other implementations of NVDs.
[0107] As mentioned above, compute instances within a customer VCN can communicate with a variety of different endpoints. These endpoints may be in the same subnet as the source compute instance, in a different subnet but in the same VCN as the source compute instance, or may include endpoints outside the VCN of the source compute instance. These communications are facilitated using the VNICs associated with the compute instances, the VCN VRs, and the gateways associated with the VCN.
[0108] Communication between two compute instances on the same subnet within a VCN is facilitated using VNICs associated with the source and destination compute instances. The source and destination compute instances may be hosted by the same host machine or different host machines. A packet originating from a source compute instance may be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. In the NVD, the packet is processed using a packet processing pipeline, which may include the execution of a VNIC associated with the source compute instance. Because the packet's destination endpoint is in the same subnet, the execution of a VNIC associated with the source compute instance forwards the packet to an NVD running a VNIC associated with the destination compute instance, which processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances may run on the same NVD (e.g., if both the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., if the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNIC can use the routing / forwarding tables stored by the NVD to determine the next hop for a packet.
[0109] When communicating a packet from a compute instance in a subnet to an endpoint in a different subnet within the same VCN, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. In the NVD, the packet is processed using a packet processing pipeline that may include running one or more VNICs and VRs associated with the VCN. For example, the NVD executes or invokes a function corresponding to a VNIC associated with the source compute instance (also referred to as executing a VNIC) as part of the packet processing pipeline. The function executed by the VNIC may include examining the VLAN tag on the packet. Because the packet's destination is outside the subnet, a VCN VR function is invoked and executed by the NVD. The VCN VR then routes the packet to the NVD executing the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards the packet to the destination compute instance. The VNICs associated with the source compute instance and the destination compute instance may run on the same NVD (e.g., if both the source compute instance and the destination compute instance are hosted by the same host machine) or may run on different NVDs (e.g., if the source compute instance and the destination compute instance are hosted by different host machines connected to different NVDs).
[0110] If the packet's destination is outside the VCN of the source compute instance, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD runs the VNIC associated with the source compute instance. Because the packet's destination endpoint is outside the VCN, the packet is processed by the VCN VR for that VCN. The NVD invokes VCN VR functions, which may result in the packet being forwarded to an NVD running the appropriate gateway associated with the VCN. For example, if the destination is an endpoint in a customer's on-premises network, the packet may be forwarded by the VCN VR to an NVD running the DRG gateway configured for the VCN. The VCN VR runs on the same NVD as the NVD running the VNIC associated with the source compute instance. The gateway may be implemented by a single NVD or may be implemented by a different NVD. The gateway may be implemented by a smart NIC, a host machine, or another NVD implementation. The packet is then processed by the gateway and forwarded to a next hop to facilitate communication of the packet to the intended destination endpoint. For example, in the embodiment shown in FIG. 2, a packet originating from compute instance 268 may be communicated from host machine 202 to NVD 210 over link 220 (using NIC 232). VNIC 276 on NVD 210 is called VNIC 276 because it is the VNIC associated with source compute instance 268. VNIC 276 is configured to inspect encapsulated information within the packet, determine a next hop for forwarding the packet to facilitate communication of the packet to the intended destination endpoint, and forward the packet to the determined next hop.
[0111] Compute instances deployed on a VCN can communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 200 may include instances within the same VCN or other VCNs (which may be customer VCNs or VCNs not belonging to the customer). Communication between endpoints hosted by CSPI 200 may be performed over physical network 218. Compute instances can also communicate with endpoints not hosted by or external to CSPI 200. Examples of these endpoints include endpoints within a customer on-premises network or data center, or public endpoints accessible over a public network such as the Internet. Communication with endpoints external to CSPI 200 may be performed over a public network (e.g., the Internet) (not shown in FIG. 2) or a private network (not shown in FIG. 2) using various communication protocols.
[0112] The architecture of CSPI 200 shown in FIG. 2 is merely exemplary and not intended to be limiting. Variations, substitutions, and modifications are possible in alternative embodiments. For example, in some implementations, CSPI 200 may have more or fewer systems or elements than those shown in FIG. 2, may combine two or more systems, or may have a different system configuration or arrangement. The systems, subsystems, and other elements shown in FIG. 2 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device).
[0113] FIG. 4 illustrates connections between host machines and an NVD to provide I / O virtualization to support multi-tenancy, according to certain embodiments. As shown in FIG. 4, a host machine 402 runs a hypervisor 404 that provides a virtualized environment. The host machine 402 runs two virtual machine instances: VM1 406, which belongs to customer / tenant #1, and VM2 408, which belongs to customer / tenant #2. The host machine 402 includes a physical NIC 410 connected to an NVD 412 via link 414. Each of the compute instances is connected to a VNIC run by the NVD 412. In the embodiment of FIG. 4, VM1 406 is connected to VNIC-VM1 420, and VM2 408 is connected to VNIC-VM2 422.
[0114] 4, NIC 410 includes two logical NICs: logical NIC A 416 and logical NIC B 418. Each virtual machine is connected to and configured to operate with its own logical NIC. For example, VM1 4 VM1 406 is connected to logical NIC A 416, and VM2 408 is connected to logical NIC B 418. The logical NICs make each tenant's virtual machine believe it owns its own host machine and NIC, even though the host machine 402 consists of only one physical NIC 410 shared by multiple tenants.
[0115] In particular embodiments, each logical NIC is assigned its own VLAN ID. Thus, logical NIC A 416 for tenant #1 is assigned a particular VLAN ID, and logical NIC B 418 for tenant #2 is assigned a different VLAN ID. When a packet is communicated from VM1 406, the hypervisor attaches a tag assigned to tenant #1 to the packet before communicating the packet from host machine 402 to NVD 412 over link 414. Similarly, when a packet is communicated from VM2 408, the hypervisor attaches a tag assigned to tenant #2 to the packet before communicating the packet from host machine 402 to NVD 412 over link 414. Thus, a packet 424 communicated from host machine 402 to NVD 412 has an associated tag 426 that identifies the particular tenant and associated VM. When a packet 424 is received on the NVD from host machine 402, the tag 426 associated with the packet is used to determine whether the packet should be processed by VNIC-VM1 420 or VNIC-VM2 422. The packet is then processed by the corresponding VNIC. The configuration shown in Figure 4 allows each tenant's compute instance to believe it owns its own host machine and NIC. The configuration shown in Figure 4 provides I / O virtualization to support multi-tenancy.
[0116] FIG. 5 is a schematic block diagram illustrating a physical network 500 according to a particular embodiment. The embodiment illustrated in FIG. 5 is constructed as a Clos network. A Clos network is a particular type of network topology designed to provide connection redundancy while maintaining high bisection bandwidth and maximum resource utilization. A Clos network is a type of non-blocking, multi-stage or multi-layer switching network, and the number of stages or layers may be 2, 3, 4, 5, etc. The embodiment illustrated in FIG. 5 is a three-layer network, including layers 1, 2, and 3. TOR switch 504 represents a layer-0 switch in the Clos network. One or more NVDs are connected to the TOR switch. The layer-0 switch is also referred to as an edge device of the physical network. The layer-0 switch is connected to a layer-1 switch, also referred to as a leaf switch. In the embodiment illustrated in FIG. 5, “n” layer-0 TOR switches are connected to “n” layer-1 switches to form a pod. Each layer-0 switch in a pod is interconnected to all layer-1 switches in the pod, but switches between pods are not connected. In a specific implementation, the two pods are referred to as blocks. Each block is served by or connected to "n" layer-2 switches (also referred to as spine switches). A physical network topology may include multiple blocks. In turn, the layer-2 switches are connected to "n" layer-3 switches (also referred to as super-spine switches). Communication of packets through the physical network 500 is typically performed using one or more layer-3 communication protocols. Typically, all layers of the physical network, except for the TOR layer, are n-way redundant, thus achieving high availability. The physical network can be scaled by specifying policies on pods and blocks to control the mutual visibility of switches in the physical network.
[0117] A characteristic of Clos networks is that the maximum hop count required to reach from one layer-0 switch to another layer-0 switch (or from an NVD connected to a layer-0 switch to another NVD connected to a layer-0 switch) is constant. For example, in a three-layer Clos network, a packet can travel a maximum of 7 hops from one NVD to another. In a four-tier Clos network, a packet requires up to nine hops to travel from one NVD to another. In this case, the source NVD and destination NVD are connected to the leaf tier of the Clos network. Similarly, in a four-tier Clos network, a packet requires up to nine hops to travel from one NVD to another. In this case, the source NVD and destination NVD are connected to the leaf tier of the Clos network. Therefore, the Clos network architecture maintains constant latency across the network, which is important for intra- and inter-datacenter communications. The Clos topology is horizontally scalable and cost-effective. The network bandwidth / throughput capacity can be easily increased by adding more switches (e.g., more leaf switches and spine switches) at each tier and by increasing the number of links between switches in adjacent tiers.
[0118] In certain embodiments, each resource in the CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource's information. This identifier can be used to manage the resource, for example, through a console or API. An exemplary syntax for a CID is as follows:
[0119] ocid1.<RESOURCE TYPE> . <realm>.[REGION].[FUTURE USE].<UNIQUE ID> is. During the ceremony, "ocid1" is a string that indicates the version of the CID.
[0120] "RESOURCE TYPE" is the type of resource (e.g., instance, volume, VCN) , subnet, user, group).
[0121] "REALM" represents the region where the resource resides. An example value is "c1" "c1" represents the government cloud region, "c2" represents the government cloud region, or "c3" represents the federal government cloud region. Each region can have its own domain name.
[0122] "REGION" represents the region that the resource belongs to. If a region does not apply to the resource, this part may be blank.
[0123] "FUTURE USE" indicates that the item is reserved for future use. "UNIQUE ID" is the unique ID part. This format is used for resources or servers. This may vary depending on the type of screw.
[0124] Loop prevention example 6 illustrates an example of a VCN 600 configured and deployed for a customer, such as an enterprise, in accordance with certain embodiments. VCN 600 includes a virtual L2 network 610 and may be implemented on Oracle® Cloud Infrastructure (OCI).
[0125] In general, VCN 600 may be a software-defined virtual version of a traditional network, including subnets, route tables, and gateways, on which various compute instances 620A, 620B-620K (collectively referred to herein as "computer instances 620" or individually as "computer instances 620") may run. For example, VCN 600 may be a virtual private network built by a customer in a cloud infrastructure. The cloud infrastructure may exist within a particular region and include all available domains in that region. Subnets are subdivisions of VCN 600. Each subnet defined in the cloud infrastructure may be within a single available domain or span all available domains in a region. At least one cloud infrastructure must be configured before launching compute instances 620. If necessary, public traffic may be shared among the subnets. Cloud connectivity options include an optional internet gateway to handle traffic, an optional IP Security (IPSec) virtual private network (VPN) connection to securely extend customer on-premises networks, or OCI's FastConnect. You can configure a private infrastructure. VCN 600 may be privately connected to another VCN so that traffic does not traverse the Internet. The CIDRs (Classless Inter-Domain Routing) of the two VCNs may not overlap.
[0126] In the example of Figure 6, a virtual L2 network 610 supports any of the virtual networking methods described above and is implemented as an overlay on a cloud infrastructure. In this case, the overlay implements an L2 protocol. Compute instances 620 may belong to and be connected through the virtual L2 network 610 using the L2 protocol. Specifically, the virtual L2 network 610 may include virtual switches, VNICs, and / or other virtual compute resources for receiving and routing frames to and from the compute instances 620. The overlay implements this virtualization on the underlying compute resources of the cloud infrastructure.
[0127] For example, compute instance 620A is hosted on a first host machine connected to an NVD via a first port. The NVD hosts a first VNIC for compute instance 620A. The first VNIC is an element of virtual L2 network 610. To send a frame to compute instance 620K, compute instance 620A includes the MAC address of compute instance 620A as the source address and the MAC address of compute instance 620K as the destination address in the frame according to the L2 protocol. These MAC addresses may be referred to as overlay MAC addresses or logical MAC addresses. The frame is received by virtual L2 network 610 via the first VNIC and sent to a second VNIC corresponding to compute instance 620K. In practice, the first host machine forwards the frame to the NVD via the first port, and the NVD routes the frame to the second host machine for compute instance 620K. Routing involves encapsulating the frame with relevant information (eg, the MAC address of the host machine, which may also be called the physical MAC address or board MAC address).
[0128] FIG. 7 illustrates an example VCN 700 including a virtual L2 network and a virtual Layer 3 (L3) network established for a customer, according to certain embodiments. In one example, the VCN 700 includes a virtual L3 network 710. The virtual L3 network 710 may be a virtual network implementing a protocol at the L3 layer of the OSI model. Multiple subnets 712A-712N (collectively referred to herein as "subnets 712" or individually as "subnets 712") may be established within the virtual network. The VCN 700 also includes a virtual L2 network 720, such as the virtual L2 network 610 of FIG. 6, that implements a protocol at the L2 layer of the OSI model. A switched virtual interface (SVI) router 730 connects the virtual L3 network 710 and the virtual L2 network 720. The virtual L2 network 720 may have a designated SVI IP address. Traffic between the subnets may pass through the SVI router 730.
[0129] In a virtual L3 network, each subnet 712 can have a contiguous range of IP version 4 (IPv4) or version 6 (IPv6) addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets in VCN 700. Subnets act as building blocks. All compute instances within a given subnet share the same route tables, security lists, and dynamic host configuration. The subnet 712 may also include a VNIC. Each compute instance 620 may reside within the subnet and be connected to a VNIC that enables network connectivity for the compute instance 620. A subnet may be either a public subnet or a private subnet. A private subnet may include a VNIC that does not have a public Internet Protocol (IP) address. A public subnet may include a VNIC that has a public IP address. A subnet may exist in a single available domain or across multiple available domains within a particular region.
[0130] In a virtual L2 network, VNICs may be attached to compute instances to enable communication between them using L2 protocols. Whether L2 or L3, VNICs may determine how compute instances 620 connect to endpoints and are hosted on NVDs (also called intelligent server adapters (ISAs)). L2 VNICs enable communication between compute instances using L2 protocols, while L3 VNICs enable communication using L3 protocols.
[0131] Each compute instance 620 may have a primary VNIC that is created during compute instance startup and is not deleted. Secondary VNICs (in the same availability domain as the primary VNIC) may be added to an existing compute instance 620 and deleted as needed.
[0132] L2 VNICs may be hosted on an NVD, which implements switch functionality. The NVD can gradually learn the source MAC addresses of compute instances based on frames received from the source, maintain an L2 forwarding table, and transmit frames based on the MAC addresses and the forwarding table. Transmitting a frame can include transmitting the actual frame (e.g., forwarding the frame) or transmitting one or more replicas of the frame (e.g., broadcasting the frame). For example, upon receiving a frame with a source MAC address from a port, the NVD can learn that it can transmit frames to the source MAC address through the port. The NVD can also cache and age out MAC addresses. The NVD can maintain static (non-aged) MAC addresses of the SVT routers 730. When it receives frames with unknown destination MAC addresses or with broadcast destination MAC addresses, the NVD can broadcast (flood) these frames.
[0133] Generally, the OSI model includes seven layers: Layer 1 (L1) is the physical layer, which is concerned with the transmission of data bits over a physical medium; L2 is the data link layer, which governs the transmission of frames between connected nodes on the physical layer; and L3 is the network layer, which describes addressing, routing, and traffic control for multi-node networks.
[0134] L2 is the broadcast MAC level network, and L3 is the segmented routing over IP networks. L2 can include two sublayers: the MAC layer authorizes devices to access and transmit on the medium, and the Logical Link Layer (LLC) handles frame traffic, including managing the communication link, identifying protocols above the network layer, and checking for errors and frame synchronization. L3 operates on IP addresses, while L2 operates on MAC addresses. MAC addresses are unique identifiers for the network adapter present in each device. IP addresses are a higher level of abstraction than MAC addresses, and can be slower. IP addresses are typically assigned via DHCP. While MAC addresses are "leased" or "assigned" by a server and can be changed, MAC addresses are fixed addresses of network adapters and do not change on a device unless the hardware adapter is changed. A frame is the protocol data unit on an L2 network. Frames have a defined structure and can be used for error detection, control plane activity, etc. The network can use some frames to control the data link itself.
[0135] In L2, unicast refers to sending a frame from one node to another single node, multicast refers to sending traffic from one node to multiple nodes, and broadcast refers to sending a frame to all nodes in the network. A broadcast domain is a logical division of a network where L2 broadcasts can reach all nodes in the network. Bridges can be used to link LAN segments at the frame level. Bridging creates separate broadcast domains on a LAN, creating VLANs. VLANs are independent logical networks that group related devices into separate network segments. The grouping of devices on a VLAN is independent of the physical placement of the devices on the LAN. In a VLAN, frames whose source and destination are in the same VLAN are forwarded only within the local VLAN.
[0136] An L2 network device is a multiport device that processes and forwards data at the data link layer using hardware addresses and MAC addresses. Frames are sent to a specific switch port based on the destination MAC address. L2 switches can automatically learn MAC addresses and build a table that can be used to selectively forward frames. For example, if a switch receives a frame from MAC address X on port 1, the switch knows that it can forward the frame from port 1 to MAC address X without having to try each available port. Because L2 information is easily obtained, frame forwarding (switching) is very fast, typically at network wire speed. Devices within the same L2 segment do not require routing to reach local peers. Therefore, L2 switching has little or no impact on network performance or bandwidth. Furthermore, L2 switches are inexpensive and easy to deploy. However, L2 switches may not apply any intelligence when forwarding frames. For example, L2 switches do not natively allow multipathing and can suffer from loop problems because L2 control protocols inherently lack loop prevention mechanisms. Additionally, L2 switches may not route frames based on IP addresses or prioritize frames sent by specific applications. In L3 networks, packets are sent to a specific next-hop IP address based on the destination IP address. L3 virtual networks (e.g., L3 VCNs) are highly scalable, natively support multipathing, and do not suffer from loops. However, many enterprise workloads still use L2 virtual networks (L2 VLANs), at least because L2 virtual networks can provide simple, inexpensive, and high-performance connectivity to hundreds or thousands of end stations.
[0137] Because L2 networks implement L2 protocols, which do not natively prevent loops, L2 networks may be susceptible to broadcast storms when loops are introduced. In contrast, L3 networks may be less susceptible to loop issues. Also, both L2 and L3 networks must include multipathing to provide redundant paths in case of link failure. To mitigate loop issues, STP can be implemented in L2 networks; however, this may hinder multipathing in L2 networks.
[0138] Specifically, STP regulates the interconnections between switches and can generate L2 networks, which are trees that define a single path to avoid loops by disabling certain links. STP is a network protocol for building a loop-free logical topology in a network. STP is required because switches in a LAN are often interconnected using redundant links to improve resilience in case one connection fails. However, this connection configuration creates switching loops that result in broadcast radiation and MAC table instability. When redundant links are used to connect switches, switching loops must be avoided. Therefore, STP is generally implemented on switches to monitor the network topology. The basic function of STP is to prevent bridge loops and the resulting broadcast radiation. Spanning tree also provides a backup link that provides fault tolerance when an active link fails. STP enables network designs that include multiple L2 links. STP creates a spanning tree that characterizes the relationships between nodes in a network of connected L2 switches or bridges. STP disables links that are not part of the spanning tree, leaving a single active path between any two network nodes, for example, between the tree root and any leaf. All other paths are forced into a standby (or blocked) state. STP-based L2 loop prevention can be achieved using various variants or extensions of STP, such as Rapid Spanning Tree Protocol (RSTP), Multiple Spanning Tree Protocol (MSTP), and VLAN Spanning Tree Protocol (VSTP).
[0139] Therefore, for L2 networks, using STP may not support optimal implementation because it prevents multipathing. Virtual L2 networks, such as virtual L2 network 720, are implemented as an overlay of an L2 network on a cloud infrastructure. Therefore, virtual L2 networks are also susceptible to loop issues. STP can also be implemented in virtual L2 networks, but would similarly result in a loss of multipathing capabilities. Embodiments of the present disclosure support loop prevention and multipathing, as further described in connection with the following figures.
[0140] 8 illustrates an example of an infrastructure supporting a virtual L2 network of a VCN, according to certain embodiments. The elements illustrated in FIG. 8 are typically part of a CSP provided by a CSP and used to provide cloud services (e.g., IaaS services) to subscribing customers. In one example, the elements are the same as or similar to elements described in the architecture diagram of FIG. 2. Similarities between the elements of the two diagrams will not be repeated here for the sake of brevity.
[0141] As shown in FIG. 8, a host machine 820A is connected to an NVD 810A, which is connected to a TOR switch 832A of a physical switch network 830. Furthermore, host machines 820A-820L (collectively referred to herein as "host 820" or individually referred to as "host 820") are connected to the NVD 810A via an Ethernet link, and the NVD 810A is connected to a TOR switch 832A via another Ethernet link. Also as shown in FIG. 8, host machines 821A-821L (collectively referred to herein as "host 821" or individually referred to as "host 821") are connected to an NVD 810K via an Ethernet link, and the NVD 810K is connected to a TOR switch 832K via an Ethernet link.
[0142] Each host 820 or 821 can run one or more compute instances for one or more customers. A customer's compute instances are configured by the customer and run by the CSPI. Each NVD 810 may be part of one or more virtual cloud networks (VCNs) hosted in the cloud by the The NVD 810 may host VNICs on the customer's compute instances. The customer's compute instances are connected via VNICs to a virtual L2 network, such as virtual L2 network 610 in FIG. 6 or L2 virtual network 720 in FIG. 7. This virtual L2 network may be implemented as an overlay on an infrastructure that includes the NVD 810, hosts 820 and 821, and switch network 830.
[0143] 8, NVD 810A includes a north port 812 that connects NVD 810A to TOR switch 832A. NVD 810A also includes a set of south ports 814A-814L, each connected to one of hosts 820A-820L. NVD 810A further hosts multiple VNICs 816A-816M, each connected to a compute instance running on host 820A.
[0144] Also in the example of FIG. 8, host 820A is connected to NVD 810A via south port 814A. Host 820A can run customer compute instance 822A. VNIC 816A hosted on NVD 810A corresponds to compute instance 822A. Compute instance 822A can belong to a virtual L2 network within the customer's VCN. Similarly, host 820L is connected to NVD 810A via south port 814L and hosts at least customer compute instance 822M, and VNIC 816M corresponds to compute instance 822M. Although FIG. 8 shows each host 820 connected to a single NVD (e.g., NVD 810A), some or all of hosts 820 may be connected to multiple NVDs.
[0145] NVD 810K may be similar to NVD 810A. Specifically, NVD 810K includes a north port 813 that connects NVD 810K to TOR switch 832K of switch network 830. NVD 810K also includes a set of south ports 815A-815L, each connected to one of hosts 821A-821L. NVD 810K also hosts multiple VNICs 817A-817M, each connected to a compute instance running on host 821A. Host 821A is connected to NVD 810K via south port 815A and hosts customer compute instance 823A. VNIC 817A hosts compute instance 823A. Similarly, host 82IL is connected to NVD 810K via south port 815L and hosts at least a customer compute instance 823M, with VNIC 817M corresponding to this compute instance 823M. Although Figure 8 shows NVD 810A and 8T0K each having the same number of ports, hosts, VNICs, and compute instances, these numbers may differ.
[0146] A frame sent from a first compute instance to a second compute instance is processed at the virtual and hardware levels. Specifically, the first compute instance 822A generates a frame containing a payload, the MAC address of the first compute instance as the source address, and the MAC address of the second compute instance as the destination address. These MAC addresses are overlay MAC addresses because they are in the virtual network. The NVD 810A receives the frame and maps it to the correct VNIC (e.g., VNIC 816A) based on the physical or virtual port (e.g., VLAN). The NVD 810A then performs packet filtering, logging, destination address lookup, and overlay encapsulation. Perform the necessary overlay functions. For example, the first NVD 810A uses the overlay and the destination MAC address from the frame to determine from its forwarding table the physical MAC address of the second host and the port to which the frame can be sent. If this information is found in the forwarding table, the NVD 810A substitutes the physical MAC address of the host 820A with its MAC address as the source and its destination MAC address with the physical MAC address of the second host as the destination in the frame, and transmits the frame over the port. If this information is not found in the forwarding table, the NVD 810A determines that the frame should be flooded, and therefore substitutes the physical MAC address of the first host 820A with its MAC address as the source and its destination MAC address with the broadcast MAC address as the destination in the frame, and broadcasts the frame. Sending the frame can include sending the actual frame or multiple replicas of the frame (which is equivalent to broadcasting the frame).
[0147] Conversely, upon receiving a frame through its north port, the NVD 810A performs any necessary packet transformations, including frame filtering, logging, source and / or destination MAC address lookup, decapsulation, and encapsulation. As a result of these transformations, the NVD 810A determines which VNIC the packet belongs to and whether the packet is legitimate. If the frame passes the forwarding table lookup process, the NVD 810A determines the appropriate port (or VNIC and VLAN) and sends the frame to the host through that port. Similarly, the VNIC corresponding to the compute instance can modify the overlay address contained in the frame by receiving and processing it. If no match is found in the forwarding table, the NVD 810A determines that the frame should be flooded to a host-facing port.
[0148] FIG. 9 illustrates an example of a loop 900 between NVDs, according to certain embodiments. The loop 900 can occur between an NVD 810A and an NVD 810K in the infrastructure described in FIG. 8 and can grow exponentially. Specifically, a frame is received by the NVD 810A and has a destination MAC address that is not included in the NVD 810A's forwarding table. In this case, the receiving NVD 810A can flood the frame (e.g., send the frame to all NVDs serving this VCN or this customer's VLAN). If the NVD 810K also does not have a forwarding table entry for this destination MAC address, the NVD 810K must also flood the frame. Flooding causes the frame to be sent back to the original sending NVD 810A, thereby forming a loop between the two NVDs 810A and 810K. Retransmitting the frame can again result in broadcasts, thereby growing the loop. Within a short period of time (eg, a few milliseconds), the growth can significantly flood the network and consume network bandwidth.
[0149] In the example of FIG. 9, frame 910 is generated by a first compute instance (e.g., compute instance 822A) executing on host 820A and transmitted to a second compute instance (e.g., compute instance 823A) executing on host 821A. This frame is an L2 frame including an L2 header and a payload. The payload includes an L2 protocol data unit (L2 PDU). The header includes a first MAC address of the first compute instance ("MAC1" in FIG. 9) as the source address and a second MAC address of the second compute instance ("MAC3" in FIG. 9) as the destination address. Nevertheless, embodiments of the present disclosure are not so limited, and this frame or other frames may be transmitted to two compute instances running on the same host, two compute instances running on different hosts connected to the same NVD, and / or two compute instances running on different hosts connected to different NVDs. In all of these cases, transmitting the frame may similarly result in loop 900.
[0150] Host 820A transmits frame 910 to NVD 810A via south port 814A. NVD 810 looks up its forwarding table and does not identify a match between the second MAC address and an entry in the forwarding table. Therefore, NVD 810A determines to flood frame 910 to include an unknown destination MAC address. Flooding may include transmitting frame 910 through all ports of NVD 810A.
[0151] Thus, host 820L receives frame 910 (e.g., a replica of frame 910) that includes the second MAC address. Because this second MAC address does not correspond to the MAC address of a compute instance hosted by host 820L, frame replica 930 is discarded.
[0152] NVD 810K receives frame 910 (a replica of this frame encapsulated with appropriate information for routing through switch network 830, e.g., by including an overlay header (denoted in the drawing as "OH"). This overlay header may include a VCN header, a source IP address, a destination IP address, a source MAC address, and a destination MAC address). Although the received frame includes a second MAC address, which corresponds to a second compute instance running on host 821A, NVD 810K has not yet learned the second MAC address and does not have an entry for the second MAC address in its forwarding table. Therefore, NVD 810K also determines that the received frame should be flooded and broadcasts the received frame through its ports. As a result, each of hosts 821 receives frame 910 (e.g., a replica of frame 910). The received frame is discarded unless its destination MAC address matches the destination MAC address of the compute instance on host 821.
[0153] Flooding by NVD 810K also involves sending frame 910 (or a replica of frame 910 encapsulated with appropriate information for routing through switch network 830) back to NVD 810A via North Port 813 and switch network 830. The retransmission of frame 910 by NVD 810A to NVD 810K and the return of frame 910 by NVD 810AK to NVD 810A creates loop 900. This loop can repeat multiple times in a short period of time, resulting in significant network bandwidth usage.
[0154] FIG. 10 illustrates an example of preventing loops between NVDs associated with a virtual L2 network, according to certain embodiments. For clarity, reference is made to NVD 810A and NVD 810K in the infrastructure described in FIG. 8. However, loop prevention may also be applied to any other loops between network resources, e.g., switches, that support the virtual L2 network. Unlike FIG. 9, each of the NVDs implements loop prevention rules that mitigate the occurrence of loops. A loop prevention rule is a rule that, when applied, prevents the occurrence of one or more loops in the network. A loop prevention rule may specify one or more parameters and one or more actions to be applied to support loop prevention. The parameters and actions may be, for example, a set of if-then statements. In one example, a loop prevention rule specifies that if a network resource receives a frame through a port, the network resource cannot send frames through the same port. Specifically, if a network resource broadcasts a frame because the frame has a destination MAC address that does not match the network resource's forwarding table, the frame cannot be broadcast through the port on which it was received.
[0155] In the example of FIG. 10, NVD 810 stores loop prevention rule 1012A. Frame 1010 is generated by a first compute instance (e.g., compute instance 822A) running on host 820A and transmitted to a second compute instance (e.g., compute instance 823A) running on host 821A. The frame is an L2 frame including a header and a payload. The payload includes an L2 PDU. The header includes a first MAC address of the first compute instance ("MAC1" in FIG. 10) as the source address and a second MAC address of the second compute instance ("MAC3" in FIG. 10) as the destination address. Nevertheless, embodiments of the present disclosure are not so limited, and the frame or other frames may be transmitted to two compute instances running on the same host, two compute instances running on different hosts connected to the same NVD, and / or two compute instances running on different hosts connected to different NVDs.
[0156] Host 820A transmits frame 1010 to NVD 810A via south port 814A. NVD 810A looks up forwarding table 1014A and does not identify a match between the second MAC address and an entry in this forwarding table. At this point, NVD 810A knows that the first MAC address is associated with south port 814A and updates forwarding table 1014A to remember this association. NVD 810A then broadcasts frame 1010 via south port 814 (excluding south port 814A) and north port 812.
[0157] Host 820L receives frame 1010 containing a second MAC address, which does not correspond to the MAC address of the compute instance hosted by host S20L, and so the frame replica is discarded.
[0158] The NVD 810K also receives frame 1010 (or a replica of frame 1010 with appropriate encapsulation) via north port 813. The NVD 810K stores loop prevention rule 1012K, which specifies that a frame received via a port cannot be sent back via the same port. In this case, the received frame contains a second MAC address, which corresponds to a second compute instance running on host 821A; however, the NVD 810K has not yet learned the second MAC address and does not have an entry for the second MAC address in its forwarding table 1014K. At this point, the NVD 810K knows that the first MAC address is associated with north port 813 and updates forwarding table 1014K to record this association. The NVD 810K also determines that the received frame should be flooded because the destination MAC address is unknown, and broadcasts the frame via ports other than north port 813 based on loop prevention rule 1012K. As a result, host 821 receives frame 1010 (or a replica thereof). As a result, each host 821 receives frame 1010 (e.g., a replica of this frame 1010). The received frame is discarded unless its destination MAC address matches the destination MAC address of a compute instance on host 821. Loop prevention rule 1012K prevents NVD 810K from sending frame 1010 back to NVD 1010A via north port 813. This prevents a loop between NVD 810A and NVD 810K.
[0159] The above-mentioned loop can occur not only with "unicast MAC addresses with unknown destinations" but also with "broadcast MAC addresses." By definition, "broadcast MAC addresses" are transmitted on all ports, so they can cause loops similar to those described above. The above-mentioned loop prevention mechanism can also be applied to such "broadcast MAC addresses."
[0160] FIG. 11 illustrates an example of a loop 1100 between an NVD and a compute instance executing on a host, according to certain embodiments. For clarity, reference is made to the NVD 810A, host 820A, and compute instance 822A in the infrastructure described in FIG. 8. However, loops may exist between the NVD and other compute instances and / or between compute instances and other NVDs. In one example, the NVD 810A transmits frame 1110 addressed to the MAC address of compute instance 822A. This frame is transmitted to host 820A via south port 814A. Upon receiving frame 1110, a software bug or software code in compute instance 822A causes compute instance 822A to replicate and send frame 1110 back to host 820A (shown as frame 1120 in FIG. 11). Host 820A then transmits frame 1120 to the NVD 810A, which received the frame via south port 814A. Thus, a loop 1100 exists, with NVD 810A receiving frames it transmitted via south port 814A.
[0161] FIG. 12 illustrates an example of preventing loops between an NVD and a compute instance running on a host, according to certain embodiments. Similarly, for clarity of explanation, reference is made to the NVD 810A, host 820A, and compute instance 822A of the infrastructure described in FIG. 8 . However, loop prevention applies equally to any other NVD. In one example, the NVD 810A implements a lightweight STP 1200 for each south port of the NVD 810A. In the diagram of FIG. 12, the lightweight STP 1200 manages the south port 814A. A similar lightweight STP can also be implemented for each of the remaining south ports.
[0162] Managing the south port 814A includes determining whether a loop (e.g., loop 1100) exists and, if so, disabling 1202 the south port 814A. Disabling 1202 the south port 814A may include storing a flag indicating the loop and deactivating the south port 814A. The flag may be stored in an entry associated with the MAC address of the compute instance 822A, for example, in the forwarding table 1014A of the NVD 810A. Deactivation may be achieved using several techniques. In one technique, the NVD 810A disconnects the south port 814A via software control. In an additional or alternative technique, the NVD 810A stops transmitting frames through the south port 814A and ignores frames received through the south port 814A. Deactivation may apply to all frames sent to or from the host 820A via the south port 814A, regardless of the particular compute instance hosted by the host 820A. Alternatively, the deactivation may apply only to frames sent to and from compute instance 822A.
[0163] To determine whether a loop exists, the lightweight STP 1200 must specify that frame 1210 be transmitted via the south port. Transmission may be triggered periodically upon startup of the compute instance 822A, upon updating the compute instance 822A, and / or when frame congestion to and / or from the NVD 810A reaches a congestion level. Frame 1210 may be a BPDU. If a loop exists, the compute instance 822A replicates the BPDU, and the resulting frame 1220 is transmitted by the host 820A to the NVD 810A. The lightweight STP 1200 detects that the received frame 1220 is a BPDU and, accordingly, determines the existence of a loop and disables the south port 814A (1202). If a loop does not exist, the frame received via the south port 814A is not a BPDU, and the lightweight STP 1200 distinguishes between the frame and a BPDU. As a result, the south port 814 remains valid.
[0164] 13-15 illustrate an example of a flow for preventing loops while enabling multipathing. The operations of the flow may be performed by an NVD. Some or all of the instructions for performing the operations of the flow may be implemented as hardware circuits and / or stored as computer-readable instructions on a non-transitory computer-readable medium of the NVD. The instructions, when implemented, represent modules containing circuitry or code that are executable by a processor of the NVD. Such instructions are used to configure the NVD to perform specific operations described herein. Each circuit or code in combination with an associated processor represents a means for performing the respective operation. While the operations are shown in a particular order, it should be understood that a particular order is not required, and one or more operations may be omitted, skipped, performed in parallel, and / or the order of the operations may be changed.
[0165] FIG. 13 shows an example of a flow for preventing loops associated with a virtual L2 network, such as virtual L2 network 610 of FIG. 6 or virtual L2 network 720 of FIG. 7. The NVD may belong to a cloud infrastructure that provides the virtual L2 network. The flow may start at operation 1302. In operation 1302, the NVD stores loop prevention rules. The loop prevention rules may indicate that frames received by the NVD through a port of the NVD will not be transmitted by the NVD through the port. In operation 1304, the NVD may maintain a forwarding table. For example, the forwarding table may include entries, each of which associates a port with a MAC address of a compute instance and / or a MAC address of a host of the compute instance. The forwarding table may be used to forward frames received by the NVD and that include a destination MAC address. Specifically, if the destination MAC address matches an entry in the forwarding table, the NVD forwards the frame based on the entry's association. Otherwise, the NVD transmits the frame through all ports except the port on which the frame was received in accordance with the loop prevention rule. In operation 1306, the NVD maintains a lightweight spanning tree for each south port. The lightweight STP can detect loops and disable south ports based on BPDU transmission and BPDU reception. In operation 1308, the NVD routes frames based on loop prevention rules, the forwarding table, and the lightweight spanning tree. In this way, frames received through one or more north ports of the NVD are not looped back through those ports if their destination MAC address does not match an entry in the forwarding table. Loops created by software bugs or by the software code of a compute instance are also detected and managed.
[0166] Figure 14 illustrates an example of a flow for preventing loops between NVDs associated with a virtual L2 network. In this case, a first NVD is connected to a second NVD via a switch network, and frames can be exchanged between the two NVDs via associated north ports. Each of the two NVDs can store loop prevention rules. This flow is described with respect to the operation of the first NVD, but it can also be applied to the operation of the second NVD.
[0167] As shown, the flow may begin at operation 1402. In operation 1402, a first NVD may receive a frame via a first port of the first NVD (e.g., the north port if the frame is received from a second NVD). The frame may originate from a second compute instance and include an L2 PDU, a first MAC address of the first compute instance as a destination address (e.g., the destination of the first frame), and a second MAC address of the second compute instance as a source address (e.g., the destination of the first frame). The first compute instance may include a second MAC address (e.g., the source of the first frame). The first compute instance and the second compute instance are members of a virtual L2 network. The first compute instance is hosted by a first host machine connected to the network interface card via a second port. At operation 1404, the first NVD may look up the forwarding table. The lookup may use the MAC address to determine whether a match exists between the MAC address from the first frame and an entry in the forwarding table. At operation 1406, the first NVD determines whether the second MAC address matches an entry in the forwarding table. If there is a match, flow proceeds to operation 1408. If there is no match, flow proceeds to operation 1420, where the first NVD updates the forwarding table by including an entry associating the second MAC address with the first port. Thus, when a subsequent frame is received and includes the second MAC address as a destination address, the frame is forwarded via the first port. At operation 1408, the first NVD determines whether the first MAC address matches an entry in the forwarding table. If there is no match, flow proceeds to operation 1410. If there is a match, flow proceeds to operation 1430. At operation 1430, the first NVD forwards the first frame to the destination (e.g., forwards the first frame to the first compute instance by transmitting the first frame through the second port). At operation 1410, the first NVD determines, based on the second MAC address, that the first frame should be transmitted through all ports of the network interface card. In other words, because the destination does not match, the first NVD determines that it needs to flood the first frame by broadcasting it through its ports. At operation 1412, the first NVD determines that its loop prevention rules prevent it from transmitting the frame using the port on which it received the frame. Therefore, this loop prevention rule prevents the first NVD from sending the first frame through the first port.In operation 1414, the NVD forwards the first frame through all ports of the network interface card except the first port based on the loop prevention rule. In other words, by applying the loop prevention rule, the first frame is broadcast through all ports of the first NVD except the first port.
[0168] FIG. 15 illustrates an example of a flow for preventing loops between an NVD and a compute instance associated with a virtual L2 network. This flow equally applies to preventing loops between an NVD and other compute instances hosted on one or more hosts connected to the NVD through the NVD's south port. The example flow may begin at operation 1502. In operation 1502, the NVD transmits a first frame through a first port (e.g., south port) of the NVD. This port is connected to the compute instance's host. The first frame may be a BPDU or may be transmitted based on lightweight STP implemented in the NVD. In operation 1504, the NVD receives a second frame through the first port. In operation 1506, the NVD determines whether the second frame is a replica of the first frame. For example, if the second frame is also a BPDU that it sent (e.g., the BPDU was received in operation 1504 and has its own MAC address as the bridge ID in the BPDU), the NVD determines that a replica exists, and operation 1506 is followed by operation 1520. Otherwise, it is determined that a replica does not exist, and operation 1506 is followed by operation 1510. In operation 1510, the NVD does not disable the first port because no loop exists. On the other hand, in operation 1520, the NVD determines that a loop exists. In operation 1522, the NVD determines a loop prevention rule that indicates one or more actions to be taken to manage the loop. For example, this rule is stored by the NVD and specifies that the port of the NVD on which the loop occurred should be disabled. In operation 1524, the NVD takes action (or (all actions) to prevent loops. For example, the NVD disables the first port.
[0169] IaaS Architecture Examples As mentioned above, IaaS (Infrastructure as a Service) is a specific type of Cloud computing. IaaS may be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider may host infrastructure elements (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may offer various services (e.g., billing, monitoring, logging, security, load balancing, clustering, etc.) associated with the infrastructure elements. Therefore, because these services can be policy-driven, IaaS users can implement policies to drive load balancing to maintain application availability and performance.
[0170] In some examples, IaaS customers can access resources and services over a wide area network (WAN), such as the Internet, and use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform, create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and install enterprise software on the VMs. Customers can use the provider's services to perform a variety of functions, including balancing network traffic, troubleshooting applications, monitoring performance, managing disaster recovery, and more.
[0171] In most cases, the cloud computing model requires the participation of a cloud provider, which can be, but does not have to be, a third-party service provider specializing in providing (e.g., offering, renting, or selling) IaaS. Alternatively, an enterprise can deploy a private cloud and become a provider of infrastructure services.
[0172] In some examples, IaaS deployment is the process of deploying a new application or a new version of an application onto a provisioned application server or the like. IaaS deployment may include the process of provisioning the server (e.g., installing libraries, daemons, etc.). IaaS deployment is often managed by the cloud provider below the hypervisor layer (e.g., server, storage, network hardware, and virtualization). Thus, customers can deploy the OS, middleware, and / or applications (e.g., self-service virtual machines (e.g., that can be spun up on demand)).
[0173] In some instances, IaaS provisioning may include obtaining the computer or virtual host to be used and installing the necessary libraries or services on the computer or virtual host. In most cases, deployment does not include provisioning, which must be performed first.
[0174] In some cases, IaaS provisioning presents two distinct challenges. First, there is the challenge of provisioning an initial set of infrastructure before you can do anything. Second, there is the challenge of consuming existing infrastructure after everything has been provisioned. One challenge is evolving the infrastructure (e.g., adding new services, modifying services, removing services). In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., which elements are required and how these elements interact) may be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which and how they work together) can be described declaratively. In some instances, once the topology is defined, workflows can be generated to create and / or manage the different elements described in the configuration files.
[0175] In some examples, the infrastructure can include many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potentially on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there may be one or more security group rules and one or more virtual machines (VMs) that are provisioned to define how the network's security is configured. Other infrastructure elements, such as load balancers, databases, etc., may also be provisioned. The infrastructure can evolve incrementally as more infrastructure elements are desired and / or added.
[0176] In some examples, continuous deployment techniques may be employed to enable deployment of infrastructure code across various virtual computing environments. The described techniques may also enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, typically many, different production environments (e.g., across a variety of different geographic locations, sometimes spanning the entire world). However, in some examples, the infrastructure for deploying the code must first be set up. In some examples, provisioning may be done manually, with provisioning tools used to provision resources and / or deployment tools used to deploy the code after the infrastructure has been provisioned.
[0177] FIG. 16 is a block diagram 1600 illustrating an example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1602 may be communicatively coupled to a secure host tenancy 1604, which may include a virtual cloud network (VCN) 1606 and a secure host subnet 1608. In some examples, the service operator 1602 may use one or more client computing devices. The one or more client computing devices may run, for example, software such as Microsoft Windows Mobile® and / or various operating systems such as iOS, Windows Phone, Android®, BlackBerry 8, and Palm OS. The device may be a handheld mobile device (e.g., iPhone®, mobile phone, iPad®, tablet, personal digital assistant (PDA) or wearable device (e.g., Google® Glass® head-mounted display) capable of running any mobile operating system and enabled for Internet, email, short message service (SMS), Blackberry® or other communication protocols. The client computing device illustratively includes a Microsoft Windows Operating System, Apple Macintosh® Operating System and and / or running various versions of the Linux operating system. The client computing device may be a general-purpose personal computer, including a personal computer and / or a laptop computer for running the client computer. The client computing devices may be, for example, workstation computers running various commercially available UNIX or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems, e.g., Google Chrome® OS. Alternatively or additionally, the client computing devices may be other electronic devices capable of communicating over VCN 1606 and / or an Internet-accessible network, such as thin-client computers, Internet-enabled gaming systems (e.g., Microsoft Xbox® game consoles with or without Kinect® gesture input devices), and / or personal messaging devices.
[0178] VCN 1606 may include a local peering gateway (LPG) 1610 that can be communicatively coupled to a secure shell (SSH) VCN 1612 via an LPG 1610 included in SSH VCN 1612. SSH VCN 1612 can include an SSH subnet 1614, and SSH VCN 1612 can be communicatively coupled to a control plane VCN 1616 via an LPG 1610 included in control plane VCN 1616. SSH VCN 1612 can also be communicatively coupled to a data plane VCN 1618 via LPG 1610. Control plane VCN 1616 and data plane VCN 1618 may be included in a service tenancy 1619, which may be owned and / or operated by the IaaS provider.
[0179] The control plane VCN 1616 may include a control plane demilitarized zone (DMZ) tier 1620 that functions as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers have a particular level of reliability and can contain security breaches. Additionally, the DMZ tier 1620 may include one or more load balancer (LB) subnets 1622, a control plane app tier 1624 that may include an app subnet 1626, and a control plane data tier 1628 that may include a database (DB) subnet 1630 (e.g., a front-end DB subnet and / or a back-end DB subnet). LB subnet 1622 included in control plane DMZ layer 1620 may be communicatively coupled to app subnet 1626 included in control plane app layer 1624 and to an Internet gateway 1634 that may be included in control plane VCN 1616, and applisubli 1626 may be communicatively coupled to DB subnet 1630, service gateway 1636, and network address translation (NAT) gateway 1638 included in control plane data layer 1628. Control plane VCN 1616 may include service gateway 1636 and NAT gateway 1638.
[0180] The control plane VCN 1616 can include a data plane mirror app layer 1640, which can include an app subnet 1626. The app subnet 1626 included in the data plane mirror app layer 1640 can include a virtual network interface controller (VNIC) 1642 on which a compute instance 1644 can run. The compute instance 1644 can communicatively couple the app subnet 1626 of the data plane mirror app layer 1640 to the app subnet 1626, which can be included in the data plane app layer 1646.
[0181] The data plane VCN 1618 may include a data plane app layer 1646, a data plane DMZ layer 1648, and a data plane data layer 1650. The data plane DMZ layer 1648 may include an app subnet 1626 of the data plane app layer 1646 and a LB subnet 1622 that may be communicatively coupled to an Internet gateway 1634 of the data plane VCN 1618. 26 may be communicatively coupled to a service gateway 1636 of the data plane VCN 1618 and a NAT gateway 1638 of the data plane VCN 1618. The data plane data layer 1650 may also include a DB subnet 1630 that may be communicatively coupled to the app subnet 1626 of the data plane app layer 1646.
[0182] The internet gateway 1634 of the control plane VCN 1616 and the internet gateway 1634 of the data plane VCN 1618 may be communicatively coupled to a metadata management service 1652, which may be communicatively coupled to the public internet 1654. The public internet 1654 may be communicatively coupled to a NAT gateway 1638 of the control plane VCN 1616 and the NAT gateway 1638 of the data plane VCN 1618. The service gateway 1636 of the control plane VCN 1616 and the service gateway 1636 of the data plane VCN 1618 may be communicatively coupled to a cloud service 1656.
[0183] In some examples, the service gateway 1636 of the control plane VCN 1616 or the data plane VCN 1618 can make application programming interface (API) calls to the cloud services 1656 without traversing the public Internet 1654. The API calls from the service gateway 1636 to the cloud services 1656 can be one-way. The service gateway 1636 can make API calls to the cloud services 1656, and the cloud services 1656 can send request data to the service gateway 1636. However, the cloud services 1656 may not initiate the API calls to the service gateway 1636.
[0184] In some examples, secure host tenancy 1604 may be directly connected to service tenancy 1619, which may be an orphan. Secure host subnet 1608 can communicate with SSH subnet 1614 through LPG 1610, which allows bidirectional communication with the orphan system. By connecting secure host subnet 1608 to SSH subnet 1614, secure host subnet 1608 can access other entities in service tenancy 1619.
[0185] The control plane VCN 1616 allows users of the service tenancy 1619 to configure or provision desired resources. The desired resources provisioned in the control plane VCN 1616 may be deployed or used in the data plane VCN 1618. In some examples, the control plane VCN 1616 may be isolated from the data plane VCN 1618, and the data plane mirror app layer 1640 of the control plane VCN 1616 can communicate with the data plane app layer 1646 of the data plane VCN 1618 via a VNIC 1642, which may be included in the data plane mirror app layer 1640 and the data plane app layer 1646.
[0186] In some examples, a user or customer of the system may make a request, such as a create, read, update, or delete (CRUD) operation, via the public internet 1654, which may communicate the request to a metadata management service 1652. The metadata management service 1652 may communicate the request to the control plane VCN 1616 via an internet gateway 1634. The request may be received by a LB subnet 1622 included in the control plane DMZ tier 1620. The LB subnet 1622 may determine that the request is valid, and in response to this determination, the LB subnet 1622 may send the request to an app subnet 1626 included in the control plane app tier 1624. If the request is validated and requires a call to the public internet 1654, the call to the public internet 1654 may be routed to the public internet 1654. The request may be sent to a NAT gateway 1638 that can make the call to the Rick Internet 1654. A memory for storing the request may be stored in the DB subnet 1630.
[0187] In some examples, the data plane mirror app layer 1640 can facilitate direct communication between the control plane VCN 1616 and the data plane VCN 1618. For example, it may be desirable for changes, updates, or other suitable modifications to a configuration to be applied to resources included in the data plane VCN 1618. The control plane VCN 1616 can communicate directly with the resources included in the data plane VCN 1618 via the VNIC 1642, allowing the changes, updates, or other suitable modifications to the configuration to be implemented.
[0188] In some embodiments, the control plane VCN 1616 and the data plane VCN 1618 may be included in the service tenancy 1619. In this case, a user or customer of the system may not own or operate either the control plane VCN 1616 or the data plane VCN 1618. Instead, an IaaS provider may own or operate the control plane VCN 1616 and the data plane VCN 1618, both of which may be included in the service tenancy 1619. This embodiment can prevent users or customers from interacting with other users' or customers' resources by enabling network isolation. This embodiment can also enable users or customers of the system to store databases privately without having to rely on the public internet 1654, which may not have the desired level of security for storage.
[0189] In another embodiment, LB subnet 1622 included in control plane VCN 1616 may be configured to receive signals from service gateway 1636. In this embodiment, control plane VCN 1616 and data plane VCN 1618 may be configured to be called by customers of the IaaS provider without calling the public internet 1654. Customers of the IaaS provider may desire this embodiment because databases used by the customers may be stored in service tenancy 1619, which is controlled by the IaaS provider and may be isolated from the public internet 1654.
[0190] 17 is a block diagram 1700 illustrating another example parameter of an IaaS architecture, according to at least one embodiment. A service operator 1702 (e.g., service operator 1602 of FIG. 16 ) may be communicatively coupled to a secure host tenancy 1704 (e.g., secure host tenancy 1604 of FIG. 16 ), which may include a virtual cloud network (VCN) 1706 (e.g., VCN 1606 of FIG. 16 ) and a secure host subnet 1708 (e.g., secure host subnet 1608 of FIG. 16 ). VCN 1706 may include a local peering gateway (LPG) 1710 (e.g., LPG 1610 of FIG. 16 ), which may be communicatively coupled to a secure shell (SSH) VCN 1712 (e.g., SSH VCN 1612 of FIG. 16 ) via an LPG 1610 included in SSH VCN 1712. SSH VCN 1712 can include SSH subnet 1714 (e.g., SSH subnet 1614 in FIG. 16 ), and SSH VCN 1712 can be communicatively coupled to control plane VCN 1716 (e.g., control plane VCN 1616 in FIG. 16 ) via LPG 1710 included in control plane VCN 1716. Control plane VCN 1716 can be included in service tenancy 1719 (e.g., service tenancy 1619 in FIG. 16 ), and data plane VCN 1718 (e.g., data plane VCN 1618 in FIG. 16 ) can be included in customer tenancy 1721, which can be owned or operated by a user or customer of the system.
[0191] The control plane VCN 1716 may include a control plane DMZ tier 1720 (e.g., the control plane DMZ tier 1620 in FIG. 16 ) that may include a LB subnet 1722 (e.g., the LB subnet 1622 in FIG. 16 ), a control plane app tier 1724 (e.g., the control plane app tier 1624 in FIG. 16 ) that may include an app subnet 1726 (e.g., the app subnet 1626 in FIG. 16 ), and a control plane data tier 1728 (e.g., the control plane data tier 1628 in FIG. 16 ) that may include a database (DB) subnet 1730 (e.g., similar to the DB subnet 1630 in FIG. 16 ). The LB subnet 1722 included in the control plane DMZ tier 1720 may be communicatively coupled to the app subnet 1726 included in the control plane app tier 1724 and to an Internet gateway 1734 (e.g., the Internet gateway 1634 in FIG. 16 ), which may be included in the control plane VCN 1716. The app subnet 1726 may be communicatively coupled to a DB subnet 1730, a service gateway 1736 (e.g., the service gateway in FIG. 16 ), and a network address translation (NAT) gateway 1738 (e.g., the NAT gateway 1638 in FIG. 16 ) included in the control plane data layer 1728. The control plane VCN 1716 may include the service gateway 1736 and the NAT gateway 1738.
[0192] The control plane VCN 1716 can include a data plane mirror app layer 1740 (e.g., data plane mirror app layer 1640 of FIG. 16 ), which can include an app subnet 1726. The app subnet 1726 included in the data plane mirror app layer 1740 can include a virtual network interface controller (VNIC) 1742 (e.g., VNIC 1642) on which a compute instance 1744 (e.g., similar to compute instance 1644 of FIG. 16 ) can run. The compute instance 1744 can facilitate communication between the app subnet 1726 of the data plane mirror app layer 1740 and the app subnet 1726, which can be included in the data plane app layer 1746 (e.g., data plane app layer 1646 of FIG. 16 ), via the VNIC 1742 included in the data plane mirror app layer 1740 and the VNIC 1742 included in the data plane app layer 1746.
[0193] An internet gateway 1734 included in the control plane VCN 1716 may be communicatively coupled to a metadata management service 1752 (e.g., metadata management service 1652 of FIG. 16 ), which may be communicatively coupled to the public internet 1754 (e.g., public internet 1654 of FIG. 16 ). The public internet 1754 may be communicatively coupled to a NAT gateway 1738 included in the control plane VCN 1716. A service gateway 1736 included in the control plane VCN 1716 may be communicatively coupled to cloud services 1756 (e.g., cloud services 1656 of FIG. 16 ).
[0194] In some examples, data plane VCN 1718 may be included in customer tenancy 1721. In this case, the IaaS provider may provide a control plane VCN 1716 for each customer, and the IaaS provider may configure a unique compute instance 1744 for each customer that is included in service tenancy 1719. Each compute instance 1744 may allow communication between the control plane VCN 1716 included in service tenancy 1719 and the data plane VCN 1718 included in customer tenancy 1721. The compute instance 1744 may allow resources provisioned in the control plane VCN 1716 included in service tenancy 1719 to be deployed or used in the data plane VCN 1718 included in customer tenancy 1721.
[0195] In another example, a customer of an IaaS provider may have a database that resides in customer tenancy 1721. In this example, control plane VCN 1716 may include data plane minor app tier 1740, which may include app subnet 1726. Data plane mirror app tier 1740 may reside in data plane VCN 1718, but data plane mirror app tier 1740 may not reside in data plane VCN 1718. That is, data plane mirror app tier 1740 may have access to customer tenancy 1721, but data plane mirror app tier 1740 may not reside in data plane VCN 1718 and may not be owned or operated by the IaaS provider's customer. Data plane mirror app tier 1740 may be configured to make calls to data plane VCN 1718, but may not be configured to make calls to any entities included in control plane VCN 1716. A customer may desire to deploy or use resources in the data plane VCN 1718 provisioned to the control plane VCN 1716, and the data plane mirror application tier 1740 may facilitate the desired deployment or other use of the customer's resources.
[0196] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 1718. In this embodiment, the customer can determine what the data plane VCN 1718 can access, and the customer can restrict access from the data plane VCN 1718 to the public internet 1754. The IaaS provider may not be able to apply filters or control access from the data plane VCN 1718 to any external network or database. Applying filters and controls to the data plane VCN 1718 included in the customer tenancy 1721 by the customer can help isolate the data plane VCN 1718 from other customers and the public internet 1754.
[0197] In some embodiments, cloud services 1756 can be called by service gateway 1736 to access services that may not reside on the public internet 1754, on the control plane VCN 1716, or on the data plane VCN 1718. The connection between cloud service 1756 and control plane VCN 1716 or data plane VCN 1718 may not be live or continuous. Cloud services 1756 may reside on a separate network owned or operated by the IaaS provider. Cloud services 1756 may be configured to receive calls from service gateway 1736 and may not be configured to receive calls from the public internet 1754. Some cloud services 1756 may be isolated from other cloud services 1756, and control plane VCN 1716 may be isolated from cloud services 1756 that may not be located in the same region as control plane VCN 1716. For example, control plane VCN 1716 may be located in “Region 1,” and cloud service “Deployment 16” may be located in “Region 1” and “Region 2.” If a call to a deployment 16 is made by a service gateway 1736 included in a control plane VCN 1716 located in Region 1, the call may be sent to the deployment 16 in Region 1. In this example, the control plane VCN 1716 or the deployment 16 in Region 1 may not be communicatively coupled with the deployment 16 in Region 2.
[0198] 18 is a block diagram 1800 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1802 (e.g., service operator 1602 of FIG. 16 ) manages a secure host tenancy 1802, which may include a virtual cloud network (VCN) 1806 (e.g., VCN 1606 of FIG. 16 ) and a secure host subnet 1808 (e.g., secure host subnet 1608 of FIG. 16 ). 16 ) via LPG 1810 included in SSH VCN 1812. SSH VCN 1812 can include an SSH subnet 1814 (e.g., SSH subnet 1614 of FIG. 16 ), which may be communicatively coupled to a control plane VCN 1816 (e.g., control plane VCN 1616 of FIG. 16 ) via LPG 1810 included in control plane VCN 1816, and may be communicatively coupled to a data plane VCN 1818 (e.g., data plane 1618 of FIG. 16 ) via LPG 1810 included in data plane VCN 1818. The control plane VCN 1816 and the data plane VCN 1818 may be included in a service tenancy 1819 (e.g., service tenant 1619 in FIG. 16).
[0199] The control plane VCN 1816 may include a control plane DMZ layer 1820 (e.g., the control plane DMZ layer 1620 of FIG. 16 ) that may include a load balancer (LB) subnet 1822 (e.g., the LB subnet 1622 of FIG. 16 ), a control plane app layer 1824 (e.g., the control plane app layer 1624 of FIG. 16 ) that may include an app subnet 1826 (e.g., similar to the app subnet 1626 of FIG. 16 ), and a control plane data layer 1828 (e.g., the control plane data layer 1628 of FIG. 16 ) that may include a DB subnet 1830. The LB subnet 1822 included in the control plane DMZ layer 1820 may be communicatively coupled to the app subnet 1826 included in the control plane app layer 1824 and to an Internet gateway 1834 (e.g., the Internet gateway 1634 of FIG. 16 ), which may be included in the control plane VCN 1816. The app subnet 1826 may be communicatively coupled to a DB subnet 1830 included in the control plane data layer 1828, and to a service gateway 1836 (e.g., the service gateway in FIG. 16 ) and a network address translation (NAT) gateway 1838 (e.g., the NAT gateway 1638 in FIG. 16 ). The control plane VCN 1816 may include the service gateway 1836 and the NAT gateway 1838.
[0200] Data plane VCN 1818 may include a data plane app layer 1846 (e.g., data plane app layer 1646 in FIG. 16 ), a data plane DMZ layer 1848 (e.g., data plane DMZ layer 1648 in FIG. 16 ), and a data plane data layer 1850 (e.g., data plane data layer 1650 in FIG. 16 ). Data plane DMZ layer 1848 may include LB subnet 1822, which may be communicatively coupled to trusted app subnet 1860 and untrusted app subnet 1862 of data plane app layer 1846 and internet gateway 1834 included in data plane VCN 1818. Trusted app subnet 1860 may be communicatively coupled to service gateway 1836 included in data plane VCN 1818, NAT gateway 1838 included in data plane VCN 1818, and DB subnet 1830 included in data plane data layer 1850. The untrusted app subnet 1862 may be communicatively coupled to a service gateway 1836 included in the data plane VCN 1818 and to a DB subnet 1830 included in the data plane data layer 1850. The data plane data layer 1850 may include a DB subnet 1830 that may be communicatively coupled to a service gateway 1836 included in the data plane VCN 1818.
[0201] The untrusted app subnet 1862 may include one or more primary VNICs 1864(1)-(N) that may be communicatively coupled to tenant virtual machines (VMs) 1866(1)-(N). Each tenant VM 1866(1)-(N) may be associated with a respective customer tenant. Each secondary VNIC 1872(1)-(N) may be communicatively coupled to a respective app subnet 1867(1)-(N), which may be included in a respective container egress VCN 1868(1)-(N), which may be included in a respective container egress VCN 1870(1)-(N). Each secondary VNIC 1872(1)-(N) may facilitate communication between the untrusted app subnet 1862 included in the data plane VCN 1818 and the app subnet included in the container egress VCN 1868(1)-(N). Each container egress VCN 1868(1)-(N) may include a NAT gateway 1838, which may be communicatively coupled to the public internet 1854 (e.g., public internet 1654 in FIG. 16 ).
[0202] The internet gateway 1834 included in the control plane VCN 1816 and the internet gateway 1834 included in the data plane VCN 1818 may be communicatively coupled to a metadata management service 1852 (e.g., metadata management system 1652 of FIG. 16 ), which may be communicatively coupled to the public internet 1854. The public internet 1854 may be communicatively coupled to a NAT gateway 1838 included in the control plane VCN 1816 and the NAT gateway 1838 included in the data plane VCN 1818. The service gateway 1836 included in the control plane VCN 1816 and the service gateway 1836 included in the data plane VCN 1818 may be communicatively coupled to cloud services 1856.
[0203] In some embodiments, data plane VCN 1818 may be consolidated into customer tenancy 1870. This consolidation may be useful or desirable for an IaaS provider's customer in some cases, such as when they may want support when running their code. Customers may provide code that, when run, may be disruptive, may communicate with other customer resources, or may cause undesirable effects. Thus, the IaaS provider can determine whether or not to run code that a customer has provided to the IaaS provider.
[0204] In some examples, an IaaS provider's customer can grant temporary network access to the IaaS provider and request a feature to be added to the data plane app layer 1846. The code to perform the feature may run in VMs 1866(1)-(N) but cannot be configured to run elsewhere on the data plane VCN 1818. Each VM 1866(1)-(N) may be connected to one customer tenancy 1870. Each container 1871(1)-(N) included in a VM 1866(1)-(N) may be configured to run code. In this case, double isolation (e.g., container 1871(1)-(N) may run code, and container 1871(1)-(N) may be included in at least one VM 1866(1)-(N) included in the untrusted app subnet 1862) may exist, which can help prevent erroneous or unwanted code from damaging the IaaS provider's network or from damaging a different customer's network. Containers 1871(1)-(N) may be communicatively coupled to customer tenancy 1870 and may be configured to send or receive data from customer tenancy 1870. Containers 1871(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 1818. Once code execution is complete, the IaaS provider can kill or discard containers 1871(I)-(N).
[0205] In some embodiments, trusted app subnet 1860 may execute code that may be owned or operated by an IaaS provider. In this embodiment, trusted app subnet 1860 may be communicatively coupled to DB subnet 1830 and configured to perform CRUD operations on DB subnet 1830. Untrusted app subnet 1862 may be communicatively coupled to DB subnet 1830, but in this embodiment, untrusted app subnet 1862 may be configured to perform read operations within DB subnet 1830. Containers 1871(1)-(N) included in each customer's VMs 1866(1)-(N) that can execute code from the customer may not be communicatively coupled to DB subnet 1830.
[0206] In other embodiments, the control plane VCN 1816 and the data plane VCN 1818 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 1816 and the data plane VCN 1818. However, there may be indirect communication by at least one method. An LPG 1810 may be established by an IaaS provider that can facilitate communication between the control plane VCN 1816 and the data plane VCN 1818. In another example, the control plane VCN 1816 or the data plane VCN 1818 can make a call to a cloud service 1856 through a service gateway 1836. For example, a call from the control plane VCN 1816 to the cloud service 1856 may include a request for a service that can communicate with the data plane VCN 1818.
[0207] 19 is a block diagram 1900 illustrating further example parameters of an IaaS architecture, according to at least one embodiment. A service operator 1902 (e.g., service operator 1602 of FIG. 16 ) may be communicatively coupled to a secure host tenancy 1904 (e.g., secure host tenancy 1604 of FIG. 16 ), which may include a virtual cloud network (VCN) 1906 (e.g., VCN 1606 of FIG. 16 ) and a secure host subnet 1908 (e.g., secure host subnet 1608 of FIG. 16 ). VCN 1906 may include an LPG 1910 (e.g., LPG 1610 of FIG. 16 ) that may be communicatively coupled to an SSH VCN 1912 (e.g., SSH VCN 1612 of FIG. 16 ) via an LPG 1910 included in SSH VCN 1912. SSH VCN 1912 can include SSH subnet 1914 (e.g., SSH subnet 1614 in FIG. 16 ), which may be communicatively coupled to control plane VCN 1916 (e.g., control plane VCN 1616 in FIG. 16 ) via LPG 1910 included in control plane VCN 1916, and may be communicatively coupled to data plane VCN 1918 (e.g., data plane 1618 in FIG. 16 ) via LPG 1910 included in data plane VCN 1918. Control plane VCN 1916 and data plane VCN 1918 may be included in service tenancy 1919 (e.g., service tenancy 1619 in FIG. 16 ).
[0208] The control plane VCN 1916 may include a control plane DMZ tier 1920 (e.g., control plane DMZ tier 1620 of FIG. 16 ) that may include a LB subnet 1922 (e.g., LB subnet 1622 of FIG. 16 ), a control plane app tier 1924 (e.g., control plane app tier 1624 of FIG. 16 ) that may include an app subnet 1926 (e.g., app subnet 1626 of FIG. 16 ), and a control plane data tier 1928 (e.g., control plane data tier 1628 of FIG. 16 ) that may include a DB subnet 1930 (e.g., DB subnet 1830 of FIG. 18 ). The LB subnet 1922 included in the control plane DMZ tier 1920 may be communicatively coupled to the app subnet 1926 included in the control plane app tier 1924 and to an Internet gateway 1934 (e.g., Internet gateway 1634 of FIG. 16 ), which may be included in the control plane VCN 1916. The app subnet 1926 may be communicatively coupled to a DB subnet 1930 included in a control plane data layer 1928, a service gateway 1936 (e.g., the service gateway in FIG. 16 ) and a network address translation (NAT) gateway 1938 (e.g., the NAT gateway 1638 in FIG. 16 ). The N1916 may include a service gateway 1936 and a NAT gateway 1938.
[0209] Data plane VCN 1918 may include a data plane app layer 1946 (e.g., data plane app layer 1646 in FIG. 16 ), a data plane DMZ layer 1948 (e.g., data plane DMZ layer 1648 in FIG. 16 ), and a data plane data layer 1950 (e.g., data plane data layer 1650 in FIG. 16 ). Data plane DMZ layer 1948 may include a trusted app subnet 1960 (e.g., trusted app subnet 1860 in FIG. 18 ) and an untrusted app subnet 1962 (e.g., untrusted app subnet 1862 in FIG. 18 ) of data plane app layer 1946 and an LB subnet 1922 that may be communicatively coupled to an Internet gateway 1934 included in data plane VCN 1918. The trusted app subnet 1960 may be communicatively coupled to a service gateway 1936 included in the data plane VCN 1918, a NAT gateway 1938 included in the data plane VCN 1918, and a DB subnet 1930 included in the data plane data layer 1950. The untrusted app subnet 1962 may be communicatively coupled to the service gateway 1936 included in the data plane VCN 1918 and the DB subnet 1930 included in the data plane data layer 1950. The data plane data layer 1950 may include a DB subnet 1930 that may be communicatively coupled to the service gateway 1936 included in the data plane VCN 1918.
[0210] The untrusted app subnet 1962 may include primary VNICs 1964(1)-(N), which may be communicatively coupled to tenant virtual machines (VMs) 1966(1)-(N) residing in the untrusted app subnet 1962. Each tenant VM 1966(1)-(N) may execute code in a respective container 1967(1)-(N), which may be communicatively coupled to an app subnet 1926, which may be included in a data plane app layer 1946, which may be included in a container egress VCN 1968. Each secondary VNIC 1972(1)-(N) may facilitate communication between the untrusted app subnet 1962, which is included in the data plane VCN 1918, and the app subnet included in the container egress VCN 1968. The container egress VCN may include a NAT gateway 1938, which may be communicatively coupled to the public internet 1954 (e.g., public internet 1654 in FIG. 16 ).
[0211] The internet gateway 1934 included in the control plane VCN 1916 and the internet gateway 1934 included in the data plane VCN 1918 may be communicatively coupled to a metadata management service 1952 (e.g., metadata management system 1652 of FIG. 16 ), which may be communicatively coupled to the public internet 1954. The public internet 1954 may be communicatively coupled to the internet gateway 1934 included in the control plane VCN 1916 and the NAT gateway 1938 included in the data plane VCN 1918. The internet gateway 1934 included in the control plane VCN 1916 and the service gateway 1936 included in the data plane VCN 1918 may be communicatively coupled to cloud services 1956.
[0212] In some instances, the pattern illustrated by the architecture of block diagram 1900 of FIG. 19 may be considered an exception to the pattern illustrated by the architecture of block diagram 1800 of FIG. 18 and may be desirable for an IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected region). The customer can access each container 1967(1)-(N) contained in each customer's VM 1966(1)-(N) in real time. The containers 1967(1)-(N) are stored in the container transport. The containers 1967(1)-(N) may be configured to call respective secondary VNICs 1972(1)-(N) included in the app subnet 1926 of the data plane app layer 1946, which may be included in the control plane VCN 1968. The secondary VNICs 1972(1)-(N) may send calls to a NAT gateway 1938, which may send calls to the public Internet 1954. In this example, the containers 1967(1)-(N) that the customer can access in real time may be isolated from the control plane VCN 1916 and may be isolated from other entities included in the data plane VCN 1918. The containers 1967(1)-(N) may also be isolated from resources of other customers.
[0213] In another example, a customer can invoke cloud service 1956 using containers 1967(1)-(N). In this example, the customer can execute code in containers 1967(1)-(N) that requests a service from cloud service 1956. Containers 1967(1)-(N) can send the request to secondary VNICs 1972(1)-(N), which can send the request to a NAT gateway that can send the request to public internet 1954. Public internet 1954 can send the request to LB subnet 1922, which is included in control plane VCN 1916, via internet gateway 1934. In response to determining that the request is valid, LB subnet 1926 can send the request to app subnet 1926, which can send the request to cloud service 1956 via service gateway 1936.
[0214] It should be noted that the illustrated IaaS architectures 1600, 1700, 1800, and 1900 may include elements other than those shown. Furthermore, the illustrated embodiments are only examples of some of the cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS system may have more or fewer elements than those shown, may combine two or more elements, or may have a different configuration or arrangement of elements.
[0215] In certain embodiments, the IaaS system described herein may include a suite of application, middleware, and database services that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example of such an IaaS system is the Oracle® Cloud Infrastructure (OCI) offered by the present applicant.
[0216] 20 illustrates an exemplary computer system 2000 upon which various embodiments may be implemented. System 2000 may be used to implement any of the computer systems described above. As shown, computer system 2000 includes a processing unit 2004 that communicates with a number of peripheral subsystems via a bus subsystem 2002. These peripheral subsystems may include a processing acceleration unit 2006, an I / O subsystem 2008, a storage subsystem 2018, and a communications subsystem 2024. Storage subsystem 2018 includes a tangible computer-readable storage medium 2022 and a system memory 2010.
[0217] Bus subsystem 2002 provides a mechanism for allowing the various components and subsystems of computer system 2000 to communicate with each other as intended. While bus subsystem 2002 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2002 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Microchannel bus, or a Serial ATA bus. These buses may include the Peripheral Component Interconnect (PCI) bus, which may be implemented as a Mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, the Multi-Core Architecture (MCA), the Bus Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus, which may be implemented as a Mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0218] Processing unit 2004, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 2000. Processing unit 2004 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 2004 may be implemented as one or more independent processing units 2032 and / or 2034, with a single-core or multi-core processor included in each processing unit. In other embodiments, processing unit 2004 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0219] In various embodiments, processing unit 2004 may execute various programs in response to program code and may maintain multiple programs or processes executing simultaneously. At any given time, some or all of the program code being executed may reside in processor 2004 and / or storage subsystem 2018. Processor 2004, through appropriate programming, may provide the various functionality described above. Computer system 2000 may further include a processing acceleration unit 2006, which may include a digital signal processor (DSP), a special purpose processor, and / or the like.
[0220] The I / O subsystem 2008 can include user interface input devices and user interface output devices. User interface input devices can include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, dials, buttons, switches, keypads, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices can also include motion sensing and / or gesture recognition devices, such as a Microsoft Kinect® motion sensor, which can provide input via a natural user interface (NUI) that utilizes gestures and voice commands, such as a Microsoft Xbox® 360 game controller. The user interface input device can control and interact with a force device. The user interface input device can also include an eye gesture recognition device, such as a Google Glass® blink detector. The Google Glass® blink detector detects a user's eye activity (e.g., "blinks" when taking a photo and / or selecting a menu) and converts the eye activity into input for input into the input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition detection device that enables a user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0221] Additionally, user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, graphics tablets, audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser range finders, and eye-tracking devices. Additionally, user interface input devices may include, for example, medical imaging devices such as computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, ultrasound emission tomography (EMC) scanners, and medical ultrasound scanners. The user interface input devices may also include audio input devices such as MIDI keyboards and electronic musical instruments.
[0222] User interface output devices may also include non-visual displays such as display subsystems, indicator lights, or audio output devices. The display subsystem may be, for example, a flat-panel device using a cathode ray tube (CRT), liquid crystal display (LCD), or plasma display, a projection device, or a touchscreen. In general, when the term "output device" is used, it is intended to include all possible types of devices and mechanisms for outputting information from computer system 2000 to a user or to another computer. For example, user interface output devices include, but are not limited to, various display devices that visually convey text, images, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, voice output devices, and modems. Computer system 2000 may include a storage subsystem 2018. The storage subsystem 2018 comprises software elements, which are illustratively located in system memory 2010. The system memory 2010 may store program instructions loadable and executable by the processing unit 2004, as well as data generated by the execution of these programs.
[0223] Depending on the configuration and type of computer system 2000, system memory 2010 may be volatile memory (e.g., random access memory (RAM)) and / or non-volatile memory (e.g., read-only memory (ROM), flash memory). Generally, RAM contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by processing unit 2004. In some implementations, system memory 2010 may be static random access memory (SRAM). In some implementations, the computer system 2000 may include a number of different types of memory, such as a basic input / output system (BIOS), which contains the basic routines that help to transfer information between elements within the computer system 2000, such as during start-up. By way of example and not limitation, the system memory 2010 may be configured to store a client application, a web browser, a mid-tier application, a relational database management system (RDBMS), Also shown are application programs 2012, program data 2014, and an operating system 2016, which may include, for example, Microsoft Windows, Apple Macintosh, and and / or various versions of the Linux operating system, various commercially available UNIX or UNIX-like operating systems (including various GNU / Linux operating systems, Google Chrome (registered trademark) (Registered trademark) OS, etc.), and / or iOS, Windows Phone, Android OS, BlackBerry (Registered trademark) Mobile devices such as Palm OS and Palm OS operating systems The system may include a file operating system.
[0224] Additionally, the storage subsystem 2018 can provide a tangible, computer-readable storage medium for storing the basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that, when executed by the processor, provide the above-described functionality may be stored in the storage subsystem 2018. These software modules or instructions are then executed by the processing unit 2004. The storage subsystem 2018 may also provide a repository for storing data used in accordance with the present disclosure.
[0225] The storage subsystem 2010 may also include a computer readable storage medium reader 2020 further connectable to a computer readable storage medium 2022. The computer readable storage medium 2022 may comprehensively represent remote, local, fixed, and / or removable storage devices, as well as storage media for temporarily and / or permanently containing, storing, transmitting, and retrieving computer readable information together with, or in combination with, the system memory 2010 as needed.
[0226] Additionally, the computer-readable storage medium 2022 containing the code or portions of code may include any suitable medium known or used in the art, including storage and communication media such as, but not limited to, volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible computer-readable media. This may also include intangible computer-readable media such as data signals, data transmissions, or other media usable to transmit the desired information and accessible by computer system 2000.
[0227] By way of example, the computer-readable storage medium 2022 may be a hard disk drive that reads from or writes to non-removable, non-volatile magnetic media, a magnetic disk drive that reads from or writes to a removable, non-volatile magnetic disk, and a CD The computer-readable storage medium 2022 may include an optical disk drive that reads from or writes to removable, non-volatile optical disks such as ROM, DVD and Blu-ray disks or other optical media. The computer-readable storage medium 2022 may include, but is not limited to, Zip drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD disks, digital video tapes, etc. The computer-readable storage medium 2022 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, solid-state RAM, dynamic RAM, etc. SSDs may include SSDs based on volatile memory such as dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 2000.
[0228] The communications subsystem 2024 provides an interface with other computer systems and networks. The communications subsystem 2024 acts as an interface for receiving data from other systems and transmitting data from the computer system 2000 to other systems. For example, the communications subsystem 2024 may enable the computer system 2000 to connect to one or more devices over the Internet. In some embodiments, the communications subsystem 2024 may be connected to a network (e.g., 3G, 4G, or EDG). Radio frequency (RF) transceiver components for accessing wireless voice and / or data networks using cellular technologies, advanced data network technologies such as IEEE 1394 (enhanced data rates for global evolution), Wi-Fi (WiFi) 602.11 family of standards or other mobile communication technologies or any combination thereof), a global positioning system (GPS) receiver component, and In some embodiments, the communications subsystem 2024 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.
[0229] Additionally, in some embodiments, the communications subsystem 2024 may receive incoming communications in the form of structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc., on behalf of one or more users who may be using the computer system 2000.
[0230] As an example, the communications subsystem 2024 may provide Twitter feeds, Facebook updates, Rich Site Summary (RSS) feeds, and other similar services. The data feeds 2026 may be configured to receive data feeds, such as web feeds, in real time from users of social networks and / or other communication services, and / or to receive real-time updates from one or more third-party sources.
[0231] The communications subsystem 2024 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 2028 of real-time events that may be continuous or may be essentially unbounded with no clear ends and / or event updates 2030. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0232] The communications subsystem 2024 may also be configured to output structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc. to one or more databases that may communicate with one or more streaming data source computers coupled to the computer system 2000.
[0233] The computer system 2000 may be one of a variety of types, including a handheld mobile device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head-mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or other data processing system.
[0234] In the foregoing description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it will be apparent that various embodiments may be practiced without these specific details. The following description provides examples only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the embodiments will provide one of ordinary skill in the art with an enabling description for practicing the embodiments. It should be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the present disclosure, as defined in the appended claims. The drawings and descriptions are not intended to be limiting. Circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other embodiments, well-known circuits, processes, algorithms, structures, and The methods and techniques may be presented without unnecessary detail so as not to obscure the embodiments. The teachings of the present disclosure may also be applied to various types of applications, such as mobile applications, non-mobile applications, desktop applications, web applications, enterprise applications, etc. Additionally, the teachings of the present disclosure are not limited to a particular operating environment (e.g., operating system, device, platform, etc.) but may be applied to multiple different operating environments.
[0235] Also, note that each embodiment is described as a process that is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. While a flowchart describes operations as sequential processes, many operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when the operation is completed, but may include additional steps not shown in the figures. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the calling function or the main function.
[0236] The words "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" or "example" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0237] The terms “machine-readable storage medium” or “computer-readable storage medium” include, but are not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instructions and / or data. Machine-readable storage media or computer-readable storage media may also include non-transitory media capable of storing data and that do not involve carrier waves and / or transitory electronic signals propagating via wireless or wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer program product may include code and / or machine-executable instructions, which may represent any combination of procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, classes, instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by transferring and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be conveyed, forwarded, or transmitted via any suitable means such as memory sharing, message passing, token passing, network transmission, etc.
[0238] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments performing the necessary operations may be stored on a machine-readable medium. A processor may perform the necessary operations. The systems illustrated in some of the figures may be provided in various configurations. In some examples, the system may be configured as a distributed system in which one or more elements of the system are distributed across one or more networks in a cloud computing system. When an element is described as being "configured" to perform a particular operation, such configuration may be achieved by, for example, designing electronic circuitry or other hardware to perform the operation, by programming or controlling electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof. This may be achieved by
[0239] While specific embodiments of the present disclosure have been described, various modifications, variations, alternative configurations, and equivalents are encompassed within the scope of the present disclosure. The embodiments of the present disclosure are not limited to operating in a particular data processing environment, but can freely operate in multiple data processing environments. Furthermore, while embodiments of the present disclosure have been described using a particular series of actions and steps, it will be apparent to those skilled in the art that the scope of the present disclosure is not limited to the series of actions and steps described. Various features and aspects of the above-described embodiments can be used individually or jointly.
[0240] Furthermore, while embodiments of the present disclosure have been described using a particular combination of hardware and software, it should be appreciated that other combinations of hardware and software are within the scope of the present disclosure. Embodiments of the present disclosure may be implemented using only hardware, only software, or a combination thereof. The various processes described in this disclosure may run on the same processor or any combination of different processors. Thus, when a component or module is described as being configured to perform a particular process, that configuration may be achieved, for example, by designing electronic circuitry to perform the process, by programming a programmable electronic circuit (such as a microprocessor) to perform the process, or a combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication. Different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0241] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and changes may be made without departing from the broad spirit and scope defined by the appended claims. Accordingly, while specific embodiments of the present disclosure have been described, these embodiments are not intended to be limiting. Various modifications and equivalents thereof are intended to be encompassed within the scope of the appended claims.
[0242] The indefinite article "a" / used in the context of describing this disclosure (particularly in the context of the claims) "An," the definite article "the" and similar references are used herein unless otherwise stated or contained in the context. Unless otherwise clearly indicated above, the terms "comprising," "having," "including ... and "containing" should be construed as an open-ended term (i.e., meaning "including, but not limited to") unless otherwise specified. The term "connected" should be construed as partly or wholly contained within, attached to, or joined together, even if there is something intervening. The recitation of ranges of values herein is intended merely as a shorthand method of referring to each individual value falling within the range, and unless otherwise stated herein, each individual value is incorporated herein as if set forth individually herein. Unless otherwise stated herein or unless the context clearly indicates otherwise, all methods described herein can be performed in any suitable order. The use of any and all examples or exemplary language (e.g., "such as") herein is intended to further clarify embodiments of the present disclosure and does not limit the scope of the disclosure, unless otherwise specified. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the disclosure.
[0243] Disjunctive language, such as the phrase "at least one of X, Y, or Z," means that an item, term, etc. may be X, Y, or Z, or any combination thereof (unless otherwise specified). For example, disjunctive language is intended to be understood in context as being generally used to indicate that any of X, Y, and / or Z may be present. Thus, such disjunctive language is not generally intended to, and does not imply, that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z be present.
[0244] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Those skilled in the art can employ such variations as appropriate, and the present disclosure may be practiced otherwise than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by this disclosure unless otherwise indicated herein.
[0245] All references cited herein, including publications, patent applications, and patents, are incorporated by reference to the same extent as if each individual reference was individually and specifically indicated to be incorporated by reference and was set forth in its entirety herein.
[0246] In the foregoing specification, aspects of the disclosure have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure may be used individually or jointly. Moreover, the embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense.< / realm>
Claims
1. 1. A method comprising: a network virtualization device receiving a first frame via a first port of the network virtualization device, the first frame including a first medium access control (MAC) address of a first compute instance that is a destination of the first frame, a second MAC address of a second compute instance that is a source of the first frame, and a Layer 2 (L2) protocol data unit (PDU), the first compute instance and the second compute instance being members of a virtual L2 network, the first compute instance being hosted by a first host machine connected to the network virtualization device via a second port of the network virtualization device; the network virtualization device determining that a loop prevention rule prevents transmission of the frame using a port on which the frame was received; determining, by the network virtualization device, based on the second MAC address, that the first frame should be transmitted through all ports of the network virtualization device; the network virtualization device transmitting the first frame through all ports of the network virtualization device except for the first port based on the loop prevention rule.
2. 2. The method of claim 1, wherein the network virtualization appliance is remote from the first host machine and connected to a network interface card of the first host machine via the second port.
3. the network virtualization appliance transmitting a bridge protocol data unit (BPDU) to the first compute instance via the second port; receiving, by the network virtualization device, the BPDU from the first compute instance via the second port in response to transmitting the BPDU; 10. The method of claim 1, further comprising: in response to receiving the BPDU, the network virtualization device determining that a loop exists between the network virtualization device and the first compute instance.
4. in response to determining that the loop exists between the network virtualization device and the first compute instance, the network virtualization device disables the second port; receiving, by the network virtualization device, a second frame including the first MAC address as a destination address; The method of claim 3 , further comprising: the network virtualization appliance preventing transmission of the second frame through the second port.
5. the network virtualization appliance sending a bridge protocol data unit (BPDU) to the first compute instance via the second port; determining, by the network virtualization device, that no BPDUs have been received from the first computing instance; the network virtualization device determining that no loop exists between the network virtualization device and the first compute instance; receiving, by the network virtualization device, a second frame including the first MAC address as a destination address; the network virtualization appliance, upon determining that no loop exists, transmits the second frame to the first compute instance via the second port; and The method of claim 1 further comprising:
6. the network virtualization device determining that the first MAC address is not included in a forwarding table of the network virtualization device; The method of claim 1 , further comprising: the network virtualization device broadcasting the frame over ports of the network virtualization device other than the first port.
7. the network virtualization device determining that the second MAC address is not included in the forwarding table; the network virtualization device updating the forwarding table by including in the forwarding table an association between at least the second MAC address and the second port; receiving, by the network virtualization device, a second frame via the second port, the second frame including the second MAC address as a destination address; 7. The method of claim 6, further comprising: the network virtualization device transmitting the second frame through the first port rather than another port of the network virtualization device based on the association in the forwarding table.
8. receiving, by the network virtualization device, a second frame including the second MAC address as a destination address; the network virtualization device determining that the first MAC address is not included in a forwarding table of the network virtualization device; The method of claim 1 , wherein the second frame is broadcast to other network virtualization devices connected to the network virtualization device via a switch network.
9. the network virtualization device is connected to other network virtualization devices via a switch network; the first frame is received from the other network virtualization device via the switch network; The method further includes determining that the first MAC address is not included in a forwarding table of the network virtualization device; The method of claim 1 , wherein the first frame is broadcast to compute instances hosted on host machines connected to the network virtualization appliance.
10. A network virtualization device, a plurality of ports, the plurality of ports including a first port connected to a first host machine hosting a first compute instance, a second port connected to a second machine hosting a second compute instance, and a third port connected to a network switch, the first compute instance and the second compute instance being members of a virtual Layer 2 (L2) network; one or more processors; one or more memories storing instructions that, when executed by the one or more processors, configure the network virtualization device to perform the following operations: The operation is transmitting a first frame including a first L2 Bridge Protocol Data Unit (BPDU) to the first compute instance via the first port; receiving a second frame from the first compute instance via the first port; To do, determining that the second frame includes the first L2 BPDU; determining, based on the first L2 BPDU of the second frame, that a loop exists between the network virtualization device and the first compute instance.
11. Execution of the instructions further configures the network virtualization device to perform the following operations: The operation is Disabling the first port based on the loop; receiving a third frame including a MAC address of the first computing instance; and preventing transmission of the third frame through the first port.
12. Execution of the instructions further configures the network virtualization device to perform the following operations: The operation is transmitting a third frame including a second L2 BPDU to the second compute instance via the second port; determining that the second L2 BPDU is not received from the second computing instance; and and determining that no loop exists between the network virtualization device and the second compute instance.
13. Execution of the instructions further configures the network virtualization device to perform the following operations: The operation is receiving a fourth frame that includes the MAC address of the second computing instance as a destination address; and transmitting the fourth frame via the second port in response to determining that no loop exists.
14. Execution of the instructions further configures the network virtualization device to perform the following operations: The operation is receiving a third frame via the second port, the third frame including a first Medium Access Control (MAC) address of the second compute instance as a source address, a second MAC address of a third compute instance as a destination address, and an L2 Protocol Data Unit (PDU), the third compute instance being a member of the virtual L2 network; determining that a loop prevention rule prevents transmission of the frame using the port on which the frame was received; determining, based on the second MAC address, that the third frame should be transmitted through all ports of the network virtualization device; and transmitting the third frame through all ports of the network virtualization device except for the second port based on the loop prevention rule.
15. When executed on a network virtualization device, the network virtualization device performs the following operations: One or more non-transitory computer-readable storage media storing instructions for causing a computer to execute The operation is receiving a first frame via a first port of the network virtualization apparatus, the first frame including a first medium access control (MAC) address of a first compute instance that is a destination of the first frame, a second MAC address of a second compute instance that is a source of the first frame, and a Layer 2 (L2) protocol data unit (PDU), the first compute instance and the second compute instance being members of a virtual L2 network, the first compute instance being hosted by a first host machine connected to the network virtualization apparatus via a second port of the network virtualization apparatus; determining that a loop prevention rule prevents transmission of the frame using the port on which the frame was received; determining, based on the second MAC address, that the first frame should be transmitted through all ports of the network virtualization device; transmitting the first frame through all ports of the network virtualization device except for the first port based on the loop prevention rule; transmitting a second frame including a bridge protocol data unit (BPDU) to the first computing instance via the second port; determining, based on receipt of the BPDU from the first compute instance, that a loop exists between the network virtualization appliance and the first compute instance.
16. The operation is determining that the first MAC address is not included in a forwarding table of the network virtualization device; and broadcasting the frame over ports of the network virtualization device other than the first port.
17. The operation is determining that the second MAC address is not included in the forwarding table; updating the forwarding table by including an association between at least the second MAC address and the second port in the forwarding table; receiving a third frame via the second port, the third frame including the second MAC address as a destination address; 17. The one or more non-transitory computer-readable storage media of claim 16, further comprising: transmitting the second frame through the first port rather than another port of the network virtualization device based on the association in the forwarding table.
18. The operations further include determining that the first MAC address is not included in a forwarding table of the network virtualization device; 16. The one or more non-transitory computer-readable storage media of claim 15, wherein the first frame is broadcast to other network virtualization appliances connected to the network virtualization appliance via a switch network.
19. The operations further include determining that the first MAC address is not included in a forwarding table of the network virtualization device; 16. The one or more non-transitory computer-readable storage media of claim 15, wherein the first frame is broadcast to other compute instances hosted on host machines connected to the network virtualization appliance.
20. The operation is Disabling the second port based on the loop; and receiving a third frame including the first MAC address as a destination address; and preventing transmission of the third frame through the second port.