Implementing dual top-of-rack switches for customer-dedicated regional clouds
The DRCC framework addresses the challenges of migrating workloads to the cloud by deploying CSP infrastructure on-premises, providing secure, scalable, and cost-effective cloud services that meet regulatory and latency requirements, enhancing cloud-scale security and scalability for traditional workloads.
Patent Information
- Application Number
- JP2025508764
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-27
- Filing Date
- 2023-07-31
- Publication Date
- 2025-09-25
AI Technical Summary
Enterprises face challenges in migrating workloads to the cloud due to misalignment between traditional application architectures and cloud architectures, leading to high costs and difficulty in achieving regulatory compliance, data residency, and latency requirements, especially for workloads that cannot be migrated to the public cloud.
A Dedicated Region Cloud at Customer (DRCC) framework that deploys a cloud service provider's infrastructure in the customer's data center, allowing enterprises to consolidate applications and database systems, providing high-performance, secure, and scalable cloud services on-premises, with isolated data and control planes, and access to public cloud services.
The DRCC framework reduces infrastructure and operational costs while meeting stringent regulatory and latency requirements, delivering cloud-scale security and scalability, and enabling incremental modernization of traditional workloads with a fully managed experience.
Smart Images

Figure 2025531668000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of the filing dates of U.S. Provisional Patent Application No. 63 / 398,134, filed August 15, 2022, U.S. Provisional Patent Application No. 63 / 381,262, filed October 27, 2022, and U.S. Provisional Patent Application No. 18 / 360,680, filed July 27, 2023, the entire contents of each of which are incorporated herein by reference. [Background technology]
[0002] background Enterprises continue to migrate their business applications and databases to the cloud to reduce the costs of purchasing, updating, and maintaining on-premises hardware and software. High-performance computer applications consistently consume all available computing power to achieve a specific outcome or result. These applications require dedicated network performance, fast storage, high computing power, and large amounts of memory—resources that are in short supply in the virtualized infrastructure that makes up today's commodity clouds.
[0003] However, enterprises typically find migrating to cloud infrastructure costly and difficult due to the inherent misalignment between traditional application architectures and cloud architectures. These challenges are compounded for workloads that cannot be migrated to the public cloud. Enterprises only have access to a small subset of cloud services on-premises, with a limited set of capabilities and performance metrics compared to those available in the public cloud.
[0004] Therefore, there is a need for a framework that provides all the functional metrics of a public cloud on-premises, allowing enterprises to reduce infrastructure and operational costs while upgrading legacy applications onto modern cloud services to meet the most stringent regulatory, data residency, and latency requirements, all while relying on the cloud service provider's infrastructure to deliver high performance and the highest levels of security. The embodiments discussed herein address these and other problems. Summary of the Invention
[0005] overview The present disclosure relates generally to cloud networks. More particularly, the present disclosure relates to a Dedicated Region Cloud at Customer (DRCC), which corresponds to a cloud service provider's (CSP) infrastructure deployed in a customer's own data center. DRCC enables enterprises to easily consolidate applications and mission-critical database systems deployed on expensive hardware on a CSP's highly available and secure infrastructure, creating operational efficiencies and modernization opportunities.
[0006] The DRCC framework delivers all the functional metrics of the public cloud on-premises, allowing enterprises to reduce infrastructure and operational costs while upgrading legacy applications to modern cloud services to meet the most stringent regulatory, data residency, and latency requirements. All of this is powered by the CSP's infrastructure, which delivers high performance and the highest levels of security. Customers gain the choice and flexibility to run all of the CSP's cloud services in their own data centers. Customers can choose from all the public cloud services the CSP offers (including VMware Cloud, Autonomous Database, Container Engine for Kubernetes, Bare Metal Servers, and Exadata Cloud Service) and pay only for the services they consume. The DRCC framework is designed to completely isolate data and customer operations from the internet, and control and data plane operations remain on-premises, helping customers meet the most stringent compliance and latency requirements. With a fully managed experience and access to the new functional metrics now available in the public cloud, the DRCC framework delivers cloud-scale security, resiliency, and scalability while supporting mission-critical workloads with tools to incrementally modernize traditional workloads.
[0007] Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media that store programs, code, or instructions that may be executed by one or more processors, etc. These exemplary embodiments are mentioned not to limit or define the disclosure, but to provide examples to aid in its understanding. Additional embodiments are also discussed and detailed descriptions are provided in the detailed description.
[0008] One aspect of the present disclosure provides a method including: communicatively coupling a first physical port of a network virtualization device (NVD) included in a data center to a first top-of-rack (TOR) switch and a second TOR switch; communicatively coupling a second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a specific TOR from a group including the first TOR and the second TOR for communication of the packet; and the NVD transmitting the packet to the specific TOR to facilitate communication of the packet to a destination host machine.
[0009] Another aspect of the present disclosure provides a computing device comprising: a processor; and memory containing instructions that, when executed by the processor, cause the computing device to: communicatively couple a first physical port of a network virtualization device (NVD) included in at least a datacenter to a first top-of-rack (TOR) switch and a second TOR switch; communicatively couple the second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a specific TOR from a group including the first TOR and the second TOR for communication of the packet; and the NVD sending the packet to the specific TOR to facilitate communication of the packet to a destination host machine.
[0010] Another aspect of the present disclosure provides a non-transitory computer-readable medium storing certain computer-executable instructions that, when executed by a processor, cause a computer system to perform operations including: communicatively coupling a first physical port of a network virtualization device (NVD) included in a datacenter to a first top-of-rack (TOR) switch and a second TOR switch; communicatively coupling a second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a specific TOR from a group including the first TOR and the second TOR for communication of the packet; and the NVD transmitting the packet to the specific TOR to facilitate communication of the packet to a destination host machine.
[0011] The features, embodiments, and advantages of the present disclosure will be better understood from the following detailed description when taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a high-level diagram of a distributed environment illustrating a virtual or overlay cloud network hosted by a cloud service provider infrastructure, according to one embodiment. [Figure 2] FIG. 2 is a simplified architectural diagram of physical components in a physical network within CSPI, according to one embodiment. [Figure 3] FIG. 1 illustrates an exemplary configuration within CSPI in which a host machine is connected to multiple network virtualization devices (NVDs), according to an embodiment. [Figure 4] FIG. 1 illustrates a connection between a host machine and an NVD to provide I / O virtualization that supports multi-tenancy, according to one embodiment. [Figure 5] FIG. 1 is a simplified block diagram of a physical network provided by CSPI, according to one embodiment. [Figure 6]FIG. 1 illustrates an arrangement of multiple TORs contained in a rack, according to at least one embodiment. [Figure 7] FIG. 10 illustrates another configuration of multiple TORs included in a rack, according to some embodiments. [Figure 8] 1 is a flowchart illustrating steps performed in providing an availability domain to a customer, according to at least one embodiment. [Figure 9] 1 illustrates an exemplary infrastructure of a DRCC, according to some embodiments. [Figure 10] 1 is a flowchart illustrating a process for providing a DRCC at a customer's on-premises location, according to some embodiments. [Figure 11A] 1 is a flowchart illustrating steps performed in transmitting a packet from a compute host included in a data center to a remote host outside the data center, according to some embodiments. [Figure 11B] 1 is a flowchart illustrating steps performed in transmitting a packet from a remote host outside a data center to a compute host included in the data center, according to some embodiments. [Figure 12] FIG. 1 illustrates another exemplary infrastructure of a DRCC, according to some embodiments. [Figure 13] 10 is a flowchart illustrating another process for providing a DRCC at a customer's on-premises location, according to some embodiments. [Figure 14] FIG. 1 illustrates an exemplary network fabric architecture for a DRCC, according to some embodiments. [Figure 15] FIG. 2 illustrates connections within the NFAB block, as well as connections between the NFAB block and multiple blocks of switches, according to some embodiments. [Figure 16] FIG. 2 illustrates an exemplary dedicated backbone network for a customer region, according to some embodiments. [Figure 17]1 is a flowchart illustrating a process for configuring a network fabric according to some embodiments. [Figure 18] FIG. 1 is a block diagram illustrating a pattern for implementing a service-based cloud infrastructure system, according to at least one embodiment. [Figure 19] FIG. 1 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, according to at least one embodiment. [Figure 20] FIG. 1 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, according to at least one embodiment. [Figure 21] FIG. 1 is a block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, according to at least one embodiment. [Figure 22] FIG. 1 is a block diagram illustrating an exemplary computer system according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The drawings and this description are not intended to be limiting. Use of the word "exemplary" herein means "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0014] Example architecture of cloud infrastructure The term cloud service is generally used to describe services made available on demand (e.g., through a subscription model) by a cloud service provider (CSP) to users or customers through the use of systems and infrastructure (cloud infrastructure) provided by the CSP. Typically, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premise servers and systems. This allows customers to use cloud services provided by CSPs without having to purchase separate hardware and software resources for the service. Cloud services are designed to provide subscribing customers with easy and scalable access to applications and computing resources without requiring the customer to invest in procuring the infrastructure used to deliver the service.
[0015] There are multiple cloud service providers offering different types of cloud services, which come in a variety of different types or models, such as Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS).
[0016] A customer may sign up for one or more cloud services offered by a CSP. A customer can be any entity, such as an individual, an organization, or a business. When a customer signs up or subscribes to a service offered by a CSP, a tenancy or account is created for the customer. This account allows the customer to access one or more registered cloud resources associated with the account.
[0017] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing service. In the IaaS model, a CSP provides infrastructure (called Cloud Service Provider Infrastructure, or CSPI) that customers can use to build their own customizable networks and deploy their resources. In this way, customer resources and networks are hosted in a distributed environment by the CSP-provided infrastructure. This differs from traditional computing, where customer resources and networks are hosted by customer-provided infrastructure.
[0018] CSPI may include interconnected high-performance compute resources, including various host machines, memory resources, and network resources, that make up a physical network, also referred to as an infrastructure network or base network. CSPI resources may be distributed across one or more data centers, which may be geographically dispersed across one or more geographic regions. These physical resources may run virtualization software to provide a virtualized distributed environment. Virtualization creates an overlay network (also known as a software-based network, software-defined network, or virtual network) on the physical network. The CSPI physical network serves as the basis for creating one or more overlay or virtual networks on the physical network. The virtual or overlay network may include one or more virtual cloud networks (VCNs). Virtual networks are realized through the use of software virtualization technologies (e.g., hypervisors, functions performed by network virtualization devices (NVDs) (e.g., smart NICs), top-of-rack (TOR) switches, smart TORs that implement one or more functions performed by the NVDs, and other mechanisms) to create a network abstraction layer that can operate on the physical network. Virtual networks can take many forms, including peer-to-peer networks, IP networks, etc. The virtual network is typically a Layer 3 IP network or a Layer 2 VLAN. This method of virtual or overlay networking is often referred to as virtual or overlay Layer 3 networking.Examples of protocols developed for virtual networks include IP-in-IP (or Generic Routing Encapsulation (GRE)), VXLAN-IETF RFC 7348 (Virtual Extensible LAN), VPN (Virtual Private Network) (e.g., MPLS Layer-3 Virtual Private Networks (RFC 4364)), VMware's NSX, and GENEVE (Generic Network Virtualization Encapsulation).
[0019] In IaaS, a CSP-provided infrastructure (CSPI) can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing service provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may also provide various services (e.g., billing, monitoring, logging, security, load balancing, clustering, etc.) associated with these infrastructure components. These services can be policy-driven, enabling IaaS users to maintain application availability and performance by implementing policies that drive load balancing. CSPI provides infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a hosted, highly available, distributed environment. CSPI delivers high-performance compute resources and capabilities, as well as storage capacity, in a flexible virtual network that can be securely accessed from various network locations, such as the customer's on-premises network. When a customer registers or subscribes to an IaaS service offered by a CSP, a tenancy is created for that customer, which is an isolated and secure partition within CSP where the customer may create, organize, and manage their cloud resources.
[0020] Customers can build their own virtual networks using compute, memory, and network resources provided by CSPI. One or more customer resources or workloads, such as compute instances, can be deployed onto these virtual networks. For example, customers can build one or more customizable private virtual networks (referred to as virtual cloud networks (VCNs)) using resources provided by CSPI. Customers can deploy one or more customer resources, such as compute instances, into the customer VCN. Compute instances can take the form of virtual machines, bare metal instances, etc. In this way, CSPI provides infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a hosted, highly available virtual environment. While customers do not manage or control the underlying physical resources provided by CSPI, they do control the operating system, storage, and deployed applications. Customers may have limited control over some networking components (e.g., firewalls).
[0021] The CSP may provide a console that enables customers and network administrators to use CSPI resources to configure, access, and manage resources deployed in the cloud. In some embodiments, the console provides a web-based user interface that can be used to access and manage the CSPI. In some implementations, the console is a web-based application provided by the CSP.
[0022] CSPI may support single-tenancy or multi-tenancy architectures. In a single-tenancy architecture, software (e.g., applications, databases) or hardware components (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenancy architecture, software or hardware components serve multiple customers or tenants. Thus, in a multi-tenancy architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenancy situation, precautions and safeguards are taken within CSPI to ensure that each tenant's data is isolated and not visible to other tenants.
[0023] In a physical network, a network endpoint ("endpoint") represents a computing device or system that is connected to the physical network and communicates with the connected network. A network endpoint of a physical network may be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints in a physical network include modems, hubs, bridges, switches, routers, and other networking devices, physical computers (or host machines), etc. Each physical device in a physical network has a fixed network address that can be used to communicate with the device. This fixed network address can be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints can include various virtual endpoints, such as virtual machines hosted by components of the physical network (e.g., hosted by a physical host machine). These virtual network endpoints are addressed by overlay addresses, such as overlay Layer 2 addresses (e.g., an overlay MAC address) and overlay Layer 3 addresses (e.g., an overlay IP address). Network overlays provide flexibility by allowing network administrators to move overlay addresses associated with network endpoints using software management (e.g., via software implementing the virtual network's control plane). Thus, unlike physical networks, in virtual networks, overlay addresses (e.g., overlay IP addresses) may be moved from one endpoint to another through the use of network management software. Because virtual networks are built on top of physical networks, communication between components of a virtual network involves both the virtual network and the underlying physical network.To facilitate such communications, CSPI components are configured to learn and store mappings from overlay addresses in the virtual network to actual physical addresses in the underlying network, and vice versa, and use these mappings to facilitate communications and encapsulate customer traffic to facilitate routing in the virtual network.
[0024] Thus, physical addresses (e.g., physical IP addresses) are associated with components of a physical network, and overlay addresses (e.g., overlay IP addresses) are associated with entities of a virtual network. Both physical and overlay IP addresses are types of real IP addresses. They are distinct from virtual IP addresses, which map to multiple real IP addresses. A virtual IP address provides a one-to-many mapping between the virtual IP address and multiple real IP addresses.
[0025] Cloud infrastructure or CSPI is physically hosted in one or more data centers in one or more regions around the world. CSPI may include physical or underlying network components as well as virtualized components (e.g., virtual networks, compute instances, virtual machines, etc.) in a virtual network built on the physical network components. In one embodiment, CSPI is organized and hosted in realms, regions, and availability domains. A region is typically a local geographic area that includes one or more data centers. Regions are generally independent of each other and may be separated by large distances, for example, across countries or continents. For example, one region may be in Australia, another in Japan, and yet another in India. CSPI resources are divided among the regions so that each region has its own independent subset of CSPI resources. Each region may provide a set of core infrastructure services and resources, such as compute resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure), storage resources (e.g., block volume storage, file storage, object storage, archive storage), networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connectivity to on-premises networks), database resources, edge networking resources (e.g., DNS), and access management and monitoring resources. Each region typically has multiple paths connecting it to other regions in the realm.
[0026] Typically, applications are deployed to the region where they are most heavily used (i.e., deployed to the infrastructure associated with that region) because using nearby resources is faster than using resources that are farther away. Applications may also be deployed to different regions for a variety of reasons, such as redundancy to mitigate the risk of region-wide events such as large-scale weather systems or earthquakes, or to meet various requirements such as jurisdictions, tax areas, and other business or societal standards.
[0027] Data centers within a region can be further organized and subdivided into availability domains (ADs). An availability domain can correspond to one or more data centers located within a region. A region can be composed of one or more availability domains. In such a distributed environment, CSPI resources are region-specific, such as a virtual cloud network (VCN), or availability domain-specific, such as a compute instance.
[0028] ADs within a region are isolated from each other, fault-tolerant, and designed to minimize the likelihood of simultaneous failures. This is achieved by ADs not sharing critical infrastructure resources such as networks, physical cables, cable paths, or cable entry points. Therefore, a failure in one AD within a region is unlikely to affect the availability of other ADs in the same region. ADs within the same region can be connected to each other via low-latency, high-bandwidth networks, enabling highly available connections to other networks (e.g., the Internet, customer on-premises networks, etc.). This allows for the creation of replica systems in multiple ADs for both high availability and disaster recovery. Cloud services use multiple ADs to ensure high availability and protect against resource failures. As the infrastructure offered by an IaaS provider grows, capacity can be added to more regions and ADs. Traffic between availability domains is typically encrypted.
[0029] In one embodiment, regions are grouped into realms. A realm is a logical collection of regions. Realms are isolated from each other and do not share any data. Regions in the same realm can communicate with each other, but regions in different realms cannot. A customer's tenancy or account using a CSP exists in a single realm and can be distributed across one or more regions belonging to that realm. Typically, when a customer signs up for an IaaS service, a tenancy or account for that customer is created in a customer-specific region (called the "home" region) within the realm. The customer can extend their tenancy across one or more other regions within the realm. A customer cannot access regions that are not in the realm in which their tenancy resides.
[0030] An IaaS provider may offer multiple realms, each corresponding to a particular set of customers or users. For example, a commercial realm may be offered for commercial customers. As another example, a realm may be offered for customers in a particular country. As yet another example, a government realm may be offered for governments, and so on. For example, a government realm may cater to a particular government and provide a higher level of security than a commercial realm. For example, OCI (Oracle Cloud Infrastructure) currently offers one realm for its commercial regions and two realms for its government cloud regions (e.g., FedRAMP and IL5 certified).
[0031] In one embodiment, an AD can be subdivided into one or more fault domains. A fault domain is a group of infrastructure resources within an AD that provides anti-affinity. Fault domains can distribute compute instances so that they do not reside on the same physical hardware within a single AD. This is known as anti-affinity. A fault domain represents a set of hardware components (computers, switches, etc.) that share a single point of failure. A compute pool is logically divided into fault domains. This ensures that a hardware failure or compute hardware maintenance event affecting one fault domain does not affect instances in other fault domains. Depending on the embodiment, the number of fault domains per AD can vary. For example, in one embodiment, each AD contains three fault domains. Fault domains act as logical data centers within an AD.
[0032] When a customer subscribes to an IaaS service, resources from CSPI are provisioned for the customer and associated with the customer's tenancy. The customer can use these provisioned resources to build private networks and deploy resources on these networks. A customer network hosted in the cloud by CSPI is called a virtual cloud network (VCN). A customer can configure one or more virtual cloud networks (VCNs) using the allocated CSPI resources. A VCN is a virtual or software-defined private network. Customer resources deployed in a customer's VCN may include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances may represent various customer workloads, such as applications, load balancers, databases, etc. Compute instances deployed in a VCN can communicate with public endpoints (“public endpoints”) accessible over a public network such as the Internet, other instances in the same VCN or other VCNs (e.g., other VCNs of the customer or VCNs not belonging to the customer), the customer's on-premises data center or network, service endpoints, and other types of endpoints.
[0033] CSPs can use CSPI to provide various services. In some cases, customers of a CSPI themselves can act as service providers and provide services using CSPI resources. Service providers may expose service endpoints characterized by identifying information (e.g., IP addresses, DNS names, and ports). Customer resources (e.g., compute instances) can consume a particular service by accessing the service endpoints exposed by that service. These service endpoints are generally publicly accessible to users over a public communications network, such as the Internet, using a public IP address associated with the endpoint. Publicly accessible network endpoints are also sometimes referred to as public endpoints.
[0034] In some embodiments, a service provider may expose a service via a service endpoint (sometimes referred to as a service endpoint). Customers of the service may then access the service using the service endpoint. In some embodiments, a given service endpoint for a service may be accessible by multiple customers intending to consume the service. In other embodiments, a customer may be given a dedicated service endpoint such that only that customer may access the service using the dedicated service endpoint.
[0035] In one embodiment, when a VCN is created, it is associated with a private overlay Classless Inter-Domain Routing (CIDR) address space (e.g., 10.0 / 16), which are broad private overlay IP addresses assigned to the VCN. A VCN includes associated subnets, route tables, and gateways. A VCN exists within a single region but can span one, more, or all availability domains in the region. A gateway is a virtual interface configured for a VCN that enables traffic to communicate between the VCN and one or more endpoints outside the VCN. One or more different types of gateways may be configured for a VCN to enable communication between different types of endpoints.
[0036] A VCN may be subdivided into one or more subnetworks, such as one or more subnets. A subnet is thus a configuration unit or subdivision that may be created within a VCN. A VCN may have one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that do not overlap with other subnets in the VCN and represent a subset of address space within the VCN's address space.
[0037] Each compute instance is associated with a virtual network interface card (VNIC) that enables the compute instance to participate in a subnet of a VCN. A VNIC is a logical representation of a physical network interface card (NIC). Generally, a VNIC is an interface between an entity (e.g., a compute instance, a service) and a virtual network. A VNIC resides in a subnet and has one or more associated IP addresses and associated security rules or policies. A VNIC corresponds to a Layer 2 port on a switch. A VNIC is associated with a compute instance and a subnet within a VCN. A VNIC associated with a compute instance enables the compute instance to be part of a subnet of a VCN and enables the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, endpoints in a different subnet of the VCN, or endpoints outside the VCN. Thus, a VNIC associated with a compute instance determines how the compute instance connects to endpoints inside and outside the VCN. A VNIC for a compute instance is created and associated with the compute instance when the compute instance is created and added to a subnet in a VCN. For a subnet containing a set of compute instances, the subnet contains VNICs corresponding to the set of compute instances, with each VNIC associated with one compute instance in the set of compute instances.
[0038] Each compute instance is assigned a private overlay IP address via the VNIC associated with that compute instance. This private overlay IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic to the compute instance. In a given subnet, all VNICs use the same route table, security list, and DHCP options. As described above, each subnet in a VCN is associated with a contiguous range of overlay IP addresses that do not overlap with other subnets in the VCN and represent a subset of address space within the VCN's address space (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24). For a VNIC on a particular subnet of a VCN, the private overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.
[0039] In some embodiments, a compute instance may be assigned a private overlay IP address and optionally additional overlay IP addresses, such as one or more public IP addresses in the case of a public subnet. These addresses are assigned to the same VNIC or multiple VNICs associated with the compute instance. However, each instance has a primary VNIC associated with an overlay private IP address that is created and assigned to the instance when the instance is launched (this primary VNIC cannot be deleted). Additional VNICs, called secondary VNICs, can be added to an existing instance in the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. Secondary VNICs can be in the same subnetwork as the primary VNIC, or in different subnetworks in the same or a different VCN.
[0040] If a compute instance resides in a public subnet, it may optionally be assigned a public IP address. A subnet can be designated as a public or private subnet at the time of creation. A private subnet means that resources (e.g., compute instances) and associated VNICs in the subnet cannot have public overlay IP addresses. A public subnet means that resources and associated VNICs in the subnet can have public IP addresses. Customers can specify a subnet to reside in a single availability domain or multiple availability domains of a region or realm.
[0041] As described above, a VCN can be subdivided into one or more subnets. In one embodiment, a virtual router (VR) (referred to as a VCN VR or simply VR) configured for a VCN enables communication between subnets in the VCN. For a subnet within a VCN, the VR represents the subnet's logical gateway, allowing that subnet (i.e., compute instances on that subnet) to communicate with endpoints on other subnets within the VCN and with other endpoints outside the VCN. A VCN VR is a logical entity configured to route traffic between VNICs in a VCN and a virtual gateway ("gateway") associated with the VCN. Gateways are described separately below with respect to FIG. 1. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is one VCN VR for a VCN, and a VCN VR potentially has a limited number of ports addressed by IP addresses, one port per subnet in the VCN. As such, a VCN VR has a different IP address for each subnet in the VCN to which the VCN VR is attached. VRs are also connected to various gateways configured for the VCN. In some embodiments, a specific overlay IP address in an overlay IP address range for a subnet is reserved for a port in the VCN VR for that subnet. For example, consider two subnets in a VCN, each with an associated address range of 10.0 / 16 and 10.1 / 16. For the first subnet in the VCN with address range 10.0 / 16, an address in this range is reserved for a port in the VCN VR for that subnet. In some cases, the first IP address in this range may be reserved for a VCN VR. For example, for a subnet with overlay IP address range 10.0 / 16, IP address 10.0.0.1 may be reserved for a port in the VCN VR for that subnet. For a second subnet in the same VCN with address range 10.1 / 16, the VCN VR may have a port with IP address 10.1.0.1 for the second subnet.A VCN VR has a different IP address for each of the VCN's subnets.
[0042] In some other embodiments, each subnet within a VCN may have its own associated VR, which is addressable by the subnet through use of a reserved or default IP address associated with the VR. The reserved or default IP address may, for example, be the first IP address in a range of IP addresses associated with the subnet. VNICs in the subnet can communicate with (e.g., send and receive packets from) the VR associated with the subnet using this default or reserved IP address. In such embodiments, the VR is the ingress / egress point for the subnet. VRs associated with a subnet within a VCN can communicate with other VRs associated with other subnets in the VCN. VRs can also communicate with gateways associated with the VCN. The VR functions for a subnet are running on or performed by one or more NVDs that perform the VNIC functions for VNICs in the subnet.
[0043] A VCN can be configured with route tables, security rules, and DHCP options. Route tables are virtual route tables for a VCN and contain rules that route traffic from subnets within the VCN to destinations outside the VCN via gateways or specially configured instances. A VCN's route tables can be customized to control how packets are forwarded / routed to the VCN. DHCP options represent configuration information that is automatically provided to instances at launch.
[0044] Security rules configured for a VCN represent the overlay firewall rules for the VCN. Security rules include ingress and egress rules and can specify the type of traffic (e.g., based on protocol and port) that can enter and exit instances in the VCN. Customers can choose whether a given rule is stateful or stateless. For example, a customer can allow ingress SSH traffic from anywhere to a set of instances by configuring a stateful ingress rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22. Security rules can be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to the resources in that group. On the other hand, a security list contains rules that apply to all resources in any subnet that uses the security list. A VCN can be provided with a default security list that contains default security rules. DHCP options configured for a VCN provide configuration information that is automatically given to instances in the VCN at launch.
[0045] In one embodiment, configuration information for a VCN is determined and stored by a VCN control plane. The VCN configuration information may include, for example, information about address ranges associated with the VCN, subnets and related information within the VCN, one or more VRs associated with the VCN, compute instances and associated VNICs in the VCN, NVDs performing various virtualized network functions (e.g., VNICs, VRs, gateways) associated with the VCN, VCN state information, and other VCN-related information. In one embodiment, a VCN distribution service publishes the configuration information stored by the VCN control plane or portions thereof to the NVD. The distributed information is stored by the NVD and may be used to update information (e.g., forwarding tables, routing tables, etc.) used to forward packets to compute instances in the VCN.
[0046] In one embodiment, VCN and subnet creation is handled by a VCN control plane (CP), and compute instance launch is handled by the compute control plane. The compute control plane is responsible for allocating physical resources for the compute instance and then calls the VCN control plane to create and associate VNICs with the compute instance. The VCN CP also sends VCN data mappings to a VCN data plane, which is configured to perform packet forwarding and routing functions. In one embodiment, a VCN VP provides a distribution service responsible for providing updates to the VCN data plane. Examples of VCN control planes are also shown in Figures 18, 19, 20, and 21 (see reference numerals 1816, 1916, 2016, and 2116) and described below.
[0047] Customers can create one or more VCNs using resources hosted by CSPI. Compute instances deployed in a customer VCN can communicate with various endpoints. These endpoints can include CSPI-hosted endpoints and endpoints external to CSPI.
[0048] Various architectures for implementing cloud-based services using CSPI are shown in Figures 1, 2, 3, 4, 5, 18, 19, 20, and 21 and are described below. Figure 1 is a high-level diagram of a distributed environment 100 illustrating an overlay or customer VCN hosted by CSPI, according to one embodiment. The distributed environment shown in Figure 1 includes multiple components of an overlay network. The distributed environment 100 shown in Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, alternatives, and modifications are possible. For example, in some embodiments, the distributed environment shown in Figure 1 may include more or fewer systems or components than those shown in Figure 1, may combine two or more systems, or may have different system configurations or arrangements.
[0049] As shown in the example of FIG. 1 , distributed environment 100 includes CSPI 101, which provides services and resources that customers can use to sign up and build their own virtual cloud networks (VCNs). In one embodiment, CSPI 101 provides IaaS services to its subscribing customers. Data centers within CSPI 101 may be organized as one or more regions. One exemplary region, “Region US” 102, is shown in FIG. 1 . A customer has configured a customer VCN 104 for region 102. The customer may deploy various compute instances into VCN 104, which may include virtual machines or bare metal instances. Example instances include applications, databases, load balancers, etc.
[0050] In the embodiment shown in FIG. 1 , customer VCN 104 includes two subnets, “Subnet 1” and “Subnet 2,” each with its own CIDR IP address range. In FIG. 1 , Subnet 1’s overlay IP address range is 10.0 / 16, and Subnet 2’s address range is 10.1 / 16. VCN virtual router 105 represents the VCN’s logical gateway, enabling communication between subnets in VCN 104 and with other endpoints outside the VCN. VCN VR 105 is configured to route traffic between VNICs in VCN 104 and the gateway associated with VCN 104. VCN VR 105 provides a port for each subnet in VCN 104. For example, VR 105 may provide a port with IP address 10.0.0.1 to Subnet 1 and a port with IP address 10.1.0.1 to Subnet 2.
[0051] Multiple compute instances may be deployed in each subnet, and the compute instances may be virtual machine instances and / or bare metal instances. The compute instances in a subnet may be hosted by one or more host machines within CSPI 101. A compute instance joins a subnet via a VNIC associated with the compute instance. For example, as shown in FIG. 1, compute instance C1 is part of subnet 1 via a VNIC associated with the compute instance. Similarly, compute instance C2 is part of subnet 1 via a VNIC associated with C2. Similarly, multiple compute instances (which may be virtual machine instances or bare metal instances) may be part of subnet 1. Each compute instance is assigned a private overlay IP address and a MAC address via its associated VNIC. For example, in Figure 1, compute instance C1 has an overlay IP address of 10.0.0.2 and a MAC address of M1, while compute instance C2 has a private overlay IP address of 10.0.0.3 and a MAC address of M2. Each compute instance in Subnet 1 (including compute instances C1 and C2) has a default route to VCN VR105 using IP address 10.0.0.1, which is the IP address of the port in VCN VR105 for Subnet 1.
[0052] Subnet2 may have multiple compute instances deployed, including virtual machine instances and / or bare metal instances. For example, as shown in FIG. 1, compute instances D1 and D2 are part of Subnet2 via VNICs associated with each of the compute instances. In the embodiment shown in FIG. 1, compute instance D1 has an overlay IP address of 10.1.0.2 and a MAC address of MM1, while compute instance D2 has a private overlay IP address of 10.1.0.3 and a MAC address of MM2. Each compute instance in Subnet2 (including compute instances D1 and D2) has a default route to VCN VR105 using IP address 10.1.0.1, which is the IP address of the port in VCN VR105 for Subnet2.
[0053] VCN A 104 may also include one or more load balancers. For example, a load balancer may be provided for a subnet and configured to load balance traffic among multiple compute instances on the subnet. A load balancer may also be provided to load balance traffic between subnets of the VCN.
[0054] A particular compute instance deployed in VCN 104 can communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints outside of CSPI 200. Endpoints hosted by CSPI 101 may include endpoints on the same subnet as the particular compute instance (e.g., communication between two compute instances in Subnet 1), endpoints on different subnets within the same VCN (e.g., communication between a compute instance in Subnet 1 and a compute instance in Subnet 2), endpoints in different VCNs in the same region (e.g., communication between a compute instance in Subnet 1 and an endpoint in a VCN in the same region 106 or 110, communication between a compute instance in Subnet 1 and an endpoint in the service network 110 in the same region), or endpoints in a VCN in a different region (e.g., communication between a compute instance in Subnet 1 and an endpoint in a VCN in a different region 108). Additionally, compute instances in a subnet hosted by CSPI 101 may communicate with endpoints not hosted by CSPI 101 (i.e., external to CSPI 101). These external endpoints include endpoints in a customer's on-premise network 116, endpoints in other hosted remote cloud networks 118, public endpoints 114 accessible over a public network such as the Internet, and other endpoints.
[0055] Communication between compute instances on the same subnet is facilitated through the use of VNICs associated with the source and destination compute instances. For example, compute instance C1 on Subnet 1 may want to send a packet to compute instance C2 on Subnet 1. For a packet originating from the source compute instance and destined for another compute instance in the same subnet, the packet is first processed by the VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance may include determining the packet's destination information from the packet header, identifying any policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the packet's next hop, performing encapsulation / decapsulation functions on the packet as needed, and forwarding / routing the packet to the next hop to facilitate communication of the packet to its intended destination. If the destination compute instance is in the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing, which then forwards the packet to the destination compute instance via the VNIC associated with the destination compute instance.
[0056] When a packet is sent from a compute instance in one subnet to an endpoint in a different subnet of the same VCN, this communication is facilitated by the VNICs and VCN VRs associated with the source and destination compute instances. For example, in Figure 1, if compute instance C1 in Subnet 1 wants to send a packet to compute instance D1 in Subnet 2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR using the default route or port 10.0.0.1 of VCN VR105. VCN VR105 is configured to route the packet to Subnet 2 using port 10.1.0.1. Once the packet is received and processed by the VNIC associated with D1, this VNIC forwards the packet to compute instance D1.
[0057] When a packet is sent from a compute instance in VCN 104 to an endpoint outside VCN 104, the communication is facilitated by the VNIC associated with the source compute instance, VCN VR 105, and a gateway associated with VCN 104. VCN 104 may have one or more types of gateways associated with it. A gateway is an interface between a VCN and another endpoint, where the other endpoint is outside the VCN. A gateway is a Layer 3 / IP layer concept that allows a VCN to communicate with endpoints outside the VCN. In this way, a gateway facilitates the flow of traffic between a VCN and other VCNs or networks. A VCN may be configured with various different types of gateways to facilitate different types of communication with different types of endpoints. Depending on the gateway, the communication may be over a public network (e.g., the Internet) or a private network. Various communication protocols may be used for these communications.
[0058] For example, compute instance C1 may want to communicate with an endpoint outside VCN 104. The packet may initially be processed by a VNIC associated with source compute instance C1. The VNIC's processing determines that the packet's destination is outside of Subnet 1 of C1. The VNIC associated with C1 may forward the packet to VCN VR105 in VCN 104. VCN VR105 then processes the packet and, as part of that processing, determines a particular gateway associated with VCN 104 as the packet's next hop based on the packet's destination. VCN VR105 may then forward the packet to the identified particular gateway. For example, if the destination is an endpoint in a customer's on-premises network, the packet may be forwarded by VCN VR105 to Dynamic Routing Gateway (DRG) gateway 122 configured for VCN 104. Forwarding the packet from the gateway to the next hop can then facilitate the packet's delivery to its ultimate destination.
[0059] A variety of different types of gateways may be configured for a VCN. Examples of gateways that may be configured for a VCN are shown in FIG. 1 and described below. Examples of gateways associated with VCNs are also shown in FIGS. 18, 19, 20, and 21 (e.g., gateways referenced by reference numbers 1834, 1836, 1838, 1934, 1936, 1938, 2034, 2036, 2038, 2134, 2136, and 2138) and described below. As shown in the embodiment of FIG. 1, a Dynamic Routing Gateway (DRG) 122 may be added to or associated with customer VCN 104 to provide a path for private network traffic communication between customer VCN 104 and another endpoint, which may be the customer's on-premises network 116, a VCN 108 in a different region of CSPI 101, or another remote cloud network 118 not hosted by CSPI 101. The customer on-premises network 116 may be a customer network built using customer resources or may be a customer data center. Access to the customer on-premises network 116 is typically highly restricted. If a customer has both the on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, they may want the on-premises network 116 and the cloud-based VCNs 104 to be able to communicate with each other. This allows the customer to build an extended hybrid environment that includes the on-premises network 116 and the customer VCNs 104 hosted by CSPI 101. This communication is enabled by the DRG 122. To enable such communication, a communication channel 124 is set up so that one endpoint of the channel resides in the customer on-premises network 116 and the other endpoint resides in CSPI 101 and is connected to the customer VCN 104. The communication channel 124 may traverse a public or private communication network, such as the Internet.A variety of different communication protocols may be used, such as IPsec VPN technology over a public communication network such as the Internet, or Oracle's FastConnect technology, which uses a private network instead of a public network. The device or equipment in the customer on-premises network 116 that forms one endpoint of the communication channel 124 is referred to as customer premises equipment (CPE), such as CPE 126 shown in Figure 1. On the CSPI 101 side, the endpoint may be a host machine running DRG 122.
[0060] In one embodiment, remote peering connections (RPCs) can be added to a DRG, allowing customers to peer one VCN with another VCN in a different region. Using such RPCs, a customer VCN 104 can connect to a VCN 108 in another region using a DRG 122. The DRG 122 can also be used to communicate with other remote cloud networks 118 not hosted by CSPI 101, such as the Microsoft Azure cloud or the Amazon AWS cloud.
[0061] 1, an Internet Gateway (IGW) 120 may be configured for customer VCN 104, allowing compute instances on VCN 104 to communicate with public endpoints 114 accessible over a public network, such as the Internet. IGW 120 is a gateway that connects a VCN to a public network, such as the Internet. IGW 120 allows a public subnet within a VCN, such as VCN 104, to directly access a public endpoint 112 on public network 114, such as the Internet, if the resources in that public subnet have public overlay IP addresses. Using IGW 120, connections can be initiated from a subnet within VCN 104 or from the Internet.
[0062] A network address translation (NAT) gateway 128 configured for a customer's VCN 104 may enable cloud resources in the customer's VCN that do not have dedicated public overlay IP addresses to access the Internet, but the NAT gateway 128 does not expose these resources to direct inbound Internet connections (e.g., L4-L7 connections). This allows private subnets within the VCN (such as Private Subnet 1 of VCN 104) to privately access public endpoints on the Internet. The NAT gateway only allows connections to be initiated from the private subnet to the public Internet, not from the Internet to the private subnet.
[0063] In some embodiments, a service gateway (SGW) 126 may be configured for customer VCN 104, providing a pathway for private network traffic between VCN 104 and service endpoints supported in service network 110. In some embodiments, service network 110 may be provided by a CSP and may offer a variety of services. One example of such a service network is Oracle's Services Network, which offers a variety of services available to customers. For example, compute instances (e.g., database systems) in a private subnet of customer VCN 104 can back up data to a service endpoint (e.g., Object Storage) without requiring a public IP address or access to the Internet. In some embodiments, a VCN may have only one SGW, and connections can be initiated only from subnets within the VCN, not from service network 110. When a VCN is peered with another VCN, resources in the other VCN typically cannot access the SGW. Additionally, resources in an on-premises network connected to a VCN with FastConnect or VPN Connect can use a service gateway configured for that VCN.
[0064] In one embodiment, SGW 126 uses the concept of a service CIDR (Classless Inter-Domain Routing) label, which is a string that represents all of the regional public IP address ranges for a service or services of interest. A customer uses the service CIDR label when configuring the SGW and associated routing rules to control traffic to the service. A customer can optionally use the service CIDR label when configuring security rules without having to adjust if the service's public IP address changes in the future.
[0065] A local peering gateway (LPG) 132 is a gateway that can be added to a customer VCN 104 to enable the VCN 104 to peer with another VCN in the same region. Peering means that the VCNs communicate using private IP addresses, but without traffic traversing a public network such as the Internet or routing traffic to the customer's on-premises network 116. In a preferred embodiment, a VCN has a separate LPG for each peering it establishes. Local peering or VCN peering is a common method used to establish network connectivity between different applications or infrastructure management functions.
[0066] Service providers (e.g., providers of services in service network 110) can provide access to services using various access models. In a public access model, services are exposed as public endpoints publicly accessible by compute instances in the customer VCN over a public network, such as the Internet, and / or they can be privately accessible through SGW 126. In a specific private access model, services are accessible as private IP endpoints in a private subnet of the customer VCN. This is called private endpoint (PE) access, and service providers can expose their services as instances in the customer's private network. Private endpoint resources represent services within the customer's VCN. Each PE appears as a VNIC in a customer-selected subnet in the customer's VCN (called a PE-VNIC and having one or more private IPs). Thus, the PE provides a way to present services in the private customer VCN subnet using VNICs. Because the endpoints are exposed as VNICs, all functionality associated with a VNIC, such as routing rules and security lists, is available to the PE VNIC.
[0067] Service providers may register their services to enable access through the PE. Providers can associate policies with services that restrict the visibility of the service to the customer tenancy. Providers can register multiple services under a single virtual IP address (VIP), especially for multi-tenant services. There may be multiple such private endpoints (in multiple VCNs) representing the same service.
[0068] Compute instances in the private subnet can then access the service using the private IP address or service DNS name of the PE VNIC. Compute instances in the customer VCN can access the service by sending traffic to the private IP address of the PE in the customer VCN. A private access gateway (PAGW) 130 is a gateway resource that can be attached to a service provider VCN (e.g., a VCN in service network 110) that acts as the ingress / egress point for all traffic to customer subnet private endpoints. The PAGW 130 allows providers to scale the number of PE connections without using their internal IP address resources. A provider need only configure one PAGW for any number of services registered in a single VCN. A provider can represent services as private endpoints in multiple VCNs for one or more customers. From the customer's perspective, the PE VNIC does not appear to be attached to the customer's instance, but rather to the service with which the customer wants to interact. Traffic destined for the private endpoint is routed to the service through the PAGW 130. These are called customer-to-service private connections (C2S connections).
[0069] The use of the PE concept also allows traffic to flow across FastConnect / IPsec links and private endpoints in the customer VCN, extending private access for services to the customer's on-premises network and data center, and allows traffic to flow between LPG 132 and PEs in the customer VCN, extending private access for services to the customer's peering VCN.
[0070] Customers can control VCN routing at the subnet level, allowing them to specify which subnets in their VCNs (such as VCN 104) use each gateway. The VCN's route tables are used to determine whether traffic is allowed from a VCN through a particular gateway. For example, in a particular instance, the route table for a public subnet in customer VCN 104 may direct non-local traffic through IGW 120. The route table for a private subnet in the same customer VCN 104 may direct traffic destined for CSP services through SGW 126. All other traffic may be directed through NAT gateway 128. Route tables only control traffic leaving a VCN.
[0071] Traffic entering the VCN via the gateway through an inbound connection is controlled using security lists associated with the VCN. All resources in a subnet use the same route table and security lists. Security lists can be used to control the specific types of traffic that can enter and exit instances in a VCN's subnets. Security list rules can include input (inbound) and output (outbound) rules. For example, an input rule can specify an allowed source address range, while an output rule can specify an allowed destination address range. Security rules can specify a specific protocol (e.g., TCP, ICMP), a specific port (e.g., 22 for SSH, 3389 for Windows RDP), etc. In some embodiments, the instance's operating system can enforce its own firewall rules in line with the security list rules. Rules can be stateful (e.g., connections are tracked and responses are automatically allowed without explicit security list rules for the response traffic) or stateless.
[0072] Access from a customer VCN (i.e., by resources or compute instances deployed in VCN 104) can be categorized as public access, private access, or dedicated access. Public access describes an access model in which public IP addresses or NATs are used to access public endpoints. Private access allows customer workloads in VCN 104 with private IP addresses (e.g., resources in a private subnet) to access services without traversing a public network such as the Internet. In one embodiment, CSPI 101 enables customer VCN workloads with private IP addresses to access (public service endpoints of) services using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer VCN and the public endpoints of the services that reside outside the customer's private network.
[0073] CSPI may also provide dedicated public access by using technologies such as FastConnect public peering, which allows customer on-premises instances to use a FastConnect connection to access one or more services in the customer VCN without traversing a public network such as the Internet. CSPI may also provide dedicated private access by using FastConnect private peering, which allows customer on-premises instances with private IP addresses to access workloads in the customer's VCN using a FastConnect connection. FastConnect is an alternative network connection using the public Internet to connect a customer's on-premises network to CSPI and its services. FastConnect provides an easy, flexible, and economical way to create a dedicated private connection with higher bandwidth options, higher reliability, and a consistent networking experience compared to Internet-based connections.
[0074] FIG. 1 and the accompanying discussion above describe various virtualized components in an exemplary virtual network. As noted above, a virtual network is built on an underlying physical or infrastructure network. FIG. 2 is a simplified architectural diagram of the physical components in a physical network within CSPI 200 that forms the basis of the virtual network, according to one embodiment. As shown, CSPI 200 provides a distributed environment including components and resources (e.g., compute, memory, and network resources) provided by a cloud service provider (CSP). These components and resources are used to provide cloud services (e.g., IaaS services) to subscribing customers, i.e., customers who have signed up for one or more services offered by the CSP. Customers are provisioned with a subset of CSPI 200's resources (e.g., compute, memory, and network resources) based on the services for which they subscribe. Customers can then build their own cloud-based (i.e., CSPI-hosted), customizable private virtual networks using the physical compute, memory, and network resources provided by CSPI 200. As previously mentioned, these customer networks are referred to as virtual cloud networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, into these Customer VCNs. The compute instances can be in the form of virtual machines, bare metal instances, etc. CSPI 200 provides infrastructure and a set of complementary cloud services that enable customers to build and run a wide range of applications and services in a hosted, highly available environment.
[0075] In the exemplary embodiment shown in FIG. 2, the physical components of CSPI 200 include one or more physical host machines or servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), and physical networks (e.g., 218) and switches of physical network 218. The physical host machines or servers can host and execute various compute instances that participate in one or more subnets of the VCN. The compute instances can include virtual machine instances and bare metal instances. For example, the various compute instances shown in FIG. 1 may be hosted by the physical host machines shown in FIG. 2. The virtual machine compute instances in the VCN may be executed by a single host machine or by multiple different host machines. Additionally, the physical host machines can host virtual host machines, container-based hosts or functions, etc. The VNICs and VCN VRs shown in FIG. 1 may be executed by the NVDs shown in FIG. 2. The gateway shown in FIG. 1 may be implemented by a host machine and / or NVD shown in FIG.
[0076] A host machine or server may run a hypervisor (also called a virtual machine monitor or VMM) that creates and makes available a virtualized environment on the host machine. Virtualization or a virtualized environment facilitates cloud-based computing. The hypervisor on a host machine may create, run, and manage one or more compute instances on the host machine. The hypervisor on a host machine enables the sharing of the host machine's physical computing resources (e.g., compute, memory, and network resources) among the various compute instances that the host machine runs.
[0077] For example, as shown in FIG. 2, host machines 202 and 208 execute hypervisors 260 and 266, respectively. These hypervisors may be implemented in software, firmware, hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that resides on a host machine's operating system (OS) and runs on the host machine's hardware processor. A hypervisor provides a virtualized environment by enabling the sharing of the host machine's physical computing resources (e.g., processing resources such as processors / cores, memory resources, and networking resources) among various virtual machine compute instances that the host machine executes. For example, in FIG. 2, hypervisor 260 may reside on the OS of host machine 202 and enable the sharing of the host machine's computing resources (e.g., processing, memory, and networking resources) among the compute instances (e.g., virtual machines) that the host machine executes. A virtual machine may have its own operating system (referred to as a guest operating system), which may be the same as or different from the host machine's OS. The operating system of a virtual machine executed by a host machine may be the same as or different from the operating system of another virtual machine executed by the same host machine. In this manner, the hypervisor allows multiple operating systems to run side by side with each other while sharing the same computing resources of the host machine. The host machines shown in Figure 2 may have the same type of hypervisor or different types of hypervisors.
[0078] A compute instance can be a virtual machine instance or a bare metal instance. In Figure 2, compute instance 268 on host machine 202 and 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.
[0079] In some cases, an entire host machine is provisioned for a single customer, and the one or more compute instances (virtual machine or bare metal instances) that it hosts all belong to that same customer. In other cases, a host machine may be shared among multiple customers (i.e., multiple tenants). In such a multi-tenancy scenario, a host machine may host virtual machine compute instances that belong to different customers. These compute instances may be part of different VCNs for different customers. In some embodiments, bare metal compute instances are hosted on bare metal servers that do not include a hypervisor. When a bare metal compute instance is provisioned, a single customer or tenant maintains control of the physical CPU, memory, and network interfaces of the host machine that hosts the bare metal instance, and the host machine is not shared with other customers or tenants.
[0080] As previously mentioned, each compute instance that is part of a VCN is associated with a VNIC, which may make the compute instance a member of a subnet of the VCN. A VNIC associated with a compute instance facilitates transmission of packets or frames to the compute instance. A VNIC is associated with the compute instance when the compute instance is created. In one embodiment, for compute instances executed by a host machine, the VNIC associated with the compute instance is implemented by an NVD connected to the host machine. For example, in FIG. 2, host machine 202 executes virtual machine compute instance 268, which is associated with VNIC 276, which is implemented by NVD 210 connected to host machine 202. In another example, host machine 206 hosts bare metal instance 272, which is associated with VNIC 280, which is implemented by NVD 212 connected to host machine 206. In yet another example, the VNIC 284 is associated with a compute instance 274 that the host machine 208 executes, and the VNIC 284 is executed by an NVD 212 that is connected to the host machine 208 .
[0081] For compute instances hosted by a host machine, the NVD connected to that host machine also runs VCN VRs corresponding to the VCNs of which the compute instances are elements. For example, in the embodiment shown in Figure 2, NVD 210 runs VCN VR 277 corresponding to the VCN of which compute instance 268 is an element. NVD 212 may also run one or more VCN VRs 283 corresponding to the VCNs corresponding to the compute instances hosted by host machines 206 and 208.
[0082] A host machine may include one or more network interface cards (NICs) that allow the host machine to connect to other devices. A NIC on a host machine may provide one or more ports (or interfaces) that allow the host machine to communicate with other devices. For example, a host machine may be connected to an NVD by one or more ports (or interfaces) provided on the host machine and the NVD. A host machine may also be connected to other devices, such as another host machine.
[0083] 2, host machine 202 is connected to NVD 210 using link 220 extending between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD 210. Host machine 206 is connected to NVD 212 using link 224 extending between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD 212. Host machine 208 is connected to NVD 212 using link 226 extending between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD 212.
[0084] The NVDs, in turn, are connected via communication links to top-of-rack (TOR) switches, which are connected to a physical network 218 (also referred to as a switch fabric). In one embodiment, the links between the host machines and the NVDs and between the NVDs and the TOR switches are Ethernet links. For example, in Figure 2, NVDs 210 and 212 are connected to TOR switches 214 and 216, respectively, using links 228 and 230. In one embodiment, links 220, 224, 226, 228, and 230 are Ethernet links. The collection of host machines and NVDs connected to a TOR may also be referred to as a rack.
[0085] The physical network 218 provides a communications fabric that allows the TOR switches to communicate with each other. The physical network 218 can be a multi-tier network. In one embodiment, the physical network 218 is a multi-tier Clos network of switches, with the TOR switches 214 and 216 representing leaf-level nodes of the multi-tier, multi-node physical switching network 218. Various Clos network configurations are possible, including, but not limited to, a 2-tier network, a 3-tier network, a 4-tier network, a 5-tier network, and generally an "n" tier network. An example of a Clos network is shown in FIG. 5 and described below.
[0086] A variety of different connection configurations are possible between host machines and NVDs, including one-to-one, many-to-one, and one-to-many configurations. In one embodiment of a one-to-one configuration, each host machine is connected to its own separate NVD. For example, in Figure 2, host machine 202 is connected to NVD 210 via NIC 232 on host machine 202. In a many-to-one configuration, multiple host machines are connected to a single NVD. For example, in Figure 2, host machines 206 and 208 are connected to the same NVD 212 via NICs 244 and 250, respectively.
[0087] In a one-to-multiple configuration, one host machine is connected to multiple NVDs. FIG. 3 shows an example of a CSPI 300 in which a host machine is connected to multiple NVDs. As shown in FIG. 3, a host machine 302 includes a network interface card (NIC) 304 including multiple ports 306 and 308. The host machine 300 is connected to a first NVD 310 via port 306 and link 320, and to a second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between the host machine 302 and the NVDs 310 and 312 may be Ethernet links. The NVD 310 is connected to a first TOR switch 314, and the NVD 312 is connected to a second TOR switch 316. The links between the NVDs 310 and 312 and the TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent layer 0 switching devices of a multi-tier physical network 318 .
[0088] 3 provides two separate physical network paths from the physical switch network 318 to the host machine 302: a first path through the TOR switch 314 to the NVD 310 and thus to the host machine 302, and a second path through the TOR switch 316 to the NVD 312 and thus to the host machine 302. These separate paths may increase the availability (referred to as high availability) of the host machine 302. If there is a problem with one of the paths (e.g., if a link on one of the paths goes down) or if there is a problem with one of the devices (e.g., if a particular NVD is not functioning), the other path may be used to communicate with the host machine 302.
[0089] In the configuration shown in Figure 3, the host machine is connected to two different NVDs using two different ports provided by the host machine's NIC. In other embodiments, the host machine may have multiple NICs allowing connection to multiple NVDs.
[0090] Referring again to Figure 2, an NVD is a physical device or component that performs one or more network and / or storage virtualization functions. An NVD may be any device that has one or more processing units (e.g., a CPU, a network processing unit (NPU), an FPGA, a packet processing pipeline, etc.), memory including cache, and ports. Various virtualization functions may be performed by software / firmware executed by one or more processing units of the NVD.
[0091] The NVD may be implemented in a variety of different forms. For example, in one embodiment, the NVD is implemented as an interface card with an embedded processor, referred to as a smart NIC or intelligent NIC. The smart NIC is a separate device from the NIC on the host machine. In Figure 2, NVDs 210 and 212 may be implemented as smart NICs connected to host machine 202 and host machines 206 and 208, respectively.
[0092] However, a smart NIC is merely one example of an implementation of an NVD. Various other implementations are possible. For example, in some other embodiments, the NVD or one or more functions performed by the NVD may be incorporated into or performed by one or more host machines, one or more TOR switches, and other components of CSPI200. For example, the NVD may be embodied in a host machine, and the functions performed by the NVD may be performed by the host machine. As another example, the NVD may be part of a TOR switch, or the TOR switch may be configured to perform the functions performed by the NVD, enabling the TOR switch to perform various complex packet transformations used in public clouds. A TOR performing the functions of an NVD may also be referred to as a smart TOR. In still other embodiments, if virtual machine (VM) instances rather than bare metal (BM) instances are provided to customers, the functions performed by the NVD may be implemented inside the hypervisor of the host machine. In some other embodiments, some of the NVD's functions may be offloaded to a centralized service running on a group of host machines.
[0093] In some embodiments, such as when implemented as a smart NIC as shown in FIG. 2, an NVD may have multiple physical ports that allow the NVD to connect to one or more host machines and one or more TOR switches. Ports on an NVD may be classified as host-side ports (also referred to as "south ports") or network-side or TOR-side ports (also referred to as "north ports"). A host-side port of an NVD is a port used to connect the NVD to a host machine. Examples of host-side ports in FIG. 2 include port 236 on NVD 210 and ports 248 and 254 on NVD 212. A network-side port of an NVD is a port used to connect the NVD to a TOR switch. Examples of network-side ports in FIG. 2 include port 256 on NVD 210 and port 258 on NVD 212. As shown in FIG. 2, the NVD 210 is connected to the TOR switch 214 by link 228, which extends from port 256 of the NVD 210 to the TOR switch 214. Similarly, the NVD 212 is connected to the TOR switch 216 by a link 230 extending from a port 258 of the NVD 212 to the TOR switch 216 .
[0094] An NVD can receive packets and frames from a host machine (e.g., packets and frames generated by a compute instance hosted by the host machine) via a host-side port, perform necessary packet processing, and then forward the packets and frames to a TOR switch via a network-side port of the NVD. An NVD can receive packets and frames from a TOR switch via a network-side port of the NVD, perform necessary packet processing, and then forward the packets and frames to a host machine via a host-side port of the NVD.
[0095] In some embodiments, there may be multiple ports and associated links between the NVD and the TOR switch. These ports and links may be aggregated to form a link aggregator group (LAG) consisting of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and the TOR switch) to be treated as a single logical link. All physical links in a given LAG may operate in full-duplex mode at the same speed. LAGs help increase the bandwidth and reliability of the connection between two endpoints. If one of the physical links in the LAG goes down, traffic is dynamically and transparently reassigned to one of the other physical links in the LAG. The aggregated physical link provides higher bandwidth than each individual link. Multiple ports associated with a LAG are treated as a single logical port. Traffic may be load-balanced across the multiple physical links in the LAG. One or more LAGs may be configured between two endpoints. The two endpoints may be, for example, between the NVD and the TOR switch, or between a host machine and the NVD.
[0096] The NVD implements or performs network virtualization functions. These functions are performed by software / firmware executed by the NVD. Examples of network virtualization functions include, but are not limited to, packet encapsulation and decapsulation functions, functions for creating VCN networks, functions for enforcing network policies such as VCN security list (firewall) functions, and functions for facilitating routing and forwarding of packets to compute instances in the VCN. In one embodiment, upon receiving a packet, the NVD is configured to execute a packet processing pipeline to process the packet and determine how to forward or route the packet. As part of this packet processing pipeline, the NVD may execute one or more virtual functions associated with the overlay network, such as running VNICs associated with cis in the VCN, running virtual routers (VRs) associated with the VCN, encapsulating and decapsulating packets to facilitate forwarding or routing in the virtual network, running a gateway (e.g., a local peering gateway), implementing security lists, network security groups, network address translation (NAT) functions (e.g., public IP to private IP translation on a per-host basis), throttling functions, and other functions.
[0097] In one embodiment, the NVD's packet processing data path may include multiple packet pipelines, each consisting of a series of packet transformation stages. In one embodiment, upon receipt of a packet, the packet is parsed and sorted into a single pipeline. The packet is then processed linearly, one stage at a time, until it is either discarded or sent out through the NVD's interface. These stages provide basic functional packet processing building blocks (e.g., header validation, throttling enforcement, new Layer 2 header insertion, L4 firewall enforcement, VCN encapsulation / decapsulation, etc.), such that new pipelines can be constructed by assembling existing stages, and new functionality can be added by creating and inserting new stages into existing pipelines.
[0098] The NVD may perform both control plane and data plane functions corresponding to the VCN's control and data planes. Examples of the VCN control plane are also shown in Figures 18, 19, 20, and 21 (see reference numbers 1816, 1916, 2016, and 2116) and described below. Examples of the VCN data plane are shown in Figures 18, 19, 20, and 21 (see reference numbers 1818, 1918, 2018, and 2118) and described below. Control plane functions include functions used to configure the network (e.g., routing and route table configuration, VNIC configuration, etc.) and control how data is forwarded. In one embodiment, a VCN control plane is provided that centrally computes and exposes all overlay-to-foundation mappings to the NVD and virtual network edge devices (e.g., various gateways such as DRGs, SGWs, and IGWs). Firewall rules may also be exposed by the same mechanism. In one embodiment, the NVD retrieves only mappings relevant to the NVD. The data plane functions include functionality for the actual routing / forwarding of packets based on the configuration set using the control plane. The VCN data plane is achieved by encapsulating customer network packets before they traverse the underlying network. The encapsulation / decapsulation functionality is achieved on the NVD. In one embodiment, the NVD is configured to intercept all network packets entering and leaving the host machine and perform network virtualization functions.
[0099] As described above, the NVD performs various virtualization functions, including VNICs and VCN VRs. The NVD may execute VNICs associated with compute instances hosted by one or more host machines connected to the VNIC. For example, as shown in FIG. 2, NVD 210 executes the functions of VNIC 276 associated with compute instance 268 hosted by host machine 202 connected to NVD 210. As another example, NVD 212 executes VNIC 280 associated with bare metal compute instance 272 hosted by host machine 206 and VNIC 284 associated with compute instance 274 hosted by host machine 208. The host machines can host compute instances that belong to different VCNs belonging to different customers, and the NVDs connected to the host machines can execute VNICs (i.e., perform VNIC-related functions) corresponding to these compute instances.
[0100] The NVDs also execute VCN virtual routers corresponding to the VCNs of the compute instances. For example, in the embodiment shown in FIG. 2, NVD 210 executes VCN VR 277 corresponding to the VCN to which compute instance 268 belongs. NVD 212 executes one or more VCN VRs 283 corresponding to one or more VCNs to which compute instances hosted by host machines 206 and 208 belong. In one embodiment, a VCN VR corresponding to a given VCN is executed by all NVDs connected to a host machine that hosts at least one compute instance belonging to that VCN. If a host machine hosts compute instances that belong to different VCNs, the NVDs connected to that host machine may execute VCN VRs corresponding to the different VCNs.
[0101] In addition to VNICs and VCN VRs, the NVD may run various software (e.g., daemons) and may include one or more hardware components that facilitate the various network virtualization functions performed by the NVD. For simplicity, these various components are grouped together as a “packet processing component” shown in FIG. 2 . For example, NVD 210 includes packet processing component 286, and NVD 212 includes packet processing component 288. For example, the packet processing component of the NVD may include a packet processor configured to interact with the NVD's ports and hardware interfaces to monitor all packets received and communicated using the NVD and to store network information. The network information may include, for example, network flow information and per-flow information (e.g., per-flow statistics) that identify the various network flows processed by the NVD. In one embodiment, network flow information may be stored per VNIC. The packet processor may perform per-packet operations as well as implement a stateful NAT and an L4 firewall (FW). As another example, the packet processing component may include a replication agent configured to replicate information stored by the NVD to one or more different replication target stores. As yet another example, the packet processing component may include a logging agent configured to perform logging functions for the NVD. The packet processing component may also include software for monitoring the performance and health of the NVD, and possibly the status and health of other components connected to the NVD.
[0102] FIG. 1 illustrates components of an exemplary virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed to the subnets, VNICs associated with the compute instances, VRs in the VCN, and a set of gateways configured for the VCN. The overlay components illustrated in FIG. 1 may be executed or hosted by one or more of the physical components illustrated in FIG. 2. For example, compute instances in a VCN may be executed or hosted by one or more host machines illustrated in FIG. 2. For compute instances hosted by a host machine, the VNICs associated with the compute instances are typically executed by an NVD connected to the host machine (i.e., the VNIC functionality is provided by an NVD connected to the host machine). The VCN VR functionality for a VCN is performed by all NVDs connected to host machines that host or execute compute instances that are part of the VCN. The gateways associated with a VCN may be executed by one or more different types of NVDs. For example, one gateway may be executed by a smart NIC, while other gateways may be executed by one or more host machines or other implementations of NVDs.
[0103] As mentioned above, a compute instance in a customer VCN can communicate with a variety of different endpoints, which can be in the same subnet as the source compute instance, in a different subnet but within the same VCN as the source compute instance, or outside the VCN of the source compute instance. These communications are facilitated through the use of VNICs associated with the compute instances, VCN VRs, and gateways associated with the VCN.
[0104] For communication between two compute instances on the same subnet in a VCN, this communication is facilitated through the use of VNICs associated with the source and destination compute instances. The source and destination compute instances may be hosted by the same host machine or different host machines. A packet originating from the source compute instance may be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. On the NVD, the packet is processed using a packet processing pipeline, which may include running the VNIC associated with the source compute instance. Because the packet's destination endpoint is in the same subnet, the VNIC associated with the source compute instance forwards the packet to the NVD running the VNIC associated with the destination compute instance. The NVD then processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances may be running on the same NVD (e.g., if both the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., if the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNICs may use routing / forwarding tables stored by the NVD to determine the next hop for a packet.
[0105] When a packet is sent from a compute instance in one subnet to an endpoint in a different subnet of the same VCN, the packet originating from the source compute instance is sent from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD processes the packet using a packet processing pipeline, which may include running one or more VNICs and VRs associated with the VCN. For example, as part of the packet processing pipeline, the NVD executes or invokes a function corresponding to the VNIC associated with the source compute instance (also referred to as VNIC execution). The function executed by the VNIC may include verifying the VLAN tag on the packet. Because the packet's destination is outside the subnet, the NVD then invokes a VCN VR function to execute it. The VCN VR then routes the packet to the NVD that executes the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances may be running on the same NVD (e.g., if both the source and destination compute instances are hosted by the same host machine) or may be running on different NVDs (e.g., if the source and destination compute instances are hosted by different host machines connected to different NVDs).
[0106] If the packet's destination is outside the VCN of the source compute instance, the packet originating from the source compute instance is sent from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD runs the VNIC associated with the source compute instance. Because the packet's destination endpoint is outside the VCN, the packet is then processed by the VCN VR for that VCN. The NVD invokes a VCN VR function, which forwards the packet to an NVD running the appropriate gateway associated with the VCN. For example, if the destination is an endpoint in a customer's on-premises network, the packet may be forwarded by the VCN VR to an NVD running a DRG gateway configured for the VCN. The VCN VR may be running on the same NVD as the NVD running the VNIC associated with the source compute instance, or it may be running on a different NVD. The gateway may be running on the NVD, which may be a smart NIC, a host machine, or another implementation of the NVD. The packet is then processed by the gateway and forwarded to a next hop that facilitates delivery of the packet to the intended destination endpoint. For example, in the embodiment shown in FIG. 2, a packet originating from compute instance 268 may be sent from host machine 202 to NVD 210 over link 220 (using NIC 232). In NVD 210, VNIC 276 is called VNIC 276 because it is associated with source compute instance 268. VNIC 276 is configured to examine information encapsulated in the packet, determine a next hop to forward the packet to facilitate delivery of the packet to the intended destination endpoint, and then forward the packet to the determined next hop.
[0107] Compute instances deployed in a VCN can communicate with a variety of different endpoints. These endpoints can include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. CSPI 200-hosted endpoints can include instances in the same VCN or other VCNs (which may be customer VCNs or VCNs not belonging to the customer). Communication between CSPI 200-hosted endpoints can be performed over physical network 218. Compute instances can also communicate with endpoints not hosted by CSPI 200 (i.e., external to CSPI 200). Examples of such endpoints include endpoints within a customer's on-premises network or data center or public endpoints accessible over a public network such as the Internet. CSPI 200's communication with external endpoints can be performed over a public network (e.g., the Internet) (not shown in FIG. 2) or over a private network (not shown in FIG. 2) using various communication protocols.
[0108] The architecture of CSPI 200 shown in FIG. 2 is illustrative only and is not intended to be limiting. Variations, substitutions, and improvements are possible in alternative embodiments. For example, in some implementations, CSPI 200 may include more or fewer systems or components than those shown in FIG. 2, may combine two or more systems, or may have different system configurations or arrangements. The systems, subsystems, and other components shown in FIG. 2 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, using hardware, or a combination thereof. Software may be stored in a non-transitory storage medium (e.g., a memory device).
[0109] FIG. 4 illustrates connections between host machines and an NVD to provide I / O virtualization to support multitenancy, according to one embodiment. As shown in FIG. 4, host machine 402 runs hypervisor 404, which provides a virtualized environment. Host machine 402 runs two virtual machine instances: VM1 406, which belongs to customer / tenant #1, and VM2 408, which belongs to customer / tenant #2. Host machine 402 includes a physical NIC 410 connected to NVD 412 via link 414. Each compute instance is associated with a VNIC that NVD 412 runs. In the embodiment of FIG. 4, VM1 406 is associated with VNIC-VM1 420, and VM2 408 is associated with VNIC-VM2 422.
[0110] 4, NIC 410 includes two logical NICs: logical NIC A 416 and logical NIC B 418. Each virtual machine is configured to associate with and work together on its own logical NIC. For example, VM1 406 is associated with logical NIC A 416, and VM2 408 is associated with logical NIC B 418. Although host machine 402 includes only one physical NIC 410 shared by multiple tenants, because of the logical NICs, each tenant's virtual machine sees itself as having its own host machine and NIC.
[0111] In one embodiment, each logical NIC is assigned its own VLAN ID. Thus, a particular VLAN ID is assigned to logical NIC A 416 for tenant #1, and a separate VLAN ID is assigned to logical NIC B 418 for tenant #2. When a packet is sent from VM1 406, the hypervisor attaches a tag assigned to tenant #1 to the packet, and the packet is then sent from host machine 402 to NVD 412 via link 414. Similarly, when a packet is sent from VM2 408, the hypervisor attaches a tag assigned to tenant #2 to the packet, and the packet is then sent from host machine 402 to NVD 412 via link 414. Thus, a packet 424 sent from host machine 402 to NVD 412 has an associated tag 426 that identifies the particular tenant and associated VM. On the NVD, for a packet 424 received from a host machine 402, a tag 426 associated with the packet is used to determine whether the packet should be processed by VNIC-VM1 420 or VNIC-VM2 422. The packet is then processed by the corresponding VNIC. According to the configuration shown in Figure 4, each tenant's compute instance can be recognized as having its own host machine and NIC. The configuration shown in Figure 4 provides I / O virtualization to support multi-tenancy.
[0112] FIG. 5 is a simplified block diagram of a physical network 500, according to one embodiment. The embodiment shown in FIG. 5 is structured as a Clos network. A Clos network is a specific type of network topology designed to provide connection redundancy while maintaining high bisection bandwidth and maximum resource utilization. A Clos network is a type of non-blocking multistage or multi-layer switching network, with possible numbers of stages or layers, such as two, three, four, or five. The embodiment shown in FIG. 5 is a three-layer network, including layers 1, 2, and 3. TOR switch 504 represents a layer 0 switch in a Clos network. One or more NVDs are connected to a TOR switch. A layer 0 switch is also referred to as an edge device of the physical network. A layer 0 switch is connected to a layer 1 switch, also referred to as a leaf switch. In the embodiment shown in FIG. 5, a set of “n” layer 0 TOR switches are connected to a set of “n” layer 1 switches, collectively forming a pod. Each layer 0 switch in a pod is interconnected to all layer 1 switches in that pod, but there are no switch connections between pods. In one embodiment, two pods are referred to as a block. Each block is served by or connected to a set of “n” layer 2 switches (sometimes referred to as spine switches). There may be multiple blocks in a physical network topology. The layer 2 switches are then connected to “n” layer 3 switches (sometimes referred to as super spine switches). Packet communication across the physical network 500 is typically performed using one or more layer 3 communication protocols. High availability is typically achieved by having n-way redundancy at all layers of the physical network except the TOR layer. To enable scaling of the physical network, policies may be specified for pods and blocks to control the visibility of switches to each other in the physical network.
[0113] A Clos network is characterized by a fixed maximum hop count for reaching another tier-0 switch (or from an NVD connected to a tier-0 switch to another NVD connected to a tier-0 switch). For example, in a three-tier Clos network, a packet requires at most seven hops to reach another NVD from one NVD, with the source and destination NVDs connected to the leaf layers of the Clos network. Similarly, in a four-tier Clos network, a packet requires at most nine hops to reach another NVD from one NVD, with the source and destination NVDs connected to the leaf layers of the Clos network. Therefore, the Clos network architecture maintains consistent latency throughout the network, which is important for intra- and inter-datacenter communications. Clos topologies are horizontally scalable and cost-effective. The network's bandwidth / throughput capacity can be easily increased by adding more switches (e.g., more leaf and spine switches) to various tiers and by increasing the number of links between switches in adjacent tiers.
[0114] In one embodiment, each resource in CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource's information and can be used when managing the resource, for example, through a console or API. An exemplary syntax for a CID is as follows: ocid1.<RESOURCE TYPE> . <realm>.[REGION][.FUTURE USE].<UNIQUE ID> where: ocid1 is a string indicating the version of the CID, RESOURCE TYPE is the type of resource (for example, instance, volume, VCN, subnet, user, group, etc.), REALM is the realm in which the resource resides (example values are "c1" for the commercial realm, "c2" for the government cloud realm, or "c3" for the federal cloud realm, etc.; each realm may have its own domain name); REGION is the region in which the resource resides (this part may be blank if region is not applicable to the resource), FUTURE USE is reserved for future use, UNIQUE ID is the unique part of the ID (the format can vary depending on the type of resource or service).
[0115] Dedicated Customer Region Cloud (DRCC) A customer-dedicated regional cloud (DRCC) corresponds to a cloud service provider's (CSP) infrastructure deployed in a customer's own data center. With a DRCC, enterprises can easily consolidate applications and mission-critical database systems deployed on expensive hardware on a CSP's highly available and secure infrastructure, creating operational efficiencies and modernization opportunities. Enterprises typically find migration to cloud infrastructure costly and difficult due to the inherent mismatch between traditional application architectures and cloud architectures. These challenges are compounded for workloads that cannot be migrated to the public cloud. Enterprises have access to only a small subset of cloud services on-premises and a limited set of capabilities and functional metrics compared to those available in the public cloud.
[0116] The DRCC framework delivers all the functional metrics of the public cloud on-premises, allowing enterprises to reduce infrastructure and operational costs while upgrading legacy applications to modern cloud services to meet the most stringent regulatory, data residency, and latency requirements. All of this is powered by the CSP's infrastructure, which delivers high performance and the highest levels of security. Customers gain the choice and flexibility to run all of the CSP's cloud services in their own data centers. Customers can choose from all the public cloud services the CSP offers (including VMware Cloud, Autonomous Database, Container Engine for Kubernetes, Bare Metal Servers, and Exadata Cloud Service) and pay only for the services they consume. The DRCC framework is designed to completely isolate data and customer operations from the internet, and control and data plane operations remain on-premises, helping customers meet the most stringent compliance and latency requirements. With a fully managed experience and access to the new functional metrics now available in the public cloud, the DRCC framework delivers cloud-scale security, resiliency, and scalability while supporting mission-critical workloads with tools to incrementally modernize traditional workloads.
[0117] The benefits of DRCC are as follows: Reduce the risk and cost of innovation by offering all public cloud services and autonomous databases on-premises Provide a framework where customers pay only for the services they consume Create a truly consistent development experience for all IaaS and PaaS applications by using the exact same tools, APIs, and SLAs available in public cloud infrastructure. Maintain complete control of all your data to meet the most stringent data privacy and latency requirements Deploy seamlessly between on-premise and public cloud without compromising functionality or development experience Consolidating workloads onto a single cloud platform allows customers to focus on their business priorities Lower the cost of running on-premise workloads with a CSP's infrastructure that offers the highest level of security
[0118] CSP infrastructure is typically hosted in regions and availability domains. A region is a local geographic area, and an availability domain is one or more data centers located within a region. Thus, a region consists of one or more availability domains. Most infrastructure resources are region-specific, such as virtual cloud networks, or availability domain-specific, such as compute instances. Furthermore, traffic between availability domains and between regions is encrypted. Availability domains are isolated from each other and fault-tolerant, with a very low probability of simultaneous failure. Because availability domains do not share infrastructure such as power or cooling, nor do they share internal availability domain networks, a failure in one availability domain within a region is unlikely to affect the availability of others in the same region.
[0119] Availability domains within the same region are connected to each other by low-latency, high-bandwidth networks, enabling highly available connections to the Internet and on-premises, and allowing replica systems to be built across multiple availability domains for both high availability and disaster recovery. Regions are independent of other regions and can be separated by large distances, even across countries or continents. Generally, you might want to deploy applications in the region that is most heavily used, because using nearby resources is faster than using resources that are farther away. However, you might want to deploy applications in different regions to mitigate the risk of a region-wide event, such as an earthquake.
[0120] A fault domain is a grouping of hardware and infrastructure within an availability domain. Each availability domain contains multiple fault domains (for example, three fault domains). Fault domains provide anti-affinity, i.e., they distribute instances so that they are not on the same physical hardware within a single availability domain. A hardware failure or compute hardware maintenance event that affects one fault domain does not affect instances in other fault domains. Additionally, the physical hardware in a fault domain has independent and redundant power sources. This ensures that a power hardware failure in one fault domain does not affect other fault domains.
[0121] According to some embodiments, in the context of a DRCC, an availability domain is provided via a rack containing multiple top-of-rack (TOR) switches and multiple host machines / servers. TOR switches are network switches used in data centers to connect servers and other network devices in the rack. The purpose of a TOR switch is to provide high-speed connectivity and efficient data transfer between devices in the rack and the larger network infrastructure. TOR switches typically have high port densities to accommodate multiple servers and devices within a single rack. They also provide Ethernet connections for these devices, allowing them to communicate with each other and with the rest of the network. TOR switches often support high-speed Ethernet standards, such as 10 Gigabit Ethernet, 25 Gigabit Ethernet, 40 Gigabit Ethernet, or 100 Gigabit Ethernet. These high-speed data transfer rates ensure efficient communication between servers and the network. TOR switches minimize latency and enable low-latency switching, which is important in data centers where applications and services running on servers require fast response times. Various TOR configurations that may be employed within a data center rack are described below with reference to FIGS. 6 and 7.
[0122] FIG. 6 illustrates a configuration of multiple TORs included in a rack, according to at least one embodiment. The multiple TOR configuration 600 illustrated in FIG. 6 corresponds to a three-TOR configuration. The rack includes three TORs (601, 602, and 603) and multiple host machines / servers. In one embodiment, fault domains are created within an availability domain (i.e., the rack) by selecting a subset of host machines from the multiple host machines. Each host machine in the selected group of host machines (i.e., the subset of host machines) is communicatively coupled to one of the TORs included in the multiple TORs. For example, as illustrated in FIG. 6, a first subset of host machines from the multiple host machines is illustrated as 605A. Each host machine / server included in 605A is communicatively coupled to a first TOR, i.e., TOR1 601. In this manner, the combination of the first subset of host machines (605A) and the first TOR (601) constitutes a first fault domain.
[0123] A second fault domain is created within the availability domain by selecting a second subset of host machines from the plurality of host machines. Each host machine in the second subset of host machines is communicatively coupled to another TOR within the plurality of TORs (i.e., a different TOR than the TOR associated with the first fault domain). Note that the first subset of host machines is separate from the second subset of host machines. For example, as shown in FIG. 6, the second subset of host machines within the plurality of host machines is designated 605B. Each host machine / server in 605B is communicatively coupled to a second TOR, i.e., TOR2 602.
[0124] Similarly, a third fault domain may be created within the availability domain by selecting a third subset of host machines from the plurality of host machines. Each host machine in the third subset of host machines is communicatively coupled to a different TOR (i.e., a different TOR than the TORs associated with the first and second fault domains). Note that the third subset of host machines is separate from both the first subset of host machines and the second subset of host machines. For example, as shown in FIG. 6, the third subset of host machines from the plurality of host machines is designated 605C. Each host machine / server included in 605C is communicatively coupled to a third TOR, i.e., TOR3 603, to form the third fault domain.
[0125] The first, second, and third subsets of servers (605A, 605B, and 605C) shown in FIG. 6 each include K servers. Of course, this does not limit the scope of the present disclosure. A particular subset of servers may include a different number of servers compared to another subset of servers. Furthermore, the rack includes multiple network virtualization devices (NVDs). Of the multiple NVDs, a first subset of NVDs is employed to connect a first subset of server / host machines to a first TOR. Similarly, of the multiple NVDs, a second subset of NVDs is employed to connect a second subset of server / host machines to a second TOR, and a third subset of NVDs is employed to connect a third subset of server / host machines to a third TOR.
[0126] Furthermore, for each fault domain (i.e., a combination of a subset of servers and a TOR switch associated with the subset of servers), a set of addresses corresponding to the host machines / servers in the first subset of servers is associated with the TOR switch associated with the first subset of servers. In this manner, the control plane is configured to forward packets destined for a particular server in the first subset of servers to the first TOR associated with the first subset of servers.
[0127] Thus, a rack configuration such as that shown in FIG. 6 provides three distinct fault domains, which have a smaller blast radius (i.e., the percentage of capacity loss in the event of a TOR switch failure) than if all servers in the rack were communicatively coupled to a single TOR switch. Specifically, a rack configuration such as that shown in FIG. 6 results in a blast radius of 33%. That is, a single TOR failure results in a 33% capacity loss. However, the probability of multiple TORs failing simultaneously within a rack is extremely low, i.e., negligible. In one embodiment, the rack configuration of FIG. 6 provides one or more fault domains that can be presented to a customer. Upon a customer request for allocation of one or more host machines in a rack, the control plane may allocate one or more host machines in one or more fault domains based on certain criteria. For example, if a customer requires high availability, the control plane may allocate host machines in different fault domains.
[0128] Referring to FIG. 7, this illustrates another configuration of multiple TORs included in a rack, according to some embodiments. Specifically, the configuration 700 shown in FIG. 7 is referred to herein as a dual TOR configuration. As shown in FIG. 7, the rack includes two TORs, TOR1 701 and TOR2 703. Additionally, the rack includes multiple server / host machines. According to some embodiments, the multiple servers are grouped into disjoint subsets of servers. For example, the multiple servers may be grouped into a first subset of servers 705A and a second subset of servers 705B.
[0129] As shown in FIG. 7, in the dual TOR configuration, a subset of servers is communicatively coupled to each of the TORs included in the rack. Specifically, a first subset of servers 705A and a second subset of servers are communicatively coupled to TOR1 701 and TOR2 703. Therefore, in this configuration, if a single TOR switch in the rack fails, no capacity loss occurs; the servers simply select the other functioning TOR to communicate data to. Note that the probability of both TORs failing simultaneously in a rack is extremely low; i.e., negligible. In contrast to the TOR configuration of FIG. 6, in which the design of the NVDs coupling servers to the TORs remains unchanged (i.e., each NVD couples a single server to a single TOR), in the configuration of FIG. 7, the NVDs are naturally configured to connect to both TORs. Therefore, in this configuration, each NVD is configured with multiple IP addresses. Of course, in the configuration of FIG. 7, the entire rack containing multiple TORs (i.e., providing redundancy) is considered a fault domain. In some embodiments, a fault domain may span multiple racks. For example, consider a switch that serves multiple racks (e.g., four racks). In this situation, the four racks can be considered fault domains, each of which may have multiple TORs providing redundancy.
[0130] FIG. 8 is a flowchart illustrating steps performed in providing an availability domain to a customer, according to at least one embodiment. The process illustrated in FIG. 8 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 8 and described below is exemplary and not intended to be limiting. While FIG. 8 depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, the steps may be performed in some different order, or some steps may be performed in parallel.
[0131] The process begins in step 801, where a control plane provides an availability domain comprising a rack. The rack includes a plurality of TOR switches and a plurality of host machines or servers. In step 803, a first fault domain is created within the availability domain. The first fault domain comprises a first TOR switch of the plurality of TOR switches and a first subset of host machines of the plurality of host machines. The first subset of host machines is communicatively coupled to the first TOR. The process then moves to step 805, where a second fault domain is created within the availability domain. The second fault domain comprises a second TOR switch of the plurality of TOR switches and a second subset of host machines of the plurality of host machines. The second subset of host machines is isolated from the first set of host machines. The second subset of host machines is communicatively coupled to a second TOR via an NVD. In this manner, one or more fault domains may be presented to the customer. Upon receiving a customer request for allocation of one or more host machines in a rack, the control plane may allocate one or more host machines to one or more fault domains based on certain criteria associated with the customer request.
[0132] 9, which illustrates an example architecture 900 of a DRCC framework that provides customers with full public cloud performance metrics, enabling them to reduce infrastructure and operational costs while upgrading legacy applications onto modern cloud services to meet the most stringent regulatory, data residency, and latency requirements.
[0133] According to some embodiments, FIG. 9 illustrates a data center 905 including a pair of TORs (i.e., TOR#1 922 and TOR#2 924), a network virtualization platform (e.g., NVD 926), and a compute host 928 (also referred to herein as a local compute host). Of course, the compute host 928 includes multiple virtual machines or bare metal instances. The NVD 926 is referred to herein as a local NVD. The compute host 928 includes a host network interface card (i.e., a host NIC). For illustrative purposes, FIG. 9 illustrates the compute host 928 as including two virtual machines, VM1 and VM2, respectively. It should be noted that the VMs are each communicatively coupled to the host NIC via one of the logical interfaces (e.g., the logical interfaces designated PF1 and PF2, respectively). Furthermore, the local NVD 926 may be located on the same chassis as the host NIC included in the compute host 928.
[0134] A compute host 928 included in a data center may be coupled to another host machine 911, referred to herein as a remote host machine. It should be understood that a remote host machine can be “any” host machine, such as (i) another host in a DRCC located behind another NVD, (ii) another host in another DRCC (e.g., in a group of DRCCs serving the same customer / organization) located behind another NVD (e.g., in a group of DRCCs serving the same customer / organization), or (iii) a host machine included in a customer’s on-premises network. Note that if a host machine is included in a customer’s on-premises network, the host machine may connect to the DRCC via a FastConnect or IPSec VPN tunnel and also connect to a host machine in the DRCC through the use of a Dynamic Routing Gateway (DRG). For purposes of illustration, the following description assumes that a remote host machine (e.g., host machine 911) is included in a DRCC and located behind another NVD (e.g., NVD 913 served by a remote TOR 915). However, the features described below are equally applicable to the other remote host machine cases described above. Furthermore, it should be appreciated that in this case (shown in FIG. 9), two host machines (i.e., local host machine 928 and remote host machine 911) may be coupled via network fabric 920. Furthermore, for convenience, NVD 913 will be referred to herein as a remote NVD.
[0135] According to some embodiments, local NVD 926 has multiple physical ports. For example, in one implementation as shown in Figure 9, local NVD 926 has two physical ports: a first physical port 927A (referred to herein as the TOR-facing port) connected to TORs 922 and 924, respectively, and a second physical port 927B (referred to herein as the host-facing port) connected to a compute host 928. Each physical port of local NVD 926 may be split into multiple logical ports. For example, as shown in Figure 9, physical port 927B is split into two host-facing logical ports, and physical port 927A is split into two TOR-facing logical ports.
[0136] Each physical port of the local NVD 926 can be flexibly divided to be represented by two logical ports (two MAC addresses and two IP addresses). For example, in FIG. 9, the overlay IP address and overlay MAC address are indicated by underscores (e.g., B1, M1), while the board IP address and MAC address are indicated without underscores (e.g., A0, M0). As shown in FIG. 9, the first physical port 927A of the local NVD 926 is associated with a first IP address (A1), a second IP address (A3), a first MAC address (M1), and a second MAC address (M3). The second physical port 927B of the local NVD 926 is associated with a first overlay IP address (B1), a second overlay IP address (C1), a first overlay MAC address (M4), and a second overlay MAC address (M6).
[0137] Of course, the limit on the number of logical ports that can be obtained by dividing a physical port (e.g., port 927A) of the NVD 926 is determined by the width of the serializer / deserializer component (i.e., SerDes component) included in the NVD chipset. In one example, each physical port of the NVD 926 may be bifurcated into four logical ports. Of course, the utilization of a gearbox component in the NVD allows for more logical ports to be obtained for each physical port of the NVD 926. The data center 905 (i.e., the DRCC) provides customers with all the functional metrics of a public cloud. Specifically, the DRCC hosts applications and data that require strict data residency, control, and security, and provides a means to keep data in a specific location for low-latency connections and data-intensive processing. Thus, customers can utilize all cloud services running directly in their own data center, rather than in a cloud region hundreds or thousands of miles away. This allows the small footprint of the DRCC (described below with reference to Figure 14) to provide organizations with the opportunity to run workloads outside of the public cloud.
[0138] Reference is now made to FIG. 10 , which is a flowchart illustrating a process for providing a DRCC, according to some embodiments. The process illustrated in FIG. 10 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 10 and described below is exemplary and not intended to be limiting. While FIG. 10 depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In some alternative embodiments, these steps may be performed in some different order, or some steps may be performed in parallel. In some implementations, the method illustrated in FIG. 10 may be performed by a cloud service to provide a DRCC to a customer.
[0139] The method begins in step 1001, where a first physical port of a network virtualization device (NVD) included in a data center is communicatively coupled to a first top-of-rack (TOR) switch and a second TOR switch, where the first and second TOR switches may be included in a rack (as shown in FIGS. 6 and 7). In step 1003, a second physical port of the NVD is communicatively coupled to a network interface card (NIC) associated with a host machine included in the data center. The second physical port provides a first logical port and a second logical port for communication between the NVD and the NIC.
[0140] The process then proceeds to step 1005, where the NVD receives the packet from the host machine via the first logical port or the second logical port. In step 1007, the NVD determines a specific TOR for communication of the packet from a group including the first TOR and the second TOR. According to some embodiments, the NVD may perform a packet forwarding mechanism, such as equal-cost multipath routing (ECMP) flow hashing, to select one of the two TORs. Further, in step 1009, the NVD transmits the packet to the specific TOR to facilitate communication to the packet's destination host machine (e.g., a host machine behind another NVD in the same rack, another host machine behind another NVD in another rack, or a host machine outside the data center (e.g., included in a customer's on-premises network)).
[0141] The following details (i) the transmission of a packet from a compute host (e.g., compute host 928) included in data center 905 to remote host 911 with reference to FIG. 11A , and (ii) the transmission of a packet from the remote host to a compute host in the data center, i.e., the return path, with reference to FIG. 11B . As noted above, overlay IP addresses and overlay MAC addresses are indicated with underscores (e.g., B1, M1), while board IP addresses and MAC addresses are indicated without underscores (e.g., A0, M0). For purposes of illustration, consider the case of a virtual machine (e.g., VM1) included in compute host 928 sending a packet to remote host 911. The steps involved in transmitting a packet are described below with reference to FIG. 11A .
[0142] FIG. 11A is a flowchart illustrating steps performed in transmitting a packet from a compute host included in a data center to a remote host. For purposes of illustration, the description provided with reference to FIGS. 11A and 11B relates to a host in a DRCC receiving and sending packets from another host in the DRCC. However, it should be understood that the features described herein are equally applicable to other cases involving a remote host (e.g., a remote host machine included in a customer's on-premises network). The process illustrated in FIG. 11A may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 11A and described below is exemplary and not intended to be limiting. While FIG. 11A depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, these steps may be performed in some different order, or some steps may be performed in parallel.
[0143] In step 1101, VM1, with an overlay IP address of B2, sends a packet (destined for remote host 911, with an overlay IP address of B100) received by NVD 926 with address B1, M4 (i.e., the overlay IP and MAC address of one of the NVD's logical ports). VM1 sends the packet to the host NIC via logical interface PF1. The host NIC, in turn, uses one of the NVD's two logical ports to send the packet. Note that, according to some embodiments, each virtual machine on a compute host is associated with a logical interface on the host NIC. These interfaces are referred to herein as virtual functions or "VFs," and each VF is associated with a VLAN-ID. Note that packets sent by a VM can use one or both of the NVD's logical ports. Whether a packet uses one or both ports, or one or the other, depends on the NVD configuration (e.g., whether the NVD is configured for active / active or active / backup operation), as well as the "state" of the NVD's interfaces (e.g., whether both interfaces are up or one interface is in down / backup mode).
[0144] In step 1103, upon receiving a packet, the NVD 926 performs a lookup operation in the VCN forwarding table. Specifically, the NVD 926 obtains information about the board IP address of the NVD serving the remote host. That is, the NVD 926 obtains information about the board IP address A100 of the remote NVD 913 serving the remote host.
[0145] In step 1105, the NVD modifies the packet's header. Specifically, the NVD 926 encapsulates information in the packet (e.g., in the VCN header) corresponding to the remote NVD 913's board IP address (A100) as the packet's intended destination and its own software interface IP address A254 as the packet's source. The NVD 926 then attempts to send the packet to the remote host.
[0146] In step 1107, the NVD 926 determines how to forward the packet to the remote NVD 1113. According to one embodiment, the NVD 926 recognizes that there are two routes for sending the packet to the remote host machine: one route through TOR#1 922 and the other route through TOR#2 924. Specifically, the NVD 926 realizes the two routes through BGP communication performed between the two TORs. The NVD 926 performs equal-cost multipath (ECMP) flow hashing to select one of the two TORs and forward the packet to the remote NVD 913. In step 1109, the packet passes through the selected TOR (e.g., TOR#1) and traverses the network fabric, eventually reaching the remote TOR that serves the remote host machine. Further, in steps 1111 and 1113, the packet passes through the remote TOR 915 to reach the remote NVD 913. The remote NVD 913 decapsulates the packet, reads the destination address (ie, B100) contained in the packet, and finally serves the packet to the remote host 911.
[0147] Reference is now made to FIG. 11B, which is a flowchart illustrating steps performed in transmitting a packet from a remote host (e.g., remote host 911 of FIG. 9) to a compute host (e.g., host 928 of FIG. 9). The process illustrated in FIG. 11B may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 11B and described below is exemplary and not intended to be limiting. While FIG. 11B depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, these steps may be performed in some different order, or some steps may be performed in parallel.
[0148] The process begins in step 1151, when a VM in remote host 911 with an overlay IP address of B100 sends a packet destined for VM1 (with an overlay IP address of B2) to NVD 913. Naturally, the description of FIG. 11B assumes that both the remote host machine and the local host machine are included in the DRCC. However, if the remote host machine is located outside the DRCC, a packet sent by such a remote host "enters" the DRCC by means of FastConnect (or IPSec VPN) or the like, and then reaches the NVD of the host machine included in the DRCC using a DRG. In step 1153, upon receiving the packet, NVD 913 performs a lookup operation in its VCN forwarding table. Specifically, NVD 913 determines that an NVD with a loopback IP address of A254 serves VM1 with an overlay address of B2.
[0149] In step 1155, the NVD 913 encapsulates information in the packet (e.g., in a VCN header) corresponding to the loopback IP address (A254) of the NVD 926 as the intended destination of the packet and its own board IP address (A100) as the source of the packet, after which, in step 1157, the NVD 913 forwards the packet to the remote TOR, which further passes the packet to the network fabric 920.
[0150] In step 1159, upon receiving the packet, the network fabric 920 performs a hashing operation (e.g., a modulo 2 operation) to select one of the routes (via TOR#1 or TOR#2) and forward the packet to the NVD 926. Of course, the hashing operation performed by a switch (e.g., in the network fabric) may be different from the hashing operation performed by the NVD.
[0151] In step 1161, the NVD 926 receives the packet. In step 1163, the NVD 926 may perform another hash operation to select one of two logical ports (e.g., the port with the overlay IP addresses of B1 and C1) to forward the packet to VM1, the packet's intended destination. Of course, as described above, the packet may arrive at either of the NVD's ports. For example, if the NVD is configured in active / active mode of operation, in one embodiment, approximately half of the flows will arrive at either port. However, if the NVD is configured in active / backup mode of operation, all of the flows will arrive at the active port (until the active port fails). In step 1165, the NVD forwards the packet to the VM on the NVD's selected logical interface.
[0152] According to some embodiments, traffic originating from local host machines and destined for remote host machines (and vice versa) is statistically load balanced (e.g., by utilizing ECMP routing) across the dual TORs 922 and 924. The advantage of using dual TORs in a DRCC architecture is failover, i.e., traffic from a failed TOR can be switched to the other (functioning) TOR. Furthermore, the architecture of FIG. 9 uses less hardware, i.e., a single local NVD (e.g., a smart NIC), as opposed to multiple smart NICs. Furthermore, the above-described mechanism for utilizing dual TORs is not limited to use only in datacenter on-premise architectures such as those described above. Rather, the concepts of utilizing dual TORs as described herein are equally applicable to any deployment, i.e., a commercial region, a DRCC, or any other type of network deployment.
[0153] Referring now to FIG. 12, another exemplary architecture 1200 of a DRCC framework is shown, providing customers with all the functionality of a public cloud. This allows customers to reduce infrastructure and operational costs while upgrading legacy applications onto modern cloud services to meet the most stringent regulatory, data residency, and latency requirements. FIG. 12 illustrates a data center 1205 that includes a host machine 1228 coupled to a remote host machine 1211 via a network fabric 1220. It should be understood that the remote host machine could be “any” host machine, such as (i) another host in a DRCC behind another NVD, (ii) another host in another DRCC behind another NVD, or (iii) a host included in a customer’s on-premises network. If the host is included in the customer’s on-premises network, it may connect to the DRCC via FastConnect or an IPSec VPN and also connect to the host 928 through the use of a Dynamic Routing Gateway (DRG). For purposes of explanation, the following description assumes that a remote host (e.g., host 1211) is included in the DRCC and is located behind another NVD (e.g., NVD 1213 served by remote TOR 1215), although the features described below are equally applicable to the other cases of remote hosts described above.
[0154] The data center 1205 includes a pair of TORs (i.e., TOR#1 1222 and TOR#2 1224), a network virtualization platform (e.g., NVD 1226), and a compute host machine 1228 (i.e., the host machine referred to herein as the local host machine). NVD 1226 is referred to herein as the local NVD. The compute host machine 1228 includes a host network interface card (i.e., a host NIC) and multiple virtual machines (VMs). For purposes of illustration, FIG. 12 illustrates the compute host 1228 as including two virtual machines, VM1 and VM2, respectively. Note that the VMs are each communicatively coupled to the host NIC via one of the logical interfaces (e.g., the logical interfaces shown as PF1 and PF2, respectively). Furthermore, the local NVD 1226 may be located on the same chassis as the host NICs included in the compute instances 1228.
[0155] According to some embodiments, the local NVD 1226 has multiple physical ports. For example, in one implementation as shown in FIG. 12, the local NVD 1226 has two physical ports: a first physical port 1227A (referred to herein as the TOR-facing port) connected to each of the TORs 1222 and 1224, and a second physical port 1227B (referred to herein as the host-facing port) connected to the compute host 1228. The physical port 1227A of the local NVD 1226 can be divided into multiple logical ports. For example, as shown in FIG. 12, the physical port 1227A is divided into two TOR-facing logical ports, one connected to each of the TORs included in the data center. The physical port 1227B is connected to the host NIC. Thus, in contrast to the DRCC implementation of FIG. 9, the DRCC implementation of FIG. 12 has a single connection from the NVD 1226 to the host NIC included in the compute instance 1228. Thus, the host NIC is associated with a single pair of IP overlay and MAC addresses (i.e., B1 and M4). This host NIC's overlay IP address can be reached via two backbone paths, i.e., different TORs (i.e., TOR#1 and TOR#2), providing TOR redundancy.
[0156] FIG. 13 is a flowchart illustrating another process for providing a DRCC, according to some embodiments. The process illustrated in FIG. 13 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 13 and described below is exemplary and not intended to be limiting. While FIG. 13 depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In an alternative embodiment, these steps may be performed in some different order, or some steps may be performed in parallel. In some implementations, the method illustrated in FIG. 13 may be performed by a cloud service provider to provide a DRCC to a customer.
[0157] The process begins in step 1301, where a first physical port of a network virtualization device (NVD) included in a data center is communicatively coupled to a first top-of-rack (TOR) switch and a second TOR switch, where the first and second TOR switches may be included in a rack (as shown in FIGS. 6 and 7). In step 1303, a second physical port of the NVD is communicatively coupled to a network interface card (NIC) associated with a host machine included in the data center.
[0158] The process then moves to step 1305, where the NVD receives the packet from the host machine via the second physical port. In step 1307, the NVD determines a specific TOR for communication of the packet from a group including the first TOR and the second TOR. According to some embodiments, the NVD may perform an equal-cost multi-path routing (ECMP) flow hash to select one of the two TORs. Further, in step 1309, the NVD transmits the packet to the specific TOR to facilitate communication of the packet to a remote host machine (e.g., remote host 1211 of FIG. 12).
[0159] Dedicated Customer Regional Clouds (DRCCs) allow cloud service providers to provide infrastructure in customers' own data centers. The DRCC framework provides all the functional metrics of a public cloud on-premises, enabling enterprises to reduce infrastructure and operational costs while upgrading legacy applications to modern cloud services to meet the most stringent regulatory, data residency, and latency requirements. All of this relies on the CSP's infrastructure. Therefore, it is desirable to design a network architecture that has a small footprint and can deliver cloud services in the customer's on-premises location of their choice. A simple solution to enable the DRCC framework is to use a network design similar to that used in commercial regions. However, the drawback of such an approach is that customers typically do not meet the power and space requirements for deploying such a network, and they may not be able to fully utilize such a network architecture. Therefore, a new network architecture with a small footprint (e.g., fewer devices, racks, etc.) is needed to deliver cloud services to customers in the customer's location of their choice.
[0160] FIG. 14 illustrates an exemplary network fabric architecture for a DRCC according to some embodiments. The DRCC network fabric architecture 1400 includes a combination of compute fabric blocks (referred to herein as CFABs 1420A and 1420B) and network fabric blocks (referred to herein as NFABs 1415). The NFABs 1415 are communicatively coupled to each of the CFAB blocks (1420A and 1420B) via multiple blocks of switches (1405 and 1410). The multiple blocks of switches are also referred to herein as layer 3 (T3) level switches. Each of the multiple blocks of switches includes a predetermined number of switches (e.g., four switches). For example, the block of switches 1405 includes four switches labeled 1405A-1405D, and the block of switches 1410 includes four switches labeled 1410A-1410D.
[0161] According to some embodiments, a compute fabric block (e.g., CFAB block 1420A) is communicatively coupled to multiple blocks of switches 1405, 1410. The compute fabric block 1420A includes a set of one or more racks (e.g., Exadata rack 1424 or compute rack 1423). Each rack in the set of one or more racks includes one or more servers configured to execute one or more customer workloads. Of course, each Exadata rack 1424 may be associated with a virtual machine cluster network 1425. The CFAB block 1420A further includes a first plurality of switches 1422 organized as first multiple levels (e.g., levels labeled CFAB T1 and CFAB T2). The first plurality of switches 1422 communicatively couple the set of one or more racks (e.g., racks 1423, 1424) to the multiple blocks of switches (1405, 1410). Specifically, the first plurality of levels associated with the first plurality of switches 1422 in the compute fabric block 1420A include (i) a first tier 1 level of switches, i.e., CFAB T1, and (ii) a first tier 2 level of switches, i.e., CFAB T2.
[0162] The first tier 1 level of switches has a first end communicatively coupled to a set of one or more racks and a second end communicatively coupled to a first tier 2 level of switches. Meanwhile, the first tier 2 level of switches, i.e., CFAB T2, connects the first tier 1 level of switches, i.e., CFAB T1, to the plurality of blocks of switches 1405, 1410. In one embodiment, the first tier 1 level of switches (CFAB T1) in the compute fabric block includes eight switches, and the first tier 2 level of switches (CFAB T2) in the compute fabric block includes four switches. Each switch in the first tier 1 level of switches in the compute fabric block is connected to each switch in the first tier 2 level of switches in the compute fabric block. Meanwhile, each switch in the first tier 2 level of switches (CFAB T2) in the compute fabric block is connected to at least one switch in each of the plurality of blocks of switches 1405, 1410.
[0163] According to some embodiments, the NFAB block 1415 is communicatively coupled to multiple blocks of switches 1405, 1410. The network fabric block 1415 includes (i) one or more edge devices 1420 and (ii) a second multiple of switches 1418 organized as a second multiple of levels. The one or more edge devices 1420 include a first edge device that provides connectivity to a first external resource. For example, the first external resource may be a public communications network (e.g., the Internet), and the first edge device may be a gateway that provides connectivity to the public communications network. The one or more edge devices 1420 may include a gateway, a backbone edge device, a metro edge device, and a route reflector. Thus, the first edge device (e.g., a gateway) enables access to the first external resource (e.g., the Internet) by workloads executed by servers included in one of a set of one or more racks included in the CFAB block 1420A.
[0164] A second plurality of switches 1418, organized as a second plurality of levels (labeled NFAB T1 and NFAB T2), communicatively couples one or more edge devices 1420 to the plurality of blocks of switches 1405, 1410. In FIG. 14, the connections between the plurality of blocks of switches and the second plurality of switches 1418 are shown as logical configuration 1416 (NFAB T1 stripe). Detailed connections in logical configuration 1416 are described below with reference to FIG. 15. According to some embodiments, the second plurality of levels associated with the second plurality of switches 1418 in the network fabric block 1415 include (i) a second tier 1 level of switches, i.e., NFAB T1, and (ii) a second tier 2 level of switches, i.e., NFAB T2. Each switch in the second tier 2 level of switches is communicatively coupled to each switch in the second tier 1 level of switches, i.e., NFAB T1. In one embodiment, the second tier 1 level of switches in the network fabric block includes eight switches and the second tier 2 level of switches in the network fabric block includes four switches.
[0165] According to some embodiments, the initial deployment of the DRCC network architecture includes deploying an NFAB block (1415), one CFAB block (1420A), and a T3 switch layer (i.e., multiple blocks of switches 1405, 1410) that interconnects the NFABs to the CFABs. Of course, additional CFAB blocks (and additional T3 layer switch blocks) may be deployed on the fly (i.e., in real time) based on customer needs. It should also be understood that the number of switches described above with reference to the first or second tier 1 level of switches in the network fabric (e.g., eight switches) is for illustrative purposes only. The number of switches included in the first or second tier 1 level may be four, sixteen, or any other number of switches, such as a variable number of switches. Similarly, the second tier 2 level of switches in the network fabric block described above includes four switches. However, it will be appreciated that this is for illustrative purposes only, and the actual number of switches at this level may be a variable number of switches (eg, half the number of switches at tier 1).
[0166] 15 illustrates connections between the NFAB block and multiple blocks of switches, as well as connections within the NFAB block, i.e., connections between a second multiple of switches organized as a second multiple of levels in the NFAB. The second multiple of levels associated with the second multiple of switches in the NFAB include (i) a second layer 1 level of switches (i.e., NFAB layer 1 1505) and (ii) a second layer 2 level of switches (i.e., NFAB layer 2 1510). In one example, the second layer 1 level of switches in the NFAB includes eight switches (labeled t1-r1 through t1-r8 in FIG. 15), and the second layer 2 level of switches in the network fabric block includes four switches (labeled t2-r1 through t2-r4 in FIG. 15).
[0167] As shown in FIG. 15, a first subset of switches included in a second Layer 1 level of switches (i.e., NFAB Layer 1) are communicatively coupled at a first end to one or more edge devices. For example, as shown in FIG. 15, switches t1-r1, t1-r2, t1-r3, and t1-r4 are each coupled to a route reflector 1520A and a VPN gateway 1520. Additionally, a second subset of switches included in the second Layer 1 level of switches (e.g., switches t1-r5, t1-r6, t1-r7, and t1-r8) are communicatively coupled at a first end to multiple blocks of switches (i.e., labeled CFAB Layer 3 in FIG. 15). Note that the second subset of switches may also be coupled to WDM metro switches, i.e., switches used to interconnect racks located in different buildings. The first and second subsets of switches included in the second Layer 1 level of switches are coupled at a second end to a second Layer 2 level of switches included in a network fabric block.
[0168] Specifically, the NFAB employs a T2 layer of switches (e.g., four switches) 1510 to provide connections between the switches in the T1 layer 1505. As shown, each of the four switches included in the T2 layer switches of the NFAB is connected to a respective T1 layer switch. In this manner, service enclaves included in different NFAB blocks are communicatively coupled to edge devices via the NFAB fabric. Thus, workloads executed by servers included in a rack of compute fabric blocks access a first external resource (e.g., the Internet) by establishing a connection to a first switch in multiple blocks of switches. 15, this connection is further routed (i) from the first switch to a second switch included in a second subset of switches in a second tier 1 level of switches (e.g., switches t1-r5, t1-r6, t1-r7, and t1-r8), (ii) from the second switch to a third switch included in a second tier 2 level of switches (e.g., one of the switches in the group of switches t1-r1, t1-r2, t1-r3, t1-r4), (iii) from the third switch to a fourth switch included in a first subset of switches in the second tier 1 level of switches (e.g., switches t1-r1, t1-r2, t1-r3, and t1-r4), and (iv) from the fourth switch to a gateway. It should be appreciated that the CFAB and JFAB blocks described above with respect to FIGS. 14 and 15 support 400G connections and operate with a power budget of 100KW. The architecture includes a total of three network racks and an optional fourth rack to support backbone and metro connections.
[0169] FIG. 16 illustrates an exemplary dedicated backbone network for a customer region, according to some embodiments. FIG. 16 illustrates multiple geographic locations (e.g., countries) in which customer DRCCs may be deployed. For example, FIG. 16 illustrates three geographic locations in which customer DRCCs are deployed: geographic location A 1602, geographic location B 1604, and geographic location C 1606. Assume that in each geographic location, a customer has a DRCC deployed in two regions within the geographic location. For example, geographic location A has two regions (region I 1602A and region II 1602B), geographic location B has two regions (region I 1604A and region II 1604B), and geographic location C has two regions (region I 1606A and region II 1606B), in each of which a customer DRCC is deployed.
[0170] Each region includes a pair of dedicated backbone routers, Router 1 and Router 2. As shown in FIG. 16, Router 1 in each region is connected in series to form a dedicated backbone ring network. Similarly, as shown in FIG. 16, Router 2 in each region is connected in series to form another dedicated backbone ring network. Naturally, the pair of dedicated backbone ring networks is dedicated to the customer; that is, no other traffic is allowed on the backbone network. The backbone network may support 10G or 100G encrypted connections and may be prepared for a single link failure, i.e., the disruption of a single backbone link included in the ring formed by Router 1 or the ring formed by Router 2, without disrupting traffic on the backbone network. Naturally, the ring topology shown in FIG. 16 is for illustrative purposes only. The backbone network topology may be a mesh, a toroid, or any other topology based on certain criteria, such as the number of DRCC regions desired by the customer and / or the customer's latency and bandwidth requirements.
[0171] Reference is now made to FIG. 17, which is a flowchart illustrating a process for configuring a network fabric, according to some embodiments. The process illustrated in FIG. 17 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, may be implemented using hardware, or a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method presented in FIG. 17 and described below is exemplary and not intended to be limiting. While FIG. 17 depicts various process steps occurring in a particular order or sequence, this is not intended to be limiting. In some alternative embodiments, these steps may be performed in some different order, or some steps may be performed in parallel. In some implementations, the method illustrated in FIG. 17 may be performed by a cloud service to provide DRCC to customers.
[0172] The process begins in step 1701, where multiple blocks of switches are provided. These correspond to blocks of switches 1405 and 1410 as shown in FIG. 14. The process then moves to step 1703, where a first compute fabric block is provided that is communicatively coupled to the multiple blocks of switches. The first compute fabric block includes (i) a set of one or more racks and (ii) a first plurality of switches organized as a first plurality of levels. Each rack in the set of one or more racks includes one or more servers configured to execute one or more customer workloads. The first plurality of switches communicatively couple the set of one or more racks to the multiple blocks of switches.
[0173] The process then moves to step 1705, where a network fabric block is provided that is communicatively coupled to the plurality of blocks of switches. The network fabric block includes (i) one or more edge devices and (ii) a second plurality of switches organized as a second plurality of levels. The first edge device provides connectivity to a first external resource. For example, the first external resource may be a public communications network (e.g., the Internet), and the first edge device is a gateway that provides connectivity to the public communications network. The first edge device enables access to the first external resource by workloads executed by servers included in one of a set of one or more racks. The second plurality of switches communicatively couple the one or more edge devices to the plurality of blocks of switches.
[0174] Exemplary Embodiments of Cloud Infrastructure As mentioned above, infrastructure as a service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, an IaaS provider may also provide various services (e.g., billing, monitoring, logging, security, load balancing, clustering, etc.) associated with these infrastructure components. Therefore, because these services can be policy-driven, IaaS users can maintain application availability and performance by implementing policies that drive load balancing.
[0175] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and use the cloud provider's services to install other elements of their application stack. For example, a user can log in to an IaaS platform to create virtual machines (VMs), install operating systems (OSs) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform a variety of functions, such as distributing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0176] In most cases, the cloud computing model requires the participation of a cloud provider, which may be, but need not be, a third-party service that specializes in providing IaaS (e.g., providing, renting, selling). An entity may also deploy a private cloud and become its own infrastructure service provider.
[0177] In some examples, IaaS deployment is the process of putting a new application or a new version of an application onto a provisioned application server, etc. It may also include the process of provisioning the server (e.g., installing libraries, daemons, etc.), which is often managed below the hypervisor layer (e.g., server, storage, network hardware, and virtualization) by the cloud provider. As such, the customer may be responsible for handling (e.g., on self-service virtual machines (which can be spun up on demand)), middleware, and / or application deployment, etc.
[0178] In some instances, IaaS provisioning may refer to obtaining computers or virtual hosts for use and even installing the necessary libraries or services on them. In most cases, deployment does not include provisioning, which may need to be performed first.
[0179] Sometimes, there are two distinct problems in IaaS provisioning. First, there is the initial challenge of provisioning an initial set of infrastructure before anything works. Second, there is the challenge of evolving the existing infrastructure after all the provisioning (e.g., adding new services, modifying services, removing services, etc.). Sometimes, these two challenges can be addressed by allowing the declarative definition of the infrastructure configuration. In other words, the infrastructure (e.g., the required components and how they interact) can be specified by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., resource dependencies and how they work together) can be declaratively described. Sometimes, once the topology is specified, workflows can be generated to configure and / or manage the various components described in the configuration files.
[0180] In some examples, the infrastructure may have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as a core network. Also, in some examples, there may be one or more security group rules provisioned to define how the network's security is configured and one or more virtual machines (VMs). Other infrastructure elements, such as load balancers, databases, etc., may also be provisioned. The infrastructure may evolve incrementally as there is a desire and / or addition of more infrastructure elements.
[0181] In some cases, employing continuous deployment techniques may enable deployment of infrastructure code across various virtual computing environments. The described techniques may also enable infrastructure management within these environments. In some instances, a service team may write code that is desired to be deployed to one or more (but often many) different production environments (e.g., across various geographic locations, possibly across the globe). However, in some instances, the infrastructure into which the code will be deployed must first be set up. In some cases, provisioning can be performed manually, a provisioning tool can be used to provision resources, and / or a deployment tool can be used to deploy the code after the infrastructure has been provisioned.
[0182] 18 is a block diagram 1800 illustrating an example pattern of an IaaS architecture according to at least one embodiment. A service operator 1802 may be communicatively coupled to a secure host tenancy 1804, which may include a virtual cloud network (VCN) 1806 and a secure host subnet 1808. In some examples, the service operator 1802 may employ one or more client computing devices, which may be portable handheld devices (e.g., iPhones, mobile phones, iPads, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google Glass head-mounted displays, etc.) running software such as Microsoft Windows Mobile and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, etc., and capable of using the Internet, email, short message service (SMS), BlackBerry, or other communications protocols. Alternatively, the client computing devices may be general-purpose personal computers, examples of which include personal computers and / or laptop computers running various versions of the Microsoft Windows, Apple Macintosh, and / or Linux operating systems. The client computing devices may also be workstation computers running any of a variety of commercially available UNIX or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS.Alternatively or additionally, the client computing device may be any other electronic device, such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device, capable of communicating over a network that can access VCN 1806 and / or the Internet.
[0183] VCN 1806 may include a local peering gateway (LPG) 1810 that may be communicatively coupled to a secure shell (SSH) VCN 1812 via an LPG 1810 included in the SSH VCN 1812. The SSH VCN 1812 may include an SSH subnet 1814, and the SSH VCN 1812 may be communicatively coupled to a control plane VCN 1816 via an LPG 1810 included in the control plane VCN 1816. The SSH VCN 1812 may also be communicatively coupled to a data plane VCN 1818 via the LPG 1810. The control plane VCN 1816 and the data plane VCN 1818 may be included in a service tenancy 1819, which may be owned and / or operated by the IaaS provider.
[0184] The control plane VCN 1816 may include a control plane demilitarized zone (DMZ) tier 1820 that serves as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers may have limited liability and help mitigate security breaches. The DMZ tier 1820 may also include one or more load balancer (LB) subnets 1822, a control plane app tier 1824 that may include an app subnet 1826, and a control plane data tier 1828 that may include a database (DB) subnet 1830 (e.g., a front-end DB subnet and / or a back-end DB subnet). LB subnet 1822 included in control plane DMZ layer 1820 is communicatively coupled to app subnet 1826 included in control plane app layer 1824 and to Internet gateway 1834, which may be included in control plane VCN 1816, and app subnet 1826 may be communicatively coupled to DB subnet 1830, service gateway 1836, and network address translation (NAT) gateway 1838, which are included in control plane data layer 1828. Control plane VCN 1816 may include service gateway 1836 and NAT gateway 1838.
[0185] Control plane VCN 1816 may include a data plane mirror app layer 1840 that may include an app subnet 1826. The app subnet 1826 included in data plane mirror app layer 1840 may include a virtual network interface controller (VNIC) 1842 that may run a compute instance 1844. The compute instance 1844 may communicatively couple the app subnet 1826 of data plane mirror app layer 1840 to the app subnet 1826 that may be included in the data plane app layer 1846.
[0186] The data plane VCN 1818 may include a data plane app layer 1846, a data plane DMZ layer 1848, and a data plane data layer 1850. The data plane DMZ layer 1848 may include an LB subnet 1822 that may be communicatively coupled to an app subnet 1826 of the data plane app layer 1846 and an internet gateway 1834 of the data plane VCN 1818. The app subnet 1826 may be communicatively coupled to a service gateway 1836 of the data plane VCN 1818 and a NAT gateway 1838 of the data plane VCN 1818. Additionally, the data plane data layer 1850 may include a DB subnet 1830 that may be communicatively coupled to the app subnet 1826 of the data plane app layer 1846.
[0187] Internet gateways 1834 of the control plane VCN 1816 and the data plane VCN 1818 may be communicatively coupled to metadata management services 1852, which may be communicatively coupled to the public internet 1854. The public internet 1854 may be communicatively coupled to NAT gateways 1838 of the control plane VCN 1816 and the data plane VCN 1818. Service gateways 1836 of the control plane VCN 1816 and the data plane VCN 1818 may be communicatively coupled to cloud services 1856.
[0188] In some examples, the service gateways 1836 of the control plane VCN 1816 and the data plan VCN 1818 can make application programming interface (API) calls to the cloud services 1856 without going over the public internet 1854. The API calls from the service gateways 1836 to the cloud services 1856 can be unidirectional. The service gateways 1836 can make API calls to the cloud services 1856, and the cloud services 1856 can send the requested data to the service gateways 1836. However, the cloud services 1856 do not have to initiate the API calls to the service gateways 1836.
[0189] In some examples, secure host tenancy 1804 may be directly connected to an otherwise isolated service tenancy 1819. Secure host subnet 1808 may communicate with SSH subnet 1814 through LPG 1810, which may allow bidirectional communication through an otherwise isolated system. Connecting secure host subnet 1808 to SSH subnet 1814 allows secure host subnet 1808 to be accessible to other entities within service tenancy 1819.
[0190] Control plane VCN 1816 may enable users of service tenancy 1819 to configure or provision desired resources. The desired resources provisioned in control plane VCN 1816 may be deployed or used in data plane VCN 1818. In some examples, control plane VCN 1816 may be isolated from data plane VCN 1818, and data plane mirror app layer 1840 of control plane VCN 1816 may communicate with data plane app layer 1846 of data plane VCN 1818 via VNIC 1842, which may be included in data plane mirror app layer 1840 and data plane app layer 1846.
[0191] In some examples, a user or customer of the system may make a request (e.g., perform a create, read, update, or delete (CRUD) operation) over public internet 1854, which may route the request to metadata management service 1852. Metadata management service 1852 may route the request to control plane VCN 1816 through internet gateway 1834. The request may be received by LB subnet 1822 included in control plane DMZ tier 1820. LB subnet 1822 may determine that the request is valid, and in response to this determination, LB subnet 1822 may send the request to app subnet 1826 included in control plane app tier 1824. If the request is validated and a call to public internet 1854 is required, the call to public internet 1854 may be sent to NAT gateway 1838, which may make the call to public internet 1854. Any memory that may be desirable for storage due to the request may be stored in DB subnet 1830.
[0192] In some examples, the data plane mirror app layer 1840 may facilitate direct communication between the control plane VCN 1816 and the data plane VCN 1818. For example, it may be desirable to apply configuration changes, updates, or other suitable modifications to resources included in the data plane VCN 1818. The control plane VCN 1816 may perform the configuration changes, updates, or other suitable modifications to the resources by communicating directly with the resources included in the data plane VCN 1818 via the VNIC 1842.
[0193] In some embodiments, the control plane VCN 1816 and the data plane VCN 1818 may be included in the service tenancy 1819. In this case, a user or customer of the system may not own or operate either the control plane VCN 1816 or the data plane VCN 1818. Alternatively, an IaaS provider may own or operate both the control plane VCN 1816 and the data plane VCN 1818, and both may be included in the service tenancy 1819. This embodiment may enable network isolation that may prevent users or customers from interacting with the resources of other users or customers. This embodiment may also enable private storage of databases by users or customers of the system without having to rely on the public internet 1854, which may not have the desired level of security for storage.
[0194] In another embodiment, LB subnet 1822 included in control plane VCN 1816 can be configured to receive signals from service gateway 1836. In this embodiment, control plane VCN 1816 and data plane VCN 1818 can be configured to be called by customers of the IaaS provider without calling the public internet 1854. Customers of the IaaS provider may desire this embodiment because the databases they use can be stored in service tenancy 1819, which is controlled by the IaaS provider and can be isolated from the public internet 1854.
[0195] 19 is a block diagram 1900 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1902 (e.g., service operator 1802 of FIG. 18 ) may be communicatively coupled to a secure host tenancy 1904 (e.g., secure host tenancy 1804 of FIG. 18 ), which may include a virtual cloud network (VCN) 1906 (e.g., VCN 1806 of FIG. 18 ) and a secure host subnet 1908 (e.g., secure host subnet 1808 of FIG. 18 ). VCN 1906 may include a local peering gateway (LPG) 1910 (e.g., LPG 1810 of FIG. 18 ), which may be communicatively coupled to a secure shell (SSH) VCN 1912 via an LPG 1810 included in the SSH VCN 1912 (e.g., SSH VCN 1812 of FIG. 18 ). SSH VCN 1912 can include SSH subnet 1914 (e.g., SSH subnet 1814 in FIG. 18 ), and SSH VCN 1912 can be communicatively coupled to control plane VCN 1916 via LPG 1910 that is included in control plane VCN 1916 (e.g., control plane VCN 1816 in FIG. 18 ). Control plane VCN 1916 can be included in service tenancy 1919 (e.g., service tenancy 1819 in FIG. 18 ), and data plane VCN 1918 (e.g., data plane VCN 1818 in FIG. 18 ) can be included in customer tenancy 1921, which can be owned or operated by a user or customer of the system.
[0196] The control plane VCN 1916 may include a control plane DMZ tier 1920 (e.g., control plane DMZ tier 1820 of FIG. 18 ) that may include a LB subnet 1922 (e.g., LB subnet 1822 of FIG. 18 ), a control plane app tier 1924 (e.g., control plane app tier 1824 of FIG. 18 ) that may include an app subnet 1926 (e.g., app subnet 1826 of FIG. 18 ), and a control plane data tier 1928 (e.g., control plane data tier 1828 of FIG. 18 ) that may include a database (DB) subnet 1930 (e.g., similar to database subnet 1830 of FIG. 18 ). LB subnet 1922 included in control plane DMZ layer 1920 is communicatively coupled to app subnet 1926 included in control plane app layer 1924 and to Internet gateway 1934 (e.g., Internet gateway 1834 in FIG. 18 ), which may be included in control plane VCN 1916, and app subnet 1926 may be communicatively coupled to DB subnet 1930, service gateway 1936 (e.g., service gateway in FIG. 18 ), and network address translation (NAT) gateway 1938 (e.g., NAT gateway 1838 in FIG. 18 ), which are included in control plane data layer 1928. Control plane VCN 1916 may comprise service gateway 1936 and NAT gateway 1938.
[0197] The control plane VCN 1916 may include a data plane mirror app layer 1940 (e.g., data plane mirror app layer 1840 of FIG. 18 ), which may include an app subnet 1926. The app subnet 1926 included in the data plane mirror app layer 1940 may include a virtual network interface controller (VNIC) 1942 (e.g., VNIC 1842) on which a compute instance 1944 (e.g., similar to compute instance 1844 of FIG. 18 ) may run. The compute instance 1944 may facilitate communication between the app subnet 1926 of the data plane mirror app layer 1940 and the app subnet 1926, which may be included in the data plane app layer 1946, via the VNIC 1942 included in the data plane mirror app layer 1940 and the VNIC 1942 included in the data plane app layer 1946 (e.g., data plane app layer 1846 of FIG. 18 ).
[0198] An internet gateway 1934 included in the control plane VCN 1916 may be communicatively coupled to a metadata management service 1952 (e.g., metadata management service 1852 of FIG. 18 ), which may be communicatively coupled to the public internet 1954 (e.g., public internet 1854 of FIG. 18 ). The public internet 1954 may be communicatively coupled to a NAT gateway 1938 included in the control plane VCN 1916. A service gateway 1936 included in the control plane VCN 1416 may be communicatively coupled to cloud services 1956 (e.g., cloud services 1856 of FIG. 18 ).
[0199] In some examples, data plane VCN 1918 may be included in customer tenancy 1921. In this case, the IaaS provider may provide a control plane VCN 1916 for each customer, and the IaaS provider may configure a unique compute instance 1944 for each customer, which may be included in service tenancy 1919. Each compute instance 1944 may enable communication between the control plane VCN 1916 in service tenancy 1919 and the data plane VCN 1918 in customer tenancy 1921. The compute instance 1944 may enable deployment or use of resources provisioned in the control plane VCN 1916 in service tenancy 1919 in the data plane VCN 1918 in customer tenancy 1921.
[0200] In another example, a customer of the IaaS provider may have a database that resides in customer tenancy 1921. In this example, control plane VCN 1916 may include data plane mirror app tier 1940, which may include app subnet 1926. Data plane mirror app tier 1940 may reside in data plane VCN 1918, but may not reside in data plane VCN 1918. That is, data plane mirror app tier 1940 may be accessible to customer tenancy 1921, but may not reside in data plane VCN 1918, and may not be owned or operated by the IaaS provider's customer. Data plane mirror app tier 1940 may be configured to make calls to data plane VCN 1918, but may not be configured to make calls to any entities included in control plane VCN 1916. A customer may desire deployment or use of resources in data plane VCN 1918 that have been provisioned in control plane VCN 1916, and data plane mirror app layer 1940 may facilitate the deployment or other use of the resources desired by the customer.
[0201] In some embodiments, an IaaS provider's customer can apply filters to data plane VCN 1918. In this embodiment, the customer can determine what data plane VCN 1918 can access, and the customer may choose to restrict access from data plane VCN 1918 to the public internet 1954. The IaaS provider may not be able to apply filters or control data plane VCN 1918's access to any external networks or databases. The customer's application of filters and controls to data plane VCN 1918 in customer tenancy 1921 can help isolate data plane VCN 1918 from other customers and the public internet 1954.
[0202] In some embodiments, cloud services 1956 can access services that may not reside on the public internet 1954, the control plane VCN 1916, or the data plane VCN 1918 through calls made by service gateway 1936. The connection between cloud services 1956 and the control plane VCN 1916 or the data plane VCN 1918 may not be live or continuous. Cloud services 1956 may reside on different networks owned or operated by the IaaS provider. Cloud services 1956 may be configured to receive calls from service gateway 1936 or may not be configured to receive calls from the public internet 1954. Some cloud services 1956 may be isolated from other cloud services 1956, and control plane VCNs 1916 may be isolated from cloud services 1956 that may not be in the same region as the control plane VCN 1916. For example, control plane VCN 1916 may be located in "Region 1," and cloud service "Deployment 13" may be located in Region 1 and Region 2. If a call to Deployment 13 is made by service gateway 1936 included in control plane VCN 1916 located in Region 1, the call may be sent to Deployment 13 in Region 1. In this example, Control Plane VCN 1916 or Deployment 13 in Region 1 may not be communicatively coupled to or in communication with Deployment 13 in Region 2.
[0203] 20 is a block diagram 2000 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 2002 (e.g., service operator 1802 of FIG. 18 ) may be communicatively coupled to a virtual cloud network (VCN) 2006 (e.g., VCN 1806 of FIG. 18 ) and a secure host tenancy 2004 (e.g., secure host tenancy 1804 of FIG. 18 ), which may include a secure host subnet 2008 (e.g., secure host subnet 1808 of FIG. 18 ). VCN 2006 may include an LPG 2010 that may be communicatively coupled to an SSH VCN 2012 via an LPG 2010 (e.g., LPG 1810 of FIG. 18 ) included in SSH VCN 2012 (e.g., SSH VCN 1812 of FIG. 18 ). SSH VCN 2012 may include SSH subnet 2014 (e.g., SSH subnet 1814 in FIG. 18 ), and SSH VCN 1812 may be communicatively coupled to control plane VCN 2016 via LPG 2010 included in control plane VCN 2016 (e.g., control plane VCN 1816 in FIG. 18 ), and may be communicatively coupled to data plane VCN 2018 via LPG 2010 included in data plane VCN 2018 (e.g., data plane VCN 1818 in FIG. 18 ). Control plane VCN 2016 and data plane VCN 2018 may be included in service tenancy 2019 (e.g., service tenancy 1819 in FIG. 18 ).
[0204] The control plane VCN 1816 may include a control plane DMZ layer 1820 (e.g., control plane DMZ layer 1820 of FIG. 18 ) that may include a load balancer (LB) subnet 1822 (e.g., LB subnet 1822 of FIG. 18 ), a control plane app layer 2024 (e.g., control plane app layer 1824 of FIG. 18 ) that may include an app subnet 2026 (e.g., similar to app subnet 1826 of FIG. 18 ), and a control plane data layer 2028 (e.g., control plane data layer 1828 of FIG. 18 ) that may include a DB subnet 2030. LB subnet 2022 included in control plane DMZ layer 2020 is communicatively coupled to app subnet 2026 included in control plane app layer 2024 and to Internet gateway 1834 (e.g., Internet gateway 1834 in FIG. 18 ), which may be included in control plane VCN 2016, and app subnet 2026 may be communicatively coupled to DB subnet 1830, service gateway 1836 (e.g., service gateway in FIG. 18 ), and network address translation (NAT) gateway 1838 (e.g., NAT gateway 1838 in FIG. 18 ), which are included in control plane data layer 1828. Control plane VCN 2016 may comprise service gateway 2036 and NAT gateway 2038.
[0205] The data plane VCN 2018 may include a data plane app layer 2046 (e.g., data plane app layer 1846 in FIG. 18 ), a data plane DMZ layer 2048 (e.g., data plane DMZ layer 1848 in FIG. 18 ), and a data plane data layer 2050 (e.g., data plane data layer 1850 in FIG. 18 ). The data plane DMZ layer 2048 may include a LB subnet 2022 that may be communicatively coupled to a trusted app subnet 2060 and an untrusted app subnet 2062 of the data plane app layer 2046 and an Internet gateway 2034 included in the data plane VCN 2018. The trusted app subnet 2060 may be communicatively coupled to a service gateway 2036 included in the data plane VCN 2018, a NAT gateway 2038 included in the data plane VCN 2018, and a DB subnet 2030 included in the data plane data layer 2050. The untrusted app subnet 2062 may be communicatively coupled to a service gateway 2036 included in the data plane VCN 2018 and a DB subnet 2030 included in the data plane data tier 2050. The data plane data tier 2050 may include a DB subnet 2030 that may be communicatively coupled to a service gateway 2036 included in the data plane VCN 2018.
[0206] Untrusted app subnet 2062 may include one or more primary VNICs 2064(1)-2064(N), which may be communicatively coupled to tenant virtual machines (VMs) 2066(1)-2066(N). Each tenant VM 2066(1)-2066(N) may be communicatively coupled to a respective app subnet 2067(1)-2067(N), which may be included in a respective container egress VCN 2068(1)-2068(N), which may be included in a respective customer tenancy 2070(1)-2070(N). Each secondary VNIC 2072(1)-2072(N) may facilitate communication between untrusted app subnet 2062, which is included in data plane VCN 2018, and the app subnets included in the respective container egress VCNs 2068(1)-2068(N). Each container egress VCN 2068(1)-2068(N) may include a NAT gateway 2038 that may be communicatively coupled to the public internet 2054 (e.g., public internet 1854 in FIG. 18).
[0207] An internet gateway 2034 included in the control plane VCN 2016 and the data plane VCN 2018 may be communicatively coupled to a metadata management service 2052 (e.g., metadata management system 1852 of FIG. 18 ), which may be communicatively coupled to the public internet 2054. The public internet 2054 may be communicatively coupled to a NAT gateway 2038 included in the control plane VCN 2016 and the data plane VCN 2018. A service gateway 2036 included in the control plane VCN 2016 and the data plane VCN 2018 may be communicatively coupled to cloud services 2056.
[0208] In some embodiments, data plane VCN 2018 may be integrated with customer tenancy 2070. This integration may be useful or desirable for an IaaS provider's customer, such as if they may want support for running code. A customer may provide code to run, which may be destructive, communicate with other customer resources, or otherwise have undesirable effects. In response, the IaaS provider may determine whether to run the code provided to it by the customer.
[0209] In some examples, a customer of an IaaS provider may grant temporary network access to the IaaS provider to request functionality provided in the data plane app layer 2046. The code that performs this functionality may be configured to run in VMs 2066(1)-2066(N) and may not be configured to run elsewhere on the data plane VCN 2018. Each VM 2066(1)-2066(N) may be connected to a single customer tenancy 2070. Each container 2071(1)-2071(N) contained in VMs 2066(1)-2066(N) may be configured to run code. In this case, double isolation may exist (e.g., containers 2071(1)-2071(N) executing code may be contained in VMs 2066(1)-2066(N) contained in at least untrusted app subnet 2062), which may help prevent erroneous or unwanted code from damaging the IaaS provider's network or a different customer's network. Containers 2071(1)-2071(N) may be communicatively coupled to customer tenancy 2070 and may be configured to send or receive data from customer tenancy 2070. Containers 2071(1)-2071(N) may be configured not to send or receive data from any other entity in data plane VCN 2018. Upon completion of code execution, the IaaS provider may disable or discard containers 2071(1)-2071(N).
[0210] In some embodiments, trusted app subnet 2060 may execute code that may be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 2060 may be communicatively coupled to DB subnet 2030 and may be configured to perform CRUD operations on DB subnet 2030. Untrusted app subnet 2062 may be communicatively coupled to DB subnet 2030, but in this embodiment, may be configured to perform read operations on DB subnet 2030. Containers 2071(1)-2071(N) included in each customer's VM 2066(1)-2066(N) and that may execute code from the customer may not be communicatively coupled to DB subnet 2030.
[0211] In other embodiments, the control plane VCN 2016 and the data plane VCN 2018 may not be directly communicatively coupled. In this embodiment, there may not be direct communication between the control plane VCN 2016 and the data plane VCN 2018. However, communication may occur indirectly in at least one manner. An IaaS provider may establish an LPG 2010 that may facilitate communication between the control plane VCN 2016 and the data plane VCN 2018. In another example, the control plane VCN 2016 or the data plane VCN 2018 may make a call to a cloud service 2056 via a service gateway 2036. For example, a call from the control plane VCN 2016 to the cloud service 2056 may include a request for a service that may communicate with the data plane VCN 2018.
[0212] 21 is a block diagram 2100 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 2102 (e.g., service operator 1802 of FIG. 18 ) may be communicatively coupled to a secure host tenancy 2104 (e.g., secure host tenancy 1804 of FIG. 18 ), which may include a virtual cloud network (VCN) 2106 (e.g., VCN 1806 of FIG. 18 ) and a secure host subnet 2108 (e.g., secure host subnet 1808 of FIG. 18 ). VCN 2106 may comprise an LPG 2110, which may be communicatively coupled to an SSH VCN 2112 via an LPG 2110 (e.g., LPG 1810 of FIG. 18 ) included in SSH VCN 2112 (e.g., SSH VCN 1812 of FIG. 18 ). SSH VCN 2112 may include SSH subnet 2114 (e.g., SSH subnet 1814 in FIG. 18 ), and SSH VCN 2112 may be communicatively coupled to control plane VCN 2116 via LPG 2110 included in control plane VCN 2116 (e.g., control plane VCN 1816 in FIG. 18 ), and may be communicatively coupled to data plane VCN 2118 via LPG 2110 included in data plane VCN 2118 (e.g., data plane VCN 1818 in FIG. 18 ). Control plane VCN 2116 and data plane VCN 2118 may be included in service tenancy 2119 (e.g., service tenancy 1819 in FIG. 18 ).
[0213] The control plane VCN 2116 may include a control plane DMZ layer 2120 (e.g., control plane DMZ layer 1820 of FIG. 18 ) that may include a LB subnet 2122 (e.g., LB subnet 1822 of FIG. 18 ), a control plane app layer 2124 (e.g., control plane app layer 1824 of FIG. 18 ) that may include an app subnet 2126 (e.g., app subnet 1826 of FIG. 18 ), and a control plane data layer 2128 (e.g., control plane data layer 1828 of FIG. 18 ) that may include a DB subnet 2130 (e.g., DB subnet 2030 of FIG. 20 ). LB subnet 2122 included in control plane DMZ layer 2120 is communicatively coupled to app subnet 2126 included in control plane app layer 2124 and to Internet gateway 2134 (e.g., Internet gateway 1834 in FIG. 18 ) that may be included in control plane VCN 2116, and app subnet 2126 may be communicatively coupled to DB subnet 2130, service gateway 2136 (e.g., service gateway in FIG. 18 ), and network address translation (NAT) gateway 2138 (e.g., NAT gateway 1838 in FIG. 18 ) that are included in control plane data layer 2128. Control plane VCN 2116 may comprise service gateway 2136 and NAT gateway 2138.
[0214] Data plane VCN 2118 may include a data plane app layer 2146 (e.g., data plane app layer 1846 in FIG. 18 ), a data plane DMZ layer 2148 (e.g., data plane DMZ layer 2148 in FIG. 18 ), and a data plane data layer 2150 (e.g., data plane data layer 1850 in FIG. 18 ). Data plane DMZ layer 2148 may include a trusted app subnet 2160 (e.g., trusted app subnet 2060 in FIG. 20 ) and a non-trusted app subnet 2162 (e.g., non-trusted app subnet 2062 in FIG. 20 ) of data plane app layer 2146 and an LB subnet 2122 that may be communicatively coupled to an Internet gateway 2134 included in data plane VCN 2118. The trusted app subnet 2160 may be communicatively coupled to a service gateway 2136 included in the data plane VCN 2118, a NAT gateway 2138 included in the data plane VCN 2118, and a DB subnet 2130 included in the data plane data layer 2150. The untrusted app subnet 2162 may be communicatively coupled to the service gateway 2136 included in the data plane VCN 2118 and the DB subnet 2130 included in the data plane data layer 2150. The data plane data layer 2150 may comprise a DB subnet 2130 that may be communicatively coupled to the service gateway 2136 included in the data plane VCN 2118.
[0215] The untrusted app subnet 2162 may include primary VNICs 2164(1)-2164(N), which may be communicatively coupled to tenant virtual machines (VMs) 2166(1)-2166(N) residing within the untrusted app subnet 2162. Each tenant VM 2166(1)-2166(N) may execute code in a respective container 2167(1)-2167(N), and may be communicatively coupled to an app subnet 2126, which may be included in a data plane app layer 2146, which may be included in a container egress VCN 2168. The secondary VNICs 2172(1)-2172(N), respectively, may facilitate communication between the untrusted app subnet 2162, which may be included in the data plane VCN 2118, and the app subnet, which may be included in the container egress VCN 2168. The container egress VCN may include a NAT gateway 2138 that may be communicatively coupled to the public internet 2154 (e.g., public internet 1854 in FIG. 18).
[0216] An internet gateway 2134 included in the control plane VCN 2116 and the data plane VCN 2118 may be communicatively coupled to a metadata management service 2152 (e.g., metadata management system 1852 of FIG. 18 ), which may be communicatively coupled to the public internet 2154. The public internet 2154 may be communicatively coupled to a NAT gateway 2138 included in the control plane VCN 2116 and the data plane VCN 2118. A service gateway 2136 included in the control plane VCN 2116 and the data plane VCN 2118 may be communicatively coupled to cloud services 2156.
[0217] In some examples, the pattern illustrated by the architecture of block diagram 2100 in FIG. 21 may be considered an exception to the pattern illustrated by the architecture of block diagram 2000 in FIG. 20 and may be desirable for customers of an IaaS provider when the IaaS provider cannot communicate directly with the customer (e.g., in an unconnected region). Each of containers 2167(1)-2167(N) contained in VMs 2166(1)-2166(N) for each customer may be accessible in real time by the customer. Each of containers 2167(1)-2167(N) may be configured to make calls to a respective secondary VNIC 2172(1)-2172(N) contained in app subnet 2126 of data plane app tier 2146, which may be included in container egress VCN 2168. Secondary VNICs 2172(1)-2172(N) can send the calls to NAT gateway 2138, which may send the calls to public Internet 2154. In this example, containers 2167(1)-2167(N) that are accessible in real time by customers may be isolated from control plane VCN 2116 and may be isolated from other entities included in data plane VCN 2118. Containers 2167(1)-2167(N) may also be isolated from resources of other customers.
[0218] In another example, a customer can use containers 2167(1)-2167(N) to invoke cloud service 2156. In this example, the customer can execute code in containers 2167(1)-2167(N) that requests a service from cloud service 2156. Containers 2167(1)-2167(N) can send the request to secondary VNICs 2172(1)-2172(N), which can send the request to a NAT gateway, which can send the request to public internet 2154. Public internet 2154 can send the request to LB subnet 2122, which is included in control plane VCN 2116, via internet gateway 2134. In response to determining that the request is valid, the LB subnet can send the request to the app subnet 2126, which can send the request to the cloud service 2156 via the service gateway 2136.
[0219] It should be understood that the IaaS architectures 1800, 1900, 2000, and 2100 depicted in the figures may include components other than those shown. Furthermore, the illustrated embodiment is merely an example of a cloud infrastructure system that may incorporate an embodiment of the present disclosure. In other embodiments, an IaaS system may include more or fewer components than those depicted, may combine two or more components, or may have a different configuration or arrangement of components.
[0220] In some embodiments, the IaaS system described herein may include a suite of application, middleware, and database services delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. Oracle Cloud Infrastructure (OCI), offered by the present assignee, is an example of such an IaaS system.
[0221] 22 illustrates an exemplary computer system 2200 upon which various embodiments of the present disclosure may be implemented. System 2200 may be adapted to implement any of the computer systems described above. As shown in the figure, computer system 2200 includes a processing unit 2204 that communicates with a number of peripheral subsystems via a bus subsystem 2202. The peripheral subsystems may include a processing acceleration unit 2206, an I / O subsystem 2208, a storage subsystem 2218, and a communication subsystem 2224. Storage subsystem 2218 includes a tangible computer-readable storage medium 2222 and a system memory 2210.
[0222] Bus subsystem 2202 provides a mechanism for allowing the various components and subsystems of computer system 2200 to communicate with each other as desired. While bus subsystem 2202 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2202 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, which may be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, a MicroChannel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0223] Processing unit 2204, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 2200. Processing unit 2204 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 2204 may be implemented as one or more independent processing units 2232 and / or 2234, each including a single-core or multi-core processor. In other embodiments, processing unit 2204 may be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.
[0224] In various embodiments, processing unit 2204 executes various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code being executed may reside in processor 2204 and / or in storage subsystem 2218. With suitable programming, processor 2204 can provide the various functions described above. Computer system 2200 may also include a processing acceleration unit 2206, which may include a digital signal processor (DSP), a special purpose processor, and / or the like.
[0225] The I / O subsystem 2208 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, an audio input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may include a motion-sensing and / or gesture-recognition device such as a Microsoft Kinect® motion sensor that enables a user to control and interact with an input device such as a Microsoft Xbox® 360 game controller through a natural user interface using gestures and voice commands. User interface input devices may also include an eye-gesture recognition device such as a Google Glass® blink detector that detects a user's eye activity (e.g., "blinking" during photography and / or menu selection) and translates the eye gesture as input to an input device (e.g., Google Glass®). The user interface input devices may also include a voice recognition sensing device that allows a user to interact with a voice recognition system (e.g., the Siri® navigator) via voice commands.
[0226] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode reader 3D scanners, 3D printers, laser range finders, and eye-tracking devices. User interface input devices may also include medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound devices. User interface input devices may also include audio input devices such as MIDI keyboards and digital musical instruments.
[0227] User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all conceivable types of devices and mechanisms for outputting information from computer system 2200 to a user or to another computer. For example, user interface output devices may include, but are not limited to, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, sound output devices, and modems.
[0228] Computer system 2200 may include a storage subsystem 2218 that includes software elements shown here as located in system memory 2210. System memory 2210 may store program instructions that can be loaded and executed by processing unit 2204, as well as data generated during the execution of these programs.
[0229] Depending on the configuration and type of computer system 2200, system memory 2210 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by the processing unit 2204. In some embodiments, system memory 2210 may include a number of different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some embodiments, ROM may typically store a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within computer system 2200, such as during start-up. Also, by way of non-limiting example, system memory 2210 illustrates application programs 2212, program data 2214, and operating system 2216, which may include client applications, a web browser, a middle-tier application, a relational database management system (RDBMS), etc. By way of example, operating system 2216 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 17 OS, and Palm® OS operating systems.
[0230] The storage subsystem 2218 may also provide a tangible computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some embodiments. The storage subsystem 2218 may store software (programs, code modules, instructions) that, when executed by a processor, provide the functionality described above. These software modules or instructions may be executed by the processing unit 2204. The storage subsystem 2218 may also provide a repository for storing data used in accordance with the present disclosure.
[0231] Storage subsystem 2200 may also include computer-readable storage medium reader 2220, which may further be connected to computer-readable storage medium 2222. Computer-readable storage medium 2222, integral with, and optionally in combination with, system memory 2210, may comprehensively represent remote, local, fixed, and / or removable storage devices, as well as storage media for temporarily and / or permanently containing, storing, transmitting, and retrieving computer-readable information.
[0232] Additionally, the computer-readable storage medium 2222 containing the code or portions of code can include any suitable medium (including storage media and communication media) known or used in the art, including, but not limited to, volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information. This can include tangible computer-readable storage media such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory, and other memory technologies; optical storage such as CD-ROM, digital versatile disk (DVD), magnetic storage such as magnetic cassettes, magnetic tape, magnetic disk storage, or other tangible computer-readable media. It can also include intangible computer-readable media such as a data signal, data transmission, or any other medium usable to transmit the desired information and accessible by computer system 2200.
[0233] By way of example, computer-readable storage medium 2222 may include hard disk drives that read from and write to non-removable, nonvolatile magnetic media, magnetic disk drives that read from and write to removable, nonvolatile magnetic disks, and optical disk drives that read from and write to removable, nonvolatile optical disks such as CD-ROMs, DVDs, Blu-Ray® disks, or other optical media. Computer-readable storage medium 2222 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD disks, digital video tapes, etc. Computer-readable storage medium 2222 may also include solid-state drives (SSDs) based on nonvolatile memory such as flash memory-based SSDs, enterprise flash drives, semiconductor ROMs, semiconductor RAMs, dynamic RAMs, static RAMs, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and volatile memory-based SSDs such as hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for computer system 2200.
[0234] The communications subsystem 2224 provides an interface to other computer systems and networks. The communications subsystem 2224 serves as an interface for transmitting and receiving data between the computer system 2200 and other systems. For example, the communications subsystem 2224 may enable the computer system 2200 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 2224 may include a wireless voice and / or data network (e.g., using cellular technology, 3G, 4G, or advanced data network technology such as EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.11 family standard, or other mobile communications technology, or any combination thereof)), a global positioning system (GPS) receiver component, and / or a radio frequency (RF) transceiver component for accessing other components. In some embodiments, the communications subsystem 2224 may provide a wired network connection (e.g., Ethernet) in addition to or as an alternative to a wireless interface.
[0235] Additionally, in some embodiments, communications subsystem 2224 can receive input information in the form of structured and / or unstructured data feeds 2226, event streams 2228, event updates 2230, etc., on behalf of one or more users who may be using computer system 2200.
[0236] As an example, the communications subsystem 2224 may be configured to receive data feeds 2226 in real time from users of other communications services, such as social networks and / or web feeds such as Twitter® feeds, Facebook® updates, RSS (Rich Site Summary) feeds, and / or real-time updates from one or more third-party sources.
[0237] The communications subsystem 2224 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 2228 of real-time events and / or event updates 2230, which may be continuous or effectively infinite with no apparent end. Examples of applications that generate continuous data include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0238] Communications subsystem 2224 may also be configured to output structured and / or unstructured data feeds 2226, event streams 2228, event updates 2230, etc. to one or more databases that may be in communication with one or more streaming data source computers coupled to computer system 2200.
[0239] The computer system 2200 can be one of a variety of types, including a portable handheld device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA, etc.), a wearable device (e.g., a Google Glass® head-mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or any other data processing system.
[0240] Due to the ever-changing nature of computers and networks, the description of the illustrated computer system 2200 is intended as an example only. Many other configurations are possible, whether containing more or fewer components than the illustrated system. For example, the use of customized hardware and / or implementation of particular elements in hardware, firmware, software (including applets), or a combination is also contemplated. Furthermore, connection to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other ways and / or methods of implementing various embodiments.
[0241] While specific embodiments of the present disclosure have been described above, various improvements, modifications, alternative configurations, and equivalents are also within the scope of the present disclosure. The embodiments of the present disclosure are not limited to operation in a particular data processing environment, but may freely operate in multiple data processing environments. Furthermore, while the embodiments of the present disclosure have been described using a particular sequence of transactions and steps, it will be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described sequence of transactions and steps. Various features and aspects of the above-described embodiments may be used individually or together.
[0242] Furthermore, while embodiments of the present disclosure have been described using particular combinations of hardware and software, it will be recognized that other combinations of hardware and software are within the scope of the present disclosure. Embodiments of the present disclosure may be implemented solely in hardware, solely in software, or using a combination thereof. The various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, when a component or module is described as being configured to perform an operation, such configuration may be achieved, for example, by designing electronic circuitry to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or any combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication. Different sets of processes may use different techniques, or the same set of processes may use different techniques at different times.
[0243] Accordingly, the specification and drawings are to be regarded as illustrative and not restrictive in any sense. However, it will be apparent that additions, differences, deletions, and other modifications and alterations may be made thereto without departing from the broad spirit and scope of the appended claims. Thus, while particular embodiments of the present disclosure have been described, they are not intended to be limiting. Various modifications and equivalents are intended to be encompassed within the scope of the following claims.
[0244] Use of the terms "a," "an," and "the," and similar referents in the context of describing embodiments of the disclosure (particularly in the context of the claims below) shall be interpreted to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" shall be interpreted as open-ended (i.e., meaning "including, but not limited to"), unless otherwise noted. The term "connected" shall be interpreted as partly or wholly contained within, attached to, or integrally connected to, even if there is something intervening. The recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value within the range, unless otherwise indicated herein, and each separate value is incorporated herein as if it were individually stated herein. Unless otherwise indicated herein or clearly contradicted by context, all methods described herein can be performed in any suitable order. The use of any and all examples or exemplary language (e.g., "such as") herein is intended only to facilitate understanding of embodiments of the disclosure and does not limit the scope of the disclosure unless otherwise stated. No language herein should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0245] Unless otherwise noted, disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in the context in which it is generally used to indicate that an item, term, etc. can be either X, Y, Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that an embodiment requires the presence of at least one of X, at least one of Y, or at least one of Z.
[0246] Preferred embodiments of the present disclosure are described herein, including the best mode known to the inventors for carrying out the disclosure. The inventors anticipate that those skilled in the art will adopt such variations as necessary, and they also intend that the present disclosure may be practiced otherwise than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Furthermore, this disclosure includes any combination of the above-described elements in all possible variations thereof unless otherwise indicated herein or otherwise clearly contradicted by context.
[0247] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference in their entirety to the same extent as if each reference were individually and specifically indicated to be incorporated by reference.
[0248] While aspects of the present disclosure have been described in the foregoing specification with reference to particular embodiments, those skilled in the art will recognize that the present disclosure is not limited thereto. Various features and aspects of the present disclosure described above may be used individually or together. Furthermore, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broad spirit and scope of the present disclosure. Accordingly, the present specification and drawings are to be regarded as illustrative, and not limiting.< / realm>
Claims
1. communicatively coupling a first physical port of a network virtualization device (NVD) included in the data center to a first top-of-rack (TOR) switch and a second TOR switch; communicatively coupling a second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a particular TOR for communication of the packet from a group including the first TOR and the second TOR; the NVD sending the packet to the particular TOR to facilitate communication of the packet to a destination host machine; A method comprising:
2. 2. The method of claim 1, wherein the first TOR is different from the second TOR, and the first TOR and the second TOR are contained in a rack in the data center.
3. The method of claim 1 or claim 2, wherein the first physical port of the NVD is associated with a first IP address, a second IP address, a first MAC address, and a second MAC address.
4. The method of claim 1 , claim 2 or claim 3 , wherein the second physical port of the NVD is associated with a first overlay IP address and a first overlay MAC address.
5. 10. The method of any one of the preceding claims, wherein the destination host machine is a remote host machine included in a customer on-premises network.
6. 10. The method of claim 1, wherein the host machine includes a plurality of virtual machines running on the host machine and each associated with a logical interface, and the packet originates from a first virtual machine of the plurality of virtual machines and is sent to the NVD via the logical interface associated with the first virtual machine.
7. 10. The method of claim 1, further comprising the NVD performing a flow hash operation to select one of the first TOR and the second TOR for communication of the packet to the destination host machine.
8. 10. The method of claim 1, wherein the NIC associated with the host machine has a single overlay IP address, and a first packet sent by the destination host machine and destined for the NIC associated with the host machine included in the data center is reachable via the first TOR and / or the second TOR.
9. 1. A computing device, comprising: a processor; When executed by the processor, communicatively coupling a first physical port of a network virtualization device (NVD) included in the data center to a first top-of-rack (TOR) switch and a second TOR switch; communicatively coupling a second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a particular TOR for communication of the packet from a group including the first TOR and the second TOR; the NVD sending the packet to the particular TOR to facilitate communication of the packet to a destination host machine; a memory containing instructions that cause the computing device to: A computing device comprising:
10. 10. The computing device of claim 9, wherein the first TOR is different from the second TOR, and the first TOR and the second TOR are contained in a rack in the data center.
11. 11. The computing device of claim 9 or claim 10, wherein the first physical port of the NVD is associated with a first IP address, a second IP address, a first MAC address, and a second MAC address.
12. 12. The computing device of claim 9, claim 10, or claim 11, wherein the second physical port of the NVD is associated with a first overlay IP address and a first overlay MAC address.
13. The computing device of any one of claims 9 to 12, wherein the destination host machine is a remote host machine included in a customer on-premises network.
14. A computing device as described in any one of claims 9 to 13, wherein the host machine includes a plurality of virtual machines running on the host machine and each associated with a logical interface, and the packet originates from a first virtual machine among the plurality of virtual machines and is sent to the NVD via the logical interface associated with the first virtual machine.
15. When executed by the processor, communicatively coupling a first physical port of a network virtualization device (NVD) included in the data center to a first top-of-rack (TOR) switch and a second TOR switch; communicatively coupling a second physical port of the NVD to a network interface card (NIC) associated with a host machine; the NVD receiving a packet from the host machine via the second physical port of the NVD; the NVD determining a particular TOR for communication of the packet from a group including the first TOR and the second TOR; the NVD sending the packet to the particular TOR to facilitate communication of the packet to a destination host machine; A non-transitory computer-readable medium that stores specific computer-executable instructions that cause a computer system to perform operations including:
16. 16. The non-transitory computer-readable medium of claim 15, wherein the first TOR is different from the second TOR, and the first TOR and the second TOR are contained in a rack in the data center.
17. 17. The non-transitory computer-readable medium of claim 15 or claim 16, wherein the first physical port of the NVD is associated with a first IP address, a second IP address, a first MAC address, and a second MAC address.
18. 18. The non-transitory computer-readable medium of claim 15, claim 16, or claim 17, wherein the host machine includes a plurality of virtual machines running on the host machine and each associated with a logical interface, and the packet originates from a first virtual machine of the plurality of virtual machines and is sent to the NVD via the logical interface associated with the first virtual machine.
19. 19. The non-transitory computer-readable medium of any one of claims 15 to 18, wherein the first physical port of the NVD is associated with a first IP address, a second IP address, a first MAC address, and a second MAC address.
20. The non-transitory computer-readable medium of any one of claims 15 to 19, wherein the destination host machine is a remote host machine included in a customer on-premises network.