Routing policy for image processing device
The GPU routing mechanism optimizes packet forwarding by determining and mapping incoming to outgoing port-links, addressing the blocking issues in ring topologies and enhancing network performance for high-performance compute applications in cloud environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2022-06-23
- Publication Date
- 2026-04-22
AI Technical Summary
Existing cloud infrastructure lacks efficient routing mechanisms for high-performance compute (HPC) applications, particularly in virtualized environments, due to the inherent blocking nature of ring topologies used for connecting GPUs, leading to degraded system performance.
Implementing a GPU routing mechanism that determines the incoming port-link of the GPU routing policy, establishing a pre-configured mapping of incoming to outgoing port-links, and forwarding packets accordingly, thereby optimizing network performance.
Enhances network performance by avoiding blocking issues in GPU connectivity across host machines, ensuring efficient packet routing and improved system performance.
Smart Images

Figure 0007850184000001 
Figure 0007850184000002 
Figure 0007850184000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit and priority under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 215,264, filed Jun. 25, 2021, and U.S. Patent Application No. 17 / 734,833, filed May 2, 2022. The entire content of the above - mentioned applications is hereby incorporated by reference in its entirety for all purposes.
[0002] Field The present disclosure relates to a framework and a routing mechanism for image - processing devices (GPUs) hosted on several host machines in a cloud environment.
Background Art
[0003] Background Organizations are continuing to migrate business applications and databases to the cloud to reduce the costs of purchasing, updating, and maintaining on - premise hardware and software. High - performance compute (HPC) applications always consume 100 percent of the available computing power to achieve a particular outcome or result. HPC applications require dedicated network performance, high - speed storage, high computing power, and large amounts of memory resources, i.e., resources that are lacking in the virtualized infrastructure that makes up today's commodity cloud.
[0004] Cloud infrastructure service providers offer newer and faster CPUs and graphics processing units (GPUs) to meet the requirements of HPC applications. Typically, a virtual topology is constructed to provision various GPUs hosted on several host machines to communicate with each other. In practice, a ring topology is used to connect various GPUs. However, a ring network is inherently blocking, thus degrading the overall system performance. The embodiments discussed herein address these and other issues concerning GPU connectivity across several host machines. [Overview of the project] [Problems that the invention aims to solve]
[0005] overview This disclosure generally relates to routing mechanisms for image processing units (GPUs) hosted on several host machines in a cloud environment. This specification describes various embodiments, including methods, systems, and non-temporary computer-readable storage media for storing programs, code, or instructions executable by one or more processors. These exemplary embodiments are mentioned not to limit or define this disclosure, but to provide examples to aid in understanding this disclosure. Additional embodiments are discussed in the detailed description section, where further explanation is provided.
[0006] One embodiment of the present disclosure is a method comprising the network device determining the incoming port-link of the network device from which a packet transmitted by a host machine's image processing unit (GPU) and received by the network device is received. The method comprises the network device identifying the outgoing port-link corresponding to the incoming port-link based on a GPU routing policy. The GPU routing policy is pre-configured before the packet is received and establishes a mapping of each incoming port-link of the network device to a unique outgoing port-link of the network device. The method comprises the network device forwarding the packet on the outgoing port-link of the network device.
[0007] One aspect of the present disclosure provides a system comprising one or more data processors and a non-temporary computer-readable storage medium containing instructions, when executed by the one or more data processors, that cause the one or more data processors to perform some or all of the methods disclosed herein.
[0008] Another aspect of the present disclosure provides a computer program product tangibly embodied in a non-temporary machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform some or all of the methods disclosed herein.
[0009] The foregoing, along with other features and embodiments, will become clearer by referring to the following specification, claims, and accompanying drawings.
[0010] The features, embodiments, and advantages of this disclosure will be better understood by reading the following detailed description with reference to the accompanying drawings. [Brief explanation of the drawing]
[0011] [Figure 1]This is a schematic diagram illustrating a distributed environment showing a virtual or overlay cloud network hosted by a cloud service provider's infrastructure, according to several embodiments. [Figure 2] This is a simplified architectural diagram of the physical components in the physical network within the CSPI, according to several embodiments. [Figure 3] This figure shows an example of a configuration within a CSPI in which a host machine is connected to multiple network virtualization devices (NVDs), according to several embodiments. [Figure 4] This figure shows connectivity between a host machine and an NVD to provide I / O virtualization that supports multi-tenancy, according to several embodiments. [Figure 5] This is a simplified block diagram of a physical network provided by CSPI, according to several embodiments. [Figure 6] This is a simplified block diagram of a cloud infrastructure incorporating a CLOS network deployment in several embodiments. [Figure 7] This figure shows example scenarios illustrating flow collisions in the cloud infrastructure of Figure 6, based on several embodiments. [Figure 8] This figure shows policy-based routing mechanisms implemented in cloud infrastructure in several embodiments. [Figure 9] This is a block diagram of a cloud infrastructure showing different types of connectivity within the cloud infrastructure, according to several embodiments. [Figure 10] This figure shows examples of rack configurations included in a cloud infrastructure, according to several embodiments. [Figure 11A] This flowchart shows the steps performed by a network device when routing packets, according to several embodiments. [Figure 11B]This is another flowchart illustrating the steps performed by a network device when routing packets, according to several embodiments. [Figure 12] This block diagram shows one pattern for implementing a cloud infrastructure as a service system, according to at least one embodiment. [Figure 13] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 14] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 15] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 16] A block diagram showing an example of a computer system according to at least one embodiment. [Modes for carrying out the invention]
[0012] Detailed explanation In the following description, specific details are provided so that some embodiments may be fully understood for illustrative purposes. However, it will be apparent that various embodiments may be carried out without these specific details. The figures and descriptions are not intended to be restrictive. The term “exemplary” is used herein to mean “serving as an example, illustration, or demonstration.” Embodiments or designs described herein as “exemplary” should not necessarily be construed as being preferable or advantageous to other embodiments or designs. Cloud infrastructure architecture examples The term "cloud service" is generally used to refer to services that can be used on demand (e.g., via a subscription model) by users or customers using the systems and infrastructure (cloud infrastructure) provided by a cloud service provider (CSP). Usually, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premises servers and systems. Thus, customers can utilize cloud services provided by a CSP without the need to separately purchase hardware and software resources for the service. Cloud services are designed to provide easy and scalable access to applications and computing resources to subscribing customers without the need to invest in acquiring the infrastructure used to provide the service.
[0013] There are several cloud service providers that offer different types of cloud services. There are various different types or models of cloud services, including software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), and so on.
[0014] Customers can subscribe to one or more cloud services provided by a CSP. A customer can be any entity such as an individual, an organization, a company, etc. When a customer subscribes or registers for a service provided by a CSP, a tenancy or account is created for that customer. Thereafter, the customer can access the one or more subscribed cloud resources associated with the account via this account.
[0015] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing service. In the IaaS model, the CSP provides the infrastructure (referred to as Cloud Service Provider Infrastructure, or CSPI), which the customer can use to build a customizable network and deploy their resources. In this way, the customer's resources and network are hosted in a distributed environment by the infrastructure provided by the CSP. This differs from traditional computing, where the customer's resources and network are hosted by the infrastructure provided by the customer.
[0016] A CSPI can include interconnected high-performance compute resources, including various host machines, memory resources, and network resources, forming a physical network also known as an underlay network. CSPI resources may be distributed across one or more data centers, which may be geographically distributed across one or more geographical regions. Virtualization software can run on these physical resources to provide a virtualized distributed environment. Virtualization creates an overlay network (also known as a software-based network, software-defined network, or virtual network) on top of the physical network. The CSPI physical network provides the foundation for creating one or more overlay or virtual networks on top of the physical network. A virtual or overlay network can include one or more virtual cloud networks (VCNs). Virtual networks are implemented using software virtualization technologies (e.g., hypervisors, functions performed by network virtualization devices (NVDs) (e.g., smart NICs), top-of-rack (TOR) switches, smart TORs implementing one or more functions performed by NVDs, and other mechanisms) to create a layer of network abstraction that can run on top of the physical network. Virtual networks can take many forms, including peer-to-peer networks and IP networks. Virtual networks are typically either Layer 3 IP networks or Layer 2 VLANs. Such virtual or overlay networking methods are often referred to as virtual or overlay Layer 3 networking.Examples of protocols developed for virtual networks include IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN - IETF RFC7348), Virtual Private Network (VPN) (for example, MPLS Layer 3 Virtual Private Network (RFC 4364)), VMware's NSX, and GENEVE (Generic Network Virtualization Encapsulation).
[0017] In the case of IaaS, the infrastructure provided by the CSP (CSPI) can be configured to deliver virtualized computing resources over a public network (e.g., the internet). In the IaaS model, the cloud computing service provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider can also provide various services associated with those infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering). Therefore, since these services may be policy-driven, IaaS users may be able to implement policies to drive load balancing and maintain application availability and performance. The CSPI provides infrastructure and a set of complementary cloud services, which enable customers to build and run a wide range of applications and services in a highly available hosted distributed environment. The CSPI provides high-performance compute resources and capabilities, as well as storage capacity, on a flexible virtual network that can be securely accessed from various network locations, including the customer's on-premises network. When a customer subscribes to or registers for IaaS services provided by a CSP, the tenancy created for that customer becomes a secure, isolated partition within the CSP where the customer can create, organize, and manage cloud resources.
[0018] Customers can build their own virtual networks using the compute, memory, and networking resources provided by CSPI. They can deploy one or more customer resources or workloads, such as compute instances, on these virtual networks. For example, a customer can build one or more customizable private virtual networks, referred to as Virtual Cloud Networks (VCNs), using the resources provided by CSPI. On a customer VCN, a customer can deploy one or more customer resources, such as compute instances. Compute instances can take the form of virtual machines, bare-metal instances, etc. Thus, CSPI provides the infrastructure and a suite of complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available virtual hosting environment. While customers do not manage or control the underlying physical resources provided by CSPI, they control the operating system, storage, and deployed applications, and, in some cases, have limited control over selected networking components (e.g., firewalls).
[0019] A CSP can provide a console that enables customers and network administrators to configure, access, and manage resources deployed to the cloud using CSPI resources. In some embodiments, the console provides a web-based user interface that can be used to access and manage CSPI. In some embodiments, the console is a web-based application provided by the CSP.
[0020] CSPI can support single-tenancy or multi-tenancy architectures. In a single-tenancy architecture, software (e.g., applications, databases) or hardware components (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenancy architecture, software or hardware components serve multiple customers or tenants. Therefore, in a multi-tenancy architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenancy scenario, CSPI employs precautions and safeguards to isolate each tenant's data and prevent it from being visible to other tenants.
[0021] In a physical network, a network endpoint ("endpoint") refers to a computing device or system that is connected to a physical network and communicates bidirectionally with the connected network. A network endpoint in a physical network may be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints in a physical network include modems, hubs, bridges, switches, routers, and other networking devices, as well as physical computers (or host machines). Each physical device in a physical network has a fixed network address that can be used to communicate with this device. This fixed network address may be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints can include various virtual endpoints, such as virtual machines hosted by components of a physical network (e.g., hosted by a physical host machine). These endpoints in a virtual network are addressed by overlay addresses, such as overlay Layer 2 addresses (e.g., overlay MAC addresses) and overlay Layer 3 addresses (e.g., overlay IP addresses). Network overlays provide flexibility by allowing network administrators to move overlay addresses associated with network endpoints using software management (for example, through software implementing the control plane of the virtual network). Therefore, unlike physical networks, virtual networks allow network management software to move overlay addresses (e.g., overlay IP addresses) from one endpoint to another. Because virtual networks are built on top of physical networks, communication between components of a virtual network involves both the virtual network and the underlying physical network.To facilitate such communication, CSPI components are configured to learn and store mappings that map virtual network overlay addresses to actual physical addresses of the underlying network, or vice versa. These mappings are then used to facilitate communication. To facilitate virtual network routing, customer traffic is encapsulated.
[0022] Therefore, physical addresses (e.g., physical IP addresses) are associated with components of a physical network, while overlay addresses (e.g., overlay IP addresses) are associated with entities in a virtual network. Both physical and overlay IP addresses are types of real IP addresses. They are distinct from virtual IP addresses, which map to multiple real IP addresses. Virtual IP addresses provide a one-to-many mapping between a virtual IP address and multiple real IP addresses.
[0023] Cloud infrastructure, or CSPI, is physically hosted in one or more data centers in one or more regions worldwide. A CSPI can include physical network or underlying network components and virtualization components of virtual networks built on top of the physical network components (e.g., virtual networks, compute instances, virtual machines, etc.). In some embodiments, a CSPI is organized and hosted in realms, regions, and availability domains. A region is typically a local geographic area containing one or more data centers. Regions are generally independent of each other and may be separated by vast distances, for example, across countries or even continents. For example, one region might be in Australia, another in Japan, and yet another in India. CSPI resources are divided across regions such that each region has its own independent subset of CSPI resources. Each region can provide a set of core infrastructure services and resources, including compute resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure), storage resources (e.g., block volume storage, file storage, object storage, archive storage), networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connectivity to on-premises networks), database resources, edge networking resources (e.g., DNS), and access management and monitoring resources. Each region generally has multiple paths connecting it to other regions within the realm.
[0024] Generally, applications are deployed in the region with the highest usage (i.e., on the infrastructure associated with that region) because using nearby resources is faster than using distant resources. Applications can also be deployed in different regions for various reasons, such as redundancy to mitigate the risk of region-wide events like large-scale weather systems or earthquakes, or to meet various requirements for legal jurisdictions, tax areas, and other business or social standards.
[0025] Data centers within a region can be further organized and subdivided into Availability Domains (ADs). An Availability Domain may correspond to one or more data centers located within a region. A region may consist of one or more Availability Domains. In such a distributed environment, CSPI resources are either region-specific, such as virtual cloud networks (VCNs), or availability domain-specific, such as compute instances.
[0026] Active Directory (AD) instances within a single region are configured to be isolated from each other, fault-tolerant, and extremely unlikely to fail simultaneously. This is achieved by ensuring that ADs do not share critical infrastructure resources such as networking, physical cables, cable routes, and cable entry points, so that the failure of one AD within a region has little impact on the availability of other ADs within the same region. ADs within the same region can be connected to each other with low-latency, high-bandwidth networks, providing highly available connectivity to other networks (e.g., the internet, customer on-premises networks, etc.) and enabling the creation of a replication system for both high availability and disaster recovery across multiple ADs. Cloud services use multiple ADs to ensure high availability and protect against resource failures. As the infrastructure provided by IaaS providers grows, more regions and ADs can be added to increase capacity. Traffic between availability domains is typically encrypted.
[0027] In some embodiments, regions are grouped into realms. A realm is a logical collection of regions. Realms are isolated from each other and do not share any data. Regions within the same realm can communicate with each other, but regions in different realms cannot. A CSP customer's tenancy or account resides in a single realm and can span one or more regions belonging to that realm. Typically, when a customer subscribes to an IaaS service, a tenancy or account is created for that customer in a customer-designated region within a realm (referred to as the "home" region). A customer can extend their tenancy to one or more other regions within a realm. A customer cannot access regions that do not reside in the realm where their tenancy resides.
[0028] An IaaS provider can offer multiple realms, each corresponding to a specific set of customers or users. For example, a commercial realm can be offered for commercial customers. Another example is a realm that can be offered for customers in a specific country. Yet another example is the provision of a government realm for a government, for instance. A government realm, for example, can be created for a specific government and may have a higher level of security than a commercial realm. For example, Oracle Cloud Infrastructure (OCI) currently offers a realm for commercial regions and two realms for government cloud regions (e.g., FedRAMP authorization and IL5 authorization).
[0029] In some embodiments, an Active Directory (AD) can be subdivided into one or more fault domains. A fault domain is a grouping of infrastructure resources within the AD to provide anti-affinity. Fault domains enable the distribution of compute instances, ensuring that compute instances do not reside on the same physical hardware within a single AD. This is known as anti-affinity. A fault domain refers to a set of hardware components (computers, switches, etc.) that share a single point of failure. A compute pool is logically divided into fault domains. Therefore, a hardware failure or compute hardware maintenance event affecting one fault domain does not affect instances in other fault domains. Depending on the embodiment, the number of fault domains in each AD may vary. For example, in some embodiments, each AD contains three fault domains. A fault domain functions as a logical data center within the AD.
[0030] When a customer subscribes to an IaaS service, resources from CSPI are provisioned to the customer and associated with the customer's tenancy. The customer can use these provisioned resources to build private networks and deploy resources on these networks. Customer networks hosted in the cloud by CSPI are referred to as Virtual Cloud Networks (VCNs). A customer can set up one or more Virtual Cloud Networks (VCNs) using the CSPI resources allocated to them. A VCN is a virtual or software-defined private network. Customer resources deployed to a customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances can represent various customer workloads such as applications, load balancers, and databases. Compute instances deployed on a VCN can communicate with publicly accessible endpoints ("public endpoints") over public networks such as the internet, communicate with other instances within the same VCN or other VCNs (for example, other VCNs of the customer or VCNs not belonging to the customer), communicate with the customer's on-premises data centers or networks, communicate with service endpoints and other types of endpoints.
[0031] A CSP can provide a variety of services using a CSPI. In some cases, the CSPI customer itself can behave like a service provider and provide services using CSPI resources. A service provider can expose service endpoints characterized by identifying information (e.g., IP address, DNS name, and port). A customer's resources (e.g., compute instances) can consume a particular service by accessing the service endpoint of that particular service exposed by the service. These service endpoints are generally publicly accessible endpoints that users can access via public communication networks such as the internet using the public IP address associated with the endpoint. Publicly accessible network endpoints are sometimes referred to as public endpoints.
[0032] In some embodiments, a service provider may expose a service through an endpoint (sometimes referred to as a service endpoint) of the service. Customers of the service can then access the service using this service endpoint. In some embodiments, the service endpoint provided for a service may be accessible to multiple customers who wish to consume that service. In other embodiments, a dedicated service endpoint may be provided to a customer so that only that customer can access the service using that dedicated service endpoint.
[0033] In some embodiments, a VCN, once created, is associated with a private overlay classless inter-domain routing (CIDR) address space, which is a range of private overlay IP addresses assigned to that VCN (e.g., 10.0 / 16). A VCN includes associated subnets, route tables, and gateways. A VCN resides within a single region but can span one, more, or all of the region's availability domains. A gateway is a virtual interface configured for a VCN that enables traffic communication between the VCN and one or more endpoints outside the VCN. One or more different types of gateways can be configured for a VCN to enable communication between different types of endpoints.
[0034] A VCN can be subdivided into one or more subnets, such as one or more subnets. Therefore, a subnet is a constituent unit or subdivision that can be created within a VCN. A VCN can have one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets within that VCN and represents a subset of the address space within the VCN's address space.
[0035] Each compute instance is associated with a virtual network interface card (VNIC), which allows the compute instance to participate in a VCN subnet. A VNIC is a logical representation of a physical network interface card (NIC). Generally, a VNIC is the interface between an entity (e.g., compute instance, service) and a virtual network. A VNIC resides in a subnet and has one or more associated IP addresses and associated security rules or policies. A VNIC is equivalent to a Layer 2 port on a switch. A VNIC connects a compute instance to a subnet within a VCN. The VNIC associated with a compute instance enables the compute instance to be part of a VCN subnet and allows the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, endpoints in different subnets within the VCN, or endpoints outside the VCN. Thus, the VNIC associated with a compute instance determines how the compute instance connects to internal and external endpoints within the VCN. A compute instance's VNIC is created when the compute instance is created and added to a subnet within the VCN, and is associated with that compute instance. If a subnet contains a set of compute instances, it contains VNICs corresponding to that set of compute instances, and each VNIC connects to a compute instance within that set of compute instances.
[0036] Each compute instance is assigned a private overlay IP address via the VNIC associated with it. This private overlay IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic to and from the compute instance. All VNICs within a given subnet use the same route table, security lists, and DHCP options. As mentioned above, each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that does not overlap with other subnets within that VCN and represents a subset of the address space of that VCN. For a VNIC on a particular subnet of a VCN, the private overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.
[0037] In some embodiments, compute instances can optionally be assigned additional overlay IP addresses in addition to their private overlay IP addresses, such as one or more public IP addresses if they are located in a public subnet. These multiple addresses can be assigned on the same VNIC or across multiple VNICs associated with the compute instance. However, each instance has a primary VNIC, which is created during instance startup and associated with the overlay private IP addresses assigned to the instance, and this primary VNIC cannot be deleted. Additional VNICs, referred to as secondary VNICs, can be added to an existing instance in the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. Secondary VNICs can be in the same VCN subnet as the primary VNIC, or in a different subnet in either the same VCN or a different VCN.
[0038] If a compute instance is in a public subnet, you can optionally assign a public IP address to that compute instance. When a subnet is created, you can specify that subnet as either a public or private subnet. A private subnet means that the resources within that subnet (e.g., compute instances) and associated VNICs cannot have public overlay IP addresses. A public subnet means that the resources within that subnet and associated VNICs can have public IP addresses. Customers can specify subnets that exist within a single availability domain or across multiple availability domains within a region or realm.
[0039] As described above, a VCN can be subdivided into one or more subnets. In some embodiments, a virtual router (VR) configured for the VCN (referred to as a VCN VR or simply VR) enables communication between subnets within the VCN. For subnets within a VCN, the VR represents the logical gateway for that subnet, enabling the subnet (i.e., compute instances on that subnet) to communicate with endpoints on other subnets within the VCN and other endpoints outside the VCN. A VCN VR is a logical entity configured to route traffic between VNICs within the VCN and virtual gateways ("gateways") associated with the VCN. Gateways will be discussed further later with reference to Figure 1. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is one VCN VR for one VCN, and this VCN VR has a potentially unlimited number of ports addressed by IP addresses, with one port for each subnet of the VCN. Thus, a VCN VR has a different IP address for each subnet of the VCN to which the VCN VR is connected. The VR is also connected to various gateways configured for the VCN. In some embodiments, a specific overlay IP address from a subnet's overlay IP address range is reserved for a port in the VCN VR for that subnet. For example, consider a VCN having two subnets, each associated with the address ranges 10.0 / 16 and 10.1 / 16, respectively. For the first subnet in the VCN having the address range 10.0 / 16, addresses from this range are reserved for a port in the VCN VR for that subnet. In some cases, the first IP address from this range may also be reserved for the VCN VR. For example, for a subnet having the overlay IP address range 10.0 / 16, the IP address 10.0.0.1 may be reserved for a port in the VCN VR for that subnet.For a second subnet within the same VCN having the address range 10.1 / 16, the VCN VR may have a port for that second subnet having the IP address 10.1.0.1. The VCN VR has a different IP address for each of the subnets within the VCN.
[0040] In some other embodiments, each subnet within a VCN may have its own associated VR, which is addressable by the subnet using a reserved or default IP address associated with that VR. The reserved or default IP address may be, for example, a first IP address from a range of IP addresses associated with that subnet. A VNIC within a subnet can use this default or reserved IP address to communicate with the VR associated with the subnet (e.g., send and receive packets). In such embodiments, the VR is the ingress / egress point for that subnet. A VR associated with a subnet within a VCN can communicate with other VRs associated with other subnets within the VCN. A VR can also communicate with gateways associated with the VCN. The VR functionality of a subnet runs on, or is performed by, one or more NVDs that perform the VNIC functionality of the VNICs within the subnet.
[0041] Route tables, security rules, and DHCP options can be configured for a VCN. The route table is the VCN's virtual route table, containing rules for routing traffic from subnets within the VCN to destinations outside the VCN, via a gateway or specially configured instance. The VCN's route table can be customized to control how packets are forwarded / routed to and from the VCN. DHCP options refer to configuration information automatically provided to an instance when it is launched.
[0042] Security rules configured for a VCN represent the VCN's overlay firewall rules. Security rules can include ingress and egress rules, specifying the types of traffic allowed to enter and exit instances within the VCN (for example, based on protocol and port). Customers can choose whether a given rule is stateful or stateless. For example, a customer can allow incoming SSH traffic to a set of instances from any location by configuring a stateful ingress rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22. Security rules can be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to resources within that group. A security list, on the other hand, contains rules that apply to all resources within a subnet that uses that security list. A VCN can be provided with a default security list along with default security rules. DHCP options configured for a VCN provide configuration information that is automatically provided to instances within the VCN when they start up.
[0043] In some embodiments, VCN configuration information is determined and stored by the VCN control plane. VCN configuration information may include, for example, the address range associated with the VCN, subnets and associated information within the VCN, one or more VRs associated with the VCN, compute instances and associated VNICs within the VCN, NVDs that perform various virtualized network functions associated with the VCN (e.g., VNICs, VRs, gateways), VCN status information, and other VCN-related information. In some embodiments, a VCN distribution service exposes the configuration information or a portion thereof stored by the VCN control plane to the NVD. The distributed information can be used by the NVD to update information it stores and uses (e.g., forwarding tables, routing tables, etc.) to forward packets to and from compute instances within the VCN.
[0044] In some embodiments, the creation of VCNs and subnets is handled by the VCN control plane (CP), and the launch of compute instances is handled by the compute control plane. The compute control plane is responsible for allocating physical resources for compute instances and then calls the VCN control plane to create VNICs and attach them to compute instances. The VCN CP also sends VCN data mappings to the VCN data plane, which is configured to perform packet forwarding and routing functions. In some embodiments, the VCN CP provides a distribution service responsible for providing updates to the VCN data plane. Examples of VCN control planes are also shown in Figures 12, 13, 14, and 15 (see references 1216, 1316, 1416, and 1516) and will be discussed later.
[0045] Customers can create one or more VCNs using resources hosted by CSPI. Compute instances deployed on a customer VCN can communicate with different endpoints. These endpoints can include endpoints hosted by CSPI and endpoints outside of CSPI.
[0046] Various different architectures for implementing cloud-based services using CSPI are shown in Figures 1, 2, 3, 4, 5, 12, 13, 14, and 15, and are described below. Figure 1 is a schematic diagram of a distributed environment 100 showing an overlay VCN or customer VCN hosted by CSPI, according to several embodiments. The distributed environment shown in Figure 1 includes multiple components within the overlay network. The distributed environment 100 shown in Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, alternative forms, and alterations are possible. For example, in some embodiments, the distributed environment shown in Figure 1 may have more or fewer systems or components than those shown in Figure 1, may combine two or more systems, or may have different configurations or arrangements of systems.
[0047] As shown in the example in Figure 1, the distributed environment 100 includes a CSPI 101 that provides services and resources that customers can subscribe to and use to build their own virtual cloud network (VCN). In some embodiments, the CSPI 101 provides IaaS services to subscribing customers. The data centers within the CSPI 101 may be organized into one or more regions. Figure 1 shows one example region, the "US region" 102. The customer has configured a customer VCN 104 for region 102. The customer can deploy various compute instances on the VCN 104, where compute instances can include virtual machines or bare metal instances. Examples of instances include applications, databases, load balancers, etc.
[0048] In the embodiment shown in Figure 1, customer VCN104 includes two subnets, namely "Subnet-1" and "Subnet-2," each subnet having its own CIDR IP address range. In Figure 1, the overlay IP address range for Subnet-1 is 10.0 / 16, and the address range for Subnet-2 is 10.1 / 16. The VCN virtual router 105 represents the logical gateway of the VCN, enabling communication between subnets of VCN104 and with other endpoints outside the VCN. The VCN VR105 is configured to route traffic between VNICs within VCN104 and gateways associated with VCN104. The VCN VR105 provides ports to each subnet of VCN104. For example, the VR105 can provide a port with IP address 10.0.0.1 to Subnet-1 and a port with IP address 10.1.0.1 to Subnet-2.
[0049] Multiple compute instances can be deployed on each subnet, where compute instances may be virtual machine instances and / or bare metal instances. Compute instances within a subnet may be hosted by one or more host machines within CSPI101. Compute instances join the subnet via the VNIC associated with the compute instance. For example, as shown in Figure 1, compute instance C1 is part of subnet-1 via the VNIC associated with the compute instance. Similarly, compute instance C2 is part of subnet-1 via the VNIC associated with C2. Similarly, multiple compute instances, which may be virtual machine instances or bare metal instances, may be part of subnet-1. Each compute instance is assigned a private overlay IP address and MAC address via its associated VNIC. For example, in Figure 1, compute instance C1 has the overlay IP address 10.0.0.2 and the MAC address M1, and compute instance C2 has the private overlay IP address 10.0.0.3 and the MAC address M2. Each compute instance in subnet-1, including compute instances C1 and C2, has a default route to VCN VR105 using IP address 10.0.0.1, which is the IP address of the port of VCN VR105 in subnet-1.
[0050] Multiple compute instances, including virtual machine instances and / or bare metal instances, can be deployed in subnet-2. For example, as shown in Figure 1, compute instances D1 and D2 are part of subnet-2 via the VNIC associated with each compute instance. In the embodiment shown in Figure 1, compute instance D1 has the overlay IP address 10.1.0.2 and the MAC address of MM1, and compute instance D2 has the private overlay IP address 10.1.0.3 and the MAC address of MM2. Each compute instance in subnet-2, including compute instances D1 and D2, has a default route to VCN VR105 using the IP address 10.1.0.1, which is the IP address of the port of VCN VR105 in subnet-2.
[0051] VCN A104 may also include one or more load balancers. For example, a load balancer may be provided for a subnet and configured to load balance traffic across multiple compute instances on the subnet. A load balancer may also be provided to load balance traffic across subnets within the VCN.
[0052] A specific compute instance deployed on VCN104 can communicate with various different endpoints. These endpoints may include those hosted by CSPI200 and those outside of CSPI200. Endpoints hosted by CSPI101 may include endpoints on the same subnet as a particular compute instance (e.g., communication between two compute instances in subnet-1), endpoints in different subnets but within the same VCN (e.g., communication between a compute instance in subnet-1 and a compute instance in subnet-2), endpoints in different VCNs within the same region (e.g., communication between a compute instance in subnet-1 and an endpoint in a VCN in the same region 106 or 110, or between a compute instance in subnet-1 and an endpoint in service network 110 in the same region), or endpoints in VCNs in different regions (e.g., communication between a compute instance in subnet-1 and an endpoint in a VCN in a different region 108). Compute instances within a subnet hosted by CSPI101 can also communicate with endpoints not hosted by CSPI101 (i.e., outside of CSPI101). These external endpoints include endpoints within the customer's on-premises network 116, endpoints within other remote cloud host networks 118, public endpoints 114 accessible via public networks such as the internet, and other endpoints.
[0053] Communication between compute instances on the same subnet is facilitated using VNICs associated with the source and destination compute instances. For example, compute instance C1 in subnet-1 may want to send a packet to compute instance C2, also in subnet-1. For a packet originating from the source compute instance and destined for another compute instance on the same subnet, the packet is first processed by the VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance may include determining the packet's destination information from the packet header, identifying policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the packet's next hop, performing any packet encapsulation / deencapsulation functions as needed, and then forwarding / routing the packet to the next hop to facilitate communication of the packet to its intended destination. If the destination compute instance is on the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing. The VNIC associated with the destination compute instance then runs and forwards the packet to the destination compute instance.
[0054] When a compute instance within a subnet communicates a packet to an endpoint in a different subnet within the same VCN, the communication is facilitated by the VNICs and VCN VR associated with the source and destination compute instances. For example, if compute instance C1 in subnet-1 in Figure 1 wants to send a packet to compute instance D1 in subnet-2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR105 using the VCN VR's default route or port 10.0.0.1. VCN VR105 is configured to route the packet to subnet-2 using port 10.1.0.1. The packet is then received and processed by the VNIC associated with D1, and the VNIC forwards the packet to compute instance D1.
[0055] When a compute instance within VCN104 communicates packets to an endpoint outside of VCN104, the communication is facilitated by the VNIC associated with the source compute instance, VCN VR105, and the gateway associated with VCN104. One or more types of gateways can be associated with VCN104. A gateway is an interface between a VCN and another endpoint, where the other endpoint is outside the VCN. A gateway is a Layer 3 / IP layer concept that enables a VCN to communicate with endpoints outside the VCN. Therefore, gateways facilitate traffic flow between a VCN and other VCNs or networks. Different types of gateways can be configured for a VCN to facilitate different types of communication with different types of endpoints. Depending on the gateway, communication may be over a public network (e.g., the internet) or over a private network. Various communication protocols can be used for these communications.
[0056] For example, compute instance C1 may want to communicate with an endpoint outside of VCN104. The packet can first be processed by the VNIC associated with source compute instance C1. The VNIC processing determines that the packet's destination is outside of C1's subnet-1. The VNIC associated with C1 can then forward the packet to VCN VR105 on VCN104. VCN VR105 then processes the packet and, as part of the processing, determines, based on the packet's destination, a specific gateway associated with VCN104 as the packet's next hop. VCN VR105 can then forward the packet to the specific identified gateway. For example, if the destination is an endpoint within the customer's on-premises network, the packet can be forwarded by VCN VR105 to the Dynamic Routing Gateway (DRG) gateway 122 configured for VCN104. The packet is then forwarded from the gateway to the next hop, facilitating the packet's communication to its final intended destination.
[0057] For a VCN, various different types of gateways can be configured. An example of a gateway that can be configured for a VCN is shown in Figure 1 and will be discussed later. Examples of gateways associated with a VCN are also shown in Figures 12, 13, 14, and 15 (for example, gateways referenced by reference numbers 1234, 1236, 1238, 1334, 1336, 1338, 1434, 1436, 1438, 1534, 1536, and 1538) and will be discussed later. As shown in the embodiment shown in Figure 1, a dynamic routing gateway (DRG) 122 can be added to or associated with a customer VCN 104 to provide a path for private network traffic communication between the customer VCN 104 and another endpoint, where the other endpoint could be the customer's on-premises network 116, a VCN 108 in a different region of CSPI 101, or another remote cloud network 118 not hosted by CSPI 101. The customer on-premises network 116 may be the customer's network or a customer data center built using the customer's resources. Access to the customer on-premises network 116 is generally strictly restricted. In the case of a customer who has both the customer on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, the customer may want their on-premises network 116 and their cloud-based VCNs 104 to be able to communicate with each other. This would allow the customer to build an enhanced hybrid environment that includes the customer's VCNs 104 hosted by CSPI 101 and their on-premises network 116. DRG 122 enables such communication. To enable such communication, a communication channel 124 is configured, where one endpoint of the channel is within the customer's on-premises network 116 and the other endpoint is within CSPI 101 and connected to the customer's VCNs 104. The communication channel 124 can be via a public communication network such as the internet or a private communication network.Various different communication protocols can be used, such as IPsec VPN technology over public communication networks like the Internet, or Oracle's FastConnect technology which uses a private network instead of a public network. A device or equipment within the customer's on-premises network 116 that forms one endpoint of communication channel 124 is referred to as customer premises equipment (CPE), such as CPE126 shown in Figure 1. On the CSPI101 side, the endpoint may be a host machine running DRG122. In some embodiments, remote peering connections (RPCs) can be added to the DRG, allowing the customer to peer one VCN with another VCN in a different region. Using such an RPC, customer VCN104 can connect to VCN108 in a different region using DRG122. It is also possible to communicate with other remote cloud networks 118 not hosted by CSPI101, such as the Microsoft Azure cloud or the Amazon AWS cloud, using DRG122.
[0058] As shown in Figure 1, an Internet Gateway (IGW) 120 can be configured in customer VCN 104, allowing compute instances on VCN 104 to communicate with public endpoints 114 accessible via a public network such as the Internet. The IGW 120 is a gateway that connects the VCN to a public network such as the Internet. The IGW 120 enables public subnets within a VCN, such as VCN 104 (where resources within the public subnet have public overlay IP addresses), to directly access public endpoints 112 on a public network such as the Internet 114. Connections can be initiated using the IGW 120 from subnets within VCN 104 or from the Internet.
[0059] A Network Address Translation (NAT) gateway 128 can be configured in the customer VCN 104, which allows cloud resources within the customer VCN that do not have dedicated public overlay IP addresses to access the internet without directly exposing those resources to incoming internet connectivity (e.g., L4-L7 connectivity). This allows private subnets within the VCN, such as private subnet-1 within VCN 104, to privately access public endpoints on the internet. With the NAT gateway, connections can only be initiated from private subnets to the public internet, and not from the internet to private subnets.
[0060] In some embodiments, a service gateway (SGW) 126 can be configured in a customer VCN 104, which provides a route for private network traffic between VCN 104 and service endpoints supported in a service network 110. In some embodiments, the service network 110 may be provided by a CSP and can provide a variety of services. An example of such a service network is Oracle's service network, which provides a variety of services that customers can use. For example, compute instances (e.g., database systems) in a private subnet of customer VCN 104 can back up data to service endpoints (e.g., object storage) without requiring a public IP address or internet access. In some embodiments, a VCN may have only one SGW, and connections can only be initiated from subnets within the VCN and not from the service network 110. When one VCN is peered with another VCN, resources in the other VCN typically cannot access the SGW. Resources in an on-premises network connected to a VCN via FastConnect or VPN Connect can also use the service gateway configured in that VCN.
[0061] In some embodiments, SGW126 uses the concept of Classless Inter-Domain Routing (CIDR) labels, where a CIDR label is a string representing all regional public IP address ranges for a service or group of services of interest. Customers use service CIDR labels when configuring SGW and associated routing rules to control traffic to services. Customers can optionally use service CIDR labels when configuring security rules without having to adjust security rules if the public IP addresses of services change in the future.
[0062] The Local Peering Gateway (LPG) 132 is a gateway that can be added to the customer VCN 104, enabling the VCN 104 to peer with other VCNs within the same region. Peering means that VCNs communicate using private IP addresses without traffic traversing public networks such as the internet and without routing traffic through the customer's on-premises network 116. In a preferred embodiment, the VCN has a separate LPG for each peering it establishes. Local peering, or VCN peering, is a common practice used to establish network connectivity between different applications or infrastructure management functions.
[0063] Service providers, such as service providers within service network 110, can provide access to their services using different access models. According to the public access model, a service may be exposed as a public endpoint, accessible publicly by compute instances within the customer VCN via a public network such as the internet, and / or privately accessible via SGW126. According to a specific private access model, a service becomes accessible as a private IP endpoint within a private subnet within the customer VCN. This is referred to as private endpoint (PE) access and allows service providers to expose their services as instances within the customer's private network. A private endpoint resource represents a service within the customer's VCN. Each PE appears as a VNIC (referred to as a PE-VNIC, having one or more private IPs) selected by the customer from a subnet within the customer VCN. Thus, a PE provides a way to present a service within the customer's private VCN subnet using a VNIC. Because the endpoint is exposed as a VNIC, all functionality associated with a VNIC, such as routing rules and security lists, becomes available on the PE VNIC.
[0064] Service providers can register those services and enable access via the PE. Providers can associate policies with services that restrict the visibility of those services to customer tenancies. Providers can register multiple services under a single virtual IP address (VIP), especially in the case of multi-tenant services. Multiple private endpoints may exist representing the same service (across multiple VCNs).
[0065] Next, compute instances within a private subnet can access the service using the PE VNIC's private IP address or service DNS name. Compute instances within a customer VCN can access the service by sending traffic to the PE's private IP address within the customer VCN. The Private Access Gateway (PAGW) 130 is a gateway resource that can be attached to a service provider VCN (for example, a VCN within service network 110) and acts as an ingress / egress point for all traffic from / to the customer subnet private endpoint. The PAGW 130 allows the provider to scale the number of PE connections without utilizing its internal IP address resources. The provider only needs to configure one PAGW for any number of services registered in a single VCN. The provider can represent a service as a private endpoint in multiple VCNs for one or more customers. From the customer's perspective, the PE VNIC appears to be attached to the service the customer wants to interact with, rather than to the customer's instance. Traffic directed to the private endpoint is routed to the service via the PAGW 130. These are referred to as customer-to-service private connections (C2S connections).
[0066] The PE concept can also be used to extend private access to the service to the customer's on-premises network and data center by enabling traffic to flow through FastConnect / IPsec links and private endpoints within the customer's VCN. Furthermore, private access to the service can be extended to the customer's peering VCN by enabling traffic to flow between LPG132 and the PE within the customer's VCN.
[0067] Customers can control VCN routing at the subnet level, allowing them to specify which subnets use which gateways within their customer VCN, such as VCN104. The VCN's route table determines whether traffic can leave the VCN through a particular gateway. For example, in a specific case, the route table for a public subnet within customer VCN104 might allow non-local traffic to be sent via IGW120. The route table for a private subnet within the same customer VCN104 might allow traffic destined for CSP services to be sent via SGW126. All remaining traffic can be sent via NAT gateway 128. The route table only controls traffic leaving the VCN.
[0068] A security list associated with a VCN is used to control traffic entering the VCN through a gateway via inbound connections. All resources within a subnet use the same route table and security list. Security lists can be used to control specific types of traffic that can enter and leave instances within a VCN subnet. Security list rules can include ingress (inbound) rules and egress (outbound) rules. For example, an ingress rule may specify an allowed source address range, and an egress rule may specify an allowed destination address range. Security rules can specify specific protocols (e.g., TCP, ICMP), specific ports (e.g., port 22 for SSH, port 3389 for Windows RDP), etc. In some embodiments, the instance's operating system may enforce its own firewall rules that match the security list rules. Rules can be stateful (e.g., connections are tracked and responses are automatically allowed without explicit security list rules for response traffic) or stateless.
[0069] Access from a customer VCN (i.e., by resources or compute instances deployed on VCN104) can be classified as public access, private access, or dedicated access. Public access refers to an access model that accesses public endpoints using public IP addresses or NAT. Private access allows customer workloads within VCN104 with private IP addresses (e.g., resources in a private subnet) to access services without traversing a public network such as the internet. In some embodiments, CSPI101 allows customer VCN workloads with private IP addresses to access services (or their public service endpoints) using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer's VCN and the public endpoint of the service that resides outside the customer's private network.
[0070] Furthermore, CSPI can provide dedicated public access using technologies such as FastConnect public peering, in which case customer on-premises instances can access one or more services within the customer VCN using FastConnect connectivity without going through public networks such as the internet. CSPI can also provide dedicated private access using FastConnect private peering, in which case customer on-premises instances with private IP addresses can access the customer VCN workloads using FastConnect connectivity. FastConnect is a network connectivity used instead of connecting a customer's on-premises network to CSPI and its services using the public internet. FastConnect provides a simple, flexible, and economical way to create dedicated private connectivity with higher bandwidth options and a more reliable and consistent networking experience compared to internet-based connectivity.
[0071] Figure 1 and the accompanying description above illustrate the various virtualization components in a virtual network example. As mentioned above, the virtual network is built on top of an underlying physical network or infrastructure network. Figure 2 shows a simplified architectural diagram of the physical components within the physical network in the CSPI200 that provides the underlay for the virtual network, according to several embodiments. As illustrated, the CSPI200 provides a distributed environment that includes components and resources (e.g., compute resources, memory resources, and networking resources) provided by a cloud service provider (CSP). These components and resources are used to deliver cloud services (e.g., IaaS services) to subscribing customers, i.e., customers who subscribe to one or more services provided by the CSP. Based on the services that the customer subscribes to, some of the CSPI200's resources (e.g., compute resources, memory resources, and networking resources) are provisioned to the customer. The customer can then use the physical compute resources, memory resources, and networking resources provided by the CSPI200 to build their own cloud-based (i.e., CSPI-hosted) customizable private virtual network. As mentioned earlier, these customer networks are referred to as Virtual Cloud Networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, to these customer VCNs. Compute instances can take the form of virtual machines, bare metal instances, etc. CSPI200 provides infrastructure and a suite of complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available host environment.
[0072] In the embodiment shown in Figure 2, the physical components of the CSPI200 include one or more physical host machines or physical servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), and a physical network (e.g., 218), as well as switches within the physical network 218. The physical host machines or servers can host and run various compute instances participating in one or more subnets of the VCN. Compute instances can include virtual machine instances and bare metal instances. For example, the various compute instances shown in Figure 1 may be hosted by the physical host machines shown in Figure 2. Virtual machine compute instances in the VCN may run on one host machine or on multiple different host machines. The physical host machines can also host virtual host machines, container-based hosts or functions, etc. The VNIC and VCN VR shown in Figure 1 may run on the NVDs shown in Figure 2. The gateway shown in Figure 1 may be run by the host machine and / or NVD shown in Figure 2.
[0073] A host machine or server can run a hypervisor (also known as a virtual machine monitor or VMM) that creates and enables a virtualized environment on the host machine. Virtualization or a virtualized environment facilitates cloud-based computing. One or more compute instances can be created, run, and managed on the host machine by a hypervisor on the host machine. The hypervisor on the host machine allows the host machine's physical computing resources (e.g., compute resources, memory resources, and networking resources) to be shared among the various compute instances running on the host machine.
[0074] For example, as shown in Figure 2, host machines 202 and 208 run hypervisors 260 and 266, respectively. These hypervisors can be implemented using software, firmware, hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that sits on top of the host machine's operating system (OS), which in turn runs on the host machine's hardware processors. A hypervisor provides a virtualized environment by allowing the host machine's physical computing resources (e.g., processing resources such as processors / cores, memory resources, and networking resources) to be shared among various virtual machine compute instances running on the host machine. For example, in Figure 2, hypervisor 260 can sit on top of the host machine 202's OS, allowing the host machine 202's computing resources (e.g., processing resources, memory resources, and networking resources) to be shared among compute instances (e.g., virtual machines) running on the host machine 202. A virtual machine can have its own operating system (referred to as a guest operating system), which may be the same as or different from the host machine's OS. The operating system of a virtual machine running on a host machine may be the same as or different from the operating system of another virtual machine running on the same host machine. Therefore, a hypervisor allows multiple operating systems to run in parallel while sharing the same computing resources of the host machine. The host machines shown in Figure 2 may have the same type of hypervisor or different types of hypervisors.
[0075] A compute instance can be a virtual machine instance or a bare metal instance. In Figure 2, compute instance 268 on host machine 202 and compute instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.
[0076] In some cases, an entire host machine may be provisioned to a single customer, and one or more compute instances (either virtual machines or bare metal instances) hosted by that host machine may all belong to that same customer. In other cases, a host machine may be shared among multiple customers (i.e., multiple tenants). In such multi-tenant scenarios, a host machine can host virtual machine compute instances belonging to different customers. These compute instances may be members of different VCNs of different customers. In some embodiments, bare metal compute instances are hosted by bare metal servers without a hypervisor. When bare metal compute instances are provisioned, a single customer or tenant maintains control of the physical CPU, memory, and network interfaces of the host machine hosting the bare metal instances, and the host machine is not shared with other customers or tenants.
[0077] As mentioned above, each compute instance that is part of a VCN is associated with a VNIC that enables the compute instance to be a member of the VCN's subnet. The VNIC associated with a compute instance facilitates the communication of packets or frames to and from the compute instance. The VNIC is associated with the compute instance when the compute instance is created. In some embodiments, for compute instances run by a host machine, the VNIC associated with that compute instance is run by an NVD connected to the host machine. For example, in Figure 2, host machine 202 runs virtual machine compute instance 268 associated with VNIC 276, and VNIC 276 is run by an NVD 210 connected to host machine 202. In another example, bare metal instance 272 hosted by host machine 206 is associated with VNIC 280, which is run by an NVD 212 connected to host machine 206. As yet another example, VNIC284 is associated with compute instance 274 running on host machine 208, and VNIC284 is run on NVD212 connected to host machine 208.
[0078] For compute instances hosted by a host machine, an NVD connected to that host machine also runs the VCN VR corresponding to the VCN of which the compute instance is a member. For example, in the embodiment shown in Figure 2, NVD210 runs VCN VR277 corresponding to the VCN of which compute instance 268 is a member. NVD212 can also run one or more VCN VR283 corresponding to the VCNs of compute instances hosted by host machines 206 and 208.
[0079] A host machine may include one or more network interface cards (NICs) that enable it to connect to other devices. The NICs on the host machine may provide one or more ports (or interfaces) that enable the host machine to connect to another device in a communicative manner. For example, a host machine can be connected to an NVD using one or more ports (or interfaces) provided on both the host machine and the NVD. A host machine can also be connected to other devices, such as another host machine.
[0080] For example, in Figure 2, host machine 202 is connected to NVD210 using a link 220 extending between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD210. Host machine 206 is connected to NVD212 using a link 224 extending between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD212. Host machine 208 is connected to NVD212 using a link 226 extending between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD212.
[0081] The NVD is further connected via communication links to a top-of-rack (TOR) switch connected to a physical network 218 (also referred to as a switch fabric). In some embodiments, the links between the host machine and the NVD, and between the NVD and the TOR switch, are Ethernet links. For example, in Figure 2, NVDs 210 and 212 are connected to TOR switches 214 and 216, respectively, using links 228 and 230. In some embodiments, links 220, 224, 226, 228, and 230 are Ethernet links. The collection of host machines and NVDs connected to the TOR is sometimes referred to as a rack.
[0082] The physical network 218 provides a communication fabric that enables TOR switches to communicate with each other. The physical network 218 can be a multi-layer network. In some embodiments, the physical network 218 is a multi-layer Clos network of switches, and TOR switches 214 and 216 represent leaf-level nodes of the multi-layer and multi-node physical switching network 218. Different Clos network configurations are possible, but are not limited to 2-layer, 3-layer, 4-layer, 5-layer, and generally "n"-layer networks. An example of a Clos network is shown in Figure 5 and will be discussed later.
[0083] Various connection configurations are possible between the host machine and the NVD, including one-to-one, many-to-one, and one-to-many configurations. In a one-to-one configuration, each host machine is connected to its own separate NVD. For example, in Figure 2, host machine 202 is connected to NVD 210 via NIC 232 of host machine 202. In a many-to-one configuration, multiple host machines are connected to a single NVD. For example, in Figure 2, host machines 206 and 208 are connected to the same NVD 212 via NICs 244 and 250, respectively.
[0084] In a one-to-many configuration, one host machine is connected to multiple NVDs. Figure 3 shows an example within a CSPI300 where a host machine is connected to multiple NVDs. As shown in Figure 3, the host machine 302 has a network interface card (NIC) 304 with multiple ports 306 and 308. The host machine 300 is connected to a first NVD 310 via port 306 and link 320, and to a second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between the host machine 302 and the NVDs 310 and 312 may be Ethernet links. The NVD 310 is further connected to a first TOR switch 314, and the NVD 312 is connected to a second TOR switch 316. The links between the NVDs 310 and 312 and the TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent layer-0 switching devices within a multi-layer physical network 318.
[0085] The configuration shown in Figure 3 provides two separate physical network paths from the physical switch network 318 to the host machine 302: a first path to the host machine 302 via the NVD 310 through the TOR switch 314, and a second path to the host machine 302 via the NVD 312 through the TOR switch 316. These separate paths provide enhanced availability (referred to as high availability) for the host machine 302. If there is a problem with one of the paths (for example, one link in the path fails) or a problem with a device (for example, a particular NVD is not functioning), the other path can be used for communication with the host machine 302.
[0086] In the configuration shown in Figure 3, the host machine is connected to two different NVDs using two different ports provided by the host machine's NIC. In other embodiments, the host machine may include multiple NICs that enable connections to multiple NVDs of the host machine.
[0087] Referring again to Figure 2, an NVD is a physical device or component that performs one or more network virtualization and / or storage virtualization functions. An NVD can be any device having one or more processing units (e.g., a CPU, a network processing unit (NPU), an FPGA, a packet processing pipeline, etc.), memory including a cache, and ports. Various virtualization functions may be performed by software / firmware running on one or more processing units of the NVD.
[0088] NVDs can be implemented in various different forms. For example, in some embodiments, an NVD is implemented as an interface card called a smart NIC or intelligent NIC with an embedded processor. A smart NIC is a separate device from the NIC on the host machine. In Figure 2, NVD210 can be implemented as a smart NIC connected to host machine 202, and NVD212 can be implemented as smart NICs connected to host machines 206 and 208.
[0089] However, the smart NIC is only one example of an NVD embodiment. Various other embodiments are possible. For example, in some other embodiments, the NVD, or one or more functions performed by the NVD, may be incorporated into or performed by one or more host machines, one or more TOR switches, and other components of the CSPI200. For example, the NVD may be incorporated into a host machine, in which case the functions performed by the NVD are performed by the host machine. As another example, the NVD may be part of a TOR switch, or the TOR switch may be configured to perform functions performed by the NVD, enabling the TOR switch to perform various complex packet translations used in public clouds. A TOR that performs the functions of the NVD may be referred to as a smart TOR. In yet another embodiment, where virtual machine (VM) instances rather than bare metal (BM) instances are provided to the customer, the functions performed by the NVD may be implemented within the hypervisor of the host machine. In some other embodiments, some of the functions of the NVD may be offloaded to a centralized service running on a fleet of host machines.
[0090] In some embodiments, such as when implemented as a smart NIC as shown in Figure 2, the NVD may have multiple physical ports that enable the NVD to connect to one or more host machines and one or more TOR switches. The ports on the NVD can be classified as host-side ports (also referred to as "south ports") or network-side or TOR-side ports (also referred to as "north ports"). The host-side ports of the NVD are the ports used to connect the NVD to a host machine. In Figure 2, examples of host-side ports include port 236 on the NVD210, and ports 248 and 254 on the NVD212. The network-side ports of the NVD are the ports used to connect the NVD to a TOR switch. In Figure 2, examples of network-side ports include port 256 on the NVD210 and port 258 on the NVD212. As shown in Figure 2, the NVD210 is connected to the TOR switch 214 using a link 228 that extends from port 256 on the NVD210 to the TOR switch 214. Similarly, the NVD212 is connected to the TOR switch 216 using a link 230 that extends from port 258 of the NVD212 to the TOR switch 216.
[0091] The NVD can receive packets and frames from the host machine (for example, packets and frames generated by compute instances hosted by the host machine) via its host-side port, perform the necessary packet processing, and then forward the packets and frames to the TOR switch via its network-side port. The NVD can also receive packets and frames from the TOR switch via its network-side port, perform the necessary packet processing, and then forward the packets and frames to the host machine via its host-side port.
[0092] In some embodiments, there may be multiple ports and associated links between the NVD and the TOR switch. These ports and links can be aggregated to form a link aggregator group (LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (for example, between the NVD and the TOR switch) to be treated as a single logical link. All physical links within a given LAG can operate in full-duplex mode at the same speed. LAGs help increase the bandwidth and reliability of the connection between the two endpoints. If one of the physical links in the LAG fails, traffic is dynamically and transparently reassigned to one of the other physical links in the LAG. Aggregated physical links provide higher bandwidth than each individual link. Multiple ports associated with a LAG are treated as a single logical port. Traffic can be load-balanced across the multiple physical links of the LAG. One or more LAGs can be configured between two endpoints. The two endpoints may be between the NVD and the TOR switch, or between a host machine and the NVD, etc.
[0093] NVD implements or performs network virtualization functions. These functions are performed by software / firmware run by NVD. Examples of network virtualization functions include, but are not limited to, packet encapsulation and deencapsulation functions, functions for creating VCN networks, functions for implementing network policies such as VCN security list (firewall) functions, and functions for facilitating the routing and forwarding of packets to and from compute instances within the VCN. In some embodiments, when NVD receives a packet, it is configured to run a packet processing pipeline that processes the packet and determines how to forward or route it. As part of this packet processing pipeline, NVD may run one or more virtual functions related to overlay networks, such as running a VNIC associated with a cis within the VCN, running a virtual router (VR) associated with the VCN, packet encapsulation and deencapsulation to facilitate forwarding or routing within the virtual network, running several gateways (e.g., a local peering gateway), implementing security lists, network security groups, network address translation (NAT) functions (e.g., translation from public IP to private IP on a per-host basis), and throttling functions.
[0094] In some embodiments, the packet processing data path within the NVD may include multiple packet pipelines, each consisting of a set of packet translation stages. In some embodiments, upon receiving a packet, it is parsed and sorted into a single pipeline. The packet is then processed stage by stage in a linear manner until it is dropped or sent out through the NVD's interface. These stages provide the basic functional packet processing building blocks (e.g., header validation, throttling, insertion of new Layer 2 headers, L4 firewall execution, VCN encapsulation / deencapsulation, etc.), thereby allowing new pipelines to be constructed by assembling existing stages, and new functionality to be added by creating new stages and inserting them into existing pipelines.
[0095] The NVD can perform both control plane and data plane functions corresponding to the VCN's control plane and data plane. Examples of the VCN control plane are also shown in Figures 12, 13, 14, and 15 (see references 1216, 1316, 1416, and 1516) and will be discussed later. Examples of the VCN data plane are shown in Figures 12, 13, 14, and 15 (see references 1218, 1318, 1418, and 1518) and will be discussed later. Control plane functions include functions used to configure the network that controls how data is forwarded (e.g., setting routes and route tables, configuring VNICs, etc.). In some embodiments, a VCN control plane is provided that centrally computes the mappings of all overlays to the infrastructure and exposes them to the NVD and various gateway virtual network edge devices such as DRG, SGW, and IGW. Firewall rules can also be exposed using the same mechanism. In some embodiments, the NVD retrieves only the mappings relevant to that NVD. The data plane function includes the ability to perform the actual routing / forwarding of packets based on the configuration set using the control plane. The VCN data plane is implemented by encapsulating customer network packets before they pass through the underlying network. The encapsulation / deencapsulation function is implemented in the NVD. In some embodiments, the NVD is configured to intercept all network packets entering and leaving the host machine and to perform network virtualization functions.
[0096] As described above, NVD performs various virtualization functions, including VNICs and VCN VRs. An NVD can run VNICs associated with compute instances hosted by one or more host machines connected to a VNIC. For example, as shown in Figure 2, NVD210 runs the functions of VNIC276 associated with compute instance 268 hosted by host machine 202 connected to NVD210. As another example, NVD212 runs VNIC280 associated with bare-metal compute instance 272 hosted by host machine 206 and VNIC284 associated with compute instance 274 hosted by host machine 208. A host machine can host compute instances belonging to different VCNs belonging to different customers, and an NVD connected to a host machine can run the VNICs corresponding to those compute instances (i.e., run the functions associated with the VNICs).
[0097] NVD also runs VCN virtual routers corresponding to the VCNs of compute instances. For example, in the embodiment shown in Figure 2, NVD210 runs VCN VR277 corresponding to the VCN to which compute instance 268 belongs. NVD212 runs one or more VCN VR283 corresponding to one or more VCNs to which compute instances hosted on host machines 206 and 208 belong. In some embodiments, the VCN VR corresponding to a VCN is run by all NVDs connected to host machines hosting at least one compute instance belonging to that VCN. If a host machine hosts compute instances belonging to different VCNs, NVDs connected to that host machine can run VCN VRs corresponding to those different VCNs.
[0098] In addition to VNICs and VCN VRs, NVDs may include one or more hardware components that run various software (e.g., daemons) and facilitate various network virtualization functions performed by NVDs. For simplicity, these various components are grouped together as “packet processing components” as shown in Figure 2. For example, NVD210 includes packet processing component 286, and NVD212 includes packet processing component 288. For example, an NVD packet processing component may include a packet processor configured to interact with the NVD’s ports and hardware interfaces to monitor all packets received by and communicated using the NVD and to store network information. Network information may include, for example, network flow information that identifies different network flows processed by the NVD, and per-flow information (e.g., per-flow statistics). In some embodiments, network flow information may be stored on a per-VNIC basis. The packet processor may perform per-packet operations as well as implement stateful NAT and L4 firewalls (FW). As another example, a packet processing component may include a replication agent configured to replicate information stored by the NVD to one or more different replication target stores. As yet another example, the packet processing component may include a logging agent configured to perform the NVD's logging function. The packet processing component may also include software that monitors the performance and health of the NVD, and possibly the status and health of other components connected to the NVD.
[0099] Figure 1 shows the components of an example virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, a VR for the VCN, and a set of gateways configured for the VCN. The overlay components shown in Figure 1 can run or host one or more of the physical components shown in Figure 2. For example, compute instances within a VCN can run or host one or more host machines shown in Figure 2. For compute instances hosted by host machines, the VNICs associated with those compute instances are typically run by NVDs connected to those host machines (i.e., the VNIC functionality is provided by the NVDs connected to those host machines). The VCN VR functionality of the VCN is run by all NVDs connected to the host machines that host or run the compute instances that are part of that VCN. Gateways associated with the VCN may run by one or more different types of NVDs. For example, some gateways may run by smart NICs, and others may run by one or more host machines or other embodiments of NVDs.
[0100] As described above, compute instances within a customer VCN can communicate with various different endpoints, which may be on the same subnet as the source compute instance, or on a different subnet but still within the same VCN, or with endpoints outside the source compute instance's VCN. These communications are facilitated using the VNIC associated with the compute instance, the VCN VR, and the gateway associated with the VCN.
[0101] Communication between two compute instances on the same subnet within a VCN is facilitated using the VNICs associated with the source and destination compute instances. The source and destination compute instances may be hosted by the same host machine or by different host machines. Packets originating from the source compute instance can be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. In the NVD, packets are processed using a packet processing pipeline, which may include the execution of the VNIC associated with the source compute instance. Because the packet's destination endpoint is within the same subnet, the execution of the VNIC associated with the source compute instance forwards the packet to the NVD running the VNIC associated with the destination compute instance, where the NVD processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances may run on the same NVD (for example, if both the source and destination compute instances are hosted by the same host machine) or on different NVDs (for example, if the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNICs can use the routing / forwarding tables stored by the NVD to determine the next hop of a packet.
[0102] When a packet is communicated from a compute instance within a subnet to an endpoint in a different subnet within the same VCN, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. In the NVD, the packet is processed using a packet processing pipeline that may include the execution of one or more VNICs, and a VR associated with the VCN. For example, as part of the packet processing pipeline, the NVD executes or invokes a function corresponding to the VNIC associated with the source compute instance (also referred to as executing the VNIC). The function executed by the VNIC may include examining the VLAN tag on the packet. Because the packet's destination is outside the subnet, the VCN VR function is then invoked and executed by the NVD. The VCN VR then routes the packet to the NVD executing the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards it to the destination compute instance. The VNICs associated with the source compute instance and the destination compute instance may run on the same NVD (for example, if both the source and destination compute instances are hosted by the same host machine) or on different NVDs (for example, if the source and destination compute instances are hosted by different host machines connected to different NVDs).
[0103] If the packet's destination is outside the VCN of the source compute instance, the packet originating from the source compute instance is communicated from the host machine hosting the source compute instance to an NVD connected to that host machine. The NVD runs the VNIC associated with the source compute instance. Because the packet's destination endpoint is outside the VCN, the packet is then processed by the VCN VR of that VCN. The NVD invokes VCN VR functionality, which may result in the packet being forwarded to an NVD running the appropriate gateway associated with the VCN. For example, if the destination is an endpoint within the customer's on-premises network, the packet may be forwarded by the VCN VR to an NVD running the DRG gateway configured for the VCN. The VCN VR may run on the same NVD running the VNIC associated with the source compute instance, or it may run on a different NVD. The gateway may run on an NVD that can be a smart NIC, a host machine, or other NVD embodiment. The packet is then processed by the gateway and forwarded to the next hop that facilitates the communication of the packet to its intended destination endpoint. For example, in the embodiment shown in Figure 2, a packet originating from compute instance 268 may be communicated from host machine 202 to NVD210 via link 220 (using NIC 232). In NVD210, VNIC 276 is invoked because it is the VNIC associated with the source compute instance 268. VNIC 276 is configured to examine the encapsulation information in the packet, determine the next hop for forwarding the packet to facilitate its communication to its intended destination endpoint, and forward the packet to the determined next hop.
[0104] Compute instances deployed on a VCN can communicate with various different endpoints. These endpoints may include endpoints hosted by CSPI200 and endpoints outside of CSPI200. Endpoints hosted by CSPI200 may include instances within the same VCN or other VCNs, which may be the customer's VCN or a VCN not belonging to the customer. Communication between endpoints hosted by CSPI200 may be performed over the physical network 218. Compute instances may also communicate with endpoints that are not hosted by CSPI200 or are outside of CSPI200. Examples of these endpoints include endpoints within the customer's on-premises network or data center, or public endpoints accessible over a public network such as the internet. Communication with endpoints outside of CSPI200 may be performed over a public network (e.g., the internet) (not shown in Figure 2) or a private network (not shown in Figure 2) using various communication protocols.
[0105] The architecture of the CSPI200 shown in Figure 2 is merely an example and is not intended to be limiting. Variations, alternative forms, and altered forms are possible in alternative embodiments. For example, in some embodiments, the CSPI200 may have more or fewer systems or components than those shown in Figure 2, may combine two or more systems, or may have different system configurations or arrangements. The systems, subsystems, and other components shown in Figure 2 can be implemented using hardware or a combination thereof, with software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored in non-temporary storage media (e.g., memory devices).
[0106] Figure 4 shows the connection between a host machine and an NVD that provides I / O virtualization to support multi-tenancy in several embodiments. As shown in Figure 4, host machine 402 runs a hypervisor 404 that provides the virtualization environment. Host machine 402 runs two virtual machine instances, namely VM1 406 belonging to customer / tenant #1 and VM2 408 belonging to customer / tenant #2. Host machine 402 includes a physical NIC 410 connected to the NVD 412 via link 414. Each compute instance is attached to a VNIC run by the NVD 412. In the embodiment of Figure 4, VM1 406 is attached to VNIC-VM1 420 and VM2 408 is attached to VNIC-VM2 422.
[0107] As shown in Figure 4, NIC410 includes two logical NICs, namely logical NIC A416 and logical NIC B418. Each virtual machine is attached to its own logical NIC and configured to operate with its own logical NIC. For example, VM1 406 is attached to logical NIC A416, and VM2 408 is attached to logical NIC B418. Although the host machine 402 contains only one physical NIC 410 shared by multiple tenants, the logical NICs allow each tenant's virtual machine to be considered to have its own host machine and NIC.
[0108] In some embodiments, each logical NIC is assigned its own VLAN ID. Thus, logical NIC A416 for tenant #1 is assigned a specific VLAN ID, and logical NIC B418 for tenant #2 is assigned a different VLAN ID. When a packet is communicated from VM1 406, the hypervisor assigns the tag assigned to tenant #1 to the packet, and then the packet is communicated from host machine 402 to NVD412 via link 414. Similarly, when a packet is communicated from VM2 408, the hypervisor assigns the tag assigned to tenant #2 to the packet, and then the packet is communicated from host machine 402 to NVD412 via link 414. Thus, the packet 424 communicated from host machine 402 to NVD412 has an associated tag 426 that identifies a specific tenant and associated VM. In NVD, for a packet 424 received from host machine 402, the tag 426 associated with the packet is used to determine whether the packet should be processed by VNIC-VM1 420 or VNIC-VM2 422. The packet is then processed by the corresponding VNIC. In the configuration shown in Figure 4, each tenant's compute instance can be considered to own its own host machine and NIC. The setup shown in Figure 4 provides I / O virtualization to support multi-tenancy.
[0109] Figure 5 shows a simplified block diagram of a physical network 500 in several embodiments. The embodiments shown in Figure 5 are structured as a Clos network. A Clos network is a specific type of network topology designed to provide connectivity redundancy while maintaining high bimodal bandwidth and maximum resource utilization. A Clos network is a type of non-blocking, multi-stage or multi-layer switching network where the number of stages or layers can be 2, 3, 4, 5, etc. The embodiments shown in Figure 5 are a three-layer network including layers 1, 2, and 3. A TOR switch 504 represents a layer-0 switch in the Clos network. One or more NVDs are connected to the TOR switch. A layer-0 switch is also referred to as an edge device in the physical network. A layer-0 switch is connected to a layer-1 switch, also referred to as a leaf switch. In the embodiments shown in Figure 5, a set of "n" layer-0 TOR switches is connected to a set of "n" layer-1 switches, together forming a pod. Each Layer-0 switch within a pod is interconnected with all Layer-1 switches within that pod, but switches between pods are not interconnected. In some embodiments, two pods are referred to as a block. Each block is serviced by or connected to a set of "n" Layer-2 switches (sometimes referred to as spine switches). A physical network topology may have several blocks, and the Layer-2 switches are connected to "n" Layer-3 switches (also referred to as superspine switches). Packet communication over the physical network is typically performed using one or more Layer 3 communication protocols. Typically, all layers of the physical network except the TOR layer are n-way redundant, thus enabling high availability. The physical network can be scaled by specifying policies for pods and blocks to control the mutual visibility of switches in the physical network.
[0110] A key feature of Clos networks is that the maximum hop count required to reach one Layer-0 switch from another Layer-0 switch (or from an NVD connected to a Layer-0 switch to another NVD connected to a Layer-0 switch) remains constant. For example, in a Layer 3 Clos network, a packet may require a maximum of 7 hops to reach one NVD from another, where both the source and target NVDs are connected to the leaf layer of the Clos network. Similarly, in a Layer 4 Clos network, a packet may require a maximum of 9 hops to reach one NVD from another, where both the source and target NVDs are connected to the leaf layer of the Clos network. Therefore, the Clos network architecture maintains consistent latency across the entire network, which is critical for communication within and between data centers. Clos topologies are horizontally scalable and cost-effective. Network bandwidth / throughput capacity can be easily increased by adding more switches to various layers (e.g., more leaf and spine switches) and increasing the number of links between switches in adjacent layers.
[0111] In some embodiments, each resource within the CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource information and can be used to manage the resource, for example, via the console or via an API. An example of CID syntax is: ocid1.<RESOURCE TYPE> . <realm>[REGION][FUTURE USE]<UNIQUE ID> The formula is as follows: "ocid1": A string indicating the CID version. "resource type": The type of resource (e.g., instance, volume, VCN, subnet, user, group, etc.). "realm": The realm in which the resource resides. Examples of values include "c1" for a commercial realm, "c2" for a government cloud realm, or "c3" for a federal government cloud realm. Each realm can have its own domain name. "region": The region where the resource resides. This field can be blank if the region is not applicable to the resource. "future use": Reserved for future use. "Unique ID": The unique part of the ID. The format may vary depending on the type of resource or service.
[0112] Figure 6 shows a block diagram of a cloud infrastructure 600 incorporating a CLOS network configuration in several embodiments. The cloud infrastructure 600 comprises multiple racks (e.g., racks 1, 610… racks M, 620). Each rack contains multiple host machines (also referred to herein as hosts). For example, rack 1, 610 contains multiple hosts (e.g., K host machines) hosts 1-A, 612 to 1-K, 614, and rack M contains K host machines, i.e., hosts MA, 622 to MK, 624. It is understood that the illustration in Figure 6 (i.e., each rack containing the same number of host machines, e.g., K host machines) is intended to be illustrative and non-limiting. For example, rack M, 620 may have more or fewer host machines compared to the number of host machines contained in rack 1, 610.
[0113] Each host machine contains multiple image processing units (GPUs). For example, as shown in Figure 6, host machine 1-A 612 contains N GPUs, e.g., GPU1, 613. Furthermore, the illustration in Figure 6 (that each host machine contains the same number of GPUs, i.e., N GPUs) is intended to be illustrative and non-limiting, i.e., it is understood that each host machine may contain a different number of GPUs. Each rack contains a top-of-rack (TOR) switch that is communicatively connected to the GPUs hosted on the host machines. For example, rack 1 610 contains a TOR switch (i.e., TOR 1) 616 communicatively connected to hosts 1-A, 612 and 1-K, 614, and rack M 620 contains a TOR switch (i.e., TOR M) 626 communicatively connected to hosts MA, 622 and MK, 624. The TOR switches shown in Figure 6 (i.e., TOR 1 616 and TOR M 626) are understood to each contain N ports used to communicatively connect the TOR switches to N GPUs hosted on each host machine contained in the rack. The connection of the TOR switches to the GPUs as shown in Figure 6 is intended to be illustrative and non-limiting. For example, in some embodiments, the TOR switches may have multiple ports, each corresponding to a GPU on each host machine, i.e., the GPUs on the host machines may be connected to a unique port of the TOR via a communication link. Furthermore, traffic received by a network device (e.g., TOR 1 616) is characterized herein as traffic received on a specific incoming port-link of the network device. For example, if GPU 1 613 on host 1-A 612 sends a data packet to TOR 1 616 (using link 617), the data packet is received on port 619 of the TOR 1 switch. TOR 1 616 characterizes this data packet as information received on the first incoming port-link of TOR. It is understood that a similar concept can be applied to all outgoing port-links of TOR.
[0114] TOR switches from each rack are communicatively connected to multiple spine switches, for example, spine switches 1, 630 and P, 640. As shown in Figure 6, TOR 1, 616 is connected to spine switch 1, 630 via two links and to spine switch P, 640 via two other links. Information transmitted from a particular TOR switch to a spine switch is referred to herein as communication via uplinks, while information transmitted from a spine switch to a TOR switch is referred to herein as communication via downlinks. According to some embodiments, the TOR switches and spine switches are connected in a CLOS network configuration (e.g., a multi-stage switching network), where each TOR switch forms a “leaf” node in the CLOS network.
[0115] In some embodiments, GPUs within a host machine perform tasks related to machine learning. In such a configuration, a single task may be executed / spread across a large number of GPUs (e.g., 64) that can be spread across multiple host machines and multiple racks. Since all of these GPUs are processing the same task (i.e., workload), they need to communicate with each other in a time-synchronized manner. Furthermore, at any given moment, the GPUs are either in compute mode or communication mode, meaning they communicate with each other at approximately the same moment. The speed of the workload is determined by the speed of the slowest GPU.
[0116] Typically, equal-cost multipath (ECMP) routing is used to route packets from a source GPU (e.g., GPU1 and 613 on host 1-A 612) to a destination GPU (e.g., GPU1 and 623 on host MA 622). In ECMP routing, if there are multiple equal-cost paths available for routing traffic from the sender to the receiver, a selection technique is used to choose a specific path. Thus, the network device receiving the traffic (e.g., a TOR switch or spine switch) uses a selection algorithm to choose the outgoing link to be used to forward the traffic in the next hop from one network device to another. Such outgoing link selection is performed at each network device along the path from sender to receiver. Hash-based selection algorithms are a widely used ECMP selection technique, where the hash can be based on a 4-tuple of the packet (e.g., source port, destination port, source IP, destination IP).
[0117] ECMP routing is a flow-aware routing technique where each flow (i.e., a stream of data packets) is hashed to a specific path over the duration of the flow. Therefore, packets within a flow are forwarded from the network device using a specific outgoing port / link. This is typically done to ensure that packets within a flow arrive in order, i.e., that packet reordering is unnecessary. However, ECMP routing is bandwidth (or throughput) unaware. In other words, TOR switches and spine switches perform statistical flow-aware (throughput-unaware) ECMP load balancing of flows on parallel links.
[0118] Standard ECMP routing (i.e., flow-only routing) has a problem where flows received by a network device via two separate incoming links can be hashed to the same outgoing link, resulting in flow collisions. For example, consider a scenario where two flows arrive via two separate incoming 100G links, and each flow is hashed to the same 100G outgoing link. In such a scenario, the incoming bandwidth is 200G, but the outgoing bandwidth is 100G, leading to congestion (i.e., a flow collision) and dropped packets. For example, Figure 7 below illustrates an exemplary scenario 700 of a flow collision.
[0119] As shown in Figure 7, there are two flows: flow 1 710 (shown as a solid line) from the first GPU on host machine host 1-A, 612, to TOR switch 616, and flow 2 720 (shown as a dashed line) from another GPU on the same host machine 612 to TOR switch 616. Note that the two flows are directed to TOR switch 616 on separate links, i.e., separate incoming port-links of TOR. Assume that all links shown in Figure 7 have a capacity (i.e., bandwidth) of 100G. If TOR switch 616 is running the ECMP routing algorithm, the two flows may be hashed to use the same outgoing port-link of TOR, for example, a port of TOR connected to link 730 which connects TOR switch 616 to spine switch 630. In this case, a collision occurs between the two flows (indicated by the "X" mark), and the packets will be dropped.
[0120] Such collision scenarios are generally problematic for all types of traffic, regardless of the protocol. For example, TCP is intelligent in that if a packet is dropped and the sender does not receive an acknowledgment for that dropped packet, the packet is retransmitted. However, the situation worsens with Remote Direct Memory Access (RDMA) type traffic. RDMA networks do not use TCP for various reasons (for example, TCP has complex logic that is not suited to low latency and high performance). RDMA networks use protocols such as RDMA over Infiniband or RDMA over converged Ethernet (RoCE). RoCE has a congestion control algorithm where, if the sender detects that a packet is congested or dropped, the sender slows down the transmission of the packet. In the case of a dropped packet, not only the dropped packet but also several packets surrounding the dropped packet are retransmitted, which further consumes available bandwidth and degrades performance.
[0121] The techniques for overcoming the aforementioned flow collision problem are described below. It is understood that the flow collision problem affects traffic from both CPUs and GPUs. However, the flow collision problem is a much greater problem for GPUs due to their stringent time synchronization requirements. Furthermore, it should be understood that the standard ECMP routing mechanism, due to its inherent characteristic of routing information in a statistically bandwidth-insensitive manner, can cause flow collision scenarios regardless of whether the network is oversubscribed or undersubscribed. A network without oversubscription is one in which the bandwidth of incoming links to a device (e.g., a TOR, spine switch) is the same as the bandwidth of outgoing links. Note that if all links have the same bandwidth capacity, the number of incoming links is the same as the number of outgoing links.
[0122] According to some embodiments, techniques for overcoming the aforementioned flow collision problem include a GPU-based policy routing mechanism (also referred to herein as a GPU-based traffic engineering mechanism) and a modified ECMP routing mechanism. Each of these techniques is described in further detail below.
[0123] Figure 8 shows a policy-based routing mechanism implemented in the network devices of the cloud infrastructure of Figure 6, according to several embodiments. Specifically, the cloud infrastructure 800 includes multiple racks, for example, racks 1 810 to M 820. Each rack includes host machines, each containing multiple GPUs. For example, rack 1 810 includes a host machine, namely host 1-A 812, and rack M 820 includes a host machine, namely host MA 822. Each rack includes a TOR switch, for example, rack 1 810 includes TOR 1 switch 814, and rack M 820 includes TOR M switch 824. The host machines in each rack are communicatively connected to the respective TOR switches in the rack. The TOR switches, namely TOR switches 1, 814 and TOR switch M 824, are further communicatively connected to spine switches, namely spine switches 830 and 840. For the purpose of explaining and illustrating policy-based routing, the cloud infrastructure 800 is shown as containing a single host machine per rack. However, it is understood that each rack in the infrastructure can have two or more host machines.
[0124] According to some embodiments, data packets from a sender to a receiver are routed within the network on a hop-by-hop basis. Routing policies are configured at each network device that matches incoming port-links to outgoing port-links. Network devices may be TOR switches or spine switches. Referring to Figure 8, two flows are shown: flow 1 from GPU 1 on host machine 812, whose intended destination is GPU 1 on host machine 822, and flow 2 from GPU N on host machine 812, whose intended destination is GPU N on host machine 822. The network devices, namely TOR 1 814, spine 1 830, TOR M 824, and spine P 840, are configured to match (or connect) incoming port-links to outgoing port-links. The matching of incoming port-links to outgoing port-links is maintained at each network device (for example, in a policy table).
[0125] Referring to Figure 8, with respect to flow 1 (i.e., the flow shown by the solid line), it can be observed that when TOR 1 814 receives a packet on link 850, it is configured to forward the received packet on the outgoing link 855. Similarly, when spine switch 830 receives a packet via link 855, it is configured to forward that packet on the outgoing link 860. Finally, when TOR M 824 receives a packet on link 860, it is configured to forward that packet on the outgoing link 865 to send it to its intended destination, namely GPU 1 on host machine 822. Similarly, with respect to flow 2 (i.e., the flow shown by the dashed line), when TOR 1 814 receives a packet on link 870, it is configured to forward the received packet on the outgoing link 875 to spine P 840. When spine switch 840 receives a packet via link 875, it is configured to forward that packet on the outgoing link 880. Ultimately, when the TOR M 824 receives a packet on link 880, it is configured to forward that packet on the outgoing link 885 to its intended destination, namely the GPU N on host machine 822.
[0126] Thus, according to the GPU policy-based routing mechanism, each network device is configured to map incoming ports / links to outgoing ports / links to avoid collisions. For example, considering the first hop flows of Flow 1 and Flow 2, TOR 1 814 receives the first data packet (corresponding to Flow 1) on link 850 and the second data packet (corresponding to Flow 2) on link 870. TOR 1 814 is configured to forward the data packet received on incoming link / port 850 to outgoing link / port 855 and the data packet received on incoming link / port 870 to outgoing link / port 875, ensuring that the first and second data packets do not collide. It is understood that in each network device within the cloud infrastructure, there is a one-to-one correspondence between incoming port-links and outgoing port-links, meaning that the mapping between incoming port-links and outgoing port-links is performed independently of the flow and / or the protocol executed by the flow. Furthermore, according to some embodiments, if an outgoing link of a particular network device fails, the network device is configured to switch its routing policy from GPU policy-based routing to standard ECMP routing, obtain a new available outgoing link (from various available outgoing links), and send the flow to that new outgoing link. It should be noted that in this case, flow collisions that cause congestion may occur.
[0127] Referring now to Figure 9, a block diagram of the cloud infrastructure 900 is shown illustrating different types of connectivity according to several embodiments. The infrastructure 900 includes multiple racks, for example, rack 1 910, rack D 920, and racks M and 930. Racks 910 and 930 include multiple host machines. For example, rack 1 910 includes multiple hosts (for example, K host machines), i.e., hosts 1-A, 912 to 1-K, 914, and rack M includes K host machines, i.e., hosts MA, 932 to MK, 934. Rack D 920 includes one or more host machines 922, each of which includes multiple CPUs, i.e., host machines 922 are non-GPU host machines. Each of racks 910, 920, and 930 includes TOR switches, i.e., TOR 1, 916, TOR D 926, and TOR M 936, respectively, which are communicably connected to the host machines in their respective racks. Furthermore, TOR switches 916, 926, and 936 are communicated to multiple spine switches, namely spine switches 940 and 950.
[0128] As shown in Figure 9, there is a first connection (i.e., connection 1, shown by the dashed line) from one GPU host (i.e., host 1-A 912) to another GPU host (i.e., host MA 932), and a second connection (i.e., connection 2, shown by the dotted line) from one GPU host (i.e., host 1-K 914) to a non-GPU host (i.e., host D 922). With respect to connection 1, data packets related to the flow are routed on a hop-by-hop basis by configuring each intermediate network device based on a GPU-based policy routing mechanism, as described above in Figure 8. Specifically, data packets from connection 1 are routed along the dashed links shown in Figure 9, i.e., from host 1-A to TOR 1, from TOR 1 to spine switch 1, from spine switch 1 to TOR M, and finally from TOR M to host MA. Each network device is configured to link its incoming link-port to its outgoing link-port.
[0129] In contrast, connection 2 is from a GPU-based host (i.e., host 1-K 914) to a non-GPU-based host (i.e., host D 922). In this case, TOR 1 916 is configured to connect the incoming link-port, i.e., the port and link (971) that receives data packets from the GPU-based host, to the outgoing link-port, i.e., the output port and link 972 connected to the output port. Thus, TOR 1 uses the outgoing link 972 to forward the data packets to spine switch 1 940. Since the data packets are destined for a non-GPU-based host, spine switch 1 940 does not utilize a policy-based routing mechanism to forward the packets to TOR D. Rather, spine switch 1 940 utilizes the ECMP routing mechanism to select one of the available links 980 to forward the packets to TOR D, and TOR D forwards the data packets to host D 922.
[0130] In some embodiments, the flow collisions described above with reference to Figure 7 are avoided by the network devices of the cloud infrastructure described herein by implementing a modified version of ECMP routing. In this case, the ECMP hash algorithm is modified so that traffic arriving through a particular incoming port-link of the network device is hashed (and directed) to the same outgoing port-link of the network device. Specifically, each network device implements a modified ECMP algorithm to determine the outgoing port-link used to forward packets, where the modified ECMP algorithm always hashes any packets received on a particular incoming port-link to the same outgoing port-link. For example, consider a case where the first and second packets are received on the first incoming port-link of the network device, and the third and fourth packets are received on the second incoming port-link of the network device. In such a situation, the network device is configured to implement a modified ECMP algorithm, in which the network device transmits the first and second packets on the network device's first outgoing port-link, the third and fourth packets on the network device's second outgoing port-link, the network device's first inbound port-link being different from the network device's second inbound port-link, and the network device's first outgoing port-link being different from the network device's second outgoing port-link. In some embodiments, the above functionality can be enabled by modifying the information stored in a forwarding information database (e.g., a forwarding table, an ECMP table, etc.).
[0131] Referring now to Figure 10, exemplary configurations of rack 1000 according to several embodiments are shown. As shown in Figure 10, rack 1000 includes two host machines, namely host machine 1010 and host machine 1020. Although rack 1000 is shown to include two host machines, it will be understood that rack 1000 may include a greater number of host machines. Each host machine includes multiple GPUs and multiple CPUs. For example, host machine 1010 includes multiple CPUs 1012 and multiple GPUs 1014, and host machine 1020 includes multiple CPUs 1022 and multiple GPUs 1024.
[0132] The host machines are communicably connected to different network fabrics via different TOR switches. For example, host machines 1010 and 1020 are communicably connected to a network fabric (referred to herein as the front-end network of rack 1000) via a TOR switch, namely TOR 1 switch 1050. The front-end network may correspond to an external network. More specifically, host machine 1010 is connected to the front-end network via a network interface card (NIC) 1030 and a network virtualization device (NVD) 1035 connected to TOR 1 switch 1050. Host machine 1020 is connected to the front-end network via NIC 1040 and NVD 1045 connected to TOR 1 switch 1050. Thus, in some embodiments, the CPU of each host machine can communicate with the front-end network via the NIC, NVD, and TOR switch. For example, the CPU 1012 of the host machine 1010 can communicate with the front-end network via the NIC 1030, NVD 1035, and TOR 1 switch 1050.
[0133] Host machines 1010 and 1020 are connected to a Quality of Service (QoS) enabled backend network. The QoS-enabled backend network is referred to herein as the backend network corresponding to the GPU cluster network, as shown in Figure 6. Host machine 1010 is connected via another NIC 1065 to a TOR 2 switch 1060 that enables host machine 1010 to communicate with the backend network. Similarly, host machine 1020 is connected via NIC 1080 to a TOR 2 switch 1060 that enables host machine 1020 to communicate with the backend network. Thus, multiple GPUs on each host machine can communicate with the backend network via NICs and TOR switches. In this way, multiple CPUs utilize separate sets of NICs (compared to the set of NICs used by the GPUs) to communicate with the frontend and backend networks, respectively.
[0134] Figure 11A shows an exemplary flowchart 1100 illustrating the steps performed by a network device when routing packets, according to several embodiments. The processes shown in Figure 11A can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system, hardware, or combination thereof. The software can be stored in a non-temporary storage medium (e.g., a memory device). The methods presented in Figure 11A and described below are intended to be illustrative and non-limiting. Figure 11A shows various processing steps occurring in a particular sequence or order, but this is not intended to be limiting. In some alternative embodiments, the steps may be performed in several different orders, or some steps may be performed in parallel.
[0135] The process begins in step 1105, where the network device receives a data packet transmitted by the host machine's graphics processing unit (GPU). In step 1110, the network device determines the incoming port / link from which the packet was received. In step 1115, the network device identifies the outgoing port / link corresponding to the incoming port / link (from which the packet was received) based on policy routing information. According to some embodiments, the policy routing information corresponds to a pre-configured GPU routing table for the network device that links each incoming port-link of the network device to a unique outgoing link-port of the network device.
[0136] Next, the process moves to step 1120, where a query is executed to determine whether the outgoing port-link is functional, for example, whether the outgoing link is active. If the response to the query is positive (i.e., the link is active), the process moves to step 1125; if the response to the query is negative (i.e., the link is in a failed / inactive state), the process moves to step 1130. In step 1125, the network device uses the outgoing port-link (identified in step 1115) to forward the received data packet to another network device. In step 1130, the network device retrieves flow information for the data packet, which may correspond to a 4-tuple associated with the packet (i.e., source port, destination port, source IP address, destination IP address). Based on the retrieved flow information, the network device uses ECMP routing to identify a new outgoing port-link, i.e., an available outgoing port-link. The process then moves to step 1135, where the network device uses the newly acquired outgoing port-link to forward the data packet received in step 1105.
[0137] Figure 11B shows another exemplary flowchart 1150 illustrating the steps performed by a network device when routing packets, according to several embodiments. The processes shown in Figure 11B can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system, hardware, or combination thereof. The software can be stored in a non-temporary storage medium (e.g., a memory device). The methods presented in Figure 11B and described below are intended to be illustrative and non-limiting. Figure 11B shows various processing steps occurring in a particular sequence or order, but this is not intended to be limiting. In some alternative embodiments, the steps may be performed in several different orders, or some steps may be performed in parallel.
[0138] The process begins in step 1155, where the network device receives a data packet sent by the host machine's graphics processing unit (GPU). In step 1160, the network device determines the flow information of the received packet. In some embodiments, the flow information may correspond to a 4-tuple associated with the packet (i.e., source port, destination port, source IP address, destination IP address). In step 1165, the network device computes the outgoing port-link by implementing a modified version of ECMP routing. The modified ECMP algorithm hashes any packet received on a particular incoming port-link so that it is always sent on the same outgoing port-link.
[0139] Next, the process moves to step 1170, where a query is executed to determine whether the outgoing port-link (determined in step 1165) is functional, for example, whether the outgoing link is active. If the response to the query is positive (i.e., the link is active), the process moves to step 1175; if the response to the query is negative (i.e., the link is in a failed / inactive state), the process moves to step 1180. In step 1175, the network device uses the outgoing port-link (identified in step 1165) to forward the received data packet to another network device. If it is determined that the identified outgoing port-link (in step 1165) is inactive, the process moves to step 1180. In step 1180, the network device implements ECMP routing (i.e., standard ECMP routing) to identify a new outgoing port-link. The process then moves to step 1185, where the network device uses the newly calculated outgoing port-link to forward the data packet received in step 1155.
[0140] It should be noted that the aforementioned technique for routing data packets originating from the host machine's GPU increases throughput by 20% in small clusters and 70% in large clusters (i.e., a three-fold improvement over the standard ECMP routing algorithm).
[0141] Cloud Infrastructure Implementation Examples As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider can also provide various services associated with the infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering). Therefore, since these services can be policy-driven, IaaS users may be able to implement policies that drive load balancing to maintain application availability and performance.
[0142] In some cases, IaaS customers can access resources and services over a wide area network (WAN), such as the internet, and install the rest of their application stack using the cloud provider's services. For example, a user can log into an IaaS platform, create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform a variety of functions, including balancing network traffic, troubleshooting applications, monitoring performance, and managing disaster recovery.
[0143] In most cases, the cloud computing model requires the participation of a cloud provider. While not mandatory, the cloud provider may be a third-party service specializing in IaaS offerings (e.g., offering, renting, or selling). An entity may also choose to deploy a private cloud and become its own provider of infrastructure services.
[0144] In some cases, IaaS deployment is the process of deploying a new application, or a new version of an application, to a prepared application server, etc. IaaS deployment may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). IaaS deployments are often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Therefore, the customer may be responsible for handling tasks such as (OS), middleware, and / or application deployment (e.g., on self-service virtual machines, which can be spun up on demand).
[0145] In some examples, IaaS provisioning may refer to acquiring the computers or virtual hosts to be used, and even installing the necessary libraries or services on those computers or virtual hosts. In most cases, deployment does not include provisioning, which must be done first.
[0146] In some cases, IaaS provisioning presents two distinct challenges. First, there's the initial challenge of provisioning an initial set of infrastructure before doing anything else. Second, there's the challenge of evolving the existing infrastructure after everything has been provisioned (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by enabling the declarative definition of infrastructure configuration. In other words, the infrastructure (e.g., which components are needed and how these components interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which and how each of them works together) can be described declaratively. In some cases, once the topology is defined, workflows can be generated to create and / or manage the different components described in the configuration files.
[0147] In some examples, infrastructure can have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs), also known as core networks (e.g., potential on-demand pools of configurable and / or shared computing resources). In some examples, there may also be one or more security group rules provisioned to define how network security is set up, and one or more virtual machines (VMs). Other infrastructure elements such as load balancers and databases may also be provisioned. As more infrastructure elements are desired and / or added, the infrastructure can evolve incrementally.
[0148] In some cases, sequential deployment techniques may be employed to enable the deployment of infrastructure code across various virtual computing environments. Furthermore, the techniques described can enable infrastructure management within these environments. In some examples, a service team may write code that is intended to be deployed to one or more, but often many, different production environments (for example, across various different geographical locations, sometimes even worldwide). However, in some examples, the infrastructure to which the code will be deployed must be set up first. In some cases, provisioning can be done manually, resources can be provisioned using provisioning tools, and / or the code can be deployed using deployment tools after the infrastructure has been provisioned.
[0149] Figure 12 is a block diagram 1200 showing an example pattern of an IaaS architecture according to at least one embodiment. A service operator 1202 can be communicably connected to a secure host tenancy 1204 which may include a virtual cloud network (VCN) 1206 and a secure host subnet 1208. In some examples, the service operator 1202 may use one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)) or wearable devices (Google Glass® head-mounted displays) that run software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, PalmOS, and are capable of using the Internet, email, short message service (SMS), Blackberry®, or other communication protocols. Alternatively, the client computing device may be a general-purpose personal computer, including, for example, personal computers and / or laptop computers running various versions of the Microsoft Windows® operating system, the Apple Macintosh® operating system, and / or the Linux® operating system. The client computing device may also be a workstation computer running various commercially available UNIX® or UNIX-like operating systems, including various GNU / Linux operating systems such as Google Chrome OS, without limitation.Alternatively, or in addition, the client computing device may be any other electronic device, such as a thin client computer, an internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device, that can communicate via a network that can access the VCN1206 and / or the Internet.
[0150] VCN1206 may include a local peering gateway (LPG) 1210, which can be communicated to a Secure Shell (SSH) VCN1212 via the LPG1210 included in the SSH VCN1212. SSH VCN1212 may include an SSH subnet 1214, which can be communicated to a control plane VCN1216 via the LPG1210 included in the control plane VCN1216. Furthermore, SSH VCN1212 can be communicated to a data plane VCN1218 via the LPG1210. The control plane VCN1216 and data plane VCN1218 may be included in a service tenancy 1219, which may be owned and / or operated by an IaaS provider.
[0151] The control plane VCN 1216 may include a control plane demilitarized zone (DMZ) layer 1220 that acts as a perimeter network (for example, the portion of the corporate network between the corporate intranet and the external network). DMZ-based servers can have limited trust and help mitigate security breaches. Furthermore, the DMZ layer 1220 may include a control plane application layer 1224 that may include one or more load balancer (LB) subnets 1222 and application subnets 1226, and a control plane data layer 1228 that may include database (DB) subnets 1230 (for example, a front-end DB subnet and / or a back-end DB subnet). The LB subnet 1222 included in the control plane DMZ layer 1220 can be communicatively connected to the application subnet 1226 included in the control plane application layer 1224 and the internet gateway 1234 which can be included in the control plane VCN 1216. The application subnet 1226 can be communicatively connected to the DB subnet 1230 included in the control plane data layer 1228, the service gateway 1236, and the network address translation (NAT) gateway 1238. The control plane VCN 1216 can include the service gateway 1236 and the NAT gateway 1238.
[0152] The control plane VCN 1216 may include a data plane mirror application layer 1240 which may include an application subnet 1226. The application subnet 1226 included in the data plane mirror application layer 1240 may include a virtual network interface controller (VNIC) 1242 which can run compute instance 1244. Compute instance 1244 can communicate with application subnet 1226 of the data plane mirror application layer 1240 which may include application subnet 1226 in the data plane application layer 1246.
[0153] The data plane VCN1218 may include a data plane application layer 1246, a data plane DMZ layer 1248, and a data plane data layer 1250. The data plane DMZ layer 1248 may include an LB subnet 1222 that can be communicatively connected to the application subnet 1226 of the data plane application layer 1246 and the internet gateway 1234 of the data plane VCN1218. The application subnet 1226 can be communicatively connected to the service gateway 1236 of the data plane VCN1218 and the NAT gateway 1238 of the data plane VCN1218. The data plane data layer 1250 may also include a DB subnet 1230 that can be communicatively connected to the application subnet 1226 of the data plane application layer 1246.
[0154] The Internet gateway 1234 of the control plane VCN1216 and the Internet gateway 1234 of the data plane VCN1218 can communicate with a metadata management service 1252, which can communicate with the public internet 1254. The public internet 1254 can communicate with the NAT gateway 1238 of the control plane VCN1216 and the NAT gateway 1238 of the data plane VCN1218. The service gateway 1236 of the control plane VCN1216 and the service gateway 1236 of the data plane VCN1218 can communicate with a cloud service 1256.
[0155] In some examples, a service gateway 1236 of the control plane VCN1216 or data plane VCN1218 can make application programming interface (API) calls to a cloud service 1256 without going through the public internet 1254. API calls from the service gateway 1236 to the cloud service 1256 can be one-way; that is, the service gateway 1236 can make an API call to the cloud service 1256, and the cloud service 1256 can send the requested data to the service gateway 1236. However, the cloud service 1256 cannot initiate an API call to the service gateway 1236.
[0156] In some examples, a secure host tenancy 1204 can connect directly to a service tenancy 1219, which would otherwise be isolated. A secure host subnet 1208 can communicate with an SSH subnet 1214 via an LPG 1210, which can enable bidirectional communication in an isolated system. By connecting the secure host subnet 1208 to the SSH subnet 1214, the secure host subnet 1208 can access other entities within the service tenancy 1219.
[0157] The control plane VCN1216 can enable users of service tenancy 1219 to set up or otherwise provision desired resources. Desired resources provisioned in the control plane VCN1216 can be deployed or otherwise used in the data plane VCN1218. In some examples, the control plane VCN1216 can be isolated from the data plane VCN1218, and the data plane mirror application layer 1240 of the control plane VCN1216 can communicate with the data plane application layer 1246 of the data plane VCN1218 via a VNIC 1242, which can be included in the data plane mirror application layer 1240 and the data plane application layer 1246.
[0158] In some examples, a system user or customer may make a request, such as a create, read, update, or delete (CRUD) operation, via the public internet 1254, and the public internet 1254 may communicate such a request to the metadata management service 1252. The metadata management service 1252 may communicate the request to the control plane VCN 1216 via the internet gateway 1234. The request may be received by the LB subnet 1222, which is included in the control plane DMZ layer 1220. The LB subnet 1222 may determine that the request is valid, and in response to this determination, the LB subnet 1222 may send the request to the application subnet 1226, which is included in the control plane application layer 1224. If the request is validated and requires a call to the public internet 1254, the call to the public internet 1254 may be sent to the NAT gateway 1238, which can make calls to the public internet 1254. Memory that may be desired to be stored by the request may be stored in the DB subnet 1230.
[0159] In some examples, the data plane mirror application layer 1240 can facilitate direct communication between the control plane VCN 1216 and the data plane VCN 1218. For example, it may be desirable to apply configuration changes, updates, or other suitable modifications to resources contained in the data plane VCN 1218. The control plane VCN 1216 can communicate directly with the resources contained in the data plane VCN 1218 via VNIC 1242, thereby enabling it to perform configuration changes, updates, or other suitable modifications to the resources.
[0160] In some embodiments, the control plane VCN1216 and data plane VCN1218 can be included in the service tenancy 1219. In this case, the system user or customer does not have to own or operate either the control plane VCN1216 or the data plane VCN1218. Instead, the IaaS provider can own or operate the control plane VCN1216 and the data plane VCN1218, and both can be included in the service tenancy 1219. This embodiment allows for network isolation, thereby preventing the user or customer from interacting with other users' resources or other customers' resources. This embodiment also allows the system user or customer to store databases privately without having to rely on the public internet 1254, which may not have the desired level of security for storage.
[0161] In another embodiment, the LB subnet 1222 included in the control plane VCN 1216 can be configured to receive signals from the service gateway 1236. In this embodiment, the control plane VCN 1216 and the data plane VCN 1218 can be configured to be invoked by the IaaS provider's customers without calling the public internet 1254. The database used by the customer can be stored in a service tenancy 1219 that can be controlled by the IaaS provider and isolated from the public internet 1254, so the IaaS provider's customers may prefer this embodiment.
[0162] Figure 13 is a block diagram 1300 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1302 (for example, service operator 1202 in Figure 12) can be communicatively connected to a secure host tenancy 1304 (for example, secure host tenancy 1204 in Figure 12) which may include a virtual cloud network (VCN) 1306 (for example, VCN1206 in Figure 12) and a secure host subnet 1308 (for example, secure host subnet 1208 in Figure 12). The VCN 1306 may include a local peering gateway (LPG) 1310 (for example, LPG1210 in Figure 12), which can be communicatively connected to a secure shell (SSH) VCN 1312 (for example, SSH VCN1212 in Figure 12) via the LPG 1310 contained within the SSH VCN 1312. SSH VCN1312 may include SSH subnet 1314 (for example, SSH subnet 1214 in Figure 12), and SSH VCN1312 may be communicably connected to control plane VCN1316 (for example, control plane VCN1216 in Figure 12) via LPG1310 included in control plane VCN1316. Control plane VCN1316 may be included in service tenancy 1319 (for example, service tenancy 1219 in Figure 12), and data plane VCN1318 (for example, data plane VCN1218 in Figure 12) may be included in customer tenancy 1321, which may be owned or operated by a user or customer of the system.
[0163] The control plane VCN1316 may include a control plane DMZ layer 1320 (for example, control plane DMZ layer 1220 in Figure 12) which can include an LB subnet 1322 (for example, LB subnet 1222 in Figure 12), a control plane application layer 1324 (for example, control plane application layer 1224 in Figure 12) which can include an application subnet 1326 (for example, application subnet 1226 in Figure 12), and a control plane data layer 1328 (for example, control plane data layer 1228 in Figure 12) which can include a database (DB) subnet 1330 (similar to DB subnet 1230 in Figure 12). The LB subnet 1322 included in the control plane DMZ layer 1320 can be communicatively connected to the application subnet 1326 included in the control plane application layer 1324 and to an internet gateway 1334 (for example, internet gateway 1234 in Figure 12), which can be included in the control plane VCN 1316. The application subnet 1326 can be communicatively connected to the DB subnet 1330 included in the control plane data layer 1328 and to a service gateway 1336 (for example, the service gateway in Figure 12) and a network address translation (NAT) gateway 1338 (for example, NAT gateway 1238 in Figure 12). The control plane VCN 1316 can include the service gateway 1336 and the NAT gateway 1338.
[0164] The control plane VCN 1316 may include a data plane mirror application layer 1340 (for example, the data plane mirror application layer 1240 in Figure 12) which may include an application subnet 1326. The application subnet 1326 included in the data plane mirror application layer 1340 may include a virtual network interface controller (VNIC) 1342 (for example, VNIC 1242) which may run a compute instance 1344 (for example, similar to compute instance 1244 in Figure 12). The compute instance 1344 can facilitate communication between the application subnet 1326 of the data plane mirror application layer 1340 and the application subnet 1326 that may be included in the data plane application layer 1346 (for example, the data plane application layer 1246 in Figure 12) via the VNIC 1342 included in the data plane mirror application layer 1340 and the VNIC 1342 included in the data plane application layer 1346.
[0165] The Internet gateway 1334 included in the control plane VCN 1316 can communicate with the metadata management service 1352 (for example, the metadata management service 1252 in Figure 12), which can communicate with the public internet 1354 (for example, the public internet 1254 in Figure 12). The public internet 1354 can communicate with the NAT gateway 1338 included in the control plane VCN 1316. The service gateway 1336 included in the control plane VCN 1316 can communicate with the cloud service 1356 (for example, the cloud service 1256 in Figure 12).
[0166] In some examples, the data plane VCN1318 can be included in customer tenancy 1321. In this case, the IaaS provider can provide a control plane VCN1316 to each customer, and the IaaS provider can set up a unique compute instance 1344 included in service tenancy 1319 for each customer. Each compute instance 1344 can enable communication between the control plane VCN1316 included in service tenancy 1319 and the data plane VCN1318 included in customer tenancy 1321. The compute instance 1344 can enable resources provisioned in the control plane VCN1316 included in service tenancy 1319 to be deployed to or otherwise used in the data plane VCN1318 included in customer tenancy 1321.
[0167] In another example, an IaaS provider's customer may have a database residing in customer tenancy 1321. In this example, control plane VCN 1316 may include a data plane mirror application layer 1340 that can include application subnet 1326. The data plane mirror application layer 1340 may reside in data plane VCN 1318, but may not reside in data plane VCN 1318. That is, the data plane mirror application layer 1340 can access customer tenancy 1321, but does not have to reside in data plane VCN 1318, nor does it have to be owned and operated by the IaaS provider's customer. The data plane mirror application layer 1340 may be configured to make calls to data plane VCN 1318, but does not have to be configured to make calls to any entity included in control plane VCN 1316. Customers may want to deploy or otherwise use resources in the data plane VCN1318 that are provisioned in the control plane VCN1316, and the data plane mirror app layer 1340 can facilitate the customer's desired deployment or other use of resources.
[0168] In some embodiments, a customer of the IaaS provider can apply filters to the data plane VCN 1318. In this embodiment, the customer can determine what the data plane VCN 1318 can access and can restrict access from the data plane VCN 1318 to the public internet 1354. The IaaS provider may not be able to apply filters to or control access from the data plane VCN 1318 to any external network or database. Applying filters and controls to the data plane VCN 1318 included in the customer tenancy 1321 can help isolate the data plane VCN 1318 from other customers and the public internet 1354.
[0169] In some embodiments, the service gateway 1336 may call a cloud service 1356 to access a service that may not reside on the public internet 1354, on the control plane VCN 1316, or on the data plane VCN 1318. The connection between the cloud service 1356 and the control plane VCN 1316 or data plane VCN 1318 may be live or not continuous. The cloud service 1356 may reside on a separate network owned or operated by the IaaS provider. The cloud service 1356 may be configured to receive calls from the service gateway 1336 and not to receive calls from the public internet 1354. Some cloud services 1356 may be isolated from other cloud services 1356, and the control plane VCN 1316 may be isolated from cloud services 1356 that may not be in the same region as the control plane VCN 1316. For example, the control plane VCN 1316 may be located in "Region 1", and the cloud service "Deployment 12" may be located in "Region 1" and "Region 2". If a call to deployment 12 is made by a service gateway 1336 included in the control plane VCN 1316 located in region 1, this call can be sent to deployment 12 in region 1. In this example, the control plane VCN 1316 or deployment 12 in region 1 may not be communicatively connected to or otherwise communicating with deployment 12 in region 2.
[0170] Figure 14 is a block diagram 1400 showing another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1402 (for example, service operator 1202 in Figure 12) can be communicatively connected to a secure host tenancy 1404 (for example, secure host tenancy 1204 in Figure 12), which may include a virtual cloud network (VCN) 1406 (for example, VCN 1206 in Figure 12) and a secure host subnet 1408 (for example, secure host subnet 1208 in Figure 12). VCN 1406 may include an LPG 1410 (for example, LPG 1210 in Figure 12), which can be communicatively connected to an SSH VCN 1412 (for example, SSH VCN 1212 in Figure 12) via an LPG 1410 included in SSH VCN 1412. SSH VCN1412 may include SSH subnet 1414 (for example, SSH subnet 1214 in Figure 12), and SSH VCN1412 may be communicably connected to control plane VCN1416 (for example, control plane VCN1216 in Figure 12) via LPG1410 included in control plane VCN1416, and to data plane VCN1418 (for example, data plane 1218 in Figure 12) via LPG1410 included in data plane VCN1418. Control plane VCN1416 and data plane VCN1418 may be included in service tenancy 1419 (for example, service tenancy 1219 in Figure 12).
[0171] The control plane VCN1416 may include a control plane DMZ layer 1420 (for example, control plane DMZ layer 1220 in Figure 12) which may include a load balancer (LB) subnet 1422 (for example, LB subnet 1222 in Figure 12), a control plane application layer 1424 (for example, control plane application layer 1224 in Figure 12) which may include an application subnet 1426 (similar to application subnet 1226 in Figure 12), and a control plane data layer 1428 (for example, control plane data layer 1228 in Figure 12) which may include a DB subnet 1430. The LB subnet 1422 included in the control plane DMZ layer 1420 can be communicatively connected to the application subnet 1426 included in the control plane application layer 1424 and to an internet gateway 1434 (for example, internet gateway 1234 in Figure 12) which can be included in the control plane VCN 1416. The application subnet 1426 can be communicatively connected to the DB subnet 1430 included in the control plane data layer 1428 and to a service gateway 1436 (for example, the service gateway in Figure 12) and a network address translation (NAT) gateway 1438 (for example, NAT gateway 1238 in Figure 12). The control plane VCN 1416 may include the service gateway 1436 and the NAT gateway 1438.
[0172] The data plane VCN1418 may include a data plane application layer 1446 (for example, the data plane application layer 1246 in Figure 12), a data plane DMZ layer 1448 (for example, the data plane DMZ layer 1248 in Figure 12), and a data plane data layer 1450 (for example, the data plane data layer 1250 in Figure 12). The data plane DMZ layer 1448 may include an LB subnet 1422 that can be communicatively connected to the trusted application subnet 1460 and the untrusted application subnet 1462 of the data plane application layer 1446, as well as the internet gateway 1434, which are included in the data plane VCN1418. The trusted application subnet 1460 can be communicatively connected to the service gateway 1436 included in the data plane VCN1418, the NAT gateway 1438 included in the data plane VCN1418, and the DB subnet 1430 included in the data plane data layer 1450. The untrusted application subnet 1462 can be communicatively connected to the service gateway 1436 included in the data plane VCN 1418 and the DB subnet 1430 included in the data plane data layer 1450. The data plane data layer 1450 may include the DB subnet 1430, which can be communicatively connected to the service gateway 1436 included in the data plane VCN 1418.
[0173] An untrusted application subnet 1462 may include one or more primary VNICs 1464(1)-(N) that can be communicatively connected to tenant virtual machines (VMs) 1466(1)-(N). Each tenant VM 1466(1)-(N) may be communicatively connected to each application subnet 1467(1)-(N) that can be included in each container egress VCN 1468(1)-(N) that can be included in each customer tenancy 1470(1)-(N). Each secondary VNIC 1472(1)-(N) can facilitate communication between the untrusted application subnet 1462 included in the data plane VCN 1418 and the application subnets included in the container egress VCN 1468(1)-(N). Each container egress VCN 1468(1)-(N) may include a NAT gateway 1438 that can be communicatively connected to the public internet 1454 (for example, the public internet 1254 in Figure 12).
[0174] The Internet gateway 1434 included in the control plane VCN1416 and the Internet gateway 1434 included in the data plane VCN1418 can communicate with a metadata management service 1452 (for example, the metadata management system 1252 in Figure 12), which can communicate with the public internet 1454. The public internet 1454 can communicate with the NAT gateway 1438 included in the control plane VCN1416 and the NAT gateway 1438 included in the data plane VCN1418. The service gateway 1436 included in the control plane VCN1416 and the service gateway 1436 included in the data plane VCN1418 can communicate with a cloud service 1456.
[0175] In some embodiments, the data plane VCN1418 can be integrated with a customer tenancy 1470. This integration may be useful or desirable for the IaaS provider's customer in several cases, such as when they may want support when executing code. The customer may provide code to be executed that may be disruptive, may communicate with other customer resources, or may cause undesirable effects. Accordingly, the IaaS provider can decide whether or not to execute the code that the customer has provided to the IaaS provider.
[0176] In some examples, an IaaS provider's customer may grant the IaaS provider temporary network access and request functionality to be attached to a dataplane tier application 1446. The code that performs the functionality can run on VMs 1466(1)-(N) and does not need to be configured to run anywhere else on the dataplane VCN 1418. Each VM 1466(1)-(N) can be connected to one customer tenancy 1470. Each container 1471(1)-(N) contained within VMs 1466(1)-(N) can be configured to run the code. In this case, there may be a double isolation (for example, containers 1471(1)-(N) run the code, and containers 1471(1)-(N) can be contained within VMs 1466(1)-(N) that are at least in an untrusted application subnet 1462). This can help prevent incorrect or otherwise undesirable code from damaging the IaaS provider's network or the networks of different customers. Containers 1471(1)-(N) can be communicatively connected to customer tenancy 1470 and can be configured to send or receive data from customer tenancy 1470. Containers 1471(1)-(N) do not need to be configured to send or receive data from any other entities in the data plane VCN 1418. When code execution is complete, the IaaS provider can terminate or otherwise dispose of containers 1471(1)-(N).
[0177] In some embodiments, a trusted application subnet 1460 can execute code that may be owned or operated by the IaaS provider. In this embodiment, the trusted application subnet 1460 may be communicatively connected to the DB subnet 1430 and configured to perform CRUD operations in the DB subnet 1430. An untrusted application subnet 1462 may be communicatively connected to the DB subnet 1430, but in this embodiment, the untrusted application subnet may be configured to perform read operations within the DB subnet 1430. Containers 1471(1)-(N), which may be included in each customer's VMs 1466(1)-(N) and can execute code from the customer, do not need to be communicatively connected to the DB subnet 1430.
[0178] In other embodiments, the control plane VCN1416 and the data plane VCN1418 do not need to be directly communicatively linked. In this embodiment, direct communication between the control plane VCN1416 and the data plane VCN1418 is not required. However, communication can occur indirectly by at least one method. The IaaS provider may establish an LPG1410 that facilitates communication between the control plane VCN1416 and the data plane VCN1418. In another example, the control plane VCN1416 or the data plane VCN1418 can make a call to the cloud service 1456 via the service gateway 1436. For example, a call from the control plane VCN1416 to the cloud service 1456 may include a request for a service that can communicate with the data plane VCN1418.
[0179] Figure 15 is a block diagram 1500 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1502 (for example, service operator 1202 in Figure 12) can be communicatively connected to a secure host tenancy 1504 (for example, secure host tenancy 1204 in Figure 12), which may include a virtual cloud network (VCN) 1506 (for example, VCN 1206 in Figure 12) and a secure host subnet 1508 (for example, secure host subnet 1208 in Figure 12). VCN 1506 may include an LPG 1510 (for example, LPG 1210 in Figure 12), and LPG 1510 can be communicatively connected to an SSH VCN 1512 (for example, SSH VCN 1212 in Figure 12) via LPG 1510 contained within the SSH VCN 1512. SSH VCN1512 may include SSH subnet 1514 (for example, SSH subnet 1214 in Figure 12), and SSH VCN1512 may be communicably connected to control plane VCN1516 (for example, control plane VCN1216 in Figure 12) via LPG1510 included in control plane VCN1516, and to data plane VCN1518 (for example, data plane 1218 in Figure 12) via LPG1510 included in data plane VCN1518. Control plane VCN1516 and data plane VCN1518 may be included in service tenancy 1519 (for example, service tenancy 1219 in Figure 12).
[0180] The control plane VCN1516 may include a control plane DMZ layer 1520 (for example, control plane DMZ layer 1220 in Figure 12) which can include an LB subnet 1522 (for example, LB subnet 1222 in Figure 12), a control plane application layer 1524 (for example, control plane application layer 1224 in Figure 12) which can include an application subnet 1526 (for example, application subnet 1226 in Figure 12), and a control plane data layer 1528 (for example, control plane data layer 1228 in Figure 12) which can include a DB subnet 1530 (for example, DB subnet 1430 in Figure 14). The LB subnet 1522 included in the control plane DMZ layer 1520 can be communicatively connected to the application subnet 1526 included in the control plane application layer 1524, and to an internet gateway 1534 (for example, internet gateway 1234 in Figure 12) which can be included in the control plane VCN 1516. The application subnet 1526 can be communicatively connected to the DB subnet 1530 included in the control plane data layer 1528, and to a service gateway 1536 (for example, the service gateway in Figure 12) and a network address translation (NAT) gateway 1538 (for example, NAT gateway 1238 in Figure 12). The control plane VCN 1516 may include the service gateway 1536 and the NAT gateway 1538.
[0181] The data plane VCN 1518 may include a data plane application layer 1546 (for example, the data plane application layer 1246 in Figure 12), a data plane DMZ layer 1548 (for example, the data plane DMZ layer 1248 in Figure 12), and a data plane data layer 1550 (for example, the data plane data layer 1250 in Figure 12). The data plane DMZ layer 1548 may include trusted application subnets 1560 (for example, the trusted application subnet 1460 in Figure 14) and untrusted application subnets 1562 (for example, the untrusted application subnet 1462 in Figure 14) of the data plane application layer 1546, which are included in the data plane VCN 1518, as well as an LB subnet 1522 that can be communicatively connected to the internet gateway 1534. A trusted application subnet 1560 can be communicatively connected to the service gateway 1536 included in the data plane VCN 1518, the NAT gateway 1538 included in the data plane VCN 1518, and the DB subnet 1530 included in the data plane data layer 1550. An untrusted application subnet 1562 can be communicatively connected to the service gateway 1536 included in the data plane VCN 1518, and the DB subnet 1530 included in the data plane data layer 1550. The data plane data layer 1550 may include a DB subnet 1530 that can be communicatively connected to the service gateway 1536 included in the data plane VCN 1518.
[0182] An untrusted application subnet 1562 may include primary VNICs 1564(1)-(N) that can communicately connect to tenant virtual machines (VMs) 1566(1)-(N) residing within the untrusted application subnet 1562. Each tenant VM 1566(1)-(N) can execute code in its respective container 1567(1)-(N) and can communicately connect to an application subnet 1526 that can be included in a dataplane application layer 1546, which can be included in a container egress VCN 1568. Each secondary VNIC 1572(1)-(N) can facilitate communication between the untrusted application subnet 1562 included in the dataplane VCN 1518 and the application subnet included in the container egress VCN 1568. The container egress VCN may include a NAT gateway 1538 that can communicately connect to the public internet 1554 (for example, the public internet 1254 in Figure 12).
[0183] The Internet gateway 1534 included in the control plane VCN1516 and the Internet gateway 1534 included in the data plane VCN1518 can communicate with a metadata management service 1552 (for example, the metadata management system 1252 in Figure 12), which can communicate with the public internet 1554. The public internet 1554 can communicate with the NAT gateway 1538 included in the control plane VCN1516 and the NAT gateway 1538 included in the data plane VCN1518. The service gateway 1536 included in the control plane VCN1516 and the service gateway 1536 included in the data plane VCN1518 can communicate with a cloud service 1556.
[0184] In some examples, the pattern shown by the architecture in block diagram 1500 of Figure 15 may be considered an exception to the pattern shown by the architecture in block diagram 1400 of Figure 14, and may be desirable for the IaaS provider's customers when the IaaS provider cannot communicate directly with the customers (e.g., in a disconnected region). Customers can access each container 1567(1)~(N) contained within each customer's VM 1566(1)~(N) in real time. Each container 1567(1)~(N) can be configured to call each secondary VNIC 1572(1)~(N) contained within the application subnet 1526 of the data plane application layer 1546, which can be included in the container egress VCN 1568. The secondary VNICs 1572(1)~(N) can send calls to a NAT gateway 1538 which can send calls to the public internet 1554. In this example, containers 1567(1)-(N), which customers can access in real time, can be isolated from the control plane VCN1516 and from other entities included in the data plane VCN1518. Containers 1567(1)-(N) can also be isolated from other customers' resources.
[0185] In another example, a customer can use containers 1567(1)-(N) to invoke cloud service 1556. In this example, the customer can execute code in containers 1567(1)-(N) to request a service from cloud service 1556. Containers 1567(1)-(N) can send this request to secondary VNICs 1572(1)-(N), which can then send the request to a NAT gateway that can send the request to the public internet 1554. The public internet 1554 can then send this request to LB subnet 1522, which is included in control plane VCN 1516, via internet gateway 1534. In response to determining that the request is valid, the LB subnet can send this request to application subnet 1526, which can then send this request to cloud service 1556 via service gateway 1536.
[0186] It should be understood that the IaaS architectures 1200, 1300, 1400, and 1500 shown in the figures may include components other than those shown. Furthermore, the embodiments shown in the figures are only examples of some of the cloud infrastructure systems that may incorporate embodiments of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than those shown in the figures, may combine two or more components, or may have different configurations or arrangements of components.
[0187] In some embodiments, the IaaS systems described herein may include a self-service, subscription-based, elastically scalable, reliable, highly available, and securely delivered suite of applications, middleware, and database service offerings to customers. An example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the assignee.
[0188] Figure 16 shows an example computer system 1600 that can implement various embodiments. Any of the computer systems described above can be implemented using system 1600. As shown in the figure, computer system 1600 includes a processing unit 1604 that communicates with a number of peripheral subsystems via a bus subsystem 1602. These peripheral subsystems may include a processing accelerator 1606, an I / O subsystem 1608, a storage subsystem 1618, and a communication subsystem 1624. The storage subsystem 1618 includes a tangible computer-readable storage medium 1622 and system memory 1610.
[0189] The bus subsystem 1602 provides a mechanism for various components and subsystems of the computer system 1600 to communicate with each other as intended. While the bus subsystem 1602 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1602 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of the various bus architectures. Examples of such architectures include the Industry Standard Architecture (ISA) bus, the Microchannel Architecture (MCA) bus, the Extended ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Interconnect (PCI) bus, which can be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0190] The processing unit 1604 can be implemented as one or more integrated circuits (for example, conventional microprocessors or microcontrollers) and controls the operation of the computer system 1600. The processing unit 1604 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, the processing unit 1604 may be implemented as one or more independent processing units 1632 and / or 1634, each containing a single-core or multi-core processor. In other embodiments, the processing unit 1604 may be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.
[0191] In various embodiments, the processing unit 1604 can execute various programs in response to program code and can maintain multiple programs or processes running simultaneously. At any given time, some or all of the program code to be executed can reside in the processor 1604 and / or the storage subsystem 1618. The processor 1604 can provide the various functionalities described above through suitable programming. The computer system 1600 may further include a processing accelerator 1606, which may include a digital signal processor (DSP), a dedicated processor, and the like.
[0192] The I / O subsystem 1608 may include user interface input devices and user interface output devices. User interface input devices may include pointing devices such as keyboards, mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, voice input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion detection and / or gesture recognition devices, such as the Microsoft Kinect® motion sensor, which allows the user to control and interact with input devices, such as Microsoft Xbox® 360 game controllers, via a natural user interface (NUI) using gestures and voice commands. User interface input devices may also include eye gesture recognition devices, such as the Google Glass® blink detector, which detects eye activity from the user (e.g., blinking when taking a picture and / or selecting a menu) and translates that eye activity into input to an input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition detection device that allows the user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0193] Furthermore, user interface input devices may also include, without limitation, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, and webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Additionally, user interface input devices may include medical imaging input devices such as computed tomography scanners, magnetic resonance imaging scanners, positron emission tomography scanners, and medical ultrasound scanners. User interface input devices may also include audio input devices such as MIDI keyboards and electronic musical instruments.
[0194] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may include flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, and touchscreens. Generally, when the term “output device” is used, it is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1600 to a user or another computer. For example, user interface output devices may include, without limitation, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.
[0195] The computer system 1600 may include a storage subsystem 1618, which may include software components located in the system memory 1610. The system memory 1610 may store program instructions that can be loaded into and executed by the processing unit 1604, and data generated by the execution of these programs.
[0196] Depending on the configuration and type of the computer system 1600, the system memory 1610 may be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that the processing unit 1604 can access immediately, and / or data and / or program modules that are currently being operated and executed by the processing unit 1604. In some embodiments, the system memory 1610 may include several different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some embodiments, ROM may typically store a basic input / output system (BIOS) containing basic routines that help transfer information between elements within the computer system 1600, such as during startup. As an example, but not an limitation, the system memory 1610 also includes application programs 1612, program data 1614, and an operating system 1616, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. For example, Operating System 1616 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 10 OS, and Palm® OS.
[0197] The storage subsystem 1618 may also provide a tangible computer-readable storage medium for storing basic programming and data structures that provide the functionality of several embodiments. The storage subsystem 1618 may store software (programs, code modules, instructions) that, when executed by a processor, provides the functionality described above. These software modules or instructions can be executed by the processing unit 1604. The storage subsystem 1618 may also provide a repository for storing data used in accordance with this disclosure.
[0198] The storage subsystem 1610 may also include a computer-readable storage medium reader 1620, which can be further connected to a computer-readable storage medium 1622. The computer-readable storage medium 1622, together with the system memory 1610, or optionally in combination with the system memory 1610, can comprehensively represent remote, local, fixed, and / or removable storage devices in addition to storage media for temporarily and / or more permanently accommodating, storing, transmitting, and retrieving computer-readable information.
[0199] Computer-readable storage medium 1622 containing code or a portion of code may also include any suitable media known or used in the art, including storage and communication media, such as volatile and non-volatile, removable and non-removable media, implemented in any way or technique for storing and / or transmitting information. This may include tangible computer-readable storage media, such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This may also include intangible computer-readable media such as data signals, data transmissions, or any other media that can be used to transmit desired information and that can be accessed by the computer system 1600.
[0200] For example, the computer-readable storage medium 1622 may include hard disk drives that read from or write to non-removable non-volatile magnetic media, magnetic disk drives that read from or write to removable non-volatile magnetic disks, and optical disk drives that read from or write to removable non-volatile optical disks such as CD-ROMs, DVDs, and Blu-Ray® discs or other optical media. The computer-readable storage medium 1622 may also include, but are not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, and digital videotapes. The computer-readable storage medium 1622 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, and static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide computer-readable instructions, data structures, program modules, and other non-volatile storage devices for data to the computer system 1600.
[0201] The communication subsystem 1624 provides an interface with other computer systems and networks. The communication subsystem 1624 acts as an interface for receiving data from other systems and transmitting data from computer system 1600 to other systems. For example, the communication subsystem 1624 can enable computer system 1600 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1624 may include radio frequency (RF) transceiver components for accessing radio voice and / or data networks (using, for example, cellular technology, 3G, 4G, or advanced data network technologies such as EDGE (enhanced data rates for global evolution), WiFi (IEEE 1602.11 family standards), or other mobile communication technologies, or any combination thereof), a Global Positioning System (GPS) receiver component, and / or other components. In some embodiments, the communication subsystem 1624 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a radio interface.
[0202] In some embodiments, the communication subsystem 1624 may also receive input communications in the form of structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., on behalf of one or more users who can use the computer system 1600.
[0203] For example, the communication subsystem 1624 can be configured to receive data feeds 1626 in real time from users of social networks and / or other communication services, such as web feeds like Twitter® feeds, Facebook® updates, and Rich Site Summary (RSS) feeds, and / or to receive real-time updates from one or more third-party sources.
[0204] Furthermore, the communication subsystem 1624 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 1628 and / or event update 1630 of real-time events that are inherently continuous or boundaryless, with no clear end. Examples of applications that generate continuous data include, for example, sensor data applications, financial stock market indicators, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.
[0205] The communication subsystem 1624 may also be configured to output structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., to one or more databases that can communicate with one or more streaming data source computers connected to the computer system 1600.
[0206] Computer system 1600 can be one of a variety of types, including handheld portable devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing systems.
[0207] Because the nature of computers and networks is constantly changing, the description of the computer system 1600 shown in the figure is intended only as a specific example. Many other configurations are possible, having more or fewer components than the system shown in the figure. For example, customized hardware may be used and / or certain elements may be implemented in hardware, firmware, software (including applets), or combinations thereof. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosures and teachings provided herein, those skilled in the art will understand other means and / or methods for implementing various embodiments.
[0208] While specific embodiments have been described, various variations, modifications, alternative structures, and equivalents are also included within the scope of this disclosure. The embodiments are not limited to operating within a particular data processing environment, but can freely operate within multiple data processing environments. Furthermore, while the embodiments have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope of this disclosure is not limited to the described set of transactions and steps. The various features and aspects of the embodiments described above may be used individually or in combination.
[0209] Furthermore, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments may be implemented using hardware alone, software alone, or a combination thereof. The various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, where a component or module is described as performing some operation, such a configuration can be realized, for example, by designing electronic circuits to perform that operation, by programming programmable electronic circuits (such as a microprocessor) to perform that operation, or by a combination thereof. Processes may communicate using various techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0210] Therefore, the specification and drawings should be considered illustrative rather than restrictive. However, it will become clear that additions, reductions, deletions, and other modifications and alterations may be made without departing from the broader intent and scope set forth in the claims. Thus, while specific embodiments of the disclosure have been described, they are not intended to be restrictive. Various forms of modifications and equivalents are within the scope of the subsequent claims.
[0211] In the context describing the embodiments disclosed (particularly in the context of the following claims), the use of the terms “a, an,” and “the,” and similar referents, should be interpreted as encompassing both singular and plural, unless otherwise indicated in this disclosure and unless the context clearly contradicts it. The terms “comprising,” “having,” “including,” and “containing,” should be interpreted as open-ended terms (i.e., “including, but not limited to”), unless otherwise specified. The term “connected,” should be interpreted as being contained, attached, or joined to one another, in part or in whole, even if there are intervening elements. Unless otherwise indicated in this specification, the enumeration of value ranges is intended merely as a way of referring to each individual value within that range individually, and each individual value is incorporated into this disclosure as if it were individually enumerated herein. Unless otherwise indicated in this disclosure and unless it is clearly inconsistent with the context, all methods described herein can be performed in any preferred order. Any examples or illustrative language provided herein (e.g., "etc.") are intended solely to clarify embodiments and, unless otherwise asserted, do not limit the scope of this disclosure. Nothing in this specification should be construed as indicating that any unclaimed element is essential to the practice of this disclosure.
[0212] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as generally used to indicate that an item, term, etc., may be X, Y, Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically otherwise specified. Therefore, such disjunctive language is not, and should not, be intended to mean that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z, respectively.
[0213] This specification describes preferred embodiments of the Disclosure, including the best known mode for carrying out the Disclosure. Those skilled in the art will be able to see variations of these preferred embodiments by reading the foregoing description. Those skilled in the art can appropriately adopt such variations, and the Disclosure may be carried out in ways other than those specifically described herein. Therefore, the Disclosure includes all variations and equivalents of the subject matter described in the claims appended herein, as permitted by applicable law. Furthermore, unless otherwise indicated herein, any combination of the elements described above in all possible variations is incorporated herein.
[0214] All references, including publications, patent applications, and patents, cited herein are incorporated by reference to the same extent as if they were included in their entirety, as is individually and specifically indicated. While aspects of the disclosure are described in the preceding specification with reference to specific embodiments, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the disclosure described above may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings should be considered illustrative rather than restrictive.< / realm>
Claims
1. It is a method, Regarding packets transmitted by the host machine's graphics processing unit (GPU) and received by network devices, The network device determines the incoming port-link of the network device from which the packet was received, The network device identifies the outgoing port-link corresponding to the incoming port-link based on the GPU routing policy, Includes, The GPU routing policy is configured in advance before receiving the packet and establishes a mapping of each incoming port-link of the network device to a unique outgoing port-link of the network device. The above method applies to the packet, The network device forwards the packet on the outgoing port link of the network device. It further includes, The aforementioned transfer step is, The network device verifies the conditions related to the outgoing port-link of the network device, In accordance with the fulfillment of the above conditions, the packet is forwarded on the outgoing port-link of the network device, It further includes, If the above conditions are not met, The network device acquires flow information related to the packet, The network device executes an equal-cost multipath algorithm to acquire a new outgoing port-link for the network device based on the flow information, The network device forwards the packet on the new outgoing port-link of the network device, Methods that further include this.
2. The method according to claim 1, wherein the condition corresponds to determining whether the outgoing port-link of the network device is active.
3. The method according to claim 1, wherein the flow information relating to the packet includes at least information relating to the source port, destination port, source IP address, and destination IP address.
4. The method according to claim 1, wherein the network device is a top-of-rack (TOR) switch.
5. The method according to claim 4, wherein the TOR and the host machine are contained in a rack, and the host machine is connected to the TOR via a plurality of links, each link being associated with a network interface card.
6. The method according to claim 4, wherein the incoming port-link of the network device is a first link connecting the host machine to the TOR, and the outgoing port-link of the network device is a second link connecting the TOR to the spine switch.
7. The method according to claim 1, wherein the packet belongs to a GPU workload.
8. A network device, Processor and Memory containing instructions, The instructions, when executed by the processor, provide the network device with at least: Regarding packets transmitted by the host machine's graphics processing unit (GPU) and received by the network device, The incoming port-link of the network device that received the packet is determined, Based on the GPU routing policy, identify the outgoing port-link corresponding to the incoming port-link, Have them do it, The GPU routing policy is configured in advance before receiving the packet and establishes a mapping of each incoming port-link of the network device to a unique outgoing port-link of the network device. When the aforementioned instruction is executed by the processor, it will instruct the network device to at least: Regarding the aforementioned packet, Further, the network device is made to forward the packet via the outgoing port-link. The conditions related to the outgoing port-link of the network device are verified, Depending on whether the above conditions are met, the packet is forwarded on the outgoing port-link of the network device. It is further configured in this way, If the above conditions are not met, The flow information related to the aforementioned packet is obtained, Based on the flow information, an equal-cost multipath algorithm is executed to obtain a new outgoing port-link for the network device. A network device further configured to forward the packets on the new outgoing port-link of the network device.
9. The network device according to claim 8, wherein the condition corresponds to determining whether the outgoing port-link of the network device is active.
10. The network device according to claim 8, wherein the flow information related to the packet includes at least information related to the source port, destination port, source IP address, and destination IP address.
11. The network device according to claim 8, wherein the network device is a top-of-rack (TOR) switch.
12. The network device according to claim 11, wherein the TOR and the host machine are contained in a rack, and the host machine is connected to the TOR via a plurality of links, each link being associated with a network interface card.
13. The network device according to claim 11, wherein the incoming port-link of the network device is a first link connecting the host machine to the TOR, and the outgoing port-link of the network device is a second link connecting the TOR to the spine switch.
14. The aforementioned packet belongs to a GPU workload, according to the network device described in claim 8.
15. A program for causing a computer system to perform the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Packet transmission device, controller, and packet transmission control method
JP2017143344A
Multipath Transmission Design
JP2019503123A
Technologies for load balancing a network
US20190044849A1
Data forwarding method and device
US20190097914A1
Methods and apparatus to configure and manage network resources for use in network-based computing
US20190230025A1