Routing policy for graphics processing units

A GPU routing policy optimizes communication between GPUs in cloud environments by determining outgoing port links, enhancing performance and overcoming the limitations of ring topologies in existing cloud infrastructure.

JP7832965B2Active Publication Date: 2026-03-18ORACLE INT CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

Existing cloud infrastructure lacks the necessary high-performance computing resources and dedicated network performance for efficient communication between GPUs across multiple host machines, leading to degraded system performance due to issues like blocking in ring topologies.

Method used

Implementing a GPU routing policy that determines the outgoing port link based on a pre-configured mapping between receiving and outgoing ports, facilitating efficient packet forwarding in a cloud environment.

Benefits of technology

Enhances system performance by optimizing communication between GPUs across multiple host machines, addressing the limitations of ring networks and improving overall network efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832965000001
    Figure 0007832965000001
  • Figure 0007832965000002
    Figure 0007832965000002
  • Figure 0007832965000003
    Figure 0007832965000003
Patent Text Reader

Abstract

Described herein is a routing mechanism for graphical processing units (GPUs) hosted on multiple host machines in a cloud environment. For a packet sent by a GPU of a host machine and received by a network device, the network device determines an inbound port link of the network device on which the packet was received. The network device obtains flow information associated with the packet, and calculates an outbound port link of the network device based on the flow information according to a hashing algorithm. The hashing algorithm is configured to hash the packet received on a particular inbound port link of the network device and transmit it on the same outbound port link of the network device. The network device forwards the packet on the outbound port link of the network device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit and priority under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 63 / 215,264, filed Jun. 25, 2021, and U.S. Non - Provisional Application No. 17 / 734,865, filed May 2, 2022. The entire contents of the foregoing applications are hereby incorporated by reference in their entirety for all purposes.

[0002] Field The present disclosure relates to frameworks and routing mechanisms for Graphics Processing Units (GPUs) hosted on multiple host machines within a cloud environment.

Background Art

[0003] Background Organizations continue to migrate business applications and databases to the cloud to reduce the costs of purchasing, updating, and maintaining on - premise hardware and software. High - Performance Computing (HPC) applications consistently consume 100% of the computing power available to achieve a particular outcome or result. HPC applications require dedicated network performance, high - speed storage, advanced computing capabilities, and large amounts of memory, but these resources are lacking in the virtualized infrastructure that makes up today's commodity clouds.

[0004] Cloud infrastructure service providers offer newer and faster CPUs and graphical processing units (GPUs) to meet the requirements of HPC applications. Typically, virtual topologies are constructed to provision and communicate with various GPUs hosted on multiple host machines. In practice, ring topologies are used to connect various GPUs. However, ring networks inherently suffer from blocking, thus degrading the overall system performance. The embodiments described herein address these and other problems associated with connecting GPUs across multiple host machines. [Overview of the project]

[0005] overview This disclosure relates, in general terms, to routing mechanisms for graphical processing units (GPUs) hosted on multiple host machines in a cloud environment. Various embodiments are described herein, including methods, systems, and non-temporary computer-readable storage media for storing programs, code, or instructions executable by one or more processors. These exemplary embodiments are mentioned not to limit or define this disclosure, but to provide examples to aid in understanding this disclosure. Additional embodiments are described in the detailed description section, which provides further details therein.

[0006] One embodiment of the present disclosure relates to a method in which, in the case of a packet transmitted by a graphical processing unit (GPU) of a host machine and received by a network device, the network device determines the receiving port link of the network device from which the packet was received; the network device identifies the outgoing port link corresponding to the receiving port link based on a GPU routing policy, the GPU routing policy being pre-configured before the packet is received and establishing a mapping between each receiving port link of the network device and a unique outgoing port link of the network device; and the network device forwards the packet on the outgoing port link of the network device.

[0007] One aspect of the present disclosure provides a system comprising one or more data processors and a non-temporary computer-readable storage medium containing instructions, which, when executed on one or more data processors, causes one or more data processors to execute some or all of the methods disclosed herein.

[0008] Another aspect of this disclosure provides a computer program product, specifically embodied in a non-temporary machine-readable storage medium, which includes instructions configured to cause one or more data processors to execute some or all of the methods disclosed herein.

[0009] The foregoing, along with other features and embodiments, will become clearer by referring to the following specification, claims, and accompanying drawings.

[0010] The features, embodiments, and advantages of this disclosure will be better understood by reading the following detailed description with reference to the accompanying drawings. [Brief explanation of the drawing]

[0011] [Figure 1]This is a high-level diagram of a delivery environment showing a virtual or overlay cloud network hosted by a cloud service provider infrastructure, according to a specific embodiment. [Figure 2] This figure shows a simplified architectural diagram of the physical components within the physical network in CSPI according to a specific embodiment. [Figure 3] This figure shows an example of a configuration within a CSPI in which a host machine is connected to multiple network virtualization devices (NVDs) according to a specific embodiment. [Figure 4] This figure shows the connection between a host machine and an NVD to provide I / O virtualization to support multi-tenancy, according to a specific embodiment. [Figure 5] This figure shows a simplified block diagram of the physical network provided by CSPI according to a specific embodiment. [Figure 6] This figure shows a simplified block diagram of a cloud infrastructure incorporating a CLOS network configuration according to a specific embodiment. [Figure 7] This figure shows an exemplary scenario illustrating flow collisions in the cloud infrastructure of Figure 6, according to a specific embodiment. [Figure 8] This diagram shows a policy-based routing mechanism implemented in a cloud infrastructure according to a specific embodiment. [Figure 9] This figure shows a block diagram of a cloud infrastructure illustrating different types of connectivity within the cloud infrastructure according to a specific embodiment. [Figure 10] This figure shows an exemplary configuration of racks included in a cloud infrastructure according to a specific embodiment. [Figure 11A] This diagram shows a flowchart illustrating the steps performed by a network device when routing packets, according to a specific embodiment. [Figure 11B]This figure shows another flowchart illustrating the steps performed by a network device when routing packets, according to a specific embodiment. [Figure 12] This block diagram shows one pattern for implementing a cloud infrastructure as a service system, according to at least one embodiment. [Figure 13] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 14] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 15] This block diagram shows another pattern for implementing cloud infrastructure as a service system, with at least one embodiment. [Figure 16] A block diagram illustrating an exemplary computer system according to at least one embodiment. [Modes for carrying out the invention]

[0012] Detailed explanation In the following description, certain details are included to provide a complete understanding of a particular embodiment for illustrative purposes. However, it will be apparent that various embodiments can be carried out without these specific details. The figures and descriptions are not intended to be limiting. The term “exemplary” is used herein to mean “serving as an example, illustration, or illustration.” Any embodiment or design described herein as “exemplary” should not necessarily be construed as being preferable or advantageous over other embodiments or designs.

[0013] Cloud infrastructure architecture examples The term "cloud service" is generally used to refer to services that can be used by a user or customer on demand (e.g., via a subscription model) using systems and infrastructure (cloud infrastructure) provided by a cloud service provider (CSP). Usually, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premises servers and systems. Thus, the customer can utilize the cloud services provided by the CSP without having to purchase hardware and software resources for the service individually. Cloud services are designed to enable subscribing customers to easily and scalably access applications and computing resources without the customer investing in the procurement of the infrastructure used for the provision of the service.

[0014] There are several cloud service providers that offer different types of cloud services. Cloud services have various different types or models, such as Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS).

[0015] A customer can subscribe to one or more cloud services provided by a CSP. The customer can be any entity such as an individual, an organization, or a company. When a customer subscribes or registers for the services provided by a CSP, a tenant or account is created for that customer. Thereafter, the customer can access the one or more subscribed cloud resources associated with that account via this account.

[0016] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing service. In the IaaS model, the CSP provides the infrastructure (called Cloud Service Provider Infrastructure or CSPI) that customers can use to build their own customizable networks and deploy their customer resources. Thus, the customer's resources and network are hosted in the delivery environment by the infrastructure provided by the CSP. This differs from traditional computing, where the customer's resources and network are hosted by the infrastructure provided by the customer.

[0017] CSPI can comprise interconnected high-performance computing resources, including various host machines, memory resources, and network resources that form a physical network also known as the underlying network or foundation network. CSPI resources can be distributed across one or more data centers that may be geographically dispersed across one or more geographical regions. Virtualization software can run on these physical resources to provide a virtualized delivery environment. Virtualization creates an overlay network (also known as a software-based network, software-defined network, or virtual network) on top of the physical network. The CSPI physical network provides the foundation for creating one or more overlay or virtual networks on top of the physical network. A virtual or overlay network may include one or more virtual cloud networks (VCNs). Virtual networks are implemented using software virtualization technologies (e.g., hypervisors, functions performed by network virtualization devices (NVDs) (e.g., smart NICs), top-of-rack (TOR) switches, smart TORs implementing one or more functions performed by NVDs, and other mechanisms) to create a network abstraction layer that can run on top of the physical network. Virtual networks can take various forms, such as peer-to-peer networks and IP networks. Virtual networks are typically either Layer 3 IP networks or Layer 2 VLANs. This virtual or overlay networking method is often referred to as a virtual or overlay Layer 3 network. Examples of protocols developed for virtual networks include IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN - IETF RFC7348), Virtual Private Networks (VPNs) (e.g., MPLS Layer 3 Virtual Private Network (RFC4364)), VMware's NSX, and GENEVE (Generic Network Virtualization Encapsulation).

[0018] In the case of IaaS, the infrastructure provided by the CSP (CSPI) can be configured to deliver virtualized computing resources over a public network (e.g., the internet). In the IaaS model, the cloud computing service provider can host infrastructure components (e.g., servers, storage, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider can also provide various services associated with these infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering). Therefore, since these services can be policy-driven, IaaS users may be able to implement policies that drive load balancing to maintain application availability and performance. The CSPI provides complementary cloud services to the infrastructure set, enabling customers to build and run a wide range of applications and services in a highly available hosted delivery environment. The CSPI provides high-performance computing resources and capabilities, as well as storage capacity, in a flexible virtual network that can be securely accessed from various network locations, such as the customer's on-premises network. When a customer subscribes to or registers for an IaaS service provided by a CSP, the tenant created for that customer becomes a secure, isolated partition within the CSP where the customer can create, organize, and manage cloud resources.

[0019] Customers can build their own virtual networks using the computing, memory, and networking resources provided by CSPI. They can deploy one or more customer resources or workloads, such as compute instances, into these virtual networks. For example, a customer can use resources provided by CSPI to build one or more customizable private virtual networks called Virtual Cloud Networks (VCNs). A customer can deploy one or more customer resources, such as compute instances, into their VCN. Computing instances can take the form of virtual machines, bare-metal instances, etc. Thus, CSPI provides complementary cloud services as part of an infrastructure set that enables customers to build and run a wide range of applications and services in a highly available virtual host environment. While customers do not manage or control the underlying physical resources provided by CSPI, they can control the operating system, storage, and deployed applications. They may also have limited control over selected network components (e.g., firewalls).

[0020] A CSP can provide a console that allows customers and network administrators to configure, access, and manage resources deployed to the cloud using CSPI resources. In certain embodiments, the console provides a web-based user interface that can be used to access and manage CSPI. In some implementations, the console is a web-based application provided by the CSP.

[0021] CSPI can support single-tenant or multi-tenant architectures. In a single-tenant architecture, software (e.g., applications, databases) or hardware components (e.g., host machines and servers) serve a single customer or tenant. In a multi-tenant architecture, software or hardware components serve multiple customers or tenants. Therefore, in a multi-tenant architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenant environment, precautions are taken to isolate each tenant's data and implement protective measures within CSPI to prevent it from being seen by other tenants.

[0022] In a physical network, a network endpoint ("endpoint") refers to a computing device or system that is connected to a physical network and communicates round-trip with the connected network. Network endpoints within a physical network can be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints within a physical network include modems, hubs, bridges, switches, routers, and other networking devices, as well as physical computers (or host machines). Each physical device within a physical network has a fixed network address that can be used to communicate with the device. This fixed network address may be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints can include various virtual endpoints, such as virtual machines hosted by components of the physical network (e.g., hosted by a physical host machine). These endpoints within a virtual network are addressed by overlay addresses, such as overlay Layer 2 addresses (e.g., overlay MAC addresses) or overlay Layer 3 addresses (e.g., overlay IP addresses). Network overlays provide flexibility by allowing network administrators to move overlay addresses associated with network endpoints using software management (e.g., through software implementing the control plane of the virtual network). Therefore, unlike physical networks, virtual networks allow network management software to move overlay addresses (e.g., overlay IP addresses) from one endpoint to another. Because virtual networks are built on top of physical networks, communication between components within a virtual network involves both the virtual network and the underlying physical network.To facilitate such communication, CSPI components are configured to learn and store mappings between overlay addresses in the virtual network and actual physical addresses in the substrate network, or vice versa. These mappings are then used to facilitate communication. Customer traffic is encapsulated and routed easily within the virtual network.

[0023] Therefore, physical addresses (e.g., physical IP addresses) are associated with components within a physical network, while overlay addresses (e.g., overlay IP addresses) are associated with entities within a virtual network. Both physical and overlay IP addresses are types of real IP addresses. They are distinct from virtual IP addresses, which map to multiple real IP addresses. Virtual IP addresses provide a one-to-many mapping between a virtual IP address and multiple real IP addresses.

[0024] Cloud infrastructure, or CSPI, is physically hosted in one or more data centers located in one or more regions around the world. CSPI may include components within a physical network or underpinning network, and virtualized components within a virtual network built on top of the physical network components (e.g., virtual networks, compute instances, virtual machines, etc.). In certain embodiments, CSPI is organized and hosted within regions, areas, and availability regions. A region is typically a localized geographical area containing one or more data centers. Regions are typically independent of each other and may be very far apart, for example, spanning countries or continents. For example, one region might be Australia, another Japan, and yet another India. CSPI resources are divided across regions such that each region has its own independent subset of CSPI resources. Each region may provide a set of core infrastructure services and resources, such as computing resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure), storage resources (e.g., block volume storage, file storage, object storage, archive storage), networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connectivity to on-premises networks), database resources, edge networking resources (e.g., DNS), and access management and monitoring resources. Typically, each region has multiple paths connecting to other regions within its territory.

[0025] Generally, applications are deployed in the region where they are used most frequently (i.e., on the infrastructure associated with that region) because using nearby resources is faster than using distant resources. Applications can also be deployed in different regions for various reasons, such as redundancy to mitigate the risks of region-wide events like large-scale weather systems or earthquakes, and to meet various requirements such as legal jurisdictions, tax areas, and other business or social standards.

[0026] Data centers within a region can be further organized and subdivided into Availability Areas (ADs). An Availability Area may correspond to one or more data centers within a region. A region can consist of one or more Availability Areas. In such a delivery environment, CSPI resources can be either region-specific, such as virtual cloud networks (VCNs), or Availability Area-specific, such as compute instances.

[0027] ADs within a region are isolated from each other, fault-tolerant, and configured to be extremely unlikely to fail simultaneously. This is achieved by ADs not sharing critical infrastructure resources such as networks, physical cables, cable paths, and cable entry points, so that a failure in one AD within a region is unlikely to affect the availability of other ADs within the same region. ADs within the same region can be interconnected by low-latency, high-bandwidth networks, providing highly available connectivity to other networks (e.g., the internet, customer on-premises networks) and enabling the construction of replicated systems across multiple ADs for both high availability and disaster recovery. Cloud services use multiple ADs to ensure high availability and protect against resource failures. As the infrastructure provided by the IaaS provider expands, additional capacity may be added to further regions and ADs. Typically, traffic between availability regions is encrypted.

[0028] In certain embodiments, regions are grouped into domains. A domain is a logical collection of regions. Domains are isolated from each other, and data is not shared. Regions within the same domain can communicate with each other, but regions in different domains cannot. A customer tenant or account with a CSP resides in a single domain and can be distributed across one or more regions belonging to that domain. Typically, when a customer subscribes to an IaaS service, their tenant or account is created in a customer-designated region within a domain (referred to as the "home" region). A customer can extend their tenant across one or more other regions within a domain. A customer cannot access regions that are not in the domain where their tenant resides.

[0029] IaaS providers can offer multiple domains, each catering to the needs of a specific customer or user group. For example, a commercial domain might be offered to commercial customers. Another example is a domain offered to a specific country for a particular domestic customer. Yet another example is a government domain, which may be offered to a government. For example, a government domain can meet the needs of a specific government and may offer a higher level of security than a commercial domain. For instance, Oracle Cloud Infrastructure (OCI) currently offers two domains: one for commercial regions and another for government cloud regions (e.g., FedRAMP certified and IL5 certified).

[0030] In certain embodiments, an Active Directory (AD) can be subdivided into one or more fault regions. A fault region is a group of infrastructure resources within the AD to provide anti-affinity. Using fault regions, compute instances can be delivered in such a way that instances do not reside on the same physical hardware within a single AD. This is known as anti-affinity. A fault region refers to a set of hardware components (computers, switches, etc.) that share a single point of failure. A compute pool is logically divided into fault regions. Therefore, a hardware failure or compute hardware maintenance event affecting one fault region does not affect instances in other fault regions. Depending on the embodiment, the number of fault regions in each AD may vary. For example, in certain embodiments, each AD contains three fault regions. Fault regions function as logical data centers within the AD.

[0031] When a customer subscribes to an IaaS service, resources from CSPI are provisioned to the customer and associated with the customer's tenant. The customer can use these provisioned resources to build private networks and deploy resources on these networks. Customer networks hosted in the cloud by CSPI are called Virtual Cloud Networks (VCNs). A customer can configure one or more Virtual Cloud Networks (VCNs) using the CSPI resources allocated to them. A VCN is a virtual network or software-defined private network. Customer resources deployed in a customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances may represent various customer workloads such as applications, load balancers, and databases. Compute instances deployed in a VCN can communicate with publicly accessible endpoints ("public endpoints") over public networks such as the internet, with other instances within the same VCN or with other VCNs (e.g., other VCNs of the customer, or VCNs not belonging to the customer), with the customer's on-premises data center or network, with service endpoints, and other types of endpoints.

[0032] A CSP can provide a variety of services using a CSPI. In some cases, the CSPI customer itself may function as a service provider, providing services using CSPI resources. A service provider may expose service endpoints characterized by identifying information (e.g., IP address, DNS name, and port). A customer's resources (e.g., computing instances) can access specific services by accessing the service endpoints exposed by the service of that particular service. These service endpoints are typically publicly accessible to users over public communication networks such as the internet, using the public IP address associated with the endpoint. Publicly accessible network endpoints are sometimes called public endpoints.

[0033] In certain embodiments, a service provider may expose a service through an endpoint (sometimes called a service endpoint) of the service. Customers of the service can then access the service using this service endpoint. In certain implementations, the service endpoint provided to a service may be accessible to multiple customers who intend to use that service. In other implementations, a dedicated service endpoint may be provided to a customer, so that only that customer can access the service using that dedicated service endpoint.

[0034] In certain embodiments, once a VCN is created, it is associated with a Private Overlay Classless Inter-Region Routing (CIDR) address space, which is a range of private overlay IP addresses assigned to the VCN (e.g., 10.0 / 16). The VCN includes associated subnets, route tables, and gateways. A VCN resides within a single region but can span one, more, or all availability regions. A gateway is a virtual interface configured for the VCN that enables traffic communication between the VCN and one or more endpoints outside the VCN. One or more different types of gateways can be configured for the VCN to enable communication with different types of endpoints.

[0035] A VCN can be subdivided into one or more subnets, such as one or more subnets. Therefore, a subnet is a unit or subdivision of configuration that can be created within a VCN. A VCN can contain one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses that do not overlap with other subnets within that VCN (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24), which represent a subset of the address space within the VCN's address space.

[0036] Each compute instance is associated with a virtual network interface card (VNIC) that allows it to join a subnet in the VCN. A VNIC is a logical representation of a physical network interface card (NIC). Generally, a VNIC is the interface between an entity (e.g., compute instance, service) and a virtual network. A VNIC resides within a subnet and has one or more associated IP addresses and associated security rules or policies. A VNIC is equivalent to a Layer 2 port on a switch. The VNIC connects the compute instance to the subnet in the VCN. The VNIC associated with a compute instance allows the compute instance to become part of a subnet in the VCN, enabling the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, endpoints on different subnets within the VCN, or endpoints outside the VCN. Therefore, the VNIC associated with a compute instance determines how the compute instance connects to endpoints inside and outside the VCN. The VNIC for a compute instance is created and associated with that compute instance when the compute instance is created and added to a subnet in the VCN. For a subnet containing a set of compute instances, the subnet includes VNICs corresponding to the set of compute instances, and each VNIC connects to a compute instance within the set of compute instances.

[0037] Each compute instance is assigned a private overlay IP address via the VNIC associated with that compute instance. This private overlay IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic to and from the compute instance. All VNICs within a given subnet use the same route table, security lists, and DHCP options. As mentioned earlier, each subnet within a VCN is associated with a contiguous range of overlay IP addresses that do not overlap with other subnets within that VCN (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) and represents a subset of the address space within the VCN's address space. For a VNIC on a particular subnet of a VCN, the private overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.

[0038] In certain embodiments, a compute instance may optionally be assigned additional overlay IP addresses in addition to its private overlay IP address, such as one or more public IP addresses if it is in a public subnet. These multiple addresses may be assigned on the same VNIC or on multiple VNICs associated with the compute instance. However, each instance has a primary VNIC, which is created when the instance is launched and is associated with the overlay private IP address assigned to the instance. This primary VNIC cannot be deleted. Additional VNICs, called secondary VNICs, can be added to an existing instance in the same availability area as the primary VNIC. All VNICs are in the same availability area as the instance. Secondary VNICs may reside in a subnet within the same VCN as the primary VNIC, or in a different subnet within the same VCN or a different VCN.

[0039] If a compute instance is located in a public subnet, it can optionally be assigned a public IP address. A subnet can be designated as either a public or private subnet when it is created. A private subnet means that resources within the subnet (e.g., compute instances) and associated VNICs cannot have public overlay IP addresses. A public subnet means that resources within the subnet and associated VNICs can have public IP addresses. Customers can specify that a subnet resides in a single availability area or spans multiple availability areas within a region or area.

[0040] As explained above, a VCN can be subdivided into one or more subnets. In certain embodiments, a virtual router (VR) configured for the VCN (referred to as a VCN VR or simply a VR) enables communication between subnets within the VCN. For subnets within a VCN, the VR represents the logical gateway for that subnet, allowing that subnet (i.e., the computing instances on that subnet) to communicate with endpoints on other subnets within the VCN and other endpoints outside the VCN. A VCN VR is a logical entity configured to route traffic between VNICs within the VCN and virtual gateways ("gateways") associated with the VCN. Gateways are further discussed below with reference to Figure 1. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is one VCN VR for a VCN, and the VCN VR has one port for each subnet of the VCN and has a potentially unlimited number of ports addressed by IP addresses. Thus, a VCN VR has a different IP address for each subnet within the VCN to which the VCN VR is connected. The VR is also connected to various gateways configured for the VCN. In certain embodiments, specific overlay IP addresses from a subnet's overlay IP address range are reserved for ports on the VCN VR for that subnet. For example, consider a VCN with two subnets whose associated address ranges are 10.0 / 16 and 10.1 / 16, respectively. For the first subnet in the VCN with address range 10.0 / 16, addresses in this range are reserved for ports on the VCN VR for that subnet. In some cases, the first IP addresses in the range may be reserved for the VCN VR. For example, for a subnet with overlay IP address range 10.0 / 16, the IP address 10.0.0.1 may be reserved for ports on the VCN VR for that subnet. For a second subnet in the same VCN with address range 10.1 / 16, the VCN VR may have a port for the second subnet with IP address 10.1.0.1. The VCN VR has a different IP address for each subnet within the VCN.

[0041] In some other embodiments, each subnet within a VCN may have its own associated VR, addressable by the subnet using a reserved or default IP address associated with the VR. The reserved or default IP address may be, for example, a first IP address from a range of IP addresses associated with that subnet. A VNIC within a subnet can use this default or reserved IP address to communicate with the VR associated with the subnet (e.g., send and receive packets). In such embodiments, the VR is the entry / exit point for that subnet. A VR associated with a subnet within a VCN can communicate with other VRs associated with other subnets within the VCN. A VR can also communicate with gateways associated with the VCN. The VR functionality of a subnet is performed on, or by, one or more NVDs that perform the VNIC functionality of the VNICs within the subnet.

[0042] You can configure route tables, security rules, and DHCP options for a VCN. The route table is the VCN's virtual route table, containing rules that route traffic from subnets within the VCN to destinations outside the VCN, via a gateway or specially configured instance. You can customize the VCN's route table to control how packets are forwarded / routed to and from the VCN. DHCP options refer to configuration information automatically provided when an instance is launched.

[0043] Security rules configured for a VCN represent overlay firewall rules for the VCN. Security rules can include ingress and egress rules, specifying the types of traffic (e.g., based on protocol and port) allowed to enter and exit instances within the VCN. Customers can choose whether specific rules are stateful or stateless. For example, a customer can configure a stateful ingress rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22 to allow incoming SSH traffic to a set of instances from any location. Security rules can be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to resources within that group. A security list, on the other hand, contains rules that apply to all resources within any subnet using the security list. A VCN may be provided with a default security list containing default security rules. DHCP options configured for a VCN provide configuration information that is automatically provided to instances within the VCN when instances are launched.

[0044] In certain embodiments, VCN configuration information is determined and stored by the VCN control plane. VCN configuration information may include, for example, information about the address range associated with the VCN, subnets and associated information within the VCN, one or more VRs associated with the VCN, computing instances and associated VNICs within the VCN, NVDs (e.g., VNICs, VRs, gateways) performing various virtualized network functions associated with the VCN, VCN status information, and other VCN-related information. In certain embodiments, a VCN distribution service exposes the configuration information stored by the VCN control plane, or a portion thereof, to the NVD. Distribution information may be used to update information stored and used by the NVD to forward packets to and from computing instances within the VCN (e.g., forwarding tables, routing tables, etc.).

[0045] In certain embodiments, the creation of VCNs and subnets is handled by the VCN control plane (CP), and the startup of compute instances is handled by the compute control plane. The compute control plane is responsible for allocating physical resources to compute instances and then calling the VCN control plane to create VNICs and connect them to the compute instances. The VCN CP also sends VCN data mappings to the VCN data plane, which is configured to perform packet forwarding and routing functions. In certain embodiments, the VCN CP provides a distribution service, which is responsible for providing updates to the VCN data plane. Examples of VCN control planes are also shown in Figures 12, 13, 14, and 15 (see references 1216, 1316, 1416, and 1516) and are described below.

[0046] Customers can create one or more VCNs using resources hosted by CSPI. Computing instances deployed in a customer's VCN can communicate with different endpoints. These endpoints may include endpoints hosted by CSPI and endpoints outside of CSPI.

[0047] Various different architectures for implementing cloud-based services using CSPI are shown in Figures 1, 2, 3, 4, 5, 12, 13, 14, and 15, and are described below. Figure 1 is a high-level diagram of a delivery environment 100 showing an overlay or customer VCN hosted by CSPI according to a particular embodiment. The delivery environment shown in Figure 1 includes multiple components within the overlay network. The delivery environment 100 shown in Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, substitutions, and modifications are possible. For example, in some implementations, the delivery environment shown in Figure 1 may have more or fewer systems or components than those shown in Figure 1, may combine two or more systems, or the configuration and arrangement of the systems may differ.

[0048] As shown in the example in Figure 1, the distribution environment 100 includes a CSPI 101 that provides services and resources that customers can use to subscribe and build a virtual cloud network (VCN). In a particular embodiment, the CSPI 101 provides IaaS services to the subscriber customer. The data centers within the CSPI 101 may be organized into one or more regions. An example of region "Region US" 102 is shown in Figure 1. The customer has configured a customer VCN 104 for region 102. The customer can deploy various computing instances on the VCN 104, which may include virtual machines or bare metal instances. Examples of instances include applications, databases, load balancers, etc.

[0049] In the embodiment shown in Figure 1, customer VCN104 comprises two subnets, namely "Subnet-1" and "Subnet-2," each subnet having its own CIDRIP address range. In Figure 1, the overlay IP address range for Subnet-1 is 10.0 / 16, and the address range for Subnet-2 is 10.1 / 16. The VCN virtual router 105 represents the logical gateway of the VCN, enabling communication between subnets of VCN104 and with other endpoints outside the VCN. The VCN VR105 is configured to route traffic between VNICs within VCN104 and gateways associated with VCN104. The VCN VR105 provides ports to each subnet of VCN104. For example, the VR105 may provide a port with IP address 10.0.0.1 to Subnet-1 and a port with IP address 10.1.0.1 to Subnet-2.

[0050] Multiple computing instances can be deployed on each subnet, in which case the computing instances may be virtual machine instances and / or bare metal instances. Computing instances within a subnet may be hosted by one or more host machines within CSPI101. Computing instances join the subnet via a VNIC associated with the computing instance. For example, as shown in Figure 1, computing instance C1 is part of Subnet-1 via a VNIC associated with the computing instance. Similarly, computing instance C2 is part of Subnet-1 via a VNIC associated with C2. Similarly, multiple computing instances (which may be virtual machine instances or bare metal instances) may be part of Subnet-1. Each computing instance is assigned a private overlay IP address and MAC address via its associated VNIC. For example, in Figure 1, computing instance C1 has the overlay IP address 10.0.0.2 and MAC address M1, while computing instance C2 has the private overlay IP address 10.0.0.3 and MAC address M2. Each compute instance in Subnet-1 (including compute instances C1 and C2) has a default route to VCN VR105 using the IP address 10.0.0.1, which is the IP address of the port on VCN VR105 in Subnet-1.

[0051] Subnet-2 can deploy multiple computing instances on it, including virtual machine instances and / or bare metal instances. For example, as shown in Figure 1, computing instances D1 and D2 are part of Subnet-2 via VNICs associated with each computing instance. In the embodiment shown in Figure 1, computing instance D1 has the overlay IP address 10.1.0.2 and MAC address MM1, while computing instance D2 has the private overlay IP address 10.1.0.3 and MAC address MM2. Each computing instance in Subnet-2, including computing instances D1 and D2, has a default route to VCN VR105 using IP address 10.1.0.1, which is the IP address of the port of VCN VR105 in Subnet-2.

[0052] VCNA104 can also include one or more load balancers. For example, a load balancer can be provided to a subnet and configured to distribute the traffic load across multiple compute instances on that subnet. A load balancer can also be provided to distribute the traffic load across subnets within a VCN.

[0053] A specific compute instance deployed on VCN104 can communicate with various different endpoints. These endpoints may include endpoints hosted by CSPI200 and endpoints outside of CSPI200. Endpoints hosted by CSPI101 may include endpoints on the same subnet as a particular compute instance (e.g., communication between two compute instances in Subnet-1), endpoints on different subnets within the same VCN (e.g., communication between a compute instance in Subnet-1 and a compute instance in Subnet-2), endpoints in different VCNs within the same region (e.g., communication between a compute instance in Subnet-1 and an endpoint in the same region 106 or 110 of the VCN, or between a compute instance in Subnet-1 and an endpoint in the same region service network 110), or endpoints in VCNs in different regions (e.g., communication between a compute instance in Subnet-1 and an endpoint in the same region 108 of the VCN). A compute instance in a subnet hosted by CSPI101 can also communicate with endpoints not hosted by CSPI101 (i.e., outside of CSPI101). These external endpoints include endpoints within the customer's on-premises network 116, endpoints within other remote cloud host networks 118, public endpoints 114 accessible via public networks such as the internet, and other endpoints.

[0054] Communication between computing instances on the same subnet is facilitated using VNICs associated with the source and destination computing instances. For example, computing instance C1 in Subnet-1 may want to send a packet to computing instance C2, also in Subnet-1. For a packet originating from the source computing instance and destined for another computing instance on the same subnet, the packet is first processed by the VNIC associated with the source computing instance. The processing performed by the VNIC associated with the source computing instance may include determining the packet's destination information from the packet header, identifying the policies (e.g., security lists) configured on the VNIC associated with the source computing instance, determining the packet's next hop, performing packet encapsulation / decapsulation functions as needed, and then forwarding / routing the packet to the next hop to facilitate communication of the packet to its intended destination. If the destination computing instance is on the same subnet as the source computing instance, the VNIC associated with the source computing instance is configured to identify the VNIC associated with the destination computing instance and forward the packet to that VNIC for processing. Next, the VNIC associated with the destination computing instance is activated, and the packet is forwarded to the destination computing instance.

[0055] When packets are communicated from a compute instance within a subnet to an endpoint in a different subnet within the same VCN, the communication is facilitated by the VNICs associated with the source and destination compute instances and VCN VRs. For example, if compute instance C1 in Subnet-1 in Figure 1 wants to send a packet to compute instance D1 in Subnet-2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR105 using the VCN VR's default route or port 10.0.0.1. VCN VR105 is configured to route the packet to Subnet-2 using port 10.1.0.1. The packet is then received and processed by the VNIC associated with D1, and the VNIC forwards the packet to compute instance D1.

[0056] When packets are communicated from a computing instance within VCN104 to an endpoint outside of VCN104, this communication is facilitated by the VNIC associated with the source computing instance, VCN VR105, and the gateway associated with VCN104. One or more types of gateways can be associated with VCN104. A gateway is an interface between the VCN and another endpoint, which is outside the VCN. A gateway is a Layer 3 / IP layer concept that allows the VCN to communicate with endpoints outside the VCN. Thus, gateways facilitate traffic flow between the VCN and other VCNs or networks. Different types of gateways can be configured for the VCN to facilitate different types of communication with different types of endpoints. Depending on the gateway, communication may take place over a public network (e.g., the internet) or over a private network. Various communication protocols can be used for these communications.

[0057] For example, computing instance C1 may want to communicate with an endpoint outside of VCN104. The packet may first be processed by the VNIC associated with source computing instance C1. VNIC processing determines that the packet's destination is outside Subnet-1 of C1. The VNIC associated with C1 can then forward the packet to VCN VR105 of VCN104. Next, VCN VR105 processes the packet and, as part of the processing, determines a specific gateway associated with VCN104 as the next hop for the packet, based on the packet's destination. VCN VR105 can then forward the packet to the specific identified gateway. For example, if the destination is an endpoint within the customer's on-premises network, the packet may be forwarded by VCN VR105 to a Dynamic Routing Gateway (DRG) gateway 122 configured for VCN104. The packet is then forwarded from the gateway to the next hop, facilitating communication of the packet to its final destination.

[0058] Various different types of gateways can be configured for a VCN. An example of a gateway that can be configured for a VCN is shown in Figure 1 and described below. Examples of gateways associated with a VCN are also shown in Figures 12, 13, 14, and 15 (for example, gateways referenced in reference numbers 1234, 1236, 1238, 1334, 1336, 1338, 1434, 1436, 1438, 1534, 1536, and 1538) and described below. As shown in the embodiment shown in Figure 1, a dynamic routing gateway (DRG) 122 can be added to or associated with a customer VCN 104, providing a path for private network traffic communication between the customer VCN 104 and another endpoint, where the other endpoint could be the customer's on-premises network 116, a VCN 108 in a different region of CSPI 101, or another remote cloud network 118 not hosted by CSPI 101. A customer's on-premises network 116 may be a customer network or customer data center built using the customer's resources. Access to the customer's on-premises network 116 is generally very restricted. For a customer who has both their on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, the customer may want to allow the on-premises network 116 and the cloud-based VCNs 104 to communicate with each other. This would allow the customer to build an enhanced hybrid environment that includes the customer's VCNs 104 hosted by CSPI 101 and the on-premises network 116. DRG 122 enables this communication. To enable such communication, a communication channel 124 is configured, with one endpoint of the channel located within the customer's on-premises network 116 and the other endpoint located within CSPI 101 and connected to the customer's VCN 104. The communication channel 124 can be via a public communication network such as the internet or a private communication network.Various different communication protocols may be used, such as IPsec VPN technology over public communication networks like the Internet, or Oracle's FastConnect technology which uses a private network instead of a public network. A device or equipment within the customer's on-premises network 116 that forms one endpoint of communication channel 124 is called customer-premises equipment (CPE), such as CPE126 shown in Figure 1. On the CSPI101 side, the endpoint may be the host machine running DRG122.

[0059] In certain embodiments, a Remote Peering Connection (RPC) can be added to the DRG, which allows a customer to peer one VCN with another VCN in a different region. Using such an RPC, a customer VCN 104 can use the DRG 122 to connect to a VCN 108 in another region. The DRG 122 may also be used to communicate with other remote cloud networks 118, such as the Microsoft Advanced Cloud or the Amazon AWS Cloud, which are not hosted by the CSPI 101.

[0060] As shown in Figure 1, an Internet Gateway (IGW) 120 can be configured for customer VCN 104, allowing computing instances on VCN 104 to communicate with public endpoints 114 accessible via a public network such as the Internet. The IGW 120 is a gateway that connects the VCN to a public network such as the Internet. The IGW 120 allows public subnets within a VCN, such as VCN 104 (where resources within the public subnet have public overlay IP addresses), to directly access public endpoints 112 on a public network 114 such as the Internet. Using the IGW 120, connections can be initiated from subnets within VCN 104 or from the Internet.

[0061] The Network Address Translation (NAT) gateway 128 can be configured for the customer's VCN 104 to allow cloud resources within the customer's VCN that do not have dedicated public overlay IP addresses to access the internet, without exposing those resources, and to direct incoming internet connections (such as L4-L7 connections). This allows private subnets within the VCN, such as private Subnet-1 within VCN 104, to have private access to public endpoints on the internet. The NAT gateway can only initiate connections from private subnets to the public internet, not from the internet to private subnets.

[0062] In certain embodiments, a Service Gateway (SGW) 126 can be configured for a customer VCN 104 and provide a path for private network traffic between VCN 104 and supported service endpoints within a Service Network 110. In certain embodiments, the Service Network 110 is provided by a CSP and can provide a variety of services. An example of such a service network is Oracle's Services Network, which provides a variety of services that customers can use. For example, a compute instance (e.g., a database system) in a private subnet of customer VCN 104 can back up data to a service endpoint (e.g., object storage) without requiring a public IP address or internet access. In certain embodiments, a VCN may have only one SGW, and connections can only be initiated from subnets within the VCN and not from the Service Network 110. If a VCN is peered with another VCN, resources in the other VCN typically cannot access the SGW. Resources in an on-premises network connected to a VCN using FastConnect or VPN Connect can also use a Service Gateway configured for that VCN.

[0063] In certain implementations, SGW126 uses the concept of service classless inter-region routing (CIDR) labels, which are strings representing all regional public IP address ranges for the service or group of services in question. Customers use service CIDR labels when configuring SGW and associated route rules to control traffic to their services. Customers can optionally utilize this when configuring security rules, without having to adjust security rules if the public IP addresses of their services change in the future.

[0064] The Local Peering Gateway (LPG) 132 can be added to the customer VCN 104 and is a gateway that allows the VCN 104 to peer with other VCNs in the same region. Peering means that VCNs communicate using private IP addresses without traffic passing through a public network such as the internet or traffic being routed through the customer's on-premises network 116. In a preferred embodiment, the VCN has a separate LPG for each peering it establishes. Local peering, or VCN peering, is a common method used to establish network connectivity between different application or infrastructure management functions.

[0065] Service providers, such as the service provider in service network 110, can provide access to their services using different access models. According to the public access model, a service may be exposed as a public endpoint accessible publicly by computing instances within the customer VCN via a public network such as the internet, or it may be privately accessible via SGW126. According to a specific private access model, the service becomes accessible as a private IP endpoint within the customer VCN's private subnet. This is called private endpoint (PE) access, and it allows service providers to expose their services as instances within the customer's private network. A private endpoint resource represents a service within the customer's VCN. Each PE appears as a VNIC (called a PE-VNIC with one or more private IPs) within the customer's VCN in a subnet selected by the customer. Thus, a PE provides a way to present a service within a private customer VCN subnet using a VNIC. Because the endpoint is exposed as a VNIC, all the functionality associated with a VNIC, such as routing rules and security lists, becomes available here on the PE VNIC.

[0066] A service provider can register a service to enable access via a Public Access Point (PE). A provider can associate policies with a service to restrict its visibility to customer tenants. A provider can register multiple services under a single virtual IP address (VIP), especially in the case of multi-tenant services. Multiple such private endpoints representing the same service may exist (within multiple VCNs).

[0067] Subsequently, compute instances within the private subnet can access the service using the PE VNIC's private IP address or service DNS name. Compute instances in the customer VCN can access the service by sending traffic to the customer VCN's PE's private IP address. The Private Access Gateway (PAGW) 130 is a gateway resource that can connect to a service provider VCN (e.g., a VCN in service network 110) and act as the entry / exit point for all traffic to and from the customer subnet private endpoint. Using PAGW 130 allows the provider to scale the number of PE connections without utilizing internal IP address resources. The provider only needs to configure one PAGW for any number of services registered in a single VCN. The provider can represent a service as a private endpoint in multiple VCNs of one or more customers. From the customer's perspective, the PE VNIC appears to be connected to the service the customer wishes to interact with, rather than to the customer's instances. Traffic destined for the private endpoint is routed to the service via PAGW 130. These are called customer-to-service (C2S) connections.

[0068] The PE concept can also be used to extend private access to a service to the customer's on-premises network and data center by allowing traffic to flow through FastConnect / IPsec links and private endpoints within the customer's VCN. Private access to the service can also be extended to the customer's peered VCN by allowing traffic to flow between LPG132 and the PE within the customer's VCN.

[0069] Customers can control routing within their VCN at the subnet level, allowing them to specify which subnets within their VCN, such as VCN104, use which gateways. The VCN's route table is used to determine whether to allow traffic from the VCN through a particular gateway. For example, in a specific example, the route table for a public subnet within customer VCN104 might send non-local traffic through IGW120. The route table for a private subnet within the same customer VCN104 might send traffic destined for CSP services through SGW126. All remaining traffic could be sent through NAT gateway 128. The route table only controls traffic leaving the VCN.

[0070] Security lists associated with a VCN are used to control traffic entering the VCN via a gateway through incoming connections. All resources within a subnet use the same route table and security list. Security lists can be used to control specific types of traffic that can enter and leave instances within a VCN subnet. Security list rules can include inbound (received) rules and outbound (transmitted) rules. For example, inbound rules may specify allowed source address ranges, while outbound rules may specify allowed destination address ranges. Security rules can specify specific protocols (e.g., TCP, ICMP), specific ports (e.g., 22 for SSH, 3389 for Windows RDP), etc. In certain implementations, the instance's operating system may enforce its own firewall rules that align with the security list rules. Rules can be stateful (e.g., connections are tracked and responses are automatically allowed without explicit security list rules for response traffic) or stateless.

[0071] Access from a customer VCN (i.e., by resources or computing instances deployed on VCN104) can be classified as public access, private access, or dedicated access. Public access refers to an access model that accesses public endpoints using public IP addresses or NAT. Private access allows customer workloads within VCN104 with private IP addresses (e.g., resources in a private subnet) to access services without going through a public network such as the internet. In certain embodiments, CSPI101 allows customer VCN workloads with private IP addresses to access services (or their public service endpoints) using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer's VCN and the public endpoint of the service located outside the customer's private network.

[0072] Furthermore, CSPI can provide dedicated public access using technologies such as FastConnect public peering, allowing a customer's on-premises instances to access one or more services within the customer's VCN using a FastConnect connection without traversing a public network such as the internet. CSPI can also provide dedicated private access using FastConnect private peering, allowing a customer's on-premises instances with private IP addresses to access the customer's VCN workloads using a FastConnect connection. FastConnect is an alternative network connectivity that uses the public internet to connect a customer's on-premises network to CSPI and its services. Compared to internet-based connectivity, FastConnect provides a simple, resilient, and economical way to create dedicated private connectivity with higher bandwidth options and a more reliable and consistent network experience.

[0073] Figure 1 and the accompanying description above illustrate the various virtualization components in a virtual network example. As described above, the virtual network is built on top of an underlying physical network or infrastructure network. Figure 2 shows a simplified architectural diagram of the physical components within the physical network within the CSPI200 that provides the foundation for the virtual network, according to a particular embodiment. As illustrated, the CSPI200 provides a delivery environment that includes components and resources (e.g., compute, memory, and networking resources) provided by a Cloud Service Provider (CSP). These components and resources are used to deliver cloud services (e.g., IaaS services) to subscriber customers, i.e., customers who subscribe to one or more services provided by the CSP. Based on the services subscribed to by the customer, a subset of the CSPI200's resources (e.g., compute, memory, and networking resources) is provisioned to the customer. The customer can then use the physical compute, memory, and networking resources provided by the CSPI200 to build their own cloud-based (i.e., CSPI-hosted) customizable private virtual network. As previously mentioned, these customer networks are called virtual cloud networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, to these customer VCNs. Compute instances can take the form of virtual machines, bare metal instances, etc. CSPI200 provides complementary cloud services in a set of infrastructure that enable customers to build and run a wide range of applications and services in a highly available hosted environment.

[0074] In the exemplary embodiment shown in Figure 2, the physical components of the CSPI200 include one or more physical host machines or physical servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), and a physical network (e.g., 218), and switches within the physical network 218. The physical host machines or servers can host and run various computing instances participating in one or more subnets of the VCN. Computing instances may include virtual machine instances and bare metal instances. For example, the various computing instances shown in Figure 1 may be hosted by the physical host machines shown in Figure 2. Virtual machine computing instances within the VCN can run on one host machine or several different host machines. The physical host machines can also host virtual host machines, container-based hosts or functions, etc. The VNIC and VCN VR shown in Figure 1 may run on the NVD shown in Figure 2. The gateway shown in Figure 1 may run on the host machines and / or NVD shown in Figure 2.

[0075] A host machine or server can run a hypervisor (also known as a virtual machine monitor or VMM) that creates and enables a virtualized environment on the host machine. Virtualization or a virtualized environment facilitates cloud-based computing. A hypervisor on a host machine allows one or more computing instances to be created, run, and managed on the host machine. A hypervisor on a host machine allows the host machine's physical computing resources (e.g., compute, memory, and networking resources) to be shared among the various computing instances running on the host machine.

[0076] For example, as shown in Figure 2, host machines 202 and 208 run hypervisors 260 and 266, respectively. These hypervisors can be implemented using software, firmware, hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that sits on top of the host machine's operating system (OS) and then runs on the host machine's hardware processor. A hypervisor provides a virtualized environment by allowing the host machine's physical computing resources (e.g., processing resources such as processors / cores, memory resources, and networking resources) to be shared among various virtual machine computing instances running on the host machine. For example, in Figure 2, hypervisor 260 can sit on top of the host machine 202's OS, enabling the host machine 202's computing resources (e.g., processing, memory, and networking resources) to be shared among computing instances (e.g., virtual machines) running on the host machine 202. Virtual machines can have their own operating systems (called guest operating systems), which may be the same as or different from the host machine's OS. The operating system of a virtual machine running on a host machine may be the same as or different from the operating system of another virtual machine running on the same host machine. Therefore, a hypervisor allows multiple operating systems to run in parallel while sharing the same computing resources on the host machine. The host machines shown in Figure 2 may have the same or different types of hypervisors.

[0077] A computing instance can be a virtual machine instance or a bare metal instance. In Figure 2, computing instance 268 on host machine 202 and computing instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.

[0078] In certain examples, an entire host machine may be provisioned to a single customer, and all one or more computing instances (virtual machines or bare metal instances) hosted by that host machine belong to that same customer. In other examples, a host machine may be shared among multiple customers (i.e., multiple tenants). In such multi-tenant scenarios, a host machine may host virtual machine computing instances belonging to different customers. These computing instances may be members of different VCNs of different customers. In certain embodiments, bare metal computing instances are hosted by bare metal servers without a hypervisor. Once a bare metal computing instance is provisioned, a single customer or tenant maintains control of the physical CPU, memory, and network interfaces of the host machine hosting the bare metal instance, and the host machine is not shared with other customers or tenants.

[0079] As mentioned above, each computing instance that is part of a VCN is associated with a VNIC that enables that computing instance to be a member of the VCN's subnet. The VNIC associated with a computing instance facilitates the communication of packets or frames to and from the computing instance. The VNIC is associated with the computing instance when the computing instance is created. In certain embodiments, for a computing instance running on a host machine, the VNIC associated with that computing instance is run on an NVD connected to the host machine. For example, in Figure 2, host machine 202 runs a virtual machine computing instance 268 associated with VNIC 276, and VNIC 276 is run on an NVD 210 connected to host machine 202. In another example, a bare metal instance 272 hosted on host machine 206 is associated with VNIC 280, which is run on an NVD 212 connected to host machine 206. In yet another example, VNIC 284 is associated with computing instance 274 running on host machine 208, and VNIC 284 is run on an NVD 212 connected to host machine 208.

[0080] For computing instances hosted by a host machine, the NVD connected to that host machine also runs the VCN VR corresponding to the VCN to which the computing instance is a member. For example, in the embodiment shown in Figure 2, NVD210 runs VCN VR277 corresponding to the VCN to which computing instance 268 is a member. NVD212 may also run one or more VCN VR283 corresponding to the VCNs to which computing instances hosted by host machines 206 and 208 are associated.

[0081] A host machine may include one or more network interface cards (NICs) that enable it to connect to other devices. The NICs on the host machine may provide one or more ports (or interfaces) that enable the host machine to connect to another device in a communicative manner. For example, the host machine can be connected to the NVD using one or more ports (or interfaces) provided on the host machine and the NVD. The host machine can also connect to other devices, such as another host machine.

[0082] For example, in Figure 2, host machine 202 is connected to NVD210 using a link 220 extending between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD210. Host machine 206 is connected to NVD212 using a link 224 extending between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD212. Host machine 208 is connected to NVD212 using a link 226 extending between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD212.

[0083] The NVDs are connected to top-of-the-rack (TOR) switches via communication links, which in turn connect to the physical network 218 (also known as the switch fabric). In certain embodiments, the links between the host machines and the NVDs, and between the NVDs and the TOR switches, are Ethernet links. For example, in Figure 2, NVDs 210 and 212 are connected to TOR switches 214 and 216, respectively, using links 228 and 230. In certain embodiments, links 220, 224, 226, 228, and 230 are Ethernet links. The collection of host machines and NVDs connected to the TOR is sometimes referred to as a rack.

[0084] The physical network 218 provides a communication fabric that allows TOR switches to communicate with each other. The physical network 218 can be a multi-layer network. In a particular implementation, the physical network 218 is a multi-layer Clos network of switches, and TOR switches 214 and 216 represent leaf-level nodes of the multi-layer and multi-node physical switching network 218. Different Clos network configurations are possible, including but not limited to 2-layer, 3-layer, 4-layer, 5-layer networks, and generally "n"-layer networks. An example of a Clos network is shown in Figure 5 and described below.

[0085] Various connection configurations are possible between the host machine and the NVD, including one-to-one, many-to-one, and one-to-many configurations. In a one-to-one configuration, each host machine is connected to its own separate NVD. For example, in Figure 2, host machine 202 is connected to NVD210 via host machine 202's NIC232. In a many-to-one configuration, multiple host machines are connected to a single NVD. For example, in Figure 2, host machines 206 and 208 are connected to the same NVD212 via NIC244 and 250, respectively.

[0086] In a one-to-many configuration, one host machine is connected to multiple NVDs. Figure 3 shows an example within CSPI300 where a host machine is connected to multiple NVDs. As shown in Figure 3, the host machine 302 has a network interface card (NIC) 304 with multiple ports 306 and 308. The host machine 300 is connected to the first NVD 310 via port 306 and link 320, and to the second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between the host machine 302 and the NVDs 310 and 312 may be Ethernet links. The NVD 310 is connected to the first TOR switch 314, while the NVD 312 is connected to the second TOR switch 316. The links between the NVDs 310 and 312 and the TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent Tier-0 switching devices within a multi-layer physical network 318.

[0087] The configuration shown in Figure 3 provides two separate physical network paths between the physical switch network 318 and the host machine 302. The first path goes through TOR switch 314 and NVD 310 to the host machine 302, and the second path goes through TOR switch 316 and NVD 312 to the host machine 302. These separate paths improve the availability of the host machine 302 (referred to as high availability). If one of the paths (e.g., one link in the path goes down) or a device (e.g., a particular NVD is not functioning) is malfunctioning, the other path can be used to communicate with the host machine 302.

[0088] In the configuration shown in Figure 3, the host machine connects to two different NVDs using two different ports provided by the host machine's NIC. In other embodiments, the host machine may include multiple NICs that enable connections to multiple NVDs of the host machine.

[0089] Referring back to Figure 2, an NVD is a physical device or component that performs one or more network and / or memory virtualization functions. An NVD can be any device with one or more processing units (e.g., a CPU, a network processing unit (NPU), an FPGA, a packet processing pipeline, etc.), memory including a cache, and ports. Various virtualization functions may be performed by software / firmware executed by one or more processing units of the NVD.

[0090] NVDs can be implemented in various different forms. For example, in certain embodiments, an NVD is implemented as an interface card called a smart NIC or an intelligent NIC with an integrated processor. A smart NIC is a separate device from the NIC on the host machine. In Figure 2, NVDs 210 and 212 may be implemented as smart NICs connected to host machines 202, 206, and 208, respectively.

[0091] However, smart NICs are just one example of an NVD implementation. Many other implementations are possible. For example, in some other implementations, the NVD, or one or more functions performed by the NVD, may be incorporated into or performed by one or more host machines, one or more TOR switches, and other components of the CSPI200. For example, the NVD may be implemented within a host machine where the functions performed by the NVD are performed by the host machine. Another example is that the NVD may be part of a TOR switch, or the TOR switch may be configured to perform functions performed by the NVD, enabling the TOR switch to perform various complex packet translations used in public clouds. A TOR that performs the functions of the NVD is sometimes called a smart TOR. In yet another implementation, virtual machine (VM) instances rather than bare metal (BM) instances are provided to the customer, and the functions performed by the NVD may be implemented within the hypervisor of the host machine. In some other implementations, some of the functions of the NVD may be offloaded to a centralized service running on a set of host machines.

[0092] In certain embodiments, for example, when implemented as a smart NIC as shown in Figure 2, the NVD may have multiple physical ports that enable connection to one or more host machines and one or more TOR switches. The ports on the NVD can be classified as host-side ports (also called “southports”) or network-side or TOR-side ports (also called “northports”). The host-side ports of the NVD are the ports used to connect the NVD to a host machine. Examples of host-side ports in Figure 2 include port 236 on the NVD210, and ports 248 and 254 on the NVD212. The network-side ports of the NVD are the ports used to connect the NVD to a TOR switch. Examples of network-side ports in Figure 2 include port 256 on the NVD210 and port 258 on the NVD212. As shown in Figure 2, the NVD210 is connected to the TOR switch 214 using a link 228 that extends from port 256 on the NVD210 to the TOR switch 214. Similarly, the NVD212 is connected to the TOR switch 216 using a link 230 that extends from port 258 of the NVD212 to the TOR switch 216.

[0093] The NVD can receive packets and frames from the host machine via its host-side port (for example, packets and frames generated by computing instances hosted by the host machine), perform the necessary packet processing, and then forward the packets and frames to the TOR switch via the NVD's network-side port. The NVD can also receive packets and frames from the TOR switch via its network-side port, perform the necessary packet processing, and then forward the packets and frames to the host machine via the NVD's host-side port.

[0094] In certain embodiments, multiple ports and associated links may exist between the NVD and the TOR switch. These ports and links can be aggregated to form a link aggregator group (LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and the TOR switch) to be treated as a single logical link. All physical links within a given LAG can operate at the same speed in full-duplex mode. LAGs help improve the bandwidth and reliability of the connection between two endpoints. If one of the physical links in the LAG goes down, traffic is dynamically and transparently reassigned to one of the other physical links in the LAG. Aggregated physical links provide higher bandwidth than individual links. Multiple ports associated with a LAG are treated as a single logical port. Traffic can be load-balanced across the multiple physical links of the LAG. One or more LAGs can be configured between two endpoints. The two endpoints could be between the NVD and the TOR switch, or between a host machine and the NVD, etc.

[0095] NVD implements or performs network virtualization functions. These functions are performed by software / firmware run by NVD. Examples of network virtualization functions include, but are not limited to, packet encapsulation and decapsulation functions, functions for creating VCN networks, functions for implementing network policies such as VCN security list (firewall) functions, and functions for facilitating the routing and forwarding of packets to and from computing instances within the VCN. In certain embodiments, upon receiving a packet, NVD is configured to run a packet processing pipeline to process the packet and determine how the packet should be forwarded or routed. As part of this packet processing pipeline, NVD can run one or more virtual functions associated with the overlay network, such as running a VNIC associated with a cis in the VCN, running a virtual router (VR) associated with the VCN, packet encapsulation and decapsulation to facilitate forwarding or routing within the virtual network, running a specific gateway (e.g., a local peering gateway), implementing security lists, network security groups, network address translation (NAT) functions (e.g., translation from public IP to private IP on a per-host basis), throttling functions, and other functions.

[0096] In certain embodiments, the packet processing data path in an NVD may comprise multiple packet pipelines, each consisting of a set of packet translation stages. In a particular implementation, upon receiving a packet, it is parsed and classified into a single pipeline. The packet is then processed linearly, one step at a time, until it is dropped or sent through the NVD's interface. These stages provide the basic functional components of packet processing (e.g., header verification, throttling enforcement, insertion of new Layer 2 headers, L4 firewall enforcement, VCN encapsulation / decapsulation, etc.), and thus new pipelines can be built by configuring existing stages, and new functionality can be added by creating new stages and inserting them into existing pipelines.

[0097] The NVD can perform both control plane and data plane functions corresponding to the VCN's control plane and data plane. Examples of the VCN control plane are shown in Figures 12, 13, 14, and 15 (see references 1216, 1316, 1416, and 1516) and are described below. Examples of the VCN data plane are shown in Figures 12, 13, 14, and 15 (see references 1218, 1318, 1418, and 1518) and are described below. Control plane functions include functions used to configure the network that controls how data is transferred (e.g., setting routes and route tables, configuring VNICs, etc.). In certain embodiments, a VCN control plane is provided that centrally calculates all overlay-to-board mappings and exposes them to the NVD and various virtual network edge devices such as gateways like DRG, SGW, and IGW. Firewall rules can also be exposed using the same mechanism. In certain embodiments, the NVD retrieves only the mappings associated with that NVD. Data plane functionality includes the ability to actually route / forward packets based on the configuration set using the control plane. The VCN data plane is implemented by encapsulating customer network packets before they pass through the substrate network. The encapsulation / decapsulation functionality is implemented in the NVD. In certain embodiments, the NVD is configured to intercept all network packets inside and outside the host machine and perform network virtualization functions.

[0098] As shown above, NVD performs various virtualization functions, including VNICs and VCN VRs. An NVD can run a VNIC associated with a computing instance hosted by one or more host machines connected to a VNIC. For example, as shown in Figure 2, NVD210 performs the functions of VNIC276 associated with computing instance 268 hosted by host machine 202 connected to NVD210. As another example, NVD212 runs VNIC280 associated with bare-metal computing instance 272 hosted by host machine 206 and runs VNIC284 associated with computing instance 274 hosted by host machine 208. A host machine can host computing instances belonging to different VCNs belonging to different customers, and an NVD connected to a host machine can run the VNIC corresponding to the computing instance (i.e., perform VNIC-related functions).

[0099] An NVD also runs a VCN virtual router corresponding to the VCN of a compute instance. For example, in the embodiment shown in Figure 2, NVD210 runs VCN VR277 corresponding to the VCN to which compute instance 268 belongs. NVD212 runs one or more VCN VR283 corresponding to one or more VCNs to which compute instances hosted by host machines 206 and 208 belong. In a particular embodiment, the VCN VR corresponding to that VCN is run by all NVDs connected to the host machine hosting at least one compute instance belonging to that VCN. If a host machine hosts compute instances belonging to different VCNs, the NVDs connected to that host machine may run VCN VRs corresponding to those different VCNs.

[0100] In addition to VNICs and VCN VRs, NVD may include one or more hardware components that run various software (e.g., daemons) and facilitate various network virtualization functions performed by NVD. For simplicity, these various components are grouped as “packet processing components” as shown in Figure 2. For example, NVD210 has packet processing component 286, and NVD212 has packet processing component 288. For example, a packet processing component of NVD may include a packet processor configured to interact with the NVD’s ports and hardware interfaces to monitor all packets received and communicated using NVD and to store network information. Network information may include, for example, network flow information that identifies different network flows handled by NVD, and per-flow information (e.g., per-flow statistics). In certain embodiments, network flow information may be stored per VNIC. The packet processor can also perform per-packet operations, as well as implement stateful NAT and L4 firewalls (FW). As another example, a packet processing component may include a replication agent configured to replicate information stored by NVD to one or more different replication target stores. As yet another example, a packet processing component may include a logging agent configured to perform the NVD's logging functions. The packet processing component may include software to monitor the NVD's performance and health, and possibly software to monitor the status and health of other components connected to the NVD.

[0101] Figure 1 shows the components of an exemplary virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, a VR for the VCN, and a set of gateways configured for the VCN. The overlay components shown in Figure 1 may run or be hosted by one or more physical components shown in Figure 2. For example, compute instances within a VCN may run or be hosted by one or more host machines shown in Figure 2. For compute instances hosted by host machines, the VNICs associated with those compute instances are typically run by NVDs connected to that host machine (i.e., VNIC functionality is provided by NVDs connected to that host machine). The VCN VR functionality of a VCN is run by all NVDs connected to the host machines hosting or running the compute instances that are part of that VCN. Gateways associated with a VCN may run by one or more different types of NVDs. For example, certain gateways may run by smart NICs, while others may run by one or more host machines or other implementations of NVDs.

[0102] As mentioned earlier, a customer VCN's compute instance can communicate with various different endpoints, which may be in the same subnet as the source compute instance, in a different subnet within the same VCN as the source compute instance, or outside the VCN of the source compute instance. These communications are facilitated using VNICs associated with the compute instance, VCN VRs, and gateways associated with the VCN.

[0103] For communication between two compute instances on the same subnet within a VCN, communication is facilitated using VNICs associated with the source and destination compute instances. The source and destination compute instances may be hosted on the same host machine or on different host machines. Packets originating from the source compute instance may be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. The NVD processes the packets using a packet processing pipeline, which may include running the VNIC associated with the source compute instance. Because the packet's destination endpoint is within the same subnet, running the VNIC associated with the source compute instance forwards the packet to the NVD running the VNIC associated with the destination compute instance, where the packet is processed and then forwarded to the destination compute instance. The VNICs associated with the source and destination compute instances may run on the same NVD (e.g., if both the source and destination compute instances are hosted on the same host machine) or on different NVDs (e.g., if the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNIC can use the routing / forwarding table stored by the NVD to determine the next hop for a packet.

[0104] When packets are communicated from a compute instance within a subnet to an endpoint in a different subnet within the same VCN, packets originating from the source compute instance are communicated from the host machine hosting the source compute instance to the NVD connected to that host machine. In the NVD, packets are processed using a packet processing pipeline. This may include the execution of one or more VNICs and VRs associated with the VCN. For example, as part of the packet processing pipeline, the NVD executes or invokes a function (also called a VNIC execution) corresponding to the VNIC associated with the source compute instance. A function executed by the VNIC may include checking the VLAN tag on the packet. Since the packet's destination is outside the subnet, the VCN VR function is then invoked and executed by the NVD. The VCN VR then routes the packet to the NVD running the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards it to the destination compute instance. The VNICs associated with a source computing instance and a destination computing instance can run on the same NVD (for example, if both the source and destination computing instances are hosted on the same host machine) or on different NVDs (for example, if the source and destination computing instances are hosted on different host machines connected to different NVDs).

[0105] If the packet's destination is outside the source computing instance's VCN, the packet originating from the source computing instance is communicated from the host machine hosting the source computing instance to the NVD connected to that host machine. The NVD runs the VNIC associated with the source computing instance. Because the packet's destination endpoint is outside the VCN, the packet is processed by the VCN's VCN VR. The NVD invokes the VCN VR function, which can forward the packet to an NVD running the appropriate gateway associated with the VCN. For example, if the destination is an endpoint within the customer's on-premises network, the packet may be forwarded by the VCN VR to an NVD running a DRG gateway configured for the VCN. The VCN VR can run on the same NVD running the VNIC associated with the source computing instance, or it can run on a different NVD. The gateway can run on an NVD, which can be a smart NIC, a host machine, or other NVD implementation. The packet is then processed by the gateway and forwarded to the next hop, facilitating the packet's communication to its intended destination endpoint. For example, in the embodiment shown in Figure 2, a packet originating from computing instance 268 may be communicated from host machine 202 to NVD210 via link 220 (using NIC 232). In NVD210, VNIC 276 is invoked because it is the VNIC associated with source computing instance 268. VNIC 276 is configured to examine the encapsulated information in the packet, determine the next hop to forward the packet with the aim of facilitating communication of the packet to its intended destination endpoint, and then forward the packet to the determined next hop.

[0106] Computing instances deployed on a VCN can communicate with various different endpoints. These endpoints may include endpoints hosted by CSPI200 and endpoints outside of CSPI200. Endpoints hosted by CSPI200 may include instances within the same VCN or other VCNs, which may be the customer's VCN or a VCN not belonging to the customer. Communication between endpoints hosted by CSPI200 may be performed over the physical network 218. Computing instances can also communicate with endpoints not hosted by CSPI200, or endpoints located outside of CSPI200. Examples of these endpoints include endpoints within the customer's on-premises network or data center, or public endpoints accessible over a public network such as the Internet. Communication with endpoints outside of CSPI200 may be performed over a public network (e.g., the Internet) (not shown in Figure 2) or a private network (not shown in Figure 2) using various communication protocols.

[0107] The architecture of the CSPI200 shown in Figure 2 is merely an example and is not intended to be limiting. Alternative embodiments are possible, and variations, substitutions, and modifications are possible. For example, in some implementations, the CSPI200 may have more or fewer systems or components than those shown in Figure 2, may combine two or more systems, or may have different configurations or arrangements of systems. The systems, subsystems, and other components shown in Figure 2 may be implemented using hardware or a combination thereof, with software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored in non-temporary storage media (e.g., memory devices).

[0108] Figure 4 shows the connection between a host machine and an NVD to provide I / O virtualization to support multi-tenancy, according to a specific embodiment. As shown in Figure 4, host machine 402 runs a hypervisor 404 that provides the virtualization environment. Host machine 402 runs two virtual machine instances: VM1 406 belonging to customer / tenant #1 and VM2 408 belonging to customer / tenant #2. Host machine 402 has a physical NIC 410 connected to the NVD 412 via link 414. Each computing instance is connected to a VNIC run by the NVD 412. In the embodiment of Figure 4, VM1 406 is connected to VNIC-VM1 420 and VM2 408 is connected to VNIC-VM2 422.

[0109] As shown in Figure 4, NIC410 has two logical NICs: logical NICA416 and logical NICB418. Each virtual machine is configured to connect to and operate on its own logical NIC. For example, VM1 406 is connected to logical NICA416, and VM2 408 is connected to logical NICB418. The host machine 402 has only one physical NIC410 that is shared by multiple tenants, but the logical NICs allow each tenant's virtual machine to perceive that it has its own host machine and NIC.

[0110] In a particular embodiment, each logical NIC is assigned its own VLAN ID. Thus, a specific VLAN ID is assigned to logical NICA416 for tenant #1, and another VLAN ID is assigned to logical NICB418 for tenant #2. When a packet is communicated from VM1 406, the tag assigned to tenant #1 is attached to the packet by the hypervisor, and the packet is communicated from host machine 402 to NVD412 via link 414. Similarly, when a packet is communicated from VM2 408, the tag assigned to tenant #2 is attached to the packet by the hypervisor, and the packet is then communicated from host machine 402 to NVD412 via link 414. Thus, a packet 424 communicated from host machine 402 to NVD412 has an associated tag 426 that identifies a specific tenant and associated VM. On the NVD, for packet 424 received from host machine 402, the tag 426 associated with the packet is used to determine whether the packet should be processed by VNIC-VM1 420 or VNIC-VM2 422. The packet is then processed by the corresponding VNIC. The configuration shown in Figure 4 allows each tenant's computing instance to recognize that it owns its own host machine and NIC. The configuration shown in Figure 4 provides I / O virtualization to support multi-tenancy.

[0111] Figure 5 shows a simplified block diagram of a physical network 500 according to a particular embodiment. The embodiment shown in Figure 5 is structured as a Clos network. A Clos network is a particular type of network topology designed to provide connectivity redundancy while maintaining high binary bandwidth and maximum resource utilization. A Clos network is a type of non-blocking multi-stage or multi-layer switching network, where the number of stages or layers can be 2, 3, 4, 5, etc. The embodiment shown in Figure 5 is a 3-layer network consisting of layers 1, 2, and 3. The TOR switch 504 represents a Tier-0 switch in the Clos network. One or more NVDs are connected to the TOR switch. Tier-0 switches are also called edge devices in the physical network. Tier-0 switches are connected to Tier-1 switches (also called leaf switches). In the embodiment shown in Figure 5, a set of "n" Tier-0 TOR switches is connected to a set of "n" Tier-1 switches, together forming a pod. Each Tier-0 switch within a pod is interconnected with all Tier-1 switches within that pod, but there are no switch connections between pods. In certain implementations, two pods are called a block. Each block is supplied by or connected to a set of n Tier-2 switches (sometimes called spine switches). Multiple blocks can exist in a physical network topology. Tier-2 switches are then connected to n Tier-3 switches (sometimes called superspine switches). Packet communication over the physical network 500 is typically performed using one or more Layer 3 communication protocols. High availability is possible because all layers of the physical network, except the TOR layer, are typically n-way redundant. To enable scaling of the physical network, policies can be specified for pods and blocks to control the mutual visibility of switches within the physical network.

[0112] A key feature of Clos networks is that the maximum number of hops required to travel from one Tier-0 switch to another (or from one NVD connected to a Tier-0 switch to another NVD connected to a Tier-0 switch) is fixed. For example, in a 3-tier Clos network, a packet requires a maximum of 7 hops to travel from one NVD to another. In this case, the source NVD and target NVD are connected to the leaf layer of the Clos network. Similarly, in a 4-tier Clos network, if the source NVD and target NVD are connected to the leaf layer of the Clos network, a packet requires a maximum of 9 hops to travel from one NVD to another. Therefore, the Clos network architecture maintains consistent latency across the entire network, which is crucial for communication within and between data centers. Clos topologies are horizontally scalable and cost-effective. Network bandwidth / throughput capacity can be easily increased by adding more switches to different layers (e.g., increasing the number of leaf and spine switches) and increasing the number of links between switches in adjacent layers.

[0113] In certain embodiments, each resource within the CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource information and can be used, for example, to manage the resource via a console or API. An example of CID syntax is as follows: ocid1.<resource type>.<domain>.[region][.future use].<unique ID> Here, ocid1: A literal string indicating the CID version. Resource type: The type of resource (e.g., instance, volume, VCN, subnet, user, group, etc.) Domain: The domain where the resource resides. Examples of values ​​include "c1" for commercial domains, "c2" for government cloud domains, and "c3" for federal government cloud domains. Each domain may have its own unique domain name. Region: The region where the resource exists. This section may be blank if the region does not apply to the resource. Future Use: Reserved for future use. Unique ID: The unique part of the ID. The format may vary depending on the type of resource or service.

[0114] Figure 6 shows a block diagram of a cloud infrastructure 600 incorporating a CLOS network configuration according to a particular embodiment. The cloud infrastructure 600 includes multiple racks (e.g., rack 1 610… rack M, 620). Each rack includes multiple host machines (also referred to herein as hosts). For example, rack 1 610 includes multiple hosts (e.g., K host machines) from host 1-A, 612 to host 1-K, 614, and rack M includes K host machines, i.e., from host MA, 622 to host MK, 624. It should be understood that the diagram in Figure 6 (i.e., each rack containing the same number of host machines, e.g., K host machines) is intended to be illustrative and not limiting. For example, rack M, 620 may contain more or fewer host machines compared to the number of host machines contained in rack 1 610.

[0115] Each host machine contains multiple graphical processing units (GPUs). For example, as shown in Figure 6, host machine 1-A612 contains N GPUs, e.g., GPU1,613. Furthermore, the diagram in Figure 6 (each host machine containing the same number of GPUs, i.e., N GPUs) is for illustrative purposes only and is not limiting; it should be understood that each host machine can contain a different number of GPUs. Each rack contains a top-of-rack (TOR) switch that is communicatively coupled to the GPUs hosted on the host machines. For example, rack 1610 contains a TOR switch (i.e., TOR1)616 communicatively coupled to hosts 1-A,612 and 1-K,614, and rack M620 contains a TOR switch (i.e., TOR M)626 communicatively coupled to hosts MA,622 and MK,624. The TOR switches shown in Figure 6 (i.e., TOR1 616 and TOR M626) are understood to each contain N ports used to connect the TOR switch to N GPUs hosted on each host machine contained in the rack. As shown in Figure 6, the coupling of the TOR switches to the GPUs is illustrative and not limiting. For example, in some embodiments, the TOR switches may each have multiple ports corresponding to a GPU on each host machine. That is, the GPUs on the host machines may be connected to specific ports of the TOR via communication links. Furthermore, traffic received by a network device (e.g., TOR1 616) is characterized herein as traffic received on a particular receiving port link of the network device. For example, when GPU1 613 on host 1-A612 sends a data packet to TOR1 616 (using link 617), the data packet is received at port 619 of the TOR1 switch. TOR1 616 characterizes this data packet as information received on a first receiving port link of the TOR. It is understood that a similar concept can be applied to all outgoing port links in TOR.

[0116] Each TOR switch from a rack is communicated to multiple spine switches, for example, spine switch 1 630 and spine switch P640. As shown in Figure 6, TOR 1, 616 is connected to spine switch 1 630 via two links and to spine switch P640 via two other links. Information transmitted from a particular TOR switch to a spine switch is referred to herein as communication via uplinks, while information transmitted from a spine switch to a TOR switch is referred to herein as communication via downlinks. According to some embodiments, the TOR switches and spine switches are connected in a CLOS network configuration (e.g., a multi-stage switching network), where each TOR switch forms a "leaf" node in the CLOS network.

[0117] In some embodiments, GPUs within a host machine perform tasks related to machine learning. In such a configuration, a single task may be executed / distributed across a large number of GPUs (e.g., 64), and further distributed across multiple host machines and multiple racks. Since all these GPUs are working on the same task (i.e., workload), they all need to communicate with each other in a time-synchronized manner. Furthermore, at any given time, the GPUs are either in compute mode or communication mode. In other words, the GPUs communicate with each other almost simultaneously. The speed of the workload is determined by the speed of the slowest GPU.

[0118] Equal-cost multipath (ECMP) routing is typically used to route packets from a source GPU (e.g., GPU1 and 613 on host 1-A612) to a destination GPU (e.g., GPU1 and 623 on host M-A622). In ECMP routing, if there are multiple equal-cost paths available for routing traffic from sender to receiver, a selection technique is used to choose a specific path. Thus, a network device receiving the traffic (e.g., a TOR switch or spine switch) uses a selection algorithm to choose the outgoing link to be used to forward the traffic from one network device on the next hop to another. Such outgoing link selection is performed at each network device in the sender-to-receiver path. Hash-based selection algorithms are a widely used ECMP selection technique, and the hash may be based on a 4-tuple of the packet (e.g., source port, destination port, source IP, destination IP).

[0119] ECMP routing is a flow-aware routing technique where each flow (i.e., a stream of data packets) is hashed to a specific path for the duration of the flow. Therefore, packets within a flow are forwarded from the network device using a specific outgoing port / link. This is typically done to ensure that packets within a flow arrive in order; in other words, packet reordering is not required. However, ECMP routing is bandwidth (or throughput) unaware. Thus, TOR switches and spine switches perform statistical flow-aware (throughput-unaware) ECMP load balancing of flows on parallel links.

[0120] Standard ECMP routing (i.e., flow-aware routing only) has a problem where flows received by a network device via two separate receive links are hashed to the same outgoing link, potentially leading to flow collisions. For example, consider a situation where two flows are received via two separate 100G receive links, and each flow is hashed to the same 100G transmit link. In such a situation, congestion (i.e., flow collisions) occurs because the receive bandwidth is 200G while the transmit bandwidth is 100G, resulting in dropped packets. For example, Figure 7, described below, illustrates an exemplary scenario of flow collision 700.

[0121] As shown in Figure 7, there are two flows: Flow 1 710, directed from the first GPU on host machine host 1-A, 612 to TOR switch 616 (shown by a solid line); and Flow 2 720, directed from another GPU on the same host machine 612 to TOR switch 616 (shown by a dashed line). Note that the two flows are directed to TOR switch 616 on separate links, i.e., separate ingress port links of TOR. Assume that all links shown in Figure 7 have a capacity (i.e., bandwidth) of 100G. If TOR switch 616 is running the ECMP routing algorithm, the two flows could be hashed and use the same outgoing port link of TOR, for example, a port of TOR connected to link 730, which connects TOR switch 616 to spine switch 630. In this case, a collision occurs between the two flows (indicated by an "X" mark), and the packet is dropped.

[0122] Such collision scenarios are generally problematic for all types of traffic, regardless of the protocol. For example, TCP is intelligent in that if a packet is dropped and the sender cannot get an acknowledgment for that dropped packet, the packet is retransmitted. However, the situation is even worse with remote direct memory access (RDMA) type traffic. RDMA networks do not use TCP for various reasons (for example, TCP has complex logic that is not suitable for low latency and high performance). RDMA networks use protocols such as RDMA over Infiniband or RDMA over converged Ethernet (RoCE). RoCE has a congestion control algorithm that slows down the transmission rate of packets when the sender identifies the occurrence of a congested or dropped packet. In the case of a dropped packet, not only the dropped packet but also several packets surrounding the dropped packet are retransmitted, further consuming available bandwidth and degrading performance.

[0123] The following describes techniques for overcoming the flow collision problem mentioned above. It is understood that flow collisions affect traffic from both CPUs and GPUs. However, due to stringent time synchronization requirements, flow collisions are a much greater problem for GPUs. Furthermore, it should be understood that, due to the inherent characteristics of routing information that are not aware of statistical bandwidth, flow collision scenarios can occur in standard ECMP routing mechanisms regardless of whether the network is oversubscribed or undersubscribed. A network without oversubscribers is one where the bandwidth of the incoming links to the equipment (e.g., TOR, spine switches) is the same as the bandwidth of the outgoing links. Note that if all links have the same bandwidth capacity, the number of incoming links will be the same as the number of outgoing links.

[0124] According to some embodiments, techniques for overcoming the aforementioned flow collision problem include a GPU-based policy routing mechanism (also referred to herein as a GPU-based traffic engineering mechanism) and a modified ECMP routing mechanism. Each of these techniques is described in more detail below.

[0125] Figure 8 shows a policy-based routing mechanism implemented in the network equipment of the cloud infrastructure shown in Figure 6, according to a specific embodiment. Specifically, the cloud infrastructure 800 includes multiple racks, for example, racks 1 810 through M820. Each rack contains host machines, each containing multiple GPUs. For example, rack 1 810 contains host machine 1-A812, and rack M820 contains host machine 1-A822. Each rack contains a TOR switch. For example, rack 1 810 contains TOR1 switch 814, and rack M820 contains TOR M switch 824. Each host machine in each rack is communicated to the respective TOR switch in the rack. The TOR switches, namely TOR switch 1 814 and TOR switch M824, are then communicated to spine switches, namely spine switch 830 and spine switch 840. For the purpose of illustrating and illustrating policy-based routing, the cloud infrastructure 800 is shown as containing a single host machine per rack. However, it is understood that each rack in the infrastructure may contain multiple host machines.

[0126] According to some embodiments, data packets from sender to receiver are routed within the network on a hop-by-hop basis. Routing policies are configured for each network device that couples an incoming port link to an outgoing port link. Network devices may be TOR switches or spine switches. Referring to Figure 8, two flows are shown: flow 1 from GPU1 on host machine 812, whose intended destination is GPU1 on host machine 822, and flow 2 from GPUN on host machine 812, whose intended destination is GPUN on host machine 822. The network devices, namely TOR1 814, spine1 830, TOR M824, and spineP840, are configured to couple (or match) incoming port links to outgoing port links. The matching of incoming and outgoing port links is maintained at each network device (for example, in a policy table).

[0127] Referring to Figure 8, with respect to flow 1 (i.e., the flow shown by the solid line), it can be observed that when TOR1 814 receives a packet on link 850, TOR is configured to forward the received packet onto the outgoing link 855. Similarly, when spine switch 830 receives a packet via link 855, it is configured to forward the packet on the outgoing link 860. Finally, when TOR M824 receives a packet on link 860, it is configured to forward the packet onto the outgoing link 865 to send the packet to its intended destination, i.e., GPU1 on host machine 822. Similarly, with respect to flow 2 (i.e., the flow shown by the dashed line), when TOR1 814 receives a packet on link 870, TOR1 is configured to forward the received packet to spine P840 on the outgoing link 875. When spine switch 840 receives a packet via link 875, it is configured to forward the packet on the outgoing link 880. Ultimately, when the TOR M824 receives a packet on link 880, it is configured to forward the packet onto the outgoing link 885 to send it to its intended destination, namely GPUN on host machine 822.

[0128] Thus, according to the GPU policy-based routing mechanism, each network device is configured to match incoming ports / links to outgoing ports / links to avoid collisions. For example, considering the first hop flows of Flow 1 and Flow 2, TOR1 814 receives the first data packet (corresponding to Flow 1) on link 850 and the second data packet (corresponding to Flow 2) on link 870. TOR1 814 is configured to forward the data packet received on incoming link / port 850 to outgoing link / port 855 and the data packet received on incoming link / port 870 to outgoing link / port 875, thus ensuring that the first and second data packets do not collide. It is understood that each network device in the cloud infrastructure has a one-to-one correspondence between incoming and outgoing port links, meaning that the mapping between incoming and outgoing port links is performed independently of the flows and / or protocols executed by the flows. Furthermore, in the event of a failure in the outgoing link of a particular network device, according to some embodiments, the network device is configured to switch its routing policy from GPU policy-based routing to standard ECMP routing, obtain a new available outgoing link (from various available outgoing links), and send the flow to that new outgoing link. Note that in this case, flow collisions may occur, leading to congestion.

[0129] Referring here to Figure 9, a block diagram of the cloud infrastructure 900 is shown illustrating different types of connectivity according to a particular embodiment. The infrastructure 900 includes multiple racks, e.g., rack 1 910, rack D920, and rack M930. Racks 910 and 930 include multiple host machines. For example, rack 1 910 includes multiple hosts (e.g., K host machines) host 1-A, 912 through host 1-K, 914, and rack M includes K host machines, i.e., host MA, 932 through host MK, 934. Rack D920 includes one or more host machines 922, each host machine 922 including multiple CPUs, i.e., host machine 922 is a non-GPU host machine. Each of racks 910, 920, and 930 includes TOR switches, i.e., TOR1, 916, TOR D926, and TOR M936, respectively, that are communicably connected to the host machines in their respective racks. Furthermore, TOR switches 916, 926, and 936 are communicatively coupled to multiple spine switches, namely spine switches 940 and 950.

[0130] As shown in Figure 9, the first connection (i.e., connection 1, indicated by the dashed line) exists from one GPU host (i.e., host 1-A912) to another GPU host (i.e., host M-A932), and the second connection (i.e., connection 2, indicated by the dotted line) exists from one GPU host (i.e., host 1-K914) to a non-GPU host (i.e., host D922). With respect to connection 1, data packets associated with the flow are routed on a hop-by-hop basis by configuring each intermediate network device based on the GPU-based policy routing mechanism described in Figure 8. Specifically, data packets from connection 1 are routed along the dashed link shown in Figure 9, i.e., from host 1-A to TOR1, from TOR1 to spine switch 1, from spine switch 1 to TOR M, and finally from TOR M to host MA. Each network device is configured to connect its receiving link port to its outgoing link port.

[0131] In contrast, connection 2 is from a GPU-based host (i.e., host 1-K914) to a non-GPU-based host (i.e., host D922). In this case, TOR1 916 is configured to connect the incoming link port, i.e., the port and link (971) that receives data packets from the GPU-based host, to the outgoing link port, i.e., the output port and the link 972 connected to that output port. Thus, TOR1 uses the outgoing link 972 to forward the data packets to spine switch 1 940. Since the data packets are destined for a non-GPU-based host, spine switch 1 940 does not use a policy-based routing mechanism to forward the packets to TOR D. Rather, spine switch 1 940 uses an ECMP routing mechanism to select one of the available links 980 to forward the packets to TOR D, which then forwards the data packets to host D922.

[0132] In some embodiments, the flow collisions described above with reference to Figure 7 are avoided by the network devices of the cloud infrastructure described herein by implementing a modified version of ECMP routing. In this case, the ECMP hash algorithm is modified so that traffic passing through a particular receiving port link of the network device is hashed (and directed) to the same outgoing port link of the network device. Specifically, each network device implements the modified ECMP algorithm to determine the outgoing port link used to forward packets, and the modified ECMP algorithm ensures that packets received through a particular receiving port link are always hashed to the same outgoing port link. For example, consider the case where the first and second packets are received on the first receiving port link of the network device, while the third and fourth packets are received on the second receiving port link of the network device. In this scenario, the network device is configured to implement a modified ECMP algorithm, and the network device transmits the first and second packets on its first outgoing port link, the third and fourth packets on its second outgoing port link, the first receiving port link is different from the second receiving port link, and the first outgoing port link is different from the second outgoing port link. In a particular implementation, the information stored in the forwarding information database (e.g., forwarding table, ECMP table, etc.) may be modified to enable the above functionality.

[0133] Referring now to Figure 10, an exemplary configuration of rack 1000 according to a particular embodiment is shown. As shown in Figure 10, rack 1000 includes two host machines, namely host machine 1010 and host machine 1020. Although rack 1000 is shown to include two host machines, it is understood that rack 1000 may include more host machines. Each host machine includes multiple GPUs and multiple CPUs. For example, host machine 1010 includes multiple CPUs 1012 and multiple GPUs 1014, while host machine 1020 includes multiple CPUs 1022 and multiple GPUs 1024.

[0134] The host machines are communicably connected to different network fabrics via different TOR switches. For example, host machines 1010 and 1020 are communicably coupled to a network fabric (referred to herein as the front-end network of rack 1000) via a TOR switch, namely TOR1 switch 1050. The front-end network may correspond to an external network. More specifically, host machine 1010 is connected to the front-end network via a network interface card (NIC) 1030 and a network virtualization device (NVD) 1035 coupled to TOR1 switch 1050. Host machine 1020 is connected to the front-end network via NIC 1040 and an NVD 1045 coupled to TOR1 switch 1050. Thus, according to some embodiments, the CPU of each host machine can communicate with the front-end network via the NIC, NVD, and TOR switch. For example, the CPU 1012 of host machine 1010 can communicate with the front-end network via NIC 1030, NVD 1035, and TOR1 switch 1050.

[0135] Host machines 1010 and 1020 are connected to the other Quality of Service (QoS) enabled backend network. The QoS-enabled backend network is referred to herein as the backend network corresponding to the GPU cluster network shown in Figure 6. Host machine 1010 is connected to a TOR2 switch 1060 via another NIC 1065, which connects host machine 1010 to the backend network. Similarly, host machine 1020 is connected to the TOR2 switch 1060 via NIC 1080, and the TOR2 switch 1060 connects host machine 1020 to the backend network. Thus, multiple GPUs on each host machine can communicate with the backend network via NICs and TOR switches. In this way, multiple CPUs utilize separate sets of NICs (compared to the set of NICs used by the GPUs) to communicate with the frontend network and the backend network, respectively.

[0136] Figure 11A shows an exemplary flowchart 1100 illustrating the steps performed by a network device when routing packets according to a particular embodiment. The processes shown in Figure 11A may be implemented by software (e.g., code, instructions, programs), hardware, or a combination thereof, performed by one or more processing units (e.g., processors, cores) of each system. The software may be stored in a non-temporary storage medium (e.g., a memory device). The methods shown in Figure 11A and described below are exemplary and not limiting. Figure 11A shows, but is not limited to, various processing steps performed in a particular flow or order. In a particular alternative embodiment, the steps may be performed in some different order, and some steps may be performed in parallel.

[0137] The process begins in step 1105, when the network device receives data packets transmitted by the host machine's graphical processing unit (GPU). In step 1110, the network device determines the receiving port / link from which the packets were received. In step 1115, the network device identifies the outgoing port / link corresponding to the receiving port / link (from which the packets were received) based on policy routing information. According to some embodiments, the policy routing information corresponds to a pre-configured GPU routing table for the network device that links each input port link of the network device to the network device's unique outgoing link port.

[0138] Next, the process proceeds to step 1120, where a query is executed to determine whether the outgoing port link is functional, for example, whether the outgoing link is active. If the response to the query is positive (i.e., the link is active), the process proceeds to step 1125; on the other hand, if the response to the query is negative (i.e., the link is in a failed / inactive state), the process proceeds to step 1130. In step 1125, the network device uses the outgoing port link (identified in step 1115) to forward the received data packet to another network device. In step 1130, the network device obtains flow information for the data packet. For example, the flow information may correspond to a 4-tuple associated with the packet (i.e., source port, destination port, source IP address, destination IP address). Based on the obtained flow information, the network device uses ECMP routing to identify a new outgoing port link, i.e., an available outgoing port link. Next, the process proceeds to step 1135, where the network device uses the newly acquired outgoing port link to forward the data packet received in step 1105.

[0139] Figure 11B shows another exemplary flowchart 1150 illustrating the steps performed by a network device when routing packets according to a particular embodiment. The processes shown in Figure 11B may be implemented by software (e.g., code, instructions, programs), hardware, or a combination thereof, performed by one or more processing units (e.g., processors, cores) of each system. The software may be stored in a non-temporary storage medium (e.g., a memory device). The methods shown in Figure 11B and described below are illustrative and not limiting. Figure 11B shows, but is not limited to, various processing steps performed in a particular flow or order. In a particular alternative embodiment, the steps may be performed in some different order, and some steps may be performed in parallel.

[0140] The process begins in step 1155, when the network device receives data packets sent by the host machine's graphical processing unit (GPU). In step 1160, the network device determines the flow information of the received packet. In some implementations, the flow information may correspond to a four-part tuple associated with the packet (i.e., source port, destination port, source IP address, destination IP address). In step 1165, the network device computes the outgoing port link by implementing a modified version of ECMP routing. The modified ECMP algorithm ensures that packets received on a particular receiving port link are always hashed and sent on the same outgoing port link.

[0141] Next, the process proceeds to step 1170, where a query is executed to determine whether the outgoing port link (determined in step 1165) is functional, for example, whether the outgoing link is active. If the response to the query is positive (i.e., the link is active), the process proceeds to step 1175; on the other hand, if the response to the query is negative (i.e., the link is in a failed / inactive state), the process proceeds to step 1180. In step 1175, the network device uses the outgoing port link (identified in step 1165) to forward the received data packet to another network device. If it is determined that the identified outgoing port link (in step 1165) is in an inactive state, the process proceeds to step 1180. In step 1180, the network device implements ECMP routing (i.e., standard ECMP routing) to identify a new outgoing port link. The process then proceeds to step 1185, where the network device uses the newly calculated outgoing port link to forward the data packet received in step 1155.

[0142] Note that the above technique for routing data packets originating from the host machine's GPU can increase throughput by 20% in small clusters and 70% in larger clusters (i.e., a 3x improvement compared to the standard ECMP routing algorithm).

[0143] Examples of cloud infrastructure As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide computing resources that are virtualized over a public network (such as the internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer)). In some cases, the IaaS provider can also provide various services that accompany these infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering). Therefore, since these services can be policy-driven, IaaS users may be able to implement policies that drive load balancing to maintain application availability and performance.

[0144] In some cases, IaaS customers can access resources and services over a wide area network (WAN), such as the internet, and install the rest of their application stack using the cloud provider's services. For example, a user can log into an IaaS platform and create virtual machines (VMs), install operating systems (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform various functions such as balancing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.

[0145] In most cases, the cloud computing model requires the participation of a cloud provider. This cloud provider does not have to be a third-party service specializing in IaaS provision (e.g., providing, renting, or selling). Entities can also choose to deploy a private cloud and become their own infrastructure service provider.

[0146] In some cases, IaaS deployment is the process of deploying a new application, or a new version of an application, to a prepared application server, etc. This may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Therefore, the customer may be responsible for handling things like the OS, middleware, and / or application deployment (e.g., self-service virtual machines, e.g., that can be spun up on demand).

[0147] In some examples, IaaS provisioning can even refer to acquiring the computers or virtual hosts to be used and installing the necessary libraries or services on them. In most cases, provisioning is not included in the deployment, so it may be necessary to perform provisioning first.

[0148] In some cases, IaaS provisioning presents two distinct challenges. First, there's the initial challenge of provisioning an initial set of infrastructure before doing anything. Second, there's the challenge of evolving existing infrastructure after everything has been provisioned (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by declaratively defining the configuration of the infrastructure. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which resources and how they all work together) can be described declaratively. In some cases, once the topology is defined, a workflow can be generated to create and / or manage the various components described in the configuration files.

[0149] In some examples, infrastructure can have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs), also known as core networks (e.g., a potentially on-demand pool of configurable and / or shared computing resources). In some examples, there may also be one or more security group rules provisioned to define how network security is configured, and one or more virtual machines (VMs). Other infrastructure elements such as load balancers and databases can also be provisioned. Infrastructure can evolve incrementally as more infrastructure elements are desired or added.

[0150] In some cases, continuous deployment techniques may be used to enable the deployment of infrastructure code across various virtual computing environments. Furthermore, the techniques described enable infrastructure management within these environments. In some examples, a service team may write code that they wish to deploy to one or more, but often many, different production environments (e.g., across various different geographical locations, and possibly worldwide). However, in some examples, the infrastructure to which the code will be deployed must first be set up. In some cases, provisioning can be done manually, resources can be provisioned using provisioning tools, and / or code can be deployed using deployment tools after the infrastructure has been provisioned.

[0151] Figure 12 is a block diagram 1200 showing an example pattern of an IaaS architecture according to at least one embodiment. The service operator 1202 can be communicatively coupled to a secure host tenant 1204 which may include a virtual cloud network (VCN) 1206 and a secure host subnet 1208. In some examples, the service operator 1202 may use one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google Glass® head-mounted displays), running software such as Microsoft Windows Mobile®, and / or various mobile operating systems such as iOS, WindowsPhone, Android, BlackBerry 8, PalmOS, as well as the internet, email, short message service (SMS), Blackberry®, or other valid communication protocols. Alternatively, the client computing device may be a general-purpose personal computer, including, for example, personal computers and / or laptop computers running various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. The client computing device may be a workstation computer running one of a variety of commercially available UNIX® or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS.Alternatively, or in addition, the client computing device may be any other electronic device, such as a thin client computer, an internet-enabled game system (e.g., a Microsoft Xbox game console with or without Kinect® gesture input), and / or a personal messaging device that can communicate via a network that has access to the VCN1206 and / or the Internet.

[0152] VCN1206 may include a local peering gateway (LPG) 1210 that can communicately connect to Secure Shell (SSH) VCN1212 via LPG 1210 included in SSHVCN1212. SSHVCN1212 may include an SSH subnet 1214, and SSHVCN1212 can communicately connect to control plane VCN1216 via LPG 1210 included in control plane VCN1216. Furthermore, SSHVCN1212 can communicately connect to data plane VCN1218 via LPG 1210. Control plane VCN1216 and data plane VCN1218 may be included in a service tenant 1219 owned and / or operated by an IaaS provider.

[0153] The control plane VCN 1216 may include a control plane demilitarized zone (DMZ) layer 1220 that functions as a perimeter network (e.g., part of the corporate network between the corporate intranet and the external network). DMZ-based servers have limited liability and can help deter breaches. Furthermore, the DMZ layer 1220 may include a control plane application layer 1224 that may include one or more load balancer (LB) subnets 1222, an application subnet 1226, and a control plane data layer 1228 which may include a database (DB) subnet 1230 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 1222 included in the control plane DMZ layer 1220 can be communicatively coupled to the application subnet 1226 included in the control plane application layer 1224 and to the Internet gateway 1234 which may be included in the control plane VCN 1216. The application subnet 1226 can be communicatively coupled to the DB subnet 1230 included in the control plane data layer 1228, as well as to the service gateway 1236 and the network address translation (NAT) gateway 1238. The control plane VCN 1216 may include the service gateway 1236 and the NAT gateway 1238.

[0154] The control plane VCN 1216 may include a data plane mirror application layer 1240 which may include an application subnet 1226. The application subnet 1226 included in the data plane mirror application layer 1240 may include a virtual network interface controller (VNIC) 1242 which can run a compute instance 1244. The compute instance 1244 can communicatively connect the application subnet 1226 of the data plane mirror application layer 1240 to an application subnet 1226 which may be included in the data plane application layer 1246.

[0155] The data plane VCN1218 may include a data plane application layer 1246, a data plane DMZ layer 1248, and a data plane data layer 1250. The data plane DMZ layer 1248 may include an LB subnet 1222 that can be communicatively coupled to the application subnet 1226 of the data plane application layer 1246 and the internet gateway 1234 of the data plane VCN1218. The application subnet 1226 may be communicatively coupled to the service gateway 1236 of the data plane VCN1218 and the NAT gateway 1238 of the data plane VCN1218. The data plane data layer 1250 may also include a DB subnet 1230 that can be communicatively coupled to the application subnet 1226 of the data plane application layer 1246.

[0156] The Internet gateway 1234 of the control plane VCN1216 and data plane VCN1218 can be communicatively coupled to a metadata management service 1252, which can be communicatively coupled to the public internet 1254. The public internet 1254 can be communicatively connected to the NAT gateway 1238 of the control plane VCN1216 and data plane VCN1218. The service gateway 1236 of the control plane VCN1216 and data plane VCN1218 can be communicatively coupled to a cloud service 1256.

[0157] In some examples, a service gateway 1236 of the control plane VCN1216 or data plane VCN1218 can make application programming interface (API) calls to a cloud service 1256 without going through the public internet 1254. API calls from service gateway 1236 to cloud service 1256 can be one-way: service gateway 1236 can make API calls to cloud service 1256, and cloud service 1256 can send the requested data to service gateway 1236. However, cloud service 1256 may not be able to initiate an API call to service gateway 1236.

[0158] In some examples, a secure host tenant 1204 can connect directly to a service tenant 1219, or it may be isolated otherwise. A secure host subnet 1208 can communicate with an SSH subnet 1214 via an LPG 1210, which can enable bidirectional communication through systems that would otherwise be isolated. Connecting the secure host subnet 1208 to the SSH subnet 1214 can give the secure host subnet 1208 access to other entities within the service tenant 1219.

[0159] The control plane VCN1216 allows users of service tenant 1219 to set up or provision desired resources. Desired resources provisioned within the control plane VCN1216 can be deployed or used within the data plane VCN1218. In some examples, the control plane VCN1216 can be separated from the data plane VCN1218, and the data plane mirror application layer 1240 of the control plane VCN1216 can communicate with the data plane application layer 1246 of the data plane VCN1218 via a VNIC 1242, which can be included in the data plane mirror application layer 1240 and the data plane application layer 1246.

[0160] In some examples, a system user or customer may make requests, such as create, read, update, or delete (CRUD) operations, via the public internet 1254, which can communicate the requests to the metadata management service 1252. The metadata management service 1252 can communicate the requests to the control plane VCN 1216 via the internet gateway 1234. This request may be received by the LB subnet 1222, which is included in the control plane DMZ layer 1220. The LB subnet 1222 may determine that the request is valid, and in response to this determination, the LB subnet 1222 may send the request to the application subnet 1226, which is included in the control plane application layer 1224. If the request is validated and a call to the public internet 1254 is required, the call to the public internet 1254 may be sent to the NAT gateway 1238, which can make calls to the public internet 1254. Memory that may be desirable to be stored by the request can be stored in the DB subnet 1230.

[0161] In some cases, the data plane mirror application layer 1240 can facilitate direct communication between the control plane VCN 1216 and the data plane VCN 1218. For example, it may be desirable to apply configuration changes, updates, or other appropriate modifications to resources contained in the data plane VCN 1218. Through VNIC 1242, the control plane VCN 1216 can communicate directly with the resources contained in the data plane VCN 1218, thereby enabling it to perform configuration changes, updates, or other appropriate modifications to the resources contained in the data plane VCN 1218.

[0162] In some embodiments, the control plane VCN1216 and data plane VCN1218 can be included in the service tenant 1219. In this case, the system's user or customer cannot own or operate either the control plane VCN1216 or the data plane VCN1218. Instead, the IaaS provider can own or operate the control plane VCN1216 and the data plane VCN1218, both of which can be included in the service tenant 1219. This embodiment can enable network isolation that can prevent a user or customer from interacting with resources of other users or other customers. This embodiment also allows the system's user or customer to store databases privately without having to rely on the public internet 1254, which may not have the desired level of security for storage.

[0163] In another embodiment, the LB subnet 1222 included in the control plane VCN 1216 may be configured to receive signals from the service gateway 1236. In this embodiment, the control plane VCN 1216 and the data plane VCN 1218 may be configured to be invoked by the IaaS provider's customers without calling the public internet 1254. The IaaS provider's customers may prefer this embodiment because the database used by the customer may be controlled by the IaaS provider and stored in a service tenant 1219 which can be isolated from the public internet 1254.

[0164] Figure 13 is a block diagram 1300 illustrating another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1302 (e.g., service operator 1202 in Figure 12) can be communicatively coupled to a secure host tenant 1304 (e.g., secure host tenant 1204 in Figure 12), which may include a virtual cloud network (VCN) 1306 (e.g., VCN1206 in Figure 12) and a secure host subnet 1308 (e.g., secure host subnet 1208 in Figure 12). The VCN 1306 may include a local peering gateway (LPG) 1310 (e.g., LPG1210 in Figure 12), which may be communicatively coupled to a secure shell (SSH) VCN 1312 (e.g., SSHVCN1212 in Figure 12) via LPG1210 contained within the SSHVCN 1312. SSHVCN1312 may include SSH subnet 1314 (e.g., SSH subnet 1214 in Figure 12), and SSHVCN1312 may be communicably coupled to control plane VCN1316 (e.g., control plane VCN1216 in Figure 12) via LPG1310 included in control plane VCN1316. Control plane VCN1316 may include service tenant 1319 (e.g., service tenant 1219 in Figure 12), and data plane VCN1318 (e.g., data plane VCN1218 in Figure 12) may include customer tenant 1321, which may be owned or operated by a user or customer of the system.

[0165] The control plane VCN1316 may include a control plane DMZ layer 1320 (e.g., control plane DMZ layer 1220 in Figure 12) which may include an LB subnet 1322 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1324 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1326 (e.g., application subnet 1226 in Figure 12), and a control plane data layer 1328 (e.g., control plane data layer 1228 in Figure 12) which may include a database (DB) subnet 1330 (e.g., similar to DB subnet 1230 in Figure 12). The LB subnet 1322 included in the control plane DMZ layer 1320 can be communicatively coupled to the application subnet 1326 included in the control plane application layer 1324 and to an internet gateway 1334 (e.g., internet gateway 1234 in Figure 12) which may be included in the control plane VCN 1316. The application subnet 1326 can be communicatively coupled to the DB subnet 1330 included in the control plane data layer 1328, as well as to a service gateway 1336 (e.g., service gateway in Figure 12) and a network address translation (NAT) gateway 1338 (e.g., NAT gateway 1238 in Figure 12). The control plane VCN 1316 may include the service gateway 1336 and the NAT gateway 1338.

[0166] The control plane VCN 1316 may include a data plane mirror application layer 1340 (e.g., data plane mirror application layer 1240 in Figure 12) which may include an application subnet 1326. The application subnet 1326 included in the data plane mirror application layer 1340 may include a virtual network interface controller (VNIC) 1342 (e.g., VNIC 1242) which may run a compute instance 1344 (e.g., similar to compute instance 1244 in Figure 12). The compute instance 1344 can facilitate communication between the application subnet 1326 of the data plane mirror application layer 1340 and the application subnet 1326, which may be included in the data plane application layer 1346 (e.g., data plane application layer 1246 in Figure 12) via the VNIC 1342 included in the data plane mirror application layer 1340 and the VNIC 1342 included in the data plane application layer 1346.

[0167] The Internet gateway 1334 included in the control plane VCN 1316 can be communicatively coupled to the metadata management service 1352 (for example, the metadata management service 1252 in Figure 12), which can be communicatively coupled to the public internet 1354 (for example, the public internet 1254 in Figure 12). The public internet 1354 can be communicatively coupled to the NAT gateway 1338 included in the control plane VCN 1316. The service gateway 1336 included in the control plane VCN 1316 can be communicatively coupled to the cloud service 1356 (for example, the cloud service 1256 in Figure 12).

[0168] In some examples, the data plane VCN1318 may be contained within the customer tenant 1321. In this case, the IaaS provider may provide a control plane VCN1316 for each customer, and the IaaS provider may set up a unique compute instance 1344 contained within the service tenant 1319 for each customer. Each compute instance 1344 may enable communication between the control plane VCN1316 contained within the service tenant 1319 and the data plane VCN1318 contained within the customer tenant 1321. The compute instance 1344 may enable resources provisioned within the control plane VCN1316 contained within the service tenant 1319 to be deployed or otherwise used within the data plane VCN1318 contained within the customer tenant 1321.

[0169] In another example, an IaaS provider's customer might have a database residing within customer tenant 1321. In this example, control plane VCN 1316 could include a data plane mirror app layer 1340 that could include app subnet 1326. The data plane mirror app layer 1340 could reside within data plane VCN 1318, but it does not have to reside within data plane VCN 1318. That is, the data plane mirror app layer 1340 can access customer tenant 1321, but it does not have to reside within data plane VCN 1318 and may be owned or operated by the IaaS provider's customer. The data plane mirror app layer 1340 may be configured to make calls to data plane VCN 1318, but it does not have to be configured to make calls to any entity contained within control plane VCN 1316. Customers may wish to deploy or otherwise use resources in the data plane VCN1318 provisioned within the control plane VCN1316, and the data plane mirror application layer 1340 can facilitate the customer's desired deployment or other use of resources.

[0170] In some embodiments, a customer of the IaaS provider can apply filters to the data plane VCN1318. In this embodiment, the customer can determine what the data plane VCN1318 can access and can restrict access from the data plane VCN1318 to the public internet 1354. The IaaS provider may not be able to apply filters or control the data plane VCN1318's access to external networks or databases. Applying customer filters and controls to the data plane VCN1318 contained in a customer tenant 1321 can help isolate the data plane VCN1318 from other customers and the public internet 1354.

[0171] In some embodiments, the cloud service 1356 can be invoked by the service gateway 1336 to access services that may not reside on the public internet 1354, the control plane VCN 1316, or the data plane VCN 1318. The connection between the cloud service 1356 and the control plane VCN 1316 or the data plane VCN 1318 may not be live or continuous. The cloud service 1356 may reside on another network owned or operated by the IaaS provider. The cloud service 1356 may be configured to receive calls from the service gateway 1336, or it may be configured not to receive calls from the public internet 1354. Some cloud services 1356 may be isolated from other cloud services 1356, and the control plane VCN 1316 may be isolated from cloud services 1356 that do not have to be in the same region as the control plane VCN 1316. For example, the control plane VCN 1316 may be located in "Region 1", and the cloud service "Deployment 7" may be located in Region 1 and "Region 2". If a call to deployment 7 is made by a service gateway 1336 included in the control plane VCN 1316 in region 1, that call may be sent to deployment 7 in region 1. In this example, the control plane VCN 1316, or deployment 7 in region 1, may not be communicatively coupled to or communicating with deployment 7 in region 2.

[0172] Figure 14 is a block diagram 1400 showing another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1402 (e.g., service operator 1202 in Figure 12) can be communicatively coupled to a secure host tenant 1404 (e.g., secure host tenant 1204 in Figure 12), which may include a virtual cloud network (VCN) 1406 (e.g., VCN1206 in Figure 12) and a secure host subnet 1408 (e.g., secure host subnet 1208 in Figure 12). VCN 1406 may include an LPG 1410 (e.g., LPG1210 in Figure 12) which can be communicatively coupled to SSHVCN 1412 (e.g., SSHVCN1212 in Figure 12) via an LPG 1410 contained within SSHVCN 1412. SSHVCN1412 may include SSH subnet 1414 (e.g., SSH subnet 1214 in Figure 12), and SSHVCN1412 may be communicatively coupled to control plane VCN1416 (e.g., control plane VCN1216 in Figure 12) via LPG1410 included in control plane VCN1416, and may be communicatively coupled to data plane VCN1418 (e.g., data plane 1218 in Figure 12) via LPG1410 included in data plane VCN1418. Control plane VCN1416 and data plane VCN1418 may be included in service tenant 1419 (e.g., service tenant 1219 in Figure 12).

[0173] The control plane VCN1416 may include a control plane DMZ layer 1420 (e.g., control plane DMZ layer 1220 in Figure 12) which may include a load balancer (LB) subnet 1422 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1424 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1426 (e.g., similar to application subnet 1226 in Figure 12), and a control plane data layer 1428 (e.g., control plane data layer 1228 in Figure 12) which may include a DB subnet 1430. The LB subnet 1422 included in the control plane DMZ layer 1420 may be communicatively coupled to the application subnet 1426 included in the control plane application layer 1424 and to an internet gateway 1434 (e.g., internet gateway 1234 in Figure 12) which may be included in the control plane VCN 1416. The application subnet 1426 may be communicatively coupled to the DB subnet 1430, service gateway 1436 (e.g., service gateway in Figure 12), and network address translation (NAT) gateway 1438 (e.g., NAT gateway 1238 in Figure 12) included in the control plane data layer 1428. The control plane VCN 1416 may include the service gateway 1436 and the NAT gateway 1438.

[0174] The data plane VCN1418 may include the data plane application layer 1446 (e.g., the data plane application layer 1246 in Figure 12), the data plane DMZ layer 1448 (e.g., the data plane DMZ layer 1248 in Figure 12), and the data plane data layer 1450 (e.g., the data plane data layer 1250 in Figure 12). The data plane DMZ layer 1448 may include the LB subnet 1422, which can be communicatively coupled to the trusted application subnet 1460 and untrusted application subnet 1462 of the data plane application layer 1446, and the internet gateway 1434 included in the data plane VCN1418. The trusted application subnet 1460 can be communicatively coupled to the service gateway 1436 included in the data plane VCN1418, the NAT gateway 1438 included in the data plane VCN1418, and the DB subnet 1430 included in the data plane data layer 1450. The untrusted application subnet 1462 can be communicatively coupled to the service gateway 1436 included in the data plane VCN 1418 and the DB subnet 1430 included in the data plane data layer 1450. The data plane data layer 1450 may include the DB subnet 1430, which can be communicatively coupled to the service gateway 1436 included in the data plane VCN 1418.

[0175] An untrusted application subnet 1462 may include one or more primary VNICs 1464(1)-(N) that can be communicatively connected to tenant virtual machines (VMs) 1466(1)-(N). Each tenant VM 1466(1)-(N) may be communicatively connected to its respective application subnet 1467(1)-(N), which may be included in each container exit VCN 1468(1)-(N), which may be included in each customer tenant 1470(1)-(N). Each secondary VNIC 1472(1)-(N) can facilitate communication between the untrusted application subnet 1462, which is included in the data plane VCN 1418, and the application subnets, which are included in the container exit VCN 1468(1)-(N). Each container exit VCN 1468(1)-(N) may include a NAT gateway 1438 that can be communicatively connected to the public internet 1454 (e.g., public internet 1254 in Figure 12).

[0176] The Internet gateway 1434, included in the control plane VCN1416 and the data plane VCN1418, can communicate with a metadata management service 1452 (for example, the metadata management system 1252 in Figure 12), which can communicate with the public internet 1454. The public internet 1454 can communicate with a NAT gateway 1438, included in the control plane VCN1416 and the data plane VCN1418. The service gateway 1436, included in the control plane VCN1416 and the data plane VCN1418, can communicate with a cloud service 1456.

[0177] In some embodiments, the data plane VCN1418 can be integrated with the customer tenant 1470. This integration may be beneficial or desirable for the IaaS provider's customer, such as when support during code execution is required. The customer may provide execution of potentially destructive code, code that may communicate with other customer resources, or code that may cause other undesirable effects. In response, the IaaS provider can decide whether to execute the code provided to the IaaS provider by the customer.

[0178] In some examples, an IaaS provider's customer may request temporary network access from the IaaS provider to add functionality to a dataplane tier application 1446. The code that performs the functionality can run on VMs 1466(1)-(N), and the code cannot be configured to run anywhere else on the dataplane VCN 1418. Each VM 1466(1)-(N) can connect to one customer tenant 1470. Each container 1471(1)-(N) contained within VMs 1466(1)-(N) may be configured to run code. In this case, a double isolation may exist (e.g., code execution in containers 1471(1)-(N), which may be contained in at least one VM 1466(1)-(N) that is in an untrusted application subnet 1462), which may help prevent incorrect or undesirable code from damaging the IaaS provider's network or another customer's network. Containers 1471(1)-(N) may be communicatively coupled to customer tenant 1470 and may be configured to send or receive data from customer tenant 1470. Containers 1471(1)-(N) may not be configured to send or receive data from any other entities in the data plane VCN1418. Once code execution is complete, the IaaS provider may terminate or otherwise destroy containers 1471(1)-(N).

[0179] In some embodiments, a trusted application subnet 1460 may execute code that may be owned or operated by the IaaS provider. In this embodiment, the trusted application subnet 1460 may be communicatively coupled to a DB subnet 1430 and configured to perform CRUD operations within the DB subnet 1430. An untrusted application subnet 1462 may be communicatively coupled to the DB subnet 1430, but in this embodiment, the untrusted application subnet may be configured to perform read operations within the DB subnet 1430. Containers 1471(1)-(N), which may be included in each customer's VM 1466(1)-(N) and capable of executing code from the customer, do not necessarily have to be communicatively coupled to the DB subnet 1430.

[0180] In other embodiments, the control plane VCN1416 and the data plane VCN1418 do not have to be directly communicatively coupled. In this embodiment, direct communication between the control plane VCN1416 and the data plane VCN1418 is not required. However, communication can be performed indirectly through at least one method. The LPG1410 may be established by an IaaS provider that can facilitate communication between the control plane VCN1416 and the data plane VCN1418. In another example, the control plane VCN1416 or the data plane VCN1418 can make a call to the cloud service 1456 via the service gateway 1436. For example, a call from the control plane VCN1416 to the cloud service 1456 may include a request for a service that can communicate with the data plane VCN1418.

[0181] Figure 15 is a block diagram 1500 showing another pattern example of an IaaS architecture according to at least one embodiment. A service operator 1502 (e.g., service operator 1202 in Figure 12) can be communicatively coupled to a secure host tenant 1504 (e.g., secure host tenant 1204 in Figure 12), which may include a virtual cloud network (VCN) 1506 (e.g., VCN1206 in Figure 12) and a secure host subnet 1508 (e.g., secure host subnet 1208 in Figure 12). VCN 1506 may include an LPG 1510 (e.g., LPG1210 in Figure 12) which can be communicatively coupled to SSHVCN 1512 (e.g., SSHVCN 1212 in Figure 12) via an LPG 1510 contained in SSHVCN 1512. SSHVCN1512 may include SSH subnet 1514 (e.g., SSH subnet 1214 in Figure 12), and SSHVCN1512 may be communicatively coupled to control plane VCN1516 (e.g., control plane VCN1216 in Figure 12) via LPG1510 included in control plane VCN1516, and may be communicatively coupled to data plane VCN1518 (e.g., data plane 1218 in Figure 12) via LPG1510 included in data plane VCN1518. Control plane VCN1516 and data plane VCN1518 may be included in service tenant 1519 (e.g., service tenant 1219 in Figure 12).

[0182] The control plane VCN1516 may include a control plane DMZ layer 1520 (e.g., control plane DMZ layer 1220 in Figure 12) which may include an LB subnet 1522 (e.g., LB subnet 1222 in Figure 12), a control plane application layer 1524 (e.g., control plane application layer 1224 in Figure 12) which may include an application subnet 1526 (e.g., application subnet 1226 in Figure 12), and a control plane data layer 1528 (e.g., control plane data layer 1228 in Figure 12) which may include a DB subnet 1530 (e.g., DB subnet 930 in Figure 9). The LB subnet 1522 included in the control plane DMZ layer 1520 can be communicatively coupled to the application subnet 1526 included in the control plane application layer 1524, and can be communicatively coupled to the internet gateway 1534 (e.g., internet gateway 1234 in Figure 12), which can be included in the control plane VCN 1516. The application subnet 1526 can be communicatively coupled to the DB subnet 1530 included in the control plane data layer 1528, and can be communicatively coupled to the service gateway 1536 (e.g., the service gateway in Figure 12) and the network address translation (NAT) gateway 1538 (e.g., NAT gateway 1238 in Figure 12). The control plane VCN 1516 can include the service gateway 1536 and the NAT gateway 1538.

[0183] The data plane VCN 1518 may include a data plane application layer 1546 (e.g., data plane application layer 1246 in Figure 12), a data plane DMZ layer 1548 (e.g., data plane DMZ layer 1248 in Figure 12), and a data plane data layer 1550 (e.g., data plane data layer 1250 in Figure 12). The data plane DMZ layer 1548 may include an LB subnet 1522, which may be communicatively coupled to a trusted application subnet 1560 (e.g., trusted application subnet 960 in Figure 9), and may be communicatively coupled to an untrusted application subnet 1562 of the data plane application layer 1546 (e.g., untrusted application subnet 962 in Figure 9) and an internet gateway 1534 included in the data plane VCN 1518. A trusted application subnet 1560 can communicately connect to a service gateway 1536 included in the data plane VCN 1518, and can communicately connect to a NAT gateway 1538 included in the data plane VCN 1518, and to a DB subnet 1530 included in the data plane data layer 1550. An untrusted application subnet 1562 can communicately connect to a service gateway 1536 included in the data plane VCN 1518 and to a DB subnet 1530 included in the data plane data layer 1550. The data plane data layer 1550 may include a DB subnet 1530 that can communicately connect to a service gateway 1536 included in the data plane VCN 1518.

[0184] An untrusted application subnet 1562 may contain primary VNICs 1564(1)-(N), which can be communicatively coupled to tenant virtual machines (VMs) 1566(1)-(N) residing within the untrusted application subnet 1562. Each tenant VM 1566(1)-(N) can execute code within its respective container 1567(1)-(N) and can be communicatively coupled to an application subnet 1526, which can be included in a dataplane application layer 1546, which can be included in a container exit VCN 1568. Each secondary VNIC 1572(1)-(N) can facilitate communication between the untrusted application subnet 1562, which is included in the dataplane VCN 1518, and the application subnet, which is included in the container exit VCN 1568. The container exit VCN may contain a NAT gateway 1538, which can be communicatively coupled to the public internet 1554 (e.g., public internet 1254 in Figure 12).

[0185] The Internet gateway 1534, included in the control plane VCN1516 and the data plane VCN1518, can communicate with a metadata management service 1552 (for example, the metadata management system 1252 in Figure 12), which can communicate with the public internet 1554. The public internet 1554 can communicate with a NAT gateway 1538, included in the control plane VCN1516 and the data plane VCN1518. The service gateway 1536, included in the control plane VCN1516 and the data plane VCN1518, can communicate with a cloud service 1556.

[0186] In some examples, the pattern shown by the architecture in block diagram 1500 of Figure 15 can be considered an exception to the pattern shown by the architecture in block diagram 1400 of Figure 14, which may be desirable for the IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected area). Each container 1567(1)-(N) contained within each customer's VM 1566(1)-(N) is accessible by the customer in real time. Each container 1567(1)-(N) may be configured to make calls to each secondary VNIC 1572(1)-(N) contained within the application subnet 1526 of the data plane application layer 1546, which may be contained within the container exit VCN 1568. The secondary VNICs 1572(1)-(N) may send calls to a NAT gateway 1538, which can send calls to the public internet 1554. In this example, the containers 1567(1)-(N), which customers can access in real time, can be isolated from the control plane VCN1516 and from other entities included in the data plane VCN1518. The containers 1567(1)-(N) may also be isolated from resources from other customers.

[0187] In another example, a customer can use containers 1567(1)-(N) to invoke cloud service 1556. In this example, the customer can execute code within containers 1567(1)-(N) to request a service from cloud service 1556. Containers 1567(1)-(N) can send this request to secondary VNICs 1572(1)-(N), which can then send the request to a NAT gateway that can send the request to the public internet 1554. The public internet 1554 can then send the request to LB subnet 1522, which is included in control plane VCN 1516, via internet gateway 1534. In response to the determination that the request is valid, the LB subnet can send the request to application subnet 1526, which can then send the request to cloud service 1556 via service gateway 1536.

[0188] It should be understood that the IaaS architectures 1200, 1300, 1400, and 1500 shown in the figures may have components other than those shown. Furthermore, the embodiments shown in the figures are only some examples of cloud infrastructure systems that may incorporate embodiments of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than those shown, may combine two or more components, or may have different configurations or arrangements of components.

[0189] In certain embodiments, the IaaS system described herein may include a suite of applications, middleware, and database service products delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the assignee.

[0190] Figure 16 shows an exemplary computer system 1600 in which various embodiments can be implemented. System 1600 can be used to implement any of the computer systems described above. As shown in the figure, computer system 1600 includes a processing unit 1604 that communicates with several peripheral subsystems via a bus subsystem 1602. These peripheral subsystems may include a processing accelerator 1606, an I / O subsystem 1608, a storage subsystem 1618, and a communication subsystem 1624. The storage subsystem 1618 includes a tangible computer-readable storage medium 1622 and system memory 1610.

[0191] The bus subsystem 1602 provides a mechanism that enables various components and subsystems of the computer system 1600 to communicate with each other as intended. Although the bus subsystem 1602 is schematically shown as a single bus, multiple buses can be utilized in alternative embodiments of the bus subsystem. The bus subsystem 1602 may be any of several types of bus structures, including a memory bus or memory controller, peripheral bus, and local bus, using any of the various bus architectures. For example, such architectures may include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. This can be implemented as a mezzanine bus manufactured according to the IEEEP1386.1 standard.

[0192] The processing unit 1604 can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 1600. One or more processors may be included in the processing unit 1604. These processors may include single-core processors or multi-core processors. In certain embodiments, the processing unit 1604 may be implemented as one or more independent processing units 1632 and / or 1634, each containing a single-core processor or a multi-core processor. In other embodiments, the processing unit 1604 may be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.

[0193] In various embodiments, the processing unit 1604 can execute various programs in response to program code and can maintain multiple concurrently running programs or processes. At any given time, some or all of the program code to be executed may reside in the processor 1604 and / or the storage subsystem 1618. Through appropriate programming, the processor 1604 can provide the various functions described above. The computer system 1600 may further include a processing accelerator 1606 which may include a digital signal processor (DSP), a dedicated processor, etc.

[0194] The I / O subsystem 1608 may include user interface input devices and user interface output devices. User interface input devices may include keyboards, pointing devices such as mice and trackballs, touchpads and touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as the Microsoft Kinect® motion sensor, which allow the user to control and interact with input devices, such as the Microsoft Xbox® 360 game controller, through a natural user interface using gestures and voice commands. User interface input devices may also include eye gesture recognition devices, such as the Google Glass® blink detector, which detects eye activity from the user (e.g., blinking while taking photos and / or selecting menus) and translates eye gestures as input to an input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition sensing device that enables the user to interact with a voice recognition system (e.g., Siri® Navigator) through voice commands.

[0195] User interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Furthermore, user interface input devices may also include medical imaging input devices such as computed tomography, magnetic resonance imaging, positional radiography, and medical ultrasound equipment. User interface input devices may also include audio input devices such as MIDI keyboards and digital musical instruments.

[0196] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may include flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touchscreens, etc. In general, the use of the term “output device” is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1600 to a user or another computer. For example, user interface output devices include, but are not limited to, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.

[0197] The computer system 1600 may include a storage subsystem 1618 having software elements that are shown to be currently located in the system memory 1610. The system memory 1610 can store program instructions that can be loaded and executed on the processing unit 1604, as well as data generated during the execution of these programs.

[0198] Depending on the configuration and type of the computer system 1600, the system memory 1610 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM) or flash memory). RAM typically contains data and / or program modules that are immediately accessible to the processing unit 1604 and / or currently operating and executing by the processing unit 1604. In some implementations, the system memory 1610 may contain several different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, the basic input / output system (BIOS), which contains basic routines that help transfer information between elements within the computer system 1600, such as during startup, may typically be stored in ROM. As an example, and not an limitation, the system memory 1610 also refers to application programs 1612, program data 1614, and the operating system 1616, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. For example, Operating System 1616 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 16 OS, and Palm® OS.

[0199] The storage subsystem 1618 may also provide a tangible, computer-readable storage medium for storing basic programming and data structures that provide the functionality of several embodiments. When executed by a processor, software (programs, code modules, instructions) that provides the functionality described above may be stored in the storage subsystem 1618. These software modules or instructions may be executed by the processing unit 1604. The storage subsystem 1618 may also provide a repository for storing data used in accordance with this disclosure.

[0200] The storage subsystem 1600 may also include a computer-readable storage medium reader 1620 that can be further connected to the computer-readable storage medium 1622. Together, and optionally in combination with the system memory 1610, the computer-readable storage medium 1622 can comprehensively represent a storage medium for temporarily and / or more permanently storing, storing, transmitting, and retrieving computer-readable information, in addition to remote, local, fixed, and / or removable storage devices.

[0201] Computer-readable storage media 1622 containing code or a portion of code may also include any suitable media known or used in the art, including, but not limited to, storage and communication media such as volatile and non-volatile, removable and non-removable media, implemented in any way or technique for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This may also include intangible computer-readable media such as any other media that can be used to transmit data signals, data transmissions, or desired information and are accessible by the computing system 1600.

[0202] As an example, the computer-readable storage medium 1622 may include a hard disk drive that reads or writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads or writes to a removable non-volatile magnetic disk, and an optical disk drive that reads or writes to a removable non-volatile optical disk such as a CD-ROM, DVD, Blu-ray® disc, or other optical medium. The computer-readable storage medium 1622 may also include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, and digital videotapes. The computer-readable storage medium 1622 may also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, and static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and associated computer-readable media can provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data for computer system 1600.

[0203] The communication subsystem 1624 provides interfaces to other computer systems and networks. The communication subsystem 1624 functions as an interface for sending and receiving data between the computer system 1600 and other systems. For example, the communication subsystem 1624 can enable the computer system 1600 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1624 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using advanced data network technologies such as cellular technology, 3G, 4G, or EDGE (Enhanced Data Rate for Global Evolution)), WiFi (IEEE 802.11 family standards, or other mobile communication technologies, or any combination thereof), Global Positioning System (GPS) receiver components, and / or other components. In some embodiments, the communication subsystem 1624 may provide, in addition to or instead of a wireless interface, a wired network connection (e.g., Ethernet).

[0204] In some embodiments, the communication subsystem 1624 may also receive input communications in the form of structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., on behalf of one or more users who can use the computer system 1600.

[0205] As an example, the communication subsystem 1624 may be configured to receive data feeds 1626 in real time from users of social networks and / or other communication services such as Twitter® feeds, Facebook® updates, and web feeds such as Rich Site Summary (RSS) feeds, as well as / or real-time updates from one or more third-party information sources.

[0206] Furthermore, the communication subsystem 1624 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1628 of real-time events and / or event updates 1630, which may be continuous or have no explicit end and may be essentially unlimited. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.

[0207] The communication subsystem 1624 may also be configured to output structured and / or unstructured data feeds 1626, event streams 1628, event updates 1630, etc., to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1600.

[0208] Computer system 1600 can be one of various types, including handheld portable devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or other data processing systems.

[0209] Due to the constantly changing nature of computers and networks, the description of the computer system 1600 shown in the figure is intended only as a specific example. Many other configurations are possible with more or fewer components than the system shown in the figure. For example, customized hardware may also be used, or certain elements may be implemented in hardware, firmware, software (including applets), or a combination thereof. Furthermore, connections to other computing devices such as network input / output devices may be used. Based on the disclosures and teachings provided herein, those skilled in the art will understand other techniques and / or methods for implementing various embodiments.

[0210] While specific embodiments have been described, various modifications, changes, alternative structures, and equivalents are also included within the scope of this disclosure. The embodiments are not limited to operation within a specific, particular data processing environment, but can freely operate within multiple data processing environments. Furthermore, while the embodiments have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope of this disclosure is not limited to the described set of transactions and steps. The various features and aspects of the embodiments described above can be used individually or in combination.

[0211] Furthermore, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments can be implemented using hardware alone, software alone, or a combination thereof. The various processes described herein can be implemented on the same processor or on any combination of different processors. Thus, where a component or module is described as being configured to perform a particular operation, such configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or by any combination thereof. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or pairs of the same process may use different techniques at different times.

[0212] Therefore, the specification and drawings should be considered illustrative, not restrictive. However, it is clear that additions, subtractions, deletions, and other modifications and changes can be made without departing from the broader intent and scope set forth in the claims. Thus, while specific embodiments of disclosure have been described, they are not intended to be limiting. Various modifications and equivalents are included in the following claims.

[0213] In the context describing the disclosed embodiments (particularly in the context of the following claims), the use of the terms “a,” “an,” “the,” and similar reference subjects should be construed to cover both singular and plural forms unless otherwise indicated herein or otherwise clearly inconsistent with the context. The terms “include,” “have,” “contain,” and “contain” should be construed as unrestricted terms (i.e., “include but not limited to”) unless otherwise specified herein. The term “connected” should be construed to be included, attached, or combined with, even if something is intervening within it. The descriptions of value ranges herein are merely intended to serve as a simplified way of referring individually to each individual value within that range unless otherwise indicated herein, and each individual value is incorporated into the specification as if it were individually described herein. All methods described herein may be performed in any appropriate order unless otherwise indicated herein or otherwise clearly inconsistent with the context. Any examples or exemplary language provided herein (e.g., "etc.") are for illustrative purposes only and, unless otherwise requested, do not limit the scope of this disclosure. Nothing in this specification should be construed as indicating that any unclaimed element is essential for the practice of this disclosure.

[0214] Disjunctive expressions, such as the phrase "at least one of X, Y, or Z," are intended to be understood in context to be generally used to indicate that an item, term, etc., can be any one of X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless otherwise specified. Therefore, such disjunctive expressions are not, and should not, be intended to mean that a particular embodiment requires the presence of at least one X, at least one Y, or at least one Z, respectively.

[0215] Preferred embodiments of the present disclosure, including the best known modes for carrying out the present disclosure, are described herein. Variations of these preferred embodiments will be apparent to those skilled in the art by reading the preceding description. Those skilled in the art should be able to adopt such variations as needed, and the present disclosure can also be carried out in ways other than those specifically described herein. Accordingly, the present disclosure includes all modifications and equivalents of the subject matter described in the claims appended herein, as permitted by applicable law. Furthermore, unless otherwise indicated herein, any combination of the elements described above in all possible modifications is incorporated herein.

[0216] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the same extent as they are incorporated herein by reference, to the extent that each reference is individually and specifically indicated to be incorporated herein by reference. While aspects of the disclosure are described in the foregoing specification with reference to specific embodiments thereof, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the foregoing disclosure can be used individually or in combination. Furthermore, embodiments can be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings should be considered illustrative and not restrictive.

Claims

1. It is a method, Regarding packets transmitted by the host machine's graphical processing unit (GPU) and received by network devices, The network device acquires flow information associated with the packet, The network device includes determining the outgoing port link of the network device according to a hash algorithm based on the flow information, wherein the hash algorithm hashs packets received on a specific receiving port link of the network device, and is configured such that packets received on a specific receiving port link of the network device are transmitted on the same outgoing port link of the network device. The aforementioned method, A method further comprising the network device forwarding the packets on the outgoing port link of the network device.

2. The aforementioned transfer is, The network device confirms the conditions related to the outgoing port link of the network device, The method according to claim 1, further comprising forwarding the packet on the outgoing port link of the network device in response to the fulfillment of the above conditions.

3. The method according to claim 2, wherein the condition corresponds to determining whether the outgoing port link of the network device is active.

4. In response to the fact that the above conditions are not met, The network device executes an equal-cost multipath algorithm to acquire a new outgoing port link for the network device based on the flow information, The method according to claim 2, further comprising the network device forwarding the packets on the new outgoing port link of the network device.

5. The method according to claim 1, wherein the network device is a top-of-rack (TOR) switch.

6. The method according to claim 1, wherein information related to hash operations performed on packets received by the network device is stored in a forwarding table database.

7. With respect to the first and second packets received on the first receiving port link of the network device, and the third and fourth packets received on the second receiving port link of the network device, the network device shall The network device transmits the first packet and the second packet on the first outgoing port link of the network device, The network device transmits the third packet and the fourth packet on the second outgoing port link of the network device. The method according to claim 1, configured to perform such a function, wherein the first receiving port link of the network device is different from the second receiving port link of the network device, and the first transmitting port link of the network device is different from the second transmitting port link of the network device.

8. The method according to claim 7, wherein the first packet, the second packet, the third packet, and the fourth packet belong to a GPU workload.

9. The method according to claim 7, wherein the first receiving port link and the second receiving port link of the network device are a first set of links connecting the host machine to a top-of-rack (TOR) switch, and the first transmitting port link and the second transmitting port link of the network device are a second set of links connecting the TOR to a spine switch.

10. Network device, Processor and A network device that, when executed by the processor, causes the network device to perform the method described in any one of claims 1 to 9.

11. A program that causes a computer system to perform the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Technologies for load balancing a network

    EP3531633A1

  • Data forwarding method and device

    US20190097914A1

  • Switch-enhanced short loop congestion notification for TCP

    US20190173776A1

  • Slice-based routing

    US20200366607A1