Routing in GPU hyper-cluster

Coupling multiple GPU clusters through multiple network devices arranged in a hierarchical structure and configuring routing policies, the problem of insufficient scalability and routing policy support in traditional GPU clusters is solved, and seamless expansion and communication of GPU hybrid clusters is achieved.

CN120153359APending Publication Date: 2025-06-13ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-18
Filing Date
2023-11-02
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional GPU clusters have limitations in scalability and routing policy support, making it difficult to support seamless communication and scaling between different generations of GPUs or GPUs operating at different speeds.

Method used

Multiple network devices arranged in a hierarchical structure couple multiple GPU clusters together, establish a mapping of the incoming port link of the network device to the unique outgoing port link by configuring routing policies, and forward packets on the outgoing port link of the network device.

Benefits of technology

It realizes the coexistence of GPU hybrid clusters in the same network architecture, supports seamless expansion and communication between different GPU clusters, and improves the expansion level and communication efficiency of GPU clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153359A_ABST
    Figure CN120153359A_ABST
Patent Text Reader

Abstract

The plurality of GPU clusters are communicatively coupled to each other via a plurality of network devices arranged in a hierarchical structure, wherein the GPU clusters include at least a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed. A routing policy is configured for each network device, wherein the configuration includes establishing a mapping of each incoming port link of the network device to a unique outgoing port link of the network device. For a data packet transmitted by the GPU of the host machine and received by the first network device, an incoming port link of the first network device on which the data packet is received is determined, and an outgoing port link corresponding to the incoming port link is identified based on the configuration. The data packet is forwarded over an outgoing port link of the network device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is a non - provisional application of and claims the benefit of each of the following provisional applications. The entire content of each of the following provisional applications is incorporated herein by reference for all purposes:

[0003] (1) U.S. Provisional Application No. 63 / 422,650, filed on November 4, 2022;

[0004] (2) U.S. Provisional Application No. 63 / 424,282, filed on November 10, 2022;

[0005] (3) U.S. Provisional Application No. 63 / 425,646, filed on November 15, 2022;

[0006] (4) U.S. Provisional Application No. 63 / 460,766, filed on April 20, 2023;

[0007] (5) U.S. Provisional Application No. 63 / 583,512, filed on September 18, 2023; Technical Field

[0008] This disclosure generally relates to cloud architectures and, more particularly, to supercluster architectures of graphics processing units (GPUs). More specifically, this disclosure relates to a network architecture that allows a hybrid cluster of GPUs (e.g., different generations of GPUs, or GPUs operating at different speeds, etc.) to co - exist within the same network architecture. The supercluster architecture allows for seamless scaling of GPUs as customer demands grow. Background Art

[0009] Organizations are increasingly moving business applications and databases to the cloud to reduce the costs of purchasing, updating, and maintaining on - premise hardware and software. High - performance computing applications continuously consume all available computing power to achieve specific outcomes or results. Such applications require dedicated network performance, fast storage, high computing power, and large amounts of memory - resources that are in short supply in the virtualized infrastructure that makes up today's commodity clouds.

[0010] Cloud infrastructure service providers supply newer and faster graphics processing units (GPUs) to meet the demands of such applications. GPU workloads are typically executed on one or more host machines. Generally, such workloads do not achieve the expected throughput levels. One factor contributing to this problem is the lack of flow entropy, e.g., equal - cost multi - path (ECMP) flow entropy. Additionally, the fact that host machines (i.e., hosts) exchange traffic without considering other hosts in their local network neighborhood exacerbates the problem.

[0011] Moreover, traditional GPU clusters typically scale in the range of 1K to 4K GPUs. The limit in scaling the number of GPUs is due to the constraints imposed by the network topology built to support the GPU cluster. The network topology built to support the GPU cluster results in a large amount of oversubscription, which poses a challenge to scaling the cluster. In addition, traditional GPU clusters impose strict restrictions on the routing policies employed within the cluster. For example, traditional GPU clusters do not support standard custom routing protocols. Further, traditional GPU clusters are built in such a way that they support one transmission speed for all GPUs in the cluster. Therefore, there is a need to build a GPU cluster that can scale to a much higher level than traditional GPU clusters and support communication between different GPU clusters operating at different transmission speeds. The embodiments discussed in this document address these and other issues. SUMMARY OF THE INVENTION

[0012] The present disclosure relates to cloud architectures and, more particularly, to a supercluster architecture for graphics processing units (GPUs). More specifically, the present disclosure relates to a network architecture that allows hybrid clusters of GPUs (e.g., different generations of GPUs, or GPUs operating at different speeds, etc.) to coexist within the same network architecture. The supercluster architecture allows for seamless scaling of GPUs with growing customer demands. Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, code, or instructions executable by one or more processors, etc. Some embodiments may be implemented by using a computer program product that includes a computer program / instructions that, when executed by a processor, cause the processor to perform any of the methods described in the present disclosure.

[0013] One aspect of the present disclosure provides a method, comprising: providing a plurality of graphics processing units (GPU) clusters communicatively coupled to each other via a plurality of network devices arranged in a hierarchy, wherein the plurality of GPU clusters includes at least a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed; configuring a routing policy for each of the plurality of network devices, wherein the configuration includes establishing a mapping of each incoming port link of the network device to a unique outgoing port link of the network device; and for a data packet transmitted by a GPU of a host machine and received by a first network device, determining the incoming port link of the first network device on which the data packet is received; identifying the outgoing port link corresponding to the incoming port link based on the configuration; and forwarding the data packet on the outgoing port link of the network device.

[0014] One aspect of the present disclosure provides a computing device including one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the computing device to perform some or all of one or more methods disclosed herein.

[0015] Another aspect of the present disclosure provides a computer program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein.

[0016] The foregoing and other features and embodiments will become more apparent when reference is made to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The features, embodiments, and advantages of the present disclosure can be better understood when the following detailed description is read with reference to the drawings.

[0018] Figure 1 is a high-level diagram of a distributed environment showing a virtual or overlay cloud network hosted by a cloud service provider infrastructure according to certain embodiments.

[0019] Figure 2 depicts a simplified architecture diagram of physical components in a physical network within a CSPI according to certain embodiments.

[0020] Figure 3 shows an example arrangement within a CSPI according to certain embodiments, where a host machine is connected to multiple network virtualization devices (NVDs).

[0021] Figure 4 depicts the connectivity between a host machine and an NVD according to certain embodiments for providing I / O virtualization to support multi-tenancy.

[0022] Figure 5 depicts a simplified block diagram of a physical network provided by a CSPI according to certain embodiments.

[0023] Fig. 6A depicts the architecture of a hybrid GPU cluster according to certain embodiments.

[0024] Figure 6B depicts another architecture of a hybrid GPU cluster according to certain embodiments.

[0025] Figure 7 illustrates an exemplary flowchart depicting steps performed when provisioning requests using a hybrid GPU cluster according to certain embodiments.

[0026] Figure 8Depicts a simplified block diagram of a cloud infrastructure incorporating a CLOS network arrangement, according to certain embodiments.

[0027] Fig. 9 Depicts a logical topology constructed without locality information, according to certain embodiments.

[0028] Fig.10 Depicts a logical topology constructed with locality information, according to certain embodiments.

[0029] Fig.11 Depicts a simplified block diagram of a hybrid GPU cluster, according to some embodiments, illustrating the locality of network components.

[0030] Fig.12 Illustrates an exemplary flowchart depicting steps performed when servicing requests using locality information, according to certain embodiments.

[0031] Fig.13 Depicts a policy - based routing mechanism implemented in a hybrid GPU cluster, according to certain embodiments.

[0032] Fig.14 Illustrates a flowchart depicting steps performed by a network device when routing data packets, according to certain embodiments.

[0033] Fig.15A Depicts the architecture of a hybrid GPU cluster, according to some embodiments, illustrating the placement of route reflectors.

[0034] Fig. 15B Illustrates a flowchart depicting steps performed by a route reflector when managing the size of an address table stored in a switch, according to certain embodiments.

[0035] Fig.16 Is a block diagram illustrating a mode for implementing an Infrastructure as a Service (IaaS) cloud system, according to at least one embodiment.

[0036] Fig.17 Is a block diagram illustrating another mode for implementing an Infrastructure as a Service (IaaS) cloud system, according to at least one embodiment.

[0037] Fig.18 Is a block diagram illustrating another mode for implementing an Infrastructure as a Service (IaaS) cloud system, according to at least one embodiment.

[0038] Fig.19 Is a block diagram illustrating another mode for implementing an Infrastructure as a Service (IaaS) cloud system, according to at least one embodiment.

[0039] Fig. 20 Is a block diagram illustrating an example computer system, according to at least one embodiment. Detailed implementation manners

[0040] In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. However, it is apparent that the various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive. The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or superior to other embodiments or designs.

[0041] Embodiments of the present disclosure relate to cloud architectures, and more particularly to a supercluster architecture of graphics processing units (GPUs). More specifically, the present disclosure relates to a network architecture that allows a hybrid cluster of GPUs (e.g., different generations of GPUs, or GPUs operating at different speeds, etc.) to coexist in the same network architecture. The supercluster architecture allows for seamless scaling of GPUs as customer demands grow.

[0042] Cloud Network Example

[0043] The term cloud service is generally used to refer to services provided by a cloud service provider (CSP) to a user or customer on demand (e.g., via a subscription model) using systems and infrastructure (cloud infrastructure) provided by the CSP. Generally, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premises servers and systems. Thus, the customer can utilize the cloud services provided by the CSP without having to purchase separate hardware and software resources for the services. Cloud services are designed to provide subscribing customers with simple, scalable access to applications and computing resources without the customer having to invest in the infrastructure for providing the services.

[0044] There are several cloud service providers that offer various types of cloud services. There are various different types or models of cloud services, including software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), etc.

[0045] A customer can subscribe to one or more cloud services provided by the CSP. A customer can be any entity, such as an individual, an organization, a business, etc. When a customer subscribes to or registers for a service provided by the CSP, a lease or account is created for that customer. The customer can then access the one or more subscribed cloud resources associated with that account via this account.

[0046] As described above, Infrastructure as a Service (IaaS) is a specific type of cloud computing service. In the IaaS model, the CSP provides the infrastructure (referred to as the cloud service provider infrastructure or CSPI), which can be used by customers to build their own customizable networks and deploy customer resources. Thus, the customer's resources and network are hosted in a distributed environment by the infrastructure provided by the CSP. This is different from traditional computing, where the customer's resources and network are hosted by the infrastructure provided by the customer.

[0047] The CSPI can include interconnected high-performance computing resources that form a physical network, including various host machines, memory resources, and network resources, which is also referred to as the substrate network or underlying network. The resources in the CSPI can be spread across one or more data centers, which can be geographically dispersed across one or more geographical regions. Virtualization software can be executed by these physical resources to provide a virtualized distributed environment. Virtualization creates an overlay network (also referred to as a software-based network, software-defined network, or virtual network) on top of the physical network. The CSPI physical network provides the underlying foundation for creating one or more overlay or virtual networks on top of the physical network. The physical network (or substrate network or underlying network) includes physical network devices such as physical switches, routers, computers, and host machines. The overlay network is a logical (or virtual) network that runs on top of the physical substrate network. A given physical network can support one or more overlay networks. Overlay networks typically use encapsulation techniques to distinguish traffic belonging to different overlay networks. The virtual or overlay network is also referred to as a Virtual Cloud Network (VCN). The virtual network is implemented using software virtualization techniques (e.g., hypervisors, virtualization functions implemented by Network Virtualization Devices (NVDs) (e.g., Smart Network Cards), Top-of-Rack (TOR) switches, Smart TORs that implement one or more functions performed by NVDs, and other mechanisms) to create a layer of network abstraction that can run on top of the physical network. The virtual network can take various forms, including peer-to-peer networks, IP networks, etc. The virtual network is typically either a Layer 3 IP network or a Layer 2 VLAN. This method of virtual or overlay networking is often referred to as virtual or overlay Layer 3 networking. Examples of protocols developed for virtual networks include IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN - IETF RFC7348), Virtual Private Network (VPN) (e.g., MPLS Layer 3 Virtual Private Network (RFC4364)), VMware's NSX, GENEVE (Generic Network Virtualization Encapsulation), etc.

[0048] For IaaS, the infrastructure provided by the CSP (CSPI) can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing service provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider can also supply various services to accompany those infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering, etc.). Thus, since these services can be policy-driven, IaaS users can be able to implement policies to drive load balancing to maintain application availability and performance. CSPI provides a collection of infrastructure and complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available, hosted, distributed environment. CSPI provides high-performance computing resources and capabilities, as well as storage capacity, in a flexible virtual network that can be securely accessed from various networked locations (such as from the customer's on-premises network). When a customer subscribes to or registers for IaaS services provided by the CSP, the lease created for that customer is a secure and isolated partition within CSPI where the customer can create, organize, and manage their cloud resources.

[0049] Customers can use the computing, memory, and networking resources provided by CSPI to build their own virtual networks. One or more customer resources or workloads, such as compute instances, can be deployed on these virtual networks. For example, customers can use the resources provided by CSPI to build one or more customizable and private virtual networks, called virtual cloud networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, on the customer VCN. Compute instances can take the form of virtual machines, bare-metal instances, etc. Thus, CSPI provides a collection of infrastructure and complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available virtual hosted environment. Customers do not manage or control the underlying physical resources provided by CSPI, but can control the operating system, storage devices, and deployed applications; and may have limited control over selected networking components (e.g., firewalls).

[0050] The CSP can provide a console that enables customers and network administrators to configure, access, and manage the resources deployed in the cloud using CSPI resources. In certain embodiments, the console provides a web-based user interface that can be used to access and manage CSPI. In certain implementations, the console is a web-based application provided by the CSP.

[0051] CSPI can support single-tenant or multi-tenant architectures. In a single-tenant architecture, software (e.g., applications, databases) or hardware components (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenant architecture, software or hardware components serve multiple customers or tenants. Thus, in a multi-tenant architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenant scenario, preventive measures are taken and protection measures are implemented in CSPI to ensure that each tenant's data is isolated and invisible to other tenants.

[0052] In a physical network, a network endpoint ("endpoint") refers to a computing device or system that is connected to a physical network and communicates back and forth with the network to which it is connected. Network endpoints in a physical network can be connected to a local area network (LAN), a wide area network (WAN), or other types of physical networks. Examples of traditional endpoints in a physical network include modems, hubs, bridges, switches, routers, and other networking devices, physical computers (or host machines), etc. Each physical device in a physical network has a fixed network address that can be used to communicate with the device. This fixed network address can be a layer 2 address (e.g., MAC address), a fixed layer 3 address (e.g., IP address), etc. In a virtualized environment or virtual network, endpoints can include various virtual endpoints, such as virtual machines hosted by components of the physical network (e.g., hosted by a physical host machine). These endpoints in a virtual network are addressed by overlay addresses, such as overlay layer 2 addresses (e.g., overlay MAC address) and overlay layer 3 addresses (e.g., overlay IP address). Network overlay enables flexibility by allowing network administrators to move around the overlay addresses associated with network endpoints using software management (e.g., via software that implements a control plane for the virtual network). Accordingly, different from a physical network, in a virtual network, an overlay address (e.g., overlay IP address) can be moved from one endpoint to another using network management software. Since a virtual network is built on top of a physical network, communication between components in a virtual network involves both the virtual network and the underlying physical network. To facilitate such communication, components of CSPI are configured to learn and store mappings that map overlay addresses in the virtual network to actual physical addresses in the underlying network, and vice versa. These mappings are then used to facilitate communication. Customer traffic is encapsulated to facilitate routing in the virtual network.

[0053] Accordingly, a physical address (e.g., a physical IP address) is associated with a component in a physical network, and an overlay address (e.g., an overlay IP address) is associated with an entity in a virtual or overlay network. A physical IP address is an IP address associated with a physical device (e.g., a network device) in a substrate or physical network. For example, each NVD has an associated physical IP address. An overlay IP address is an overlay address associated with an entity in an overlay network, such as an overlay address associated with a compute instance in a customer's Virtual Cloud Network (VCN). Two different customers or tenants (each with its own private VCN) can potentially use the same overlay IP address in their VCNs without knowing about each other. Both physical IP addresses and overlay IP addresses are types of real IP addresses. These addresses are separate from virtual IP addresses. A virtual IP address is typically a single IP address that represents or maps to multiple real IP addresses. A virtual IP address provides a one-to-many mapping between the virtual IP address and multiple real IP addresses. For example, a load balancer can use a VIP to map or represent multiple servers, each with its own real IP address.

[0054] A cloud infrastructure or CSPI is physically hosted in one or more data centers in one or more regions of the world. The CSPI can include components in a physical or substrate network and virtualized components (e.g., virtual networks, compute instances, virtual machines, etc.) in a virtual network built on top of the physical network components. In some embodiments, the CSPI is organized and hosted in realms, regions, and availability domains. A region is generally a local geographic area that contains one or more data centers. Regions are generally independent of each other and can be far apart, e.g., spanning countries or even continents. For example, a first region can be in Australia, another in Japan, another in India, and so on. CSPI resources are partitioned across regions such that each region has its own independent subset of CSPI resources. Each region can provide a core set of infrastructure services and resources, such as compute resources (e.g., bare metal servers, virtual machines, containers, and associated infrastructure, etc.); storage resources (e.g., block volume storage, file storage, object storage, archival storage); networking resources (e.g., Virtual Cloud Network (VCN), load balancing resources, connection to an on-premises network), database resources; edge networking resources (e.g., DNS); and access management and monitoring resources, etc. Each region generally has multiple paths connecting it to other regions in the realm.

[0055] Generally, an application is deployed in the region where it is most frequently used (i.e., deployed on the infrastructure associated with that region), because using nearby resources is faster than using distant resources. An application can also be deployed in different regions for various reasons, such as redundancy to mitigate the risk of region-wide events (such as large weather systems or earthquakes), to meet different requirements such as legal jurisdictions, tax domains, and other commercial or social criteria.

[0056] Data centers within a region can be further organized and subdivided into Availability Domains (ADs). An Availability Domain can correspond to one or more data centers located within the region. A region can consist of one or more Availability Domains. In such a distributed environment, CSPI resources are either region-specific, such as a Virtual Cloud Network (VCN), or Availability Domain-specific, such as a compute instance.

[0057] ADs within a region are isolated from each other, fault-tolerant, and configured such that it is highly unlikely for them to fail simultaneously. This is achieved by ADs not sharing critical infrastructure resources (such as networking, physical cables, cable paths, cable entry points, etc.), such that a failure at one AD within a region is unlikely to affect the availability of other ADs within the same region. ADs within the same region can be connected to each other via a low-latency, high-bandwidth network, which enables providing highly available connectivity to other networks (e.g., the Internet, a customer's on-premises network, etc.) and building replicated systems across multiple ADs to achieve both high availability and disaster recovery. Cloud services use multiple ADs to ensure high availability and prevent resource failures. As the infrastructure provided by an IaaS provider grows, more regions and ADs, as well as additional capacity, can be added. Traffic between Availability Domains is typically encrypted.

[0058] In some embodiments, regions are grouped into realms. A realm is a logical collection of regions. Realms are isolated from each other and do not share any data. Regions within the same realm can communicate with each other, but regions in different realms cannot. A customer's lease or account with a CSP exists within a single realm and can be spread across one or more regions belonging to that realm. Typically, when a customer subscribes to an IaaS service, a lease or account for that customer is created in the region (referred to as the "primary" region) specified by the customer within the realm. A customer can extend the customer's lease to one or more other regions within the realm. A customer cannot access regions that are not within the realm where the customer's lease resides.

[0059] IaaS providers can offer multiple realms, each catering to a specific set of customers or users. For example, a business realm can be offered for business customers. As another example, a realm can be offered for a specific country for the customers within that country. As yet another example, a government realm can be offered for a government, etc. For example, a government realm can cater to a specific government and can have a higher security level than a business realm. For example, Oracle Cloud Infrastructure (OCI) currently offers realms for commercial regions and two realms for government cloud regions (e.g., FedRAMP authorized and IL5 authorized).

[0060] In some embodiments, an AD can be subdivided into one or more fault domains. A fault domain is a grouping of infrastructure resources within an AD to provide anti-affinity. Fault domains allow the distribution of compute instances such that these instances do not reside on the same physical hardware within a single AD. This is referred to as anti-affinity. A fault domain refers to a collection of hardware components (computers, switches, etc.) that share a single point of failure. The compute pool is logically divided into fault domains. Thus, a hardware failure or a compute hardware maintenance event affecting one fault domain does not affect the instances in other fault domains. Depending on the embodiment, the number of fault domains for each AD can vary. For example, in some embodiments, each AD contains three fault domains. Fault domains act as logical data centers within an AD.

[0061] When a customer subscribes to an IaaS service, resources from the CSPI are provisioned to the customer and associated with the customer's tenancy. The customer can use these provisioned resources to build private networks and deploy resources on these networks. The customer network hosted by the CSPI in the cloud is referred to as a Virtual Cloud Network (VCN). The customer can use the CSPI resources allocated to the customer to set up one or more Virtual Cloud Networks (VCNs). A VCN is a virtual or software-defined private network. The customer resources deployed in the customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances can represent various customer workloads such as applications, load balancers, databases, etc. The compute instances deployed on a VCN can communicate with public accessible endpoints ("public endpoints") via a public network such as the Internet, with other instances in the same VCN or other VCNs (e.g., other VCNs of the customer or VCNs that do not belong to the customer), with the customer's on-premises data center or network, and with service endpoints and other types of endpoints.

[0062] CSPs can use CSPIs to provide various services. In some cases, the customers of CSPIs can themselves act like service providers and use CSPI resources to provide services. A service provider can expose a service endpoint, which is characterized by identification information (e.g., IP address, DNS name, and port). A customer's resources (e.g., compute instances) can use a particular service by accessing the service endpoint exposed by the service for that particular service. These service endpoints are generally endpoints that are publicly accessible to users via a public communication network such as the Internet using the public IP address associated with the endpoint. A publicly accessible network endpoint is sometimes also referred to as a public endpoint.

[0063] In some embodiments, a service provider can expose a service via an endpoint for the service (sometimes referred to as a service endpoint). A customer of the service can then use this service endpoint to access the service. In some implementations, the service endpoint provided for a service can be accessed by multiple customers who intend to consume the service. In other implementations, a dedicated service endpoint can be provided for a customer such that only that customer can use the dedicated service endpoint to access the service.

[0064] In some embodiments, when a VCN is created, it is associated with a private overlay Classless Inter-Domain Routing (CIDR) address space, which is a range of private overlay IP addresses assigned to the VCN (e.g., 10.0 / 16). A VCN includes associated subnets, a routing table, and a gateway. A VCN resides within a single region but can span one or more or all of the availability domains of that region. A gateway is a virtual interface configured for the VCN and enables communication of traffic between the VCN and one or more endpoints external to the VCN. One or more different types of gateways can be configured for the VCN to enable communication to and from different types of endpoints.

[0065] A VCN can be subdivided into one or more sub-networks, such as one or more subnets. Thus, a subnet is a configured unit or subdivision that can be created within a VCN. A VCN can have one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that do not overlap with other subnets in the VCN and represent a subset of the address space within the VCN's address space.

[0066] Each compute instance is associated with a virtual network interface card (VNIC), which enables the compute instance to participate in a subnet of a VCN. A VNIC is a logical representation of a physical network interface card (NIC). Generally speaking, a VNIC is an interface between an entity (e.g., a compute instance, a service) and a virtual network. A VNIC exists within a subnet, has one or more associated IP addresses, and associated security rules or policies. A VNIC is equivalent to a Layer 2 port on a switch. A VNIC is attached to a compute instance and a subnet within a VCN. The VNIC associated with a compute instance makes the compute instance part of a subnet of a VCN and enables the compute instance to communicate (e.g., send and receive data packets) with endpoints on the same subnet as the compute instance, with endpoints in different subnets within the VCN, or with endpoints outside the VCN. Thus, the VNIC associated with a compute instance determines how the compute instance connects to endpoints inside and outside the VCN. When a compute instance is created and added to a subnet within a VCN, a VNIC for the compute instance is created and associated with that compute instance. For a subnet that includes a set of compute instances, the subnet contains VNICs corresponding to that set of compute instances, and each VNIC is attached to a compute instance within that set of compute instances.

[0067] A private overlay IP address is assigned to each compute instance via the VNIC associated with the compute instance. This private overlay network IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic to and from the compute instance. All VNICs within a given subnet use the same routing table, security list, and DHCP options. As described above, each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24), which do not overlap with other subnets in the VCN and represent a subset of the address space within the VCN's address space. For a VNIC on a particular subnet of a VCN, the private overlay IP address assigned to the VNIC is an address from the contiguous range of overlay IP addresses assigned to the subnet.

[0068] In some embodiments, in addition to the private secondary IP address, a compute instance can optionally be assigned additional secondary IP addresses, such as, for example, one or more public IP addresses if in a public subnet. These addresses are assigned either on the same VNIC or on multiple VNICs associated with the compute instance. However, each instance has a primary VNIC, which is created during instance launch and is associated with the secondary private IP address assigned to the instance - this primary VNIC cannot be deleted. Additional VNICs, called secondary VNICs, can be added to an existing instance in the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. A secondary VNIC can be in a subnet in the same VCN as the primary VNIC, or in a different subnet in the same VCN or a different VCN.

[0069] If a compute instance is in a public subnet, it can optionally be assigned a public IP address. When creating a subnet, the subnet can be designated as either a public subnet or a private subnet. A private subnet means that resources (e.g., compute instances) in the subnet and the associated VNICs cannot have public secondary IP addresses. A public subnet means that resources and associated VNICs in the subnet can have public IP addresses. A customer can specify that the subnet exists in a single availability domain or across multiple availability domains in a region or realm.

[0070] As described above, a VCN can be subdivided into one or more subnets. In some embodiments, a virtual router (VR) configured for the VCN (referred to as the VCN VR or simply the VR) enables communication between the subnets of the VCN. For subnets within a VCN, the VR represents the logical gateway for that subnet, which enables the subnet (i.e., compute instances on that subnet) to communicate with endpoints on other subnets within the VCN as well as endpoints outside the VCN. The VCN VR is a logical entity that is configured to route traffic between VNICs in the VCN and virtual gateways ("gateways") associated with the VCN. Below regarding Figure 1Describe the gateway further. VCN VR is a Layer 3 / IP layer concept. In one embodiment, there is a VCN VR for a VCN, where the VCN VR has a potentially unlimited number of ports addressed by IP addresses, with one port for each subnet of the VCN. In this way, the VCN VR has a different IP address for each subnet in the VCN to which the VCN VR is attached. The VR is also connected to various gateways configured for the VCN. In some embodiments, a specific overlay IP address within the overlay IP address range for a subnet is reserved for the port of the VCN VR for that subnet. For example, consider a VCN with two subnets, and the associated address ranges are 10.0 / 16 and 10.1 / 16 respectively. For the first subnet in the VCN with the address range 10.0 / 16, the addresses within this range are reserved for the ports of the VCN VR for that subnet. In some cases, the first IP address within the range can be reserved for the VCN VR. For example, for a subnet with an overlay IP address range of 10.0 / 16, the IP address 10.0.0.1 can be reserved for the port of the VCN VR for that subnet. For the second subnet in the same VCN with the address range 10.1 / 16, the VCN VR can have a port for the second subnet with the IP address 10.1.0.1. The VCN VR has a different IP address for each subnet in the VCN.

[0071] In some other embodiments, each subnet within a VCN can have its own associated VR, which can be addressed by the subnet using a reserved or default IP address associated with the VR. For example, the reserved or default IP address can be the first IP address within the IP address range associated with that subnet. The VNICs within the subnet can use this default or reserved IP address to communicate (e.g., send and receive data packets) with the VR associated with the subnet. In such an embodiment, the VR is the ingress / egress point for that subnet. The VRs associated with subnets within a VCN can communicate with other VRs associated with other subnets within the VCN. The VR can also communicate with the gateways associated with the VCN. The VR functionality for a subnet runs on or is performed by one or more NVDs that perform VNIC functionality for the VNICs within the subnet.

[0072] A route table, security rules, and DHCP options can be configured for the VCN. The route table is a virtual route table for the VCN and includes rules for routing traffic from subnets within the VCN to destinations outside the VCN via gateways or specially configured instances. The route table for the VCN can be customized to control how data packets are forwarded / routed into and out of the VCN. The DHCP options refer to the configuration information automatically provided to an instance when it is launched.

[0073] The security rules configured for a VCN represent an overlay firewall rule for the VCN. The security rules can include ingress and egress rules and specify the types of traffic (e.g., based on protocol and port) that are allowed to flow in and out of the instances within the VCN. The customer can choose whether a given rule is stateful or stateless. For example, the customer can allow incoming SSH traffic to a set of instances from anywhere by setting a stateful ingress rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22. The security rules can be implemented using network security groups or security lists. A network security group consists of a collection of security rules that apply only to the resources within that group. On the other hand, a security list includes rules that apply to all resources in any subnet that uses that security list. A default security list with default security rules can be provided for the VCN. The DHCP options configured for the VCN provide the configuration information that is automatically provided to the instances within the VCN when they are launched.

[0074] In some embodiments, the configuration information for a VCN is determined and stored by the VCN control plane. For example, the configuration information for a VCN can include information about: the address range associated with the VCN, the subnets within the VCN and associated information, one or more VRs associated with the VCN, the compute instances in the VCN and associated VNICs, the NVDs (e.g., VNICs, VRs, gateways) that perform various virtualized network functions associated with the VCN, the status information for the VCN, and other VCN-related information. In some embodiments, the VCN distribution service publishes the configuration information stored by the VCN control plane or a portion thereof to the NVD. The distributed information can be used to update the information (e.g., forwarding tables, routing tables, etc.) stored and used by the NVD to forward data packets to and from the compute instances in the VCN.

[0075] In some embodiments, the creation of the VCN and subnets is handled by the VCN control plane (CP) and the launching of compute instances is handled by the compute control plane. The compute control plane is responsible for allocating physical resources for the compute instances and then calls the VCN control plane to create a VNIC and attach it to the compute instance. The VCN CP also sends the VCN data map to the VCN data plane that is configured to perform packet forwarding and routing functions. In some embodiments, the VCN CP provides a distribution service that is responsible for providing updates to the VCN data plane. Examples of the VCN control plane are also depicted in Fig.12 、 Fig.13 、 Fig.14 and Figure 15 (see reference numerals 1216, 1316, 1416, and 1516) and are described below.

[0076] Customers can create one or more VCNs using resources hosted by CSPI. Compute instances deployed on the customer VCN can communicate with different endpoints. These endpoints can include endpoints hosted by CSPI and endpoints external to CSPI.

[0077] Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 12-16 Various different architectures for implementing cloud-based services using CSPI are depicted in ,

[0077] , Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 12-16 and are described below. Figure 1 FIG. Figure 1 is a high-level diagram of a distributed environment 100 showing an overlay or customer VCN hosted by CSPI according to certain embodiments. Figure 1 The distributed environment depicted in FIG. Figure 1 includes multiple components in an overlay network. Figure 1 The distributed environment 100 depicted in FIG. Figure 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, alternatives, and modifications are possible. For example, in some implementations, Figure 1 the distributed environment depicted in FIG. Figure 1 can have more or fewer systems or components than shown in Figure 1 FIG. Figure 1 , can combine two or more systems, or can have a different system configuration or arrangement.

[0078] As shown in the example depicted in Figure 1 FIG. Figure 1 , the distributed environment 100 includes CSPI 101 that provides services and resources that customers can subscribe to and use to build their virtual cloud network (VCN). In certain embodiments, CSPI 101 provides IaaS services to subscribing customers. Data centers within CSPI 101 can be organized into one or more regions. Figure 1 An example region "Region US" 102 is shown in FIG. Figure 1 . The customer has configured a customer VCN c / o Oracle International Corporation for region 102. The customer can deploy various compute instances on VCN 104, where the compute instances can include virtual machines or bare metal instances. Examples of instances include applications, databases, load balancers, etc.

[0079] In Figure 1 the embodiment depicted in FIG. Figure 1 , the customer VCN 104 includes two subnets, namely, "Subnet-1" and "Subnet-2", each with its own CIDR IP address range. In Figure 1Among them, the covered IP address range of Subnet-1 is 10.0 / 16, and the address range of Subnet-2 is 10.1 / 16. The VCN virtual router 105 represents the logical gateway for the VCN, which enables communication between the subnets of VCN104 and other endpoints outside the VCN. The VCN VR 105 is configured to route traffic between the VNICs in VCN 104 and the gateways associated with VCN 104. The VCN VR 105 provides ports for each subnet of VCN 104. For example, VR 105 can provide a port with the IP address 10.0.0.1 for Subnet-1 and a port with the IP address 10.1.0.1 for Subnet-2.

[0080] Multiple computing instances can be deployed on each subnet, where the computing instances can be virtual machine instances and / or bare metal instances. The computing instances in the subnet can be hosted by one or more host machines within CSPI 101. The computing instances participate in the subnet via the VNIC associated with the computing instance. For example, as Figure 1 shown, the computing instance C1 becomes part of Subnet-1 via the VNIC associated with the computing instance. Similarly, the computing instance C2 becomes part of Subnet-1 via the VNIC associated with C2. In a similar manner, multiple computing instances (which can be virtual machine instances or bare metal instances) can be part of Subnet-1. Via its associated VNIC, each computing instance is assigned a private covered IP address and a MAC address. For example, in Figure 1 it is shown that the covered IP address of the computing instance C1 is 10.0.0.2 and the MAC address is M1, while the private covered IP address of the computing instance C2 is 10.0.0.3 and the MAC address is M2. Each computing instance in Subnet-1 (including computing instances C1 and C2) has a default route to the VCN VR 105 using the IP address 10.0.0.1, which is the IP address of the port of the VCN VR 105 for Subnet-1.

[0081] Multiple computing instances can be deployed on Subnet-2, including virtual machine instances and / or bare metal instances. For example, as Figure 1 shown, the computing instances D1 and D2 become part of Subnet-2 via the VNICs associated with the respective computing instances. In the Figure 1 embodiment shown, the covered IP address of the computing instance D1 is 10.1.0.2 and the MAC address is MM1, while the private covered IP address of the computing instance D2 is 10.1.0.3 and the MAC address is MM2. Each computing instance in Subnet-2 (including computing instances D1 and D2) has a default route to the VCN VR 105 using the IP address 10.1.0.1, which is the IP address of the port of the VCN VR 105 for Subnet-2.

[0082] VCN A 104 may also include one or more load balancers. For example, a load balancer may be provided for a subnet and configured to load balance traffic across multiple compute instances on the subnet. A load balancer may also be provided to load balance traffic across subnets within a VCN.

[0083] A particular compute instance deployed on VCN 104 may communicate with a variety of different endpoints. These endpoints may include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 101 may include: endpoints on the same subnet as the particular compute instance (e.g., communication between two compute instances in Subnet-1); endpoints on different subnets but within the same VCN (e.g., communication between a compute instance in Subnet-1 and a compute instance in Subnet-2); endpoints in different VCNs within the same region (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN within the same Region 106 or 110, communication between a compute instance in Subnet-1 and an endpoint in a service point 110 within the same region); or endpoints in VCNs in different regions (e.g., communication between a compute instance in Subnet-1 and an endpoint in a VCN in a different Region 108). Compute instances in subnets hosted by CSPI 101 may also communicate with endpoints not hosted by CSPI 101 (i.e., external to CSPI 101). These external endpoints include endpoints in the customer's on-premises network 116, endpoints in other remote cloud-hosted networks 118, public endpoints 114 accessible via a public network such as the Internet, and other endpoints.

[0084] Use the VNICs associated with the source compute instance and the destination compute instance to facilitate communication between compute instances on the same subnet. For example, a compute instance C1 in subnet-1 may want to send a packet to a compute instance C2 in subnet-1. For a packet originating from a source compute instance and destined for another compute instance in the same subnet, the packet is first processed by the VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance can include determining the destination information of the packet from the packet header, identifying any policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the next hop for the packet, performing any packet encapsulation / decapsulation functions as needed, and then forwarding / routing the packet to the next hop with the aim of facilitating the communication of the packet to its intended destination. When the destination compute instance is in the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing. The VNIC associated with the destination compute instance then processes the packet and forwards the packet to the destination compute instance.

[0085] For packets to be transmitted from a compute instance in a subnet to an endpoint in a different subnet within the same VCN, communication is facilitated through the VNICs associated with the source and destination compute instances and the VCN VR. For example, if Figure 1 compute instance C1 in subnet-1 in wants to send a packet to compute instance D1 in subnet-2, then the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR 105 using the default route or port 10.0.0.1 of the VCN VR. VCN VR 105 is configured to route the packet to subnet-2 using port 10.1.0.1. Then, the VNIC associated with D1 receives and processes the packet and the VNIC forwards the packet to compute instance D1.

[0086] For packets to be sent from a compute instance in the VCN 104 to an endpoint outside the VCN 104, the communication is facilitated by the VNIC associated with the source compute instance, the VCN VR 105, and the gateway associated with the VCN 104. One or more types of gateways can be associated with the VCN 104. A gateway is an interface between the VCN and another endpoint, where that other endpoint is outside the VCN. A gateway is a layer 3 / IP layer concept and enables the VCN to communicate with endpoints outside the VCN. Thus, the gateway facilitates the flow of traffic between the VCN and other VCNs or networks. Various different types of gateways can be configured for the VCN to facilitate different types of communication with different types of endpoints. Depending on the gateway, the communication can occur over a public network (e.g., the Internet) or over a private network. Various communication protocols can be used for these communications.

[0087] For example, compute instance C1 may want to communicate with an endpoint outside the VCN 104. The packet can first be processed by the VNIC associated with the source compute instance C1. The VNIC processing determines that the destination of the packet is outside C1's subnet-1. The VNIC associated with C1 can forward the packet to the VCN VR 105 for the VCN 104. The VCN VR 105 then processes the packet and, as part of the processing, determines a specific gateway associated with the VCN 104 as the next hop for the packet based on the destination of the packet. The VCN VR 105 can then forward the packet to the specific identified gateway. For example, if the destination is an endpoint within the customer's on-premises network, then the packet can be forwarded by the VCN VR 105 to the Dynamic Routing Gateway (DRG) gateway 122 configured for the VCN 104. The packet can then be forwarded from the gateway to the next hop to facilitate the delivery of the packet to its final intended destination.

[0088] Various different types of gateways can be configured for the VCN. Examples of gateways that can be configured for the VCN are depicted in Figure 1 and described below. Examples of gateways associated with the VCN are also depicted in Fig.12 、 Fig.13 、 Fig.14 and FIG. 15 (e.g., gateways referenced by reference numerals 1234, 1236, 1238, 1334, 1336, 1338, 1434, 1436, 1438, 1534, 1536, and 1538) and are described as follows. As Figure 1As shown in the embodiments depicted, a Dynamic Routing Gateway (DRG) 122 can be added to or associated with a customer VCN 104 and provide a path for private network traffic communication between the customer VCN 104 and another endpoint, where the other endpoint can be the customer's on-premises network 116, a VCN 108 in a different region of the CSPI 101, or another remote cloud network 118 not hosted by the CSPI 101. The customer on-premises network 116 can be a customer network or customer data center built using the customer's resources. Access to the customer on-premises network 116 is generally very restricted. For customers who have both a customer on-premises network 116 and one or more VCNs 104 deployed or hosted by the CSPI 101 in the cloud, the customer may want their on-premises network 116 and their cloud-based VCN 104 to be able to communicate with each other. This enables the customer to build an extended hybrid environment that includes the customer's VCN 104 hosted by the CSPI 101 and their on-premises network 116. The DRG 122 enables this communication. To enable such communication, a communication channel 124 is set up, where one endpoint of the channel is in the customer on-premises network 116 and the other endpoint is in the CSPI 101 and connected to the customer VCN 104. The communication channel 124 can be through a public communication network (such as the Internet) or a private communication network. Various different communication protocols can be used, such as IPsec VPN technology over a public communication network (such as the Internet), Oracle's FastConnect technology that uses a private network instead of a public network, etc. The device or equipment that forms one endpoint of the communication channel 124 in the customer on-premises network 116 is referred to as customer-premises equipment (CPE), such as Figure 1 the CPE 126 depicted in

[0089] In some embodiments, a Remote Peering Connection (RPC) can be added to the DRG, which allows the customer to peer one VCN with another VCN in a different region. Using this RPC, the customer VCN 104 can be connected to the VCN 108 in another region using the DRG 122. The DRG 122 can also be used to communicate with other remote cloud networks 118 not hosted by the CSPI 101 (such as the Microsoft Azure cloud, the Amazon AWS cloud, etc.).

[0090] As Figure 1As shown, an Internet Gateway (IGW) 120 can be configured for the customer VCN 104, which enables computing instances on the VCN 104 to communicate with public endpoints 114 that can be accessed via a public network such as the Internet. The IGW 120 is a gateway that connects the VCN to a public network such as the Internet. The IGW 120 enables public subnets within the VCN (such as VCN 104), where resources in the public subnets have public covering IP addresses, to directly access public endpoints 112 on the public network 114 (such as the Internet). Using the IGW 120, connections can be initiated from subnets within the VCN 104 or from the Internet.

[0091] A Network Address Translation (NAT) Gateway 128 can be configured for the customer's VCN 104 and enables cloud resources in the customer's VCN that do not have dedicated public covering IP addresses to access the Internet and do so without exposing those resources to direct incoming Internet connections (e.g., L4-L7 connections). This enables private subnets within the VCN (such as Private Subnet-1 in VCN 104) to privately access public endpoints on the Internet. In a NAT gateway, connections to the public Internet can only be initiated from the private subnets and not from the Internet to the private subnets.

[0092] In some embodiments, a Service Gateway (SGW) 126 can be configured for the customer VCN 104 and provides a path for private network traffic between the VCN 104 and service endpoints supported in the service network 110. In some embodiments, the service network 110 can be provided by a CSP and can provide various services. An example of such a service network is Oracle's service network, which provides various services available to customers. For example, a computing instance (e.g., a database system) in a private subnet of the customer VCN 104 can back up data to a service endpoint (e.g., Object Storage) without a public IP address or access to the Internet. In some embodiments, a VCN can have only one SGW, and the connection can only be initiated from subnets within the VCN and not from the service network 110. If a VCN is peered with another, resources in the other VCN generally cannot access the SGW. Resources in an on-premises network connected to the VCN using FastConnect or VPN Connect can also use the service gateway configured for that VCN.

[0093] In some embodiments, the SGW 126 uses the concept of a service Classless Inter-Domain Routing (CIDR) label, which is a string representing all the regional public IP address ranges for a service or group of services of interest. Customers use the service CIDR label when they configure the SGW and associated routing rules to control traffic to the service. If the public IP address of the service changes in the future, then customers can optionally use it when configuring security rules without having to adjust them.

[0094] A Local Peering Gateway (LPG) 132 is a gateway that can be added to a customer VCN 104 and enables the VCN 104 to peer with another VCN in the same region. Peering means that the VCNs communicate using private IP addresses and traffic does not need to cross a public network (such as the Internet) or route traffic through the customer's on-premises network 116. In the preferred embodiment, the VCN has a separate LPG for each peer it establishes. Local peering or VCN peering is a common practice for establishing network connectivity between different applications or infrastructure management functions.

[0095] Service providers (such as providers of services in the service network 110) can provide access to services using different access models. According to the public access model, the service can be exposed as a public endpoint that can be publicly accessed by compute instances in the customer VCN via a public network (such as the Internet), and / or can be privately accessed via the SGW 126. According to a specific private access model, the service can be accessed as a private IP endpoint in a private subnet within the customer's VCN. This is referred to as Private Endpoint (PE) access and enables the service provider to expose its service as an instance within the customer's private network. A private endpoint resource represents a service within the customer's VCN. Each PE appears as a VNIC (referred to as a PE-VNIC, having one or more private IPs) in a subnet selected by the customer within the customer's VCN. Thus, the PE provides a way to present the service in a private customer VCN subnet using a VNIC. Since the endpoint is exposed as a VNIC, all the features associated with the VNIC (such as routing rules, security lists, etc.) can now be used for the PE VNIC.

[0096] Service providers can register their services to enable access through the PE. The provider can associate policies with the service, which limits the visibility of the service to the customer's tenancy. The provider can register multiple services under a single Virtual IP Address (VIP), especially for multi-tenant services. There can be multiple such private endpoints (in multiple VCNs) representing the same service.

[0097] Compute instances in the private subnet can then access the service using the private IP address of the PE VNIC or the service DNS name. Compute instances in the customer VCN can access the service by sending traffic to the private IP address of the PE in the customer VCN. The Private Access Gateway (PAGW) 130 is a gateway resource that can be attached to a service provider VCN (e.g., a VCN in the service network 110), which serves as the ingress / egress point for all traffic to / from the private endpoints of the customer subnets. The PAGW 130 enables the provider to scale the number of PE connections without utilizing its internal IP address resources. The provider only needs to configure one PAGW for any number of services registered in a single VCN. The provider can represent a service as private endpoints in multiple VCNs of one or more customers. From the customer's perspective, the PE VNIC is not attached to the customer's instance but appears to be attached to the service with which the customer wishes to interact. Traffic destined for the private endpoint is routed to the service via the PAGW 130. These are referred to as customer-to-service private connections (C2S connections).

[0098] By allowing traffic to flow through the FastConnect / IPsec link and the private endpoints in the customer VCN, the PE concept can also be used to extend private access for services to the customer's on-premises networks and data centers. By allowing traffic to flow between the LPG 132 and the PE in the customer's VCN, private access to the service can also be extended to the customer's peered VCNs.

[0099] The customer can control routing within the VCN at the subnet level, so the customer can specify which subnets in the customer's VCN (such as VCN 104) use each gateway. The routing table of the VCN is used to decide whether to allow traffic to leave the VCN through a specific gateway. For example, in a specific instance, the routing table for the public subnet within the customer VCN 104 can send non-local traffic through the IGW 120. The routing table for the private subnet within the same customer VCN 104 can send traffic destined for the CSP service through the SGW 126. All remaining traffic can be sent via the NAT gateway 128. The routing table only controls traffic flowing out of the VCN.

[0100] The security list associated with the VCN is used to control the traffic entering the VCN via the gateway through the inbound connection. All resources in the subnet use the same route table and security list. The security list can be used to control the specific type of traffic allowed to and from the instances in the subnet of the VCN. The security list rules can include ingress (inbound) and egress (outbound) rules. For example, the ingress rule can specify the allowed source address range, while the egress rule can specify the allowed destination address range. The security rules can specify a specific protocol (e.g., TCP, ICMP), a specific port (e.g., 22 for SSH, 3389 for Windows RDP), etc. In some embodiments, the operating system of the instance can enforce its own firewall rules that comply with the security list rules. The rules can be stateful (e.g., tracking the connection and automatically allowing the response without an explicit security list rule for the response traffic) or stateless.

[0101] Access from the customer VCN (i.e., through resources or compute instances deployed on VCN 104) can be classified as public access, private access, or dedicated access. Public access refers to an access model that uses a public IP address or NAT to access a public endpoint. Private access enables customer workloads with private IP addresses in VCN 104 (e.g., resources in a private subnet) to access services without traversing a public network such as the Internet. In some embodiments, CSPI 101 enables customer VCN workloads with private IP addresses to access the public service endpoints of services using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer's VCN and the public endpoints of the services residing outside the customer's private network.

[0102] In addition, CSPI can provide dedicated public access using technologies such as FastConnect public peering, where customer on-premises instances can use FastConnect connections to access one or more services in the customer VCN without traversing a public network such as the Internet. CSPI can also provide dedicated private access using FastConnect private peering, where customer on-premises instances with private IP addresses can use FastConnect connections to access the customer's VCN workloads. FastConnect is a network connectivity alternative to using the public Internet to connect the customer's on-premises network to CSPI and its services. Compared with Internet-based connections, FastConnect provides a simple, flexible, and cost-effective way to create dedicated and private connections with higher bandwidth options and a more reliable and consistent network experience.

[0103] Figure 1The above and accompanying description describes various virtualized components in an example virtual network. As described above, the virtual network is built on an underlying physical or substrate network. Figure 2 FIG. 2 depicts a simplified architecture diagram of physical components in a physical network within CSPI 200 that provides the underlying for a virtual network according to certain embodiments. As shown, CSPI 200 provides a distributed environment that includes components and resources (e.g., computing, memory, and network resources) provided by a cloud service provider (CSP). These components and resources are used to provide cloud services (e.g., IaaS services) to subscribing customers (i.e., customers who have subscribed to one or more services provided by the CSP). Based on the services subscribed to by the customer, a subset of the resources of CSPI 200 (e.g., computing, memory, and network resources) are provisioned for the customer. The customer can then use the physical computing, memory, and networking resources provided by CSPI 200 to build their own cloud-based (i.e., CSPI-hosted) customizable and private virtual network. As previously indicated, these customer networks are referred to as virtual cloud networks (VCNs). The customer can deploy one or more customer resources, such as computing instances, on these customer VCNs. The computing instances can be in the form of virtual machines, bare metal instances, etc. CSPI 200 provides an infrastructure and a collection of complementary cloud services that enable the customer to build and run a wide range of applications and services in a highly available hosted environment.

[0104] In Figure 2 the example embodiment depicted in FIG. 2, the physical components of CSPI 200 include one or more physical host machines or physical servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), and a physical network (e.g., 218), as well as switches in the physical network 218. The physical host machines or servers can host and execute various computing instances that participate in one or more subnets of the VCN. The computing instances can include virtual machine instances and bare metal instances. For example, Figure 1 the various computing instances depicted in FIG. 2 can be hosted by Figure 2 the physical host machines depicted in FIG. 2. The virtual machine computing instances in the VCN can be executed by one host machine or multiple different host machines. The physical host machines can also host virtual host machines, container-based hosts, functions, etc. Figure 1 The VNICs and VCN VRs depicted in FIG. 2 can be executed by Figure 2 the NVDs depicted in FIG. 2. Figure 1 The gateways depicted in FIG. 2 can be executed by Figure 2 the host machines and / or NVDs described in FIG. 2.

[0105] A host machine or server can execute a hypervisor (also known as a virtual machine monitor or VMM) that creates and enables a virtualized environment on the host machine. Virtualization or the virtualized environment facilitates cloud-based computing. One or more computing instances can be created, executed, and managed by the hypervisor on the host machine. The hypervisor on the host machine enables the physical computing resources of the host machine (e.g., computing, memory, and network resources) to be shared among the various computing instances executed by the host machine.

[0106] For example, as Figure 2 depicted in, host machines 202 and 208 execute hypervisors 260 and 266 respectively. These hypervisors can be implemented using software, firmware, hardware, or a combination thereof. Generally, a hypervisor is a process or software layer that resides above the operating system (OS) of the host machine, and the OS in turn executes on the hardware processor of the host machine. The hypervisor provides a virtualized environment by enabling the physical computing resources of the host machine (e.g., processing resources such as processors / cores, memory resources, networking resources) to be shared among the various virtual machine computing instances executed by the host machine. For example, in Figure 2 , hypervisor 260 can reside above the OS of host machine 202 and enable the computing resources of host machine 202 (e.g., processing, memory, and networking resources) to be shared among the computing instances (e.g., virtual machines) executed by host machine 202. A virtual machine can have its own operating system (referred to as a guest operating system), which can be the same as or different from the OS of the host machine. The operating system of a virtual machine executed by a host machine can be the same as or different from the operating system of another virtual machine executed by the same host machine. Thus, the hypervisor enables multiple operating systems to be executed simultaneously while sharing the same computing resources of the host machine. Figure 2 The host machines depicted in can have the same or different types of hypervisors.

[0107] A computing instance can be a virtual machine instance or a bare-metal instance. In Figure 2 , computing instances 268 on host machine 202 and computing instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare-metal instance provided to a customer.

[0108] In some cases, an entire host machine can be provisioned to a single customer, and one or more compute instances (either virtual machines or bare metal instances) hosted by that host machine all belong to the same customer. In other cases, the host machine can be shared among multiple customers (i.e., multiple tenants). In such a multi-tenant scenario, the host machine can host virtual machine compute instances belonging to different customers. These compute instances can be members of different VCNs of different customers. In some embodiments, bare metal compute instances are hosted by bare metal servers without a hypervisor. When provisioning a bare metal compute instance, a single customer or tenant maintains control over the physical CPUs, memory, and network interfaces of the host machine hosting the bare metal instance, and the host machine is not shared with other customers or tenants.

[0109] As previously described, each compute instance that is part of a VCN is associated with a VNIC that enables the compute instance to be a member of a subnet of the VCN. The VNIC associated with a compute instance facilitates the communication of data packets or frames to and from the compute instance. The VNIC is associated with the compute instance when the compute instance is created. In some embodiments, for a compute instance executed by a host machine, the VNIC associated with the compute instance is executed by an NVD connected to the host machine. For example, in Figure 2 the embodiment depicted in, host machine 202 executes virtual machine compute instance 268 associated with VNIC 276, and VNIC 276 is executed by NVD 210 connected to host machine 202. As another example, bare metal instance 272 hosted by host machine 206 is associated with VNIC 280 executed by NVD 212 connected to host machine 206. As yet another example, VNIC 284 is associated with compute instance 274 executed by host machine 208, and VNIC 284 is executed by NVD 212 connected to host machine 208.

[0110] For a compute instance hosted by a host machine, the NVD connected to that host machine also executes a VCN VR corresponding to the VCN of which the compute instance is a member. For example, in the embodiment depicted in Figure 2 , NVD 210 executes VCN VR 277 corresponding to the VCN of which compute instance 268 is a member. NVD 212 can also execute one or more VCN VRs 283 corresponding to the VCNs corresponding to the compute instances hosted by host machines 206 and 208.

[0111] The host machine may include one or more network interface cards (NICs) that enable the host machine to connect to other devices. The NICs on the host machine may provide one or more ports (or interfaces) that enable the host machine to communicate and connect to another device. For example, the host machine may connect to an NVD using one or more ports (or interfaces) provided on the host machine and on the NVD. The host machine may also connect to other devices (such as another host machine).

[0112] For example, in Figure 2 , the host machine 202 uses the link 220 to connect to the NVD 210, and the link 220 extends between the port 234 provided by the NIC 232 of the host machine 202 and the port 236 of the NVD 210. The host machine 206 uses the link 224 to connect to the NVD 212, and the link 224 extends between the port 246 provided by the NIC 244 of the host machine 206 and the port 248 of the NVD 212. The host machine 208 uses the link 226 to connect to the NVD 212, and the link 226 extends between the port 252 provided by the NIC 250 of the host machine 208 and the port 254 of the NVD 212.

[0113] The NVDs are in turn connected to top-of-rack (TOR) switches via communication links, and these switches are connected to the physical network 218 (also referred to as the switch fabric). In some embodiments, the links between the host machines and the NVDs and between the NVDs and the TOR switches are Ethernet links. For example, in Figure 2 , the NVDs 210 and 212 use the links 228 and 230 respectively to connect to the TOR switches 214 and 216. In some embodiments, the links 220, 224, 226, 228, and 230 are Ethernet links. The set of host machines and NVDs connected to the TOR is sometimes referred to as a rack.

[0114] The physical network 218 provides a communication fabric that enables the TOR switches to communicate with each other. The physical network 218 may be a multi-layer network. In some implementations, the physical network 218 is a multi-layer Clos network of switches, where the TOR switches 214 and 216 represent the leaf-level nodes of the multi-layer and multi-node physical switching network 218. Different Clos network configurations are possible, including but not limited to 2-layer networks, 3-layer networks, 4-layer networks, 5-layer networks, and general "n"-layer networks. Examples of Clos networks are depicted in Figure 5 and described below.

[0115] There may be various different connection configurations between the host machines and the NVDs, such as one-to-one configurations, many-to-one configurations, one-to-many configurations, etc. In a one-to-one configuration implementation, each host machine is connected to its own separate NVD. For example, in Figure 2 In the figure, the host machine 202 is connected to the NVD 210 via the NIC 232 of the host machine 202. In a multi-to-one configuration, multiple host machines are connected to one NVD. For example, in Figure 2 the figure, the host machines 206 and 208 are respectively connected to the same NVD 212 via NICs 244 and 250.

[0116] In a one-to-many configuration, one host machine is connected to multiple NVDs. Figure 3 An example within the CSPI 300 is shown, where a host machine is connected to multiple NVDs. As Figure 3 shown in the figure, the host machine 302 includes a network interface card (NIC) 304, which includes multiple ports 306 and 308. The host machine 300 is connected to the first NVD 310 via port 306 and link 320, and is connected to the second NVD 312 via port 308 and link 322. Ports 306 and 308 can be Ethernet ports and the links 320 and 322 between the host machine 302 and the NVDs 310 and 312 can be Ethernet links. The NVD 310 is further connected to the first TOR switch 314 and the NVD 312 is connected to the second TOR switch 316. The links between the NVDs 310 and 312 and the TOR switches 314 and 316 can be Ethernet links. The TOR switches 314 and 316 represent layer 0 switching devices in the multi-layer physical network 318.

[0117] Figure 3 The arrangement depicted in the figure provides two separate physical network paths between the physical switch network 318 and the host machine 302: the first path passes through the TOR switch 314 to the NVD 310 and then to the host machine 302, and the second path passes through the TOR switch 316 to the NVD 312 and then to the host machine 302. The separate paths provide enhanced availability (referred to as high availability) for the host machine 302. If there is a problem with one of the paths (e.g., a link in one of the paths is broken) or one of the devices (e.g., a particular NVD is not operating), then the other path can be used for communication with the host machine 302.

[0118] In Figure 3 the configuration depicted in the figure, the host machine uses two different ports provided by the NIC of the host machine to connect to two different NVDs. In other embodiments, the host machine can include multiple NICs that enable the host machine to connect to multiple NVDs.

[0119] Referring back to Figure 2, an NVD is a physical device or component that performs one or more network and / or storage virtualization functions. An NVD can be any device having one or more processing units (e.g., CPU, network processing unit (NPU), FPGA, packet processing pipeline, etc.), memory (including caches), and ports. Various virtualization functions can be performed by software / firmware executed by one or more processing units of the NVD.

[0120] The NVD can be implemented in various different forms. For example, in some embodiments, the NVD is implemented as an interface card called a smartNIC or a smart NIC with an on-board embedded processor. A smartNIC is a device independent of the NIC on the host machine. In Figure 2 , the NVDs 210 and 212 can be implemented as smartNICs respectively connected to the host machine 202 and the host machines 206 and 208.

[0121] However, the smartNIC is just one example of an NVD implementation. Various other implementations are possible. For example, in some other embodiments, the NVD or one or more functions performed by the NVD can be incorporated into or performed by one or more host machines, one or more TOR switches, and other components of the CSPI 200. For example, the NVD can be implemented in a host machine, where the functions performed by the NVD are performed by the host machine. As another example, the NVD can be part of a TOR switch, or the TOR switch can be configured to perform the functions performed by the NVD, which enables the TOR switch to perform various complex packet conversions for a public cloud. A TOR that performs the functions of an NVD is sometimes referred to as a smart TOR. In other embodiments that provide virtual machine (VM) instances rather than bare metal (BM) instances to customers, the functions performed by the NVD can be implemented inside the hypervisor of the host machine. In some other embodiments, some of the functions of the NVD can be offloaded to a centralized service running on a cluster of host machines.

[0122] In certain embodiments, such as when implemented as a smartNIC as shown in Figure 2 , the NVD can include multiple physical ports that enable it to connect to one or more host machines and one or more TOR switches. The ports on the NVD can be classified as host-facing ports (also referred to as "south ports") or network-facing or TOR-facing ports (also referred to as "north ports"). The host-facing ports of the NVD are the ports used to connect the NVD to the host machine. Figure 2 Examples of the host-facing ports in Figure 2 Examples of network-facing ports include port 256 on NVD 210 and port 258 on NVD 212. As Figure 2 shown, NVD 210 is connected to TOR switch 214 using link 228 that extends from port 256 of NVD 210 to TOR switch 214. Similarly, NVD 212 is connected to TOR switch 216 using link 230 that extends from port 258 of NVD 212 to TOR switch 216.

[0123] The NVD receives packets and frames from the host machine via the host-facing port (e.g., packets and frames generated by a compute instance hosted by the host machine), and after performing necessary packet processing, can forward the packets and frames to the TOR switch via the network-facing port of the NVD. The NVD can receive packets and frames from the TOR switch via the network-facing port of the NVD, and after performing necessary packet processing, can forward the packets and frames to the host machine via the host-facing port of the NVD.

[0124] In some embodiments, there can be multiple ports and associated links between the NVD and the TOR switch. These ports and links can be aggregated to form a link aggregation group (referred to as a LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and the TOR switch) to be treated as a single logical link. All physical links within a given LAG can operate at the same speed in full-duplex mode. LAG helps increase the bandwidth and reliability of the connection between two endpoints. If one of the physical links within the LAG fails, then traffic will be dynamically and transparently re-assigned to one of the other physical links within the LAG. The aggregated physical links deliver a higher bandwidth than each individual link. The multiple ports associated with the LAG are treated as a single logical port. Traffic can be load balanced across the multiple physical links of the LAG. One or more LAGs can be configured between two endpoints. These two endpoints can be located between the NVD and the TOR switch, between the host machine and the NVD, and so on.

[0125] The NVD implements or executes network virtualization functions. These functions are executed by the software / firmware executed by the NVD. Examples of network virtualization functions include, but are not limited to: packet encapsulation and decapsulation functions; functions for creating VCN networks; functions for implementing network policies, such as VCN security list (firewall) functionality; functions for facilitating packet routing and forwarding to and from compute instances in the VCN; and so on. In some embodiments, upon receiving a packet, the NVD is configured to execute a packet processing pipeline to process the packet and determine how to forward or route the packet. As part of this packet processing pipeline, the NVD may execute one or more virtual functions associated with the overlay network, such as executing a VNIC associated with a compute instance in the VCN, executing a virtual router (VR) associated with the VCN, encapsulating and decapsulating packets to facilitate forwarding or routing in the virtual network, executing certain gateways (e.g., local peer gateways), implementing security lists, network security groups, network address translation (NAT) functionality (e.g., translating public IPs to private IPs on a per-host basis), throttling functions, and other functions.

[0126] In some embodiments, the packet processing data path in the NVD may include multiple packet pipelines, each pipeline consisting of a series of packet transformation stages. In some implementations, upon receiving a packet, the packet is parsed and classified into a single pipeline. The packet is then processed linearly, stage by stage, until the packet is either discarded or sent out through an interface of the NVD. These stages provide basic functional packet processing building blocks (e.g., validating headers, enforcing throttling, inserting new layer 2 headers, enforcing L4 firewalls, VCN encapsulation / decapsulation, etc.) so that new pipelines can be constructed by combining existing stages, and new functionality can be added by creating new stages and inserting them into existing pipelines.

[0127] The NVD may execute control plane and data plane functions corresponding to the control plane and data plane of the VCN. Examples of the VCN control plane are also depicted in Fig.12 , Fig.13 , Fig.14 and FIG. 15 (see reference numerals 1216, 1316, 1416, and 1516) and are described below. Examples of the VCN data plane are in Fig.12 , Fig.13 , Fig.14Depicted in FIGS. 15 (see reference numerals 1218, 1318, 1418, and 1518) and described below. Control plane functions include functions for configuring the network for how control data is forwarded (e.g., setting up routing and routing tables, configuring VNICs, etc.). In some embodiments, a VCN control plane is provided that centrally computes all overlay-to-substrate mappings and publishes them to the NVD and virtual network edge devices (such as various gateways, such as DRG, SGW, IGW, etc.). Firewall rules can also be published using the same mechanism. In some embodiments, the NVD only obtains the mappings relevant to that NVD. Data plane functions include the function of actually routing / forwarding data packets based on the configuration set using the control plane. The VCN data plane is implemented by encapsulating the customer's network packets before they pass through the substrate network. Encapsulation / decapsulation functionality is implemented on the NVD. In some embodiments, the NVD is configured to intercept all network packets going in and out of the host machine and perform network virtualization functions.

[0128] As indicated above, the NVD performs various virtualization functions, including VNIC and VCN VR. The NVD can perform the VNIC associated with a computing instance hosted by one or more host machines connected to the VNIC. For example, as Figure 2 depicted, the NVD 210 performs the functionality of the VNIC 276 associated with the computing instance 268 hosted by the host machine 202 connected to the NVD 210. As another example, the NVD 212 performs the VNIC 280 associated with the bare-metal computing instance 272 hosted by the host machine 206 and performs the VNIC 284 associated with the computing instance 274 hosted by the host machine 208. The host machine can host computing instances belonging to different VCNs (belonging to different customers), and the NVD connected to the host machine can perform the VNIC corresponding to the computing instance (i.e., perform VNIC-related functionality).

[0129] The NVD also performs the VCN virtual router corresponding to the VCN of the computing instance. For example, in the Figure 2 embodiment depicted, the NVD 210 performs the VCN VR 277 corresponding to the VCN to which the computing instance 268 belongs. The NVD 212 performs one or more VCN VRs 283 corresponding to one or more VCNs to which the computing instances hosted by the host machines 206 and 208 belong. In some embodiments, the VCN VR corresponding to the VCN is performed by all NVDs connected to the host machine hosting at least one computing instance belonging to that VCN. If the host machine hosts computing instances belonging to different VCNs, then the NVD connected to that host machine can perform the VCN VRs corresponding to those different VCNs.

[0130] In addition to VNICs and VCN VRs, an NVD can execute various software (e.g., daemons) and includes one or more hardware components that facilitate the various network virtualization functions performed by the NVD. For simplicity, these various components are grouped together as the Figure 2 "packet processing components" shown in. For example, NVD 210 includes packet processing component 286 and NVD 212 includes packet processing component 288. For example, the packet processing components for an NVD can include a packet processor that is configured to interact with the ports and hardware interfaces of the NVD to monitor all packets received by and transmitted using the NVD and store network information. The network information can include, for example, network flow information identifying different network flows handled by the NVD and per-flow information (e.g., per-flow statistics) for each flow. In some embodiments, the network flow information can be stored on a per-VNIC basis. The packet processor can perform per-packet manipulation and implement stateful NAT and L4 firewall (FW). As another example, the packet processing components can include a replication agent configured to copy information stored by the NVD to one or more different replication target repositories. As yet another example, the packet processing components can include a logging agent configured to perform the logging function of the NVD. The packet processing components can also include software for monitoring the performance and health of the NVD and also potentially the state and health of other components connected to the NVD.

[0131] Figure 1 Shows the components of an example virtual or overlay network, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, a VR for the VCN, and a collection of gateways configured for the VCN. Figure 1 The overlay components depicted in can be executed or hosted by one or more of the Figure 2 physical components depicted in. For example, a compute instance in a VCN can be executed or hosted by one or more of the Figure 2 host machines depicted in. For a compute instance hosted by a host machine, the VNIC associated with that compute instance is typically executed by an NVD connected to that host machine (i.e., VNIC functionality is provided by the NVD connected to that host machine). The VCN VR functionality for a VCN is executed by all NVDs connected to the host machine that hosts or executes the compute instances that are part of that VCN. Gateways associated with a VCN can be executed by one or more different types of NVDs. For example, some gateways can be executed by smartNICs, while other gateways can be executed by one or more host machines or other implementations of NVDs.

[0132] As described above, the compute instances in the customer VCN can communicate with various different endpoints, where the endpoints can be within the same subnet as the source compute instance, in different subnets but within the same VCN as the source compute instance, or endpoints located outside the VCN of the source compute instance. The VNIC associated with the compute instance, the VCN VR, and the gateways associated with the VCN are used to facilitate these communications.

[0133] For communication between two compute instances on the same subnet in the VCN, the VNICs associated with the source and destination compute instances are used to facilitate the communication. The source and destination compute instances can be hosted by the same host machine or different host machines. Packets originating from the source compute instance can be forwarded from the host machine hosting the source compute instance to the NVD connected to that host machine. At the NVD, the packets are processed using the packet processing pipeline, which can include the execution of the VNIC associated with the source compute instance. Since the destination endpoint for the packet is within the same subnet, the execution of the VNIC associated with the source compute instance causes the packet to be forwarded to the NVD that executes the VNIC associated with the destination compute instance, and then the NVD processes the packet and forwards it to the destination compute instance. The VNICs associated with the source and destination compute instances can be executed on the same NVD (e.g., when the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., when the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNIC can use the routing / forwarding table stored by the NVD to determine the next hop for the packet.

[0134] For packets to be transferred from a compute instance in a subnet to an endpoint in a different subnet within the same VCN, the packet originating from the source compute instance is transferred from the host machine hosting the source compute instance to the NVD connected to that host machine. At the NVD, the packet is processed using a packet processing pipeline, which can include the execution of one or more VNICs and the VR associated with the VCN. For example, as part of the packet processing pipeline, the NVD executes or invokes the functionality corresponding to the VNIC associated with the source compute instance (also referred to as executing the VNIC). The functionality executed by the VNIC can include examining the VLAN tag on the packet. Since the destination of the packet is outside the subnet, the NVD then calls and executes the VCN VR functionality. The VCN VR then routes the packet to the NVD that executes the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packet and forwards the packet to the destination compute instance. The VNICs associated with the source and destination compute instances can be executed on the same NVD (e.g., when the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., when the source and destination compute instances are hosted by different host machines connected to different NVDs).

[0135] If the destination for the packet is outside the VCN of the source compute instance, then the packet originating from the source compute instance is transferred from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD executes the VNIC associated with the source compute instance. Since the destination endpoint of the packet is outside the VCN, the packet is then processed by the VCN VR for that VCN. The NVD calls the VCN VR functionality, which causes the packet to be forwarded to the NVD that executes the appropriate gateway associated with the VCN. For example, if the destination is an endpoint within the customer's on-premises network, then the packet can be forwarded by the VCN VR to the NVD that executes the DRG gateway configured for the VCN. The VCN VR can be executed on the same NVD as the NVD that executes the VNIC associated with the source compute instance, or by a different NVD. The gateway can be executed by an NVD, which can be a smartNIC, a host machine, or other NVD implementation. The packet is then processed by the gateway and forwarded to the next hop, which facilitates the transfer of the packet to its intended destination endpoint. For example, in Figure 2In the illustrated embodiment, data packets originating from compute instance 268 can be transferred from host machine 202 to NVD 210 via link 220 (using NIC 232). At NVD 210, VNIC 276 is invoked because it is the VNIC associated with source compute instance 268. VNIC 276 is configured to inspect the information encapsulated in the data packet and determine the next hop for forwarding the data packet, with the aim of facilitating the transfer of the data packet to its intended destination endpoint, and then forwarding the data packet to the determined next hop.

[0136] Compute instances deployed on a VCN can communicate with a variety of different endpoints. These endpoints can include endpoints hosted by CSPI 200 and endpoints external to CSPI 200. Endpoints hosted by CSPI 200 can include instances in the same VCN or other VCNs, which can be the customer's VCNs or VCNs that do not belong to the customer. Communication between endpoints hosted by CSPI 200 can be performed via physical network 218. Compute instances can also communicate with endpoints that are not hosted by CSPI 200 or are external to CSPI 200. Examples of these endpoints include endpoints or data centers within the customer's on-premises network, or public endpoints accessible via a public network (such as the Internet). Communication with endpoints external to CSPI 200 can be performed via a public network (e.g., the Internet) Figure 2 (not shown in ) or a private network ( Figure 2 (not shown in ).

[0137] Figure 2 The architecture of CSPI 200 depicted in is merely an example and is not intended to be limiting. In alternative embodiments, variations, substitutions, and modifications are possible. For example, in some implementations, CSPI 200 can have more or fewer systems or components than Figure 2 the systems or components shown in, two or more systems can be combined, or it can have a different system configuration or arrangement. Figure 2 The systems, subsystems, and other components depicted in can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, using hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., a memory device).

[0138] Figure 4 depicts the connection between a host machine and an NVD according to certain embodiments for providing I / O virtualization to support multi-tenancy. As Figure 4As depicted, host machine 402 executes hypervisor 404 that provides a virtualized environment. Host machine 402 executes two virtual machine instances, VM1 406 belonging to customer / tenant #1 and VM2 408 belonging to customer / tenant #2. Host machine 402 includes physical NIC 410 connected to NVD 412 via link 414. Each computing instance is attached to a VNIC executed by NVD 412. In Figure 4 the embodiment of, VM1 406 is attached to VNIC-VM1 420 and VM2 408 is attached to VNIC-VM2 422.

[0139] As Figure 4 shown, NIC 410 includes two logical NICs, logical NIC A 416 and logical NIC B 418. Each virtual machine is attached to its own logical NIC and is configured to work with its own logical NIC. For example, VM1 406 is attached to logical NIC A 416 and VM2 408 is attached to logical NIC B 418. Although host machine 402 includes only one physical NIC 410 shared by multiple tenants, due to the logical NICs, each tenant's virtual machine believes they have their own host machine and network card.

[0140] In some embodiments, each logical NIC is assigned its own VLAN ID. Thus, a specific VLAN ID is assigned to logical NIC A 416 for tenant #1, and a separate VLAN ID is assigned to logical NIC B 418 for tenant #2. When a data packet is transmitted from VM1 406, the label assigned to tenant #1 is attached to the data packet by the hypervisor, and then the data packet is transmitted from host machine 402 to NVD 412 via link 414. In a similar manner, when a data packet is transmitted from VM2 408, the label assigned to tenant #2 is attached to the data packet by the hypervisor, and then the data packet is transmitted from host machine 402 to NVD 412 via link 414. Accordingly, data packet 424 transmitted from host machine 402 to NVD 412 has an associated label 426 that identifies the specific tenant and the associated VM. At the NVD, for data packet 424 received from host machine 402, the label 426 associated with the data packet is used to determine whether the data packet is to be processed by VNIC-VM1 420 or by VNIC-VM2 422. The data packet is then processed by the corresponding VNIC. Figure 4 The configuration described in enables each tenant's computing instance to believe they have their own host machine and NIC. Figure 4 The setup described in provides I / O virtualization to support multi-tenancy.

[0141] Figure 5Depicts a simplified block diagram of a physical network 500 according to certain embodiments. Figure 5 The embodiment depicted in Figure 5 is structured as a Clos network. A Clos network is a particular type of network topology that is designed to provide connection redundancy while maintaining high bisection bandwidth and maximum resource utilization. A Clos network is a non-blocking, multi-stage or multi-layer switching network, where the number of stages or layers can be two, three, four, five, etc. Figure 5 The embodiment depicted in Figure 5 is a three-layer network, including Layer 1, Layer 2, and Layer 3. The TOR switch 504 represents the Layer 0 switch in the Clos network. One or more NVDs are connected to the TOR switch. The Layer 0 switch is also referred to as an edge device of the physical network. The Layer 0 switch is connected to the Layer 1 switches, which are also referred to as leaf switches. In Figure 5 the embodiment depicted in Figure 5 , a set of "n" Layer 0 TOR switches is connected to a set of "n" Layer 1 switches and forms a pod. Each Layer 0 switch in a pod is interconnected to all the Layer 1 switches in that pod, but there is no switch connectivity between pods. In certain implementations, two pods are referred to as a block. Each block is served or connected to by a set of "n" Layer 2 switches (sometimes referred to as spine switches). There can be several blocks in the physical network topology. The Layer 2 switches are in turn connected to "n" Layer 3 switches (sometimes referred to as super spine switches). Communication of data packets over the physical network 500 is typically performed using one or more Layer 3 communication protocols. Generally, all layers of the physical network (except the TOR layer) are n-way redundant, thus allowing for high availability. Policies can be specified for pods and blocks to control the visibility of switches to each other in the physical network, thereby enabling the scaling of the physical network.

[0142] A characteristic of a Clos network is that the maximum number of hops from one Layer 0 switch to another Layer 0 switch (or from an NVD connected to a Layer 0 switch to another NVD connected to a Layer 0 switch) is fixed. For example, in a three-layer Clos network, a data packet requires a maximum of seven hops to reach another NVD, where the source and destination NVDs are connected to the leaf layer of the Clos network. Similarly, in a four-layer Clos network, a data packet requires a maximum of nine hops to reach another NVD, where the source and destination NVDs are connected to the leaf layer of the Clos network. Thus, the Clos network architecture maintains a consistent latency throughout the network, which is important for communication within and between data centers. The Clos topology is horizontally scalable and cost-effective. The bandwidth / throughput capacity of the network can be easily increased by adding more switches at each layer (e.g., more leaf switches and backbone switches) and by increasing the number of links between switches in adjacent layers.

[0143] In some embodiments, each resource within CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the information for the resource and can be used to manage the resource, e.g., via the console or through an API. An example syntax for the CID is:

[0144] ocid1.<RESOURCE TYPE>. <realm>.[REGION][.FUTURE USE]. <uniqueid>

[0145] Among them,

[0146] ocid1: A text string indicating the version of the CID;

[0147] resource type: The type of the resource (e.g., instance, volume, VCN, subnet, user, group, etc.);

[0148] realm: The realm where the resource is located. Example values are "c1" for the commercial realm, "c2" for the government cloud realm, or "c3" for the federal government cloud realm, etc. Each realm can have its own domain name;

[0149] region: The region where the resource is located. If this region is not applicable to the resource, then this part may be empty;

[0150] future use: Reserved for future use.

[0151] uniqueID: The unique part of the ID. The format can vary depending on the type of the resource or service.

[0152] Hybrid GPU Architecture (GPU Super Cluster)

[0153] Fig. 6A Depicts the architecture of a hybrid general - purpose processing unit (GPU) cluster according to some embodiments. The hybrid GPU cluster includes multiple GPU clusters that are communicatively coupled to each other via a hierarchy of switches (e.g., a CLOS architecture of switches arranged in a hierarchical manner). Note that the CLOS architecture enables the GPU to be scaled to a much larger extent than traditional GPU clusters. The hybrid GPU architecture enables multiple GPU clusters to co - exist in the same network architecture - referred to herein as the GPU architecture. In some embodiments, the operating speed of at least some GPU clusters is different from that of some other GPU clusters included in the hybrid GPU architecture. For example, the hybrid GPU cluster can include one or more GPU clusters operating at 100G and one or more GPU clusters operating at 400G. Fig. 6A The hybrid GPU cluster architecture herein is also referred to as the GPU super - cluster architecture.

[0154] The GPU super - cluster architecture 600 includes multiple blocks that host multiple GPU clusters. For example, as Fig. 6A As shown, the GPU supercluster architecture 600 includes "K" blocks (e.g., block 1, 605 to block K, 625), where each block is configured to host a specific GPU cluster. The first block (e.g., block 1 605) hosts the first GPU cluster 631, and the Kth block (i.e., block K 625) hosts the second GPU cluster 633. Note that each GPU cluster contained within a block is hosted on one or more racks. For example, the first GPU cluster 631 includes a first set of GPUs hosted on one or more racks (e.g., rack 1 (619A) to rack K (619B)). Similarly, the second GPU cluster 633 includes a second set of GPUs hosted on one or more racks (e.g., rack 1 (629A) to rack M (629B)).

[0155] According to some embodiments, the first GPU cluster 631 is a GPU cluster operating at a first speed (e.g., 100G). More specifically, the first GPU cluster 631 includes one or more GPUs (i.e., the first set of GPUs), where the network links from the GPUs operate at the first speed. In other words, in the host machine, the GPU cards are paired with network interface cards (NICs) that operate at the first speed. Similarly, the second GPU cluster 633 is a GPU cluster operating at a second speed (e.g., 400G). Specifically, the second GPU cluster 633 includes one or more GPUs (i.e., the second set of GPUs), where the second speed is different from the first speed.

[0156] The CLOS architecture includes a hierarchy or hierarchy of switches, including first-level switches (T1), second-level switches (T2), and third-level switches (T3). In one implementation, the first-level switches and the second-level switches are each contained within a block of the supercluster architecture. For example, as Fig. 6A shown, block 1 605 includes a plurality of first-level switches 617A - 617B, and block K 625 includes a plurality of first-level switches 628A - 628B. In a similar manner, block 1 605 includes a plurality of second-level switches 615A - 615B, and block K 625 includes a plurality of second-level switches 627A - 627B. The network architecture also includes third-level switches. In some embodiments, the switches in the third level can be partitioned to form groups of third-level switches, e.g., the switches in the group labeled 601 (i.e., switches 613A to 613B) and the switches in the group labeled 621 (i.e., switches 623A to 623B). Note that the switches in the third level are also referred to as upper-level switches herein. Each group of switches in the third-level switches communicatively couples the blocks as described below.

[0157] In some embodiments, the hierarchical / level arrangement of the switches is as follows: (i) A layer 3 (or upper-level switch) having ports operating at a second speed (e.g., 400G). Thus, these switches (e.g., switches 613A, 613B, 623A, and 623B) can support 100G, 200G, or 400G switches attached thereto. (ii) A layer 2 (or mid-level switch, e.g., switches 615A, 615B, 627A, and 627B) configured to support 400G connections (either as 4x100G connections (i.e., multiple 100G connections) or 1x400G connections (i.e., a single 400G connection)); and (iii) A layer 1 (or lower-level switch, e.g., switches 617A, 617B, 628A, and 628B), where the type of switch to be deployed is selected based on the type of GPU cluster to which these switches are expected to be connected. For example, if a rack of GPUs is desired to be connected, where each GPU operates at 100G speed (e.g., the GPUs in cluster 1 631), then the selected layer 1 switch is a 100G switch, and if a rack of GPUs is desired to be connected, where each GPU operates at 400G speed (e.g., the GPUs in cluster 633), then the selected layer 1 switch is a 400G switch.

[0158] As previously described, the architecture of the GPU supercluster is arranged in multiple blocks, where each block supports a GPU cluster operating at a specific speed. The GPU supercluster includes at least a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed. In each block, a layer 1 switch (at one end) is communicatively coupled to a GPU cluster hosted on one or more racks and (at the other end) is communicatively coupled to a layer 2 switch. In turn, the layer 2 switch communicatively couples the layer 1 switch to the layer 3 switch. The layer 3 switches are configured to couple different blocks together, i.e., they communicatively couple different GPU clusters.

[0159] Reference Fig. 6A , block 1 605 includes K Layer 1 switches (e.g., K = 64 switches), such as switches labeled 617A - 617B, where each Layer 1 switch has M upstream ports (i.e., ports facing the Layer 2 switches) and M downstream ports (i.e., ports facing the GPU racks), e.g., M = 32 ports. Each upstream and downstream port of these Layer 1 switches operates at a first speed (e.g., 100G) (since these switches are coupled to a GPU cluster 631 operating at a first speed of 100G). The downstream ports of the Layer 1 switches are coupled to the GPU racks. For example, a particular Layer 1 switch in block 1 605 (e.g., switch 617A) is coupled to a rack of GPUs (619A), while switch 617B is coupled to rack 619B. The M downstream ports of switch 617A are communicatively coupled to the rack (619A), which includes P GPUs (e.g., P = M / 2), where each GPU is connected to the switch via two links (i.e., using two ports of switch 617A), thus providing a pair of 2x100G connections.

[0160] The M upstream ports of switch 617A (e.g., M = 32 ports) are communicatively coupled to M Layer 2 switches (e.g., switches 615A, 615B). Note that the Layer 2 switches have a dimension of P downstream ports and P upstream ports (e.g., P = 16), where each port operates at a second speed of, for example, 400G (i.e., different from the operating speed of the upstream ports of the Layer 1 switches). For example, as Fig. 6A shown, switch 615A has P = 16 downstream ports labeled 630. In this case, each downstream port of the Layer 2 switch (e.g., switch 615A) operating at 400G is (e.g., optically) split into multiple (e.g., four) 100G ports (referred to herein as sub - ports). Thereafter, each sub - port operating at 100G can be communicatively coupled to the corresponding upstream port of the Layer 1 switch (operating at 100G). It should be recognized that the splitting of the downstream ports of the Layer 2 switch can be achieved via an optical module connector that has a splitting function capable of evenly distributing the large bandwidth at one end to several low - speed connections at the other end. Each of the P upstream ports of switch 615A (Layer 2) operating at 400G is communicatively coupled to P = 16 Layer 3 switches. For example, as Fig. 6A shown, the P = 16 upstream ports of switch 615A are communicatively coupled to the first Layer 3 switch / upper - layer switch in each group, e.g., switch 613A (in upper layer 1) and switch 623A (in upper layer P).

[0161] Now turn to block K 625, which includes M layer 2 switches (e.g., switches 627A, 627B), each switch having dimensions of P upstream ports and P downstream ports (e.g., P = M / 2). Each of these switches has ports that operate at a second speed. The upstream ports of the layer 2 switches in block K 625 are communicatively coupled to the layer 3 switches in a manner similar to the coupling of the layer 2 and layer 3 switches of block 1 605. However, the P downstream ports of the layer 2 switches are communicatively coupled to P layer 1 switches. Specifically, compared to the coupling described above for the downstream port 630 of layer 2 switch 615A (of block 1), the downstream ports of the layer 2 switches in block K do not require any form of splitting.

[0162] According to some embodiments, the coupling of the downstream ports of the layer 1 switches in block K 625 to the GPU clusters 633 hosted on racks 629A - 629B can take one of two forms. In one implementation, as Fig. 6A shown, for a particular layer 1 switch (e.g., switch 628A), the M downstream ports of the switch are connected to the M GPUs hosted in the rack. Thus, there is a one-to-one correspondence between the downstream ports of the switch and the GPUs hosted in the rack. Note that this coupling provides a single connection 632 between the GPUs hosted in the rack and the downstream ports of switch 628A. In another implementation 650 (as Figure 6B shown), to provide more flexibility to the host, link aggregation techniques (e.g., link bonding) are used, where, for example, a single 400G interface of the ports of switch 628A is presented as two different links (labeled 640 in Figure 6B ) to an application running on a GPU hosted on rack 629A. Thus, the above-described implementation provides the functionality of presenting either a 1x400G link (i.e., a single connection) or a 2x200G link (i.e., bonded or multiple connections) to an application running on a GPU.

[0163] Therefore, as described above, Fig. 6A and 6B 's (one or more) supercluster architectures provide ultra-high performance under enhanced scaling. The multi-layer CLOS topology provides a non-blocking network architecture that can scale to tens of thousands of GPUs. Note that while traditional GPU clusters may be suitable for a few rows in a single room of a data center, large-scale superclusters can span multiple rooms within a building (i.e., data halls), or even multiple adjacent buildings within a data center complex. The cable distance between two GPUs may be longer, which may cause some data packets to cross these data halls and result in slightly higher latency. Refer to Figure 9-12 Discuss the techniques that may lead to potentially higher latency for cancellation. Additionally, it should be recognized that the architectures referenced above Fig. 6A and 6B are merely examples and are not intended to limit the scope of the present disclosure. In alternative embodiments, variations, substitutions, and modifications can be made. For example, in some implementations, the architecture can have more or fewer systems or components than those shown in Fig. 6A and 6B , can combine two or more systems, or can have a different configuration (e.g., switch dimensions, number of hierarchical layers in a CLOS topology, number of blocks, type (i.e., speed) of GPU clusters, etc.) or arrangement of systems. Fig. 6A and 6B The systems, subsystems, and other components depicted in

[0164] Figure 7 can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, using hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 7 The processing depicted in Figure 7 can be implemented in software executed by one or more processing units (e.g., processors, cores) of the corresponding system, in hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 7 The methods presented in Figure 7 and described below are intended to be illustrative and non-limiting.

[0165] The processing begins at step 705, where a network architecture is provided, and the network architecture includes multiple GPU clusters. Each GPU cluster includes one or more GPUs hosted on one or more racks. The multiple GPU clusters include at least a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed (different from the first speed). In some implementations, the provision of the network architecture in step 705 can include sub-steps 710 - 720 as described below.

[0166] In step 710, multiple blocks are instantiated, where each block includes one or more racks to host GPUs belonging to a specific GPU cluster. Note that a single GPU cluster is contained within a block. Thus, the first block can host a first GPU cluster operating at a first speed, and the second block can host a second GPU cluster operating at a second speed. In step 715, multiple switches are arranged in a hierarchical manner (i.e., a CLOS architecture). In some embodiments, the hierarchy of switches can correspond to a three-tier switch architecture. In this case, switches belonging to layer 1 and layer 2 can be provided in each block of the network architecture, as Fig. 6A depicted. In step 720, multiple sets of upper-tier switches can be provided. Note that the layer switches can correspond to the layer 3 switches in the CLOS architecture. The layer 3 switches communicatively couple different blocks included in the network architecture. In step 725, the control plane of the network architecture can receive a request. The request can correspond to a customer's request to execute a workload. In response to receiving the request, in step 730, the control plane can allocate one or more GPUs (based on constraints associated with the request) from multiple GPU clusters to execute the workload.

[0167] Figure 8 FIG. depicts a block diagram of a cloud infrastructure 800 including a CLOS network switch arrangement according to certain embodiments. The cloud infrastructure 800 includes multiple blocks, e.g., block 1 605 - block K 625. Each block can include multiple racks, where each rack hosts multiple host machines (also referred to herein as hosts). For simplicity, Figure 8 the blocks depicted in are illustrated as including a single rack. However, note that a block can include multiple racks. Block 1 includes rack 619A, which is depicted as including two host machines, namely, host 1-A 812 and host 1-B 814. Similarly, block K 625 includes rack 629A, which is depicted as including two other host machines, namely, host 2-A 822 and host 2-B 824. It should be recognized that Figure 8 the illustration in (i.e., each rack includes two host machines) is illustrative and not restrictive. For example, the cloud infrastructure can include more than two racks, where each rack can include more than two host machines. Also, note that each rack is not limited to having the same number of hosts. Rather, a rack can have a greater or fewer number of host machines compared to the number of host machines contained in another rack.

[0168] Each host machine includes multiple graphics processing units (GPUs). For example, host machine 1-A 812 includes N GPUs, e.g., GPU 1, 813. Also, it should be recognized that Figure 8 The illustration that each host machine includes the same number of GPUs (i.e., N GPUs) is intended to be illustrative and non - limiting, i.e., each host machine may include a different number of GPUs. Each rack is associated with a Layer 1 switch (also referred to herein as a top - of - rack (TOR) switch), which is communicatively coupled to the GPUs hosted on the host machines within the rack. For example, Rack 1 619A is associated with TOR switch 617A, which is communicatively coupled to host machines Host 1 - A, 812 and Host 1 - B, 814, while Rack 2 629A is associated with TOR switch 627A, which is communicatively coupled to host machines Host 2 - A, 622 and Host 2 - K, 624. It should be recognized that the TOR switches depicted in FIG. 6 (i.e., switches 617A and 627A) each include N ports that are used to communicatively couple the TOR switch to the N GPUs hosted on each host machine included in the corresponding rack. As Figure 8 The coupling of the TOR switches depicted in is intended to be illustrative and non - limiting. For example, in some embodiments, a TOR switch may have multiple ports, each corresponding to a GPU on each host machine, i.e., the GPUs on a host machine may be connected via communication links to a unique port of the TOR.

[0169] The TOR switch associated with each rack is communicatively coupled to one or more Layer 2 switches in each block. For example, switch 617A is communicatively coupled to Layer 2 switches 615A - 615B, while switch 627A is communicatively coupled to switches 623A - 623B. Similar to the Fig. 6A configuration, the Layer 2 switches are communicatively coupled to upper - layer switches 613A - 613B and 623A - 623B, as Figure 8 shown in. Information transmitted from a particular Layer 1 switch to a Layer 2 switch (or from a Layer 2 switch to a Layer 3 switch) is referred to herein as communication via an uplink, while information transmitted from a Layer 1 switch to a host machine (or from a Layer 3 switch to a Layer 2 switch) is referred to herein as communication via a downlink. According to some embodiments, Figure 8 the hierarchical switches form a CLOS network arrangement (e.g., a multi - stage switching network), where each Layer 1 switch can be considered to form a "leaf" node in the CLOS network.

[0170] According to some embodiments, a GPU included in a host machine performs tasks related to machine learning. In such a setup, a single task can be executed / distributed across a large number of GPUs, which can be distributed across multiple host machines, across multiple racks, and / or across different blocks. Since all these GPUs are working on the same task (i.e., workload), they all need to communicate with each other in a time-synchronized manner. Additionally, at any given time, a GPU is in either a compute mode or a communication mode, i.e., the GPUs communicate with each other roughly at the same time. The speed of the workload is determined by the speed of the slowest GPU.

[0171] Typically, to route a data packet from a source GPU to a destination GPU, Equal-Cost Multi-Path (ECMP) routing is used. In ECMP routing, when there are multiple equivalent paths available for routing traffic from a sender to a receiver, selection techniques are used to select a specific path. Correspondingly, at a network device (e.g., a TOR switch) receiving traffic, a selection algorithm is used to select an outgoing link to be used for forwarding the traffic from the network device to a subsequent device. Such outgoing link selection occurs at each network device in the path from the sender to the receiver. Hash-based selection is a widely used ECMP selection technique, where the hash can be based on, for example, the 4-tuple of the data packet (e.g., source port, destination port, source IP, destination IP).

[0172] ECMP routing is a flow-aware routing technique, where each flow (i.e., a data packet flow) is hashed to the same path for the duration of the flow. Thus, the data packets in a flow are forwarded from a network device using a specific outgoing port / link. This is generally to ensure that the data packets in a flow arrive in order, i.e., there is no need to reorder the data packets. However, ECMP routing is bandwidth (or throughput) insensitive. In other words, the switch performs statistical flow-aware (throughput-insensitive) ECMP load balancing across parallel links.

[0173] In standard ECMP routing (i.e., flow-aware routing only), one problem is that flows received by a network device through two separate incoming links may be hashed to the same outgoing link, resulting in flow conflicts. For example, consider a scenario where two flows come in through two separate incoming 100G links, and each flow is hashed to the same 100G outgoing link. This situation causes congestion (i.e., flow conflict) and results in data packets being dropped because the incoming bandwidth is 200G, but the outgoing bandwidth is 100G. As Figure 8 As shown, there are two flows: flow 1841 directed from the first GPU of host machine host 1-A 812 to switch 617A, and flow 2843 directed from another GPU on host machine host 1-B 814 to switch 617A. Note that both of these flows are directed to the switch on separate links. For ease of explanation, assume that Figure 8 the links depicted in Figure 8 have the same capacity (i.e., bandwidth) of 100G. In the case where the switch 617A executes the ECMP routing algorithm, the two flows may be hashed to use the same outgoing link of the switch, e.g., link 850 that connects the switch 617A to a higher-level switch (e.g., switch 615A). In this case, there is a conflict between the two flows (marked with "X"), which results in packets being dropped.

[0174] Regardless of the protocol, this conflict scenario is generally problematic for all types of traffic. For example, TCP is intelligent in that when a packet is dropped and the sender does not receive an acknowledgment for that dropped packet, the packet is retransmitted. However, for traffic of the Remote Direct Memory Access (RDMA) type, the situation deteriorates. There are various reasons for RDMA networks not to use TCP (e.g., the performance of TCP is not high). RDMA networks use protocols such as RDMA over Infiniband or RDMA over converged Ethernet (RoCE). In RoCE, there is a congestion control algorithm where when the sender identifies the occurrence of congestion or dropped packets, the sender slows down the transmission of packets. For dropped packets, not only the dropped packet but also several packets around the dropped packet are retransmitted, which further consumes the available bandwidth and results in poor performance.

[0175] Due to strict time synchronization requirements, the flow conflict problem is for Fig. 6A is a key issue for a supercluster GPU architecture. For example, as previously mentioned, GPUs can execute machine learning tasks (i.e., workloads) where all GPUs communicate with each other in a time-synchronized manner. For machine learning tasks and other types of tasks, a logical topology (e.g., ring topology, tree topology, etc.) is constructed for the host machine to enable communication between GPUs. The GPUs are interconnected using the logical topology, which can be multi-level or multi-dimensional. In some embodiments, to execute a workload, an application constructs a virtual (or logical) topology to interconnect the GPUs. Typically, such an application does not know the underlying physical topology of the host machine and thus attempts to construct the logical topology in a random (i.e., arbitrary) manner. This randomly constructed logical topology causes the GPU host machines to exchange traffic without considering which other (one or more) GPU hosts are in its local network neighborhood. As a result, the likelihood of traffic congestion increases, leading to poor GPU throughput. For example, consider a scenario where a pair of host machines is required to execute a certain machine learning task. In this case, if a random selection is made for the pair of hosts, where one host machine is in a first local neighborhood (e.g., a first rack) and the other host machine is in another local neighborhood (different from the first local neighborhood), e.g., a second rack, then the execution of the machine learning task will incur a certain amount of latency (e.g., the delay that occurs in the communication between the first host machine and the second host machine), and may increase the likelihood of traffic congestion. In contrast, if the pair of host machines selected to execute the machine learning task is in the same local neighborhood (e.g., in the same rack or the same block), then it should be recognized that the communication between the host machines will incur minimal latency and increase the likelihood of avoiding traffic congestion.

[0176] Techniques for overcoming the above problems are described below. Specifically, the techniques described herein utilize the hierarchical locality information of GPUs during the construction of the logical topology, thereby avoiding unnecessary traffic congestion. Moreover, embodiments of the present disclosure allow customers to reduce the latency of their applications to services by "placing" their (one or more) workloads on nearby host machines. In addition, customers can use the locality information and place their workloads in a certain way to obtain higher anti-affinity, thereby achieving higher resilience by reducing the shared fate of resources.

[0177] According to some embodiments, the host machine does not know the physical topology of the network, i.e., a particular host machine does not know the physical location of other host machines in the network. For example, referring to Figure 8 , the host machine 1-A 812 is unaware that the host machine 1-B 814 is actually contained in the same rack (i.e., rack 619A) and positioned behind the same TOR switch (i.e., switch 617A). However, the network control plane knows the overall physical topology of the host machines. In one embodiment, the network control plane publishes such locality information (e.g., hierarchical locality information identifying at least the rack including the host machine, the block hosting the rack, etc.) to the host machines to achieve traffic locality and avoid unnecessary traffic congestion. Doing so has a significant impact on the performance of GPU workloads, as shown below with reference to Fig. 9 and Fig.10 as shown.

[0178] According to some embodiments, the network control plane utilizes the Instance Metadata Service (IMDS) to publish (and store) metadata information (e.g., hierarchical locality information) to the host machines. Such metadata information can be published to the (one or more) host machines via a Network Virtualization Device (NVD) associated with the (one or more) host machines. It should be recognized that the hierarchical locality information can include metadata indicating the rack identifier of the rack including the host machine and the block identifier of the block hosting the rack. It should be recognized that each host machine can query the IMDS to obtain metadata information associated with the host machine. As described below, the published locality information is used to construct an optimal logical topology to achieve higher GPU workload throughput.

[0179] Turning to Fig. 9 , the figure depicts a logical topology constructed without considering the locality information of the host machines according to certain embodiments. Fig. 9 The logical topology depicted in Fig. 9 corresponds to a scenario with eight host machines, namely, host machine 1-A 901, host machine 1-B 903, host machine 2-A 905, host machine 2-B 907, host machine 3-A 909, host machine 3-B 911, host machine 4-A 913, and host machine 4-B 915. As Fig. 9 shown, note that host machines 1-A 901 and 1-B 903 are located in the same rack, i.e., positioned behind the same TOR switch. Similarly, the pairs of host machines: (host machines 2-A 905 and 2-B 907), (host machines 3-A 909 and 3-B 911), and (host machines 4-A 913 and 4-B 915) are respectively contained in other racks. This is represented as "TOR local traffic" in Fig. 9 is represented as "block local traffic".

[0180] Fig. 9 The logical topology depicted in is a ring topology, which is constructed without hierarchical locality information, i.e., the ring topology is constructed in a random (arbitrary) manner. As Fig. 9 shown, the ring is constructed such that host 1-A is directly connected to host 4-B via the link labeled 941. In addition, host 4-B is connected to host 3-B via the link labeled 942, host 3-B is connected to host 4-A via the link labeled 943, host 4-A is connected to host 3-A via the link labeled 944, host 3-A is connected to host 2-B via the link labeled 945, host 2-B is connected to host 2-A via the link labeled 946, host 2-A is connected to host 1-B via the link labeled 947, and host 1-B is connected to host 1-A via the link labeled 948.

[0181] In the Fig. 9 way shown, the logical topology is prone to network flow conflicts due to ECMP traffic distribution. Note that although host 4-A 913 and host 4-B 915 are located behind the same TOR switch, i.e., included in the same rack, the traffic originating from host 4-B and destined for host 4-A will follow the following route: the traffic first routes from host 4-B to host 3-B (i.e., via virtual link 942), and then from host 3-B to host 4-A (i.e., via virtual link 943). Therefore, the traffic destined for a destination host (e.g., host 4-A) located in the same rack as the original host (i.e., host 4-B) is unnecessarily routed to a host machine (i.e., host 3-B) that is not only outside the rack but also in a completely different block (compared to the source host machine and the destination host machine). In other words, the traffic that could have been locally routed within block 930 is routed to a host machine outside the block and then only rerouted back into block 930. This is due to Fig. 9 the arbitrarily constructed logical topology in, which leads to an increased likelihood of flow conflicts and thus reduces the throughput of the GPU workload (e.g., longer latency, jitter loss, etc.).

[0182] Now turning to Fig.10 , which describes another logical topology constructed considering the hierarchical locality information of the host machines according to certain embodiments. Fig.10 The logical topology depicted in corresponds to the same eight host machines depicted in Fig. 9 . Fig.10 The logical topology depicted in is a ring topology constructed using hierarchical locality information, i.e., the ring topology is constructed based on hierarchical locality information obtained from, for example, IMDS. According to some embodiments, the logical topology may be constructed by a configuration host machine, which may be Fig.10 One of the host machines depicted in . Fig.10 As shown in , the ring is constructed in such a way that: host 1-A is directly connected to host 4-B via a link labeled 1041. In addition, host 4-B is connected to host 4-A via a link labeled 1042, host 4-A is connected to host 3-B via a link labeled 1043, host 3-B is connected to host 3-A via a link labeled 1044, host 3-A is connected to host 2-B via a link labeled 1045, host 2-B is connected to host 2-A via a link labeled 1046, host 2-A is connected to host 1-B via a link labeled 1047, and host 1-B is connected to host 1-A via a link labeled 1048.

[0183] Fig.10 The logical topology depicted in avoids traffic between two host machines located in the same rack (i.e., positioned behind the same TOR switch) from being unnecessarily routed outside the rack and / or block. For example, consider Fig. 9 In the same example of , where host 4-B intends to send traffic to host 4-A, that traffic can be routed via link 1042 (through a single hop in the virtualization layer). This is similar to Fig. 9 This is in contrast to the situation depicted in , where traffic is routed outside the block (i.e., to host 3-B) and then rerouted back into the block (i.e., back to host 4-A). Fig.11 Describes the above Fig. 9 and Fig.10 1100 of eight nodes considered in the example of FIG. 1101 . Specifically, host 1-A and host 1-B are depicted as being contained in the same rack 1105A (connected by layer 1 switch 1102A), host 2-A and host 2-B are depicted as being contained in the same rack 1105B (connected by layer 1 switch 1102B), host 3-A and host 3-B are depicted as being contained in the same rack 1105C (connected by layer 1 switch 1102C), and host 4-A and host 4-B are depicted as being contained in the same rack 1107A (connected by layer 1 switch 1104A). Note that layer 1 switches 1102A, 1102B, and 1102C are located in the same block 920, while layer 1 switch 1104A is located in another block 930. In this way, constructing a logical topology based on hierarchical locality information reduces the likelihood of flow conflicts, thereby increasing the throughput of GPU workloads.

[0184] Fig.12 FIG. 1200 is an exemplary flowchart depicting steps performed when supplying requests using hierarchical locality information according to some embodiments. Fig.12 The processing depicted therein may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, in hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). Fig.12 The methods presented herein and described below are intended to be illustrative and non-limiting. Although Fig.12 various processing steps are depicted as occurring in a particular order or sequence, this is not intended to be limiting. In certain alternative embodiments, these steps may be performed in some different order, or some steps may also be performed in parallel.

[0185] The processing begins at step 1205, where for each of a plurality of host machines (e.g., host machines included in a GPU cluster), the hierarchical locality information of the host machine is stored. The hierarchical locality information of the host machine includes information identifying, for example, the rack including the host machine, the block in which the rack is located, etc. According to some embodiments, an instance metadata service may be used to store the hierarchical locality information in the host machine via a network virtualization device (NVD) associated with the host machine. Note that the hierarchical locality information may correspond to information indicating, for example, an identifier of the rack including the host machine (i.e., rack ID), an identifier of the TOR switch associated with the rack, a block identifier of the block in which the rack is located (i.e., block ID), etc.

[0186] At step 1210, the control plane (e.g., from a customer) receives a request to execute a workload. Note that the workload corresponds to one or more processes to be performed using a GPU associated with a host machine. The processing then moves to step 1215, where one or more host machines available for executing the workload are identified from among the plurality of host machines. The identification of available host machines may be performed in several ways. For example, according to one embodiment, the control plane may maintain the current load (i.e., processing workload) handled by the host machines. Based on the capacity of each host machine and the amount of the current load handled by the host machine, the control plane may select one or more host machines available for executing the customer's workload. Additionally, according to another embodiment, the control plane may pre-allocate a certain number of host machines for each customer. Additionally, the available host machines may be determined from among the certain number of host machines.

[0187] According to some embodiments, the request received in step 1210 may include one or more constraints. For example, the request may include a first constraint associated with a latency threshold, i.e., the customer may expect to execute the workload by keeping the latency below a predetermined threshold. The second constraint may correspond to an anti-affinity constraint. Such a constraint corresponds to the customer's expectation that the host machines have a certain degree of availability, i.e., at least some of the host machines selected for executing the workload need to be located in different racks and / or different blocks. Customers typically adopt such a constraint to address rack failure issues. Additionally, the customer's request may include a constraint on the type of GPUs required for executing the workload. For example, the request may indicate that the customer expects a first number of GPUs operating at a first speed (e.g., 100G) and a second number of GPUs operating at a second speed (e.g., 400G). The control plane may consider these constraints when allocating / identifying a certain number of host machines for the customer.

[0188] Then, the process moves to step 1220, where the hierarchical locality information of each of the one or more host machines identified in step 1215 is obtained. Note that the locality information (of each host machine) may be stored, for example, in the corresponding host machine by the instance metadata service. At step 1225, the process identifies the link information of the one or more host machines. It should be recognized that the link information of the one or more host machines corresponds to the logical topology formed by the one or more host machines (e.g., the logical topology depicted in Fig.10 ). Thereafter, the process moves to step 1230, where the hierarchical locality information and the link information of the one or more host machines are provided in response to the request from the customer. According to some embodiments, after obtaining the hierarchical locality information and the link information, the customer may select a subset of the one or more host machines to execute the workload (step 1240). Note that the selection of the host machines may be performed based on one or more constraints associated with the workload.

[0189] As previously mentioned, ECMP routing is a flow-aware routing technique where each flow (i.e., a packet flow) is hashed to a specific path for the duration of the flow. Thus, the packets in the flow are forwarded from the network device using a specific outgoing port / link. This is typically to ensure that the packets in the flow arrive in order, i.e., there is no need to reorder the packets. However, ECMP routing is bandwidth (or throughput) insensitive. In other words, the switch performs statistical flow-aware (throughput-insensitive) ECMP load balancing of the flows across parallel links. In standard ECMP routing (i.e., only flow-aware routing), one problem is that flows received by the network device through two separate incoming links may be hashed to the same outgoing link, resulting in flow conflicts.

[0190] The routing techniques for overcoming the above flow conflict problems are described below (referred to herein as the GPU-based routing policy mechanism or the GPU-based traffic routing mechanism). It should be recognized that the flow conflict problem affects traffic from both the CPU and the GPU. However, due to the strict time synchronization requirements, the flow conflict problem is a greater problem for the GPU. In addition, it should be recognized that due to the nature of the standard ECMP routing mechanism to route information in a statistically bandwidth-insensitive manner, it will trigger flow conflict scenarios regardless of whether the network is oversubscribed or undersubscribed.

[0191] Now turning to Fig.13 , which depicts a GPU-based policy routing mechanism implemented in a hybrid GPU cluster according to certain embodiments. For convenience and illustration, Fig.13 the architecture depicted in Figure 8 is the same as the architecture depicted in Fig.13 . According to some embodiments, packets from a sender to a receiver are routed hop-by-hop in the network architecture. A routing policy is configured at each network device that binds (or matches) an incoming port link of the (network device) to an outgoing port link of the (network device). The network device can be any one of the switches included in the hierarchy of switches (i.e., layer 1, 2, or 3). Referring to Fig.13 , which depicts two flows: flow 1 from GPU 1 on host machine 812, whose intended destination is GPU 1 on host machine 822, and flow 2 from GPU N on host machine 814, whose intended destination is GPU N on host machine 824. In one implementation, all network devices (i.e., the switches included in the network architecture) are configured to bind (or match) the incoming port link to the outgoing port link. The matching of the incoming port link to the outgoing port link is maintained at each network device (e.g., in a policy table).

[0192] Referring to Fig.13 , it can be observed that for Flow 1 (i.e., the flow depicted by the solid line), when the first - layer switch (617A) receives a data packet on link 841, the first - layer switch is configured to forward the received data packet on the outgoing link 850. Similarly, when the second - layer switch (615A) receives a data packet via link 850, it is configured to forward the data packet on the outgoing link 851 to the layer - 3 switch (613A). Further, the layer - 3 switch (613A) is configured to forward the data packet received on link 851 to the outgoing link 852 in order to transmit the data packet to the layer - 2 switch 623A (which is included in another block). The switch 623A then forwards the data packet on its outgoing link 853 to the switch 627A, and the switch 627A then forwards the data packet on the outgoing link 854 for delivery to its intended destination. Note that the intermediate network devices forward the route of Flow 2 (i.e., the flow depicted by the dashed line) in a manner similar to the above (referring to Flow 1) for delivery to its intended destination. Note that, compared with Figure 8 the scenario depicted (using ECMP routing) in Fig.13 the scenario depicted avoids routing conflicts.

[0193] Thus, in this way, in one embodiment of the GPU - based policy routing mechanism, Fig.14 each network device in the network architecture of

[0194] Fig.14 is configured to bind the incoming port / link to the outgoing port / link to avoid conflicts. It should be recognized that at each network device within the cloud infrastructure, there is a one - to - one correspondence between the incoming port link and the outgoing port link, i.e., the mapping between the incoming port link and the outgoing port link is performed independently of the flow and / or the protocol by which the flow is executed. Additionally, in the case where an outgoing link of a particular network device fails, according to some embodiments, the network device is configured to switch its routing policy from GPU - based policy routing to standard ECMP routing to obtain a new available output link (from various available output links) and send the flow to the new output link. Note that in this case, flow conflicts may occur, resulting in congestion. Fig.14 The processing depicted in Fig.14 can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, in hardware, or a combination thereof. The software can be stored on a non - transitory storage medium (e.g., on a memory device). Fig.14 Describes various processing steps that occur in a specific order or sequence, but this is not intended to be limiting. In some alternative embodiments, these steps may be performed in a different order, or some steps may be performed in parallel.

[0195] The process begins at step 1405, where multiple GPU clusters are provided in a network architecture, and these clusters are communicatively coupled to each other via multiple network devices (e.g., switches) arranged in a hierarchical manner. It should be recognized that the first GPU cluster among the multiple GPU clusters operates at a first speed (e.g., 100G), while the second GPU cluster among the multiple GPU clusters operates at a second speed different from the first speed (e.g., 400G). The process in step 1410 performs a pre-configuration step, where a routing policy is configured for each network device included in the network architecture. Note that the routing policy corresponds to mapping the incoming port link of the network device to the outgoing port link.

[0196] In step 1415, a network device (e.g., the first network device) receives a data packet transmitted by a graphics processing unit (GPU) of a host machine. In step 1420, the network device determines the incoming port / link of the received data packet. In step 1425, the network device identifies the outgoing port / link corresponding to the incoming port / link (of the received data packet) based on the policy routing information (i.e., the routing information pre-configured in step 1410). According to some embodiments, the policy routing information corresponds to a pre-configured GPU routing table of the network device, which binds each incoming port link of the network device to a unique outgoing link port of the network device.

[0197] Then, the process moves to step 1430, where a query is executed to determine whether the outgoing port link is in a working state, e.g., whether the outgoing link is active. If the response to the query is affirmative (i.e., the link is active), then the process moves to step 1435, otherwise, if the response to the query is negative (i.e., the link is in a faulty state / inactive state), then the process moves to step 1450.

[0198] In step 1435, the network device forwards the received data packet to another network device using the egress port link (identified in step 1425). Additionally, the processing in step 1440 performs another query to determine whether the data packet has reached its intended destination. If the response to the query is affirmative, then the processing simply terminates (step 1460). Otherwise, if the response to the query in step 1440 is negative, then the processing moves to step 1445, where the next network device (i.e., the next-hop network device on the path from the source host machine to the destination host machine) processes the data packet. Specifically, the next network device repeats steps 1420, 1425, 1430, 1435, and 1440.

[0199] If the response to the query in step 1430 is negative, then in step 1450, the network device obtains the flow information of the data packet. For example, the flow information may correspond to a quadruple associated with the data packet (i.e., source port, destination port, source IP address, destination IP address). Based on the obtained flow information, the network device uses ECMP routing to identify a new egress port link, i.e., an available egress port link. Then, the processing moves to step 1455, where the network device forwards the data packet using the newly obtained egress port link (in step 1450). Thereafter, the processing loops back to step 1440 to repeat the processing until the data packet is delivered to its intended destination.

[0200] According to some embodiments of the present disclosure, in another implementation of the GPU-based policy routing framework, only a subset of the network devices in the network architecture are configured to bind the ingress port / link to the egress port / link. In this implementation, the subset of network devices that implement the GPU-based policy routing mechanism corresponds to the switches included at the layer 1 and layer 2 levels of the switch hierarchy. For example, refer to Fig.13 , switches 617A, 627A (included in layer 1) and switches 615A, 615B, 623A, 623B (included in layer 2) implement policy-based routing. In this embodiment, the switches at layer 1 and layer 2 levels bind incoming uplinks / ports to the outgoing uplinks / ports of the switches. Note that in this embodiment, the switches included at the layer 3 level do not implement GPU policy-based routing. Instead, the switches at this level can implement standard routing protocols (e.g., ECMP routing) to route traffic. Additionally, it should be recognized that in both embodiments of GPU policy-based routing, when a particular switch receives a packet to be forwarded, the switch determines whether the next hop of the packet is on one of the switch's downlinks. If so, then in some embodiments, the packet will be transmitted without policy-based routing, i.e., forwarded using a standard routing protocol. If the packet is to be forwarded on an uplink, then the switch can use policy-based routing - i.e., depending on the layer in which the switch is included.

[0201] According to some embodiments, Fig. 6A The deployment of a GPU supercluster architecture such as 6 or 6B poses challenges to the management of the global address space (e.g., the MAC addresses of each GPU included in multiple GPU clusters). Specifically, with the deployment of large-scale GPU clusters, each layer 1 switch needs to manage and maintain (i.e., store) a MAC address table (e.g., a forwarding table) that includes the addresses of each GPU included in the cluster. Doing so may cause the forwarding table of the switch to overflow. There are two issues associated with this scenario: the switch is associated with the control plane (i.e., the plane where BGP and other routing protocols reside) and the data plane (i.e., the plane where the forwarding table resides). The problem of maintaining an efficient address space exists in both the control plane and the data plane.

[0202] It should be recognized that a forwarding table is needed to store MAC addresses (i.e., for overlaying the customer network) as well as IP addresses (i.e., for the underlying physical network). Since the storage space in the forwarding table is limited, the problem is more severe in the scenario of deploying Fig. 6A a GPU supercluster architecture such as 6 or 6B. In other words, the problem that arises due to the limited storage space of the forwarding table is how to limit the size of the forwarding table so that the network can be expanded without having to expand the size of the forwarding table. Techniques that can be used in the control plane and the data plane respectively are described below to provide a mechanism for efficiently managing the storage space of the forwarding table.

[0203] Fig.15A Depicts the architecture of a hybrid GPU cluster according to some embodiments, which illustrates the placement of route reflectors. For ease of explanation, Fig.15A The architecture is depicted as including two blocks, namely, a first block 1501 and a second block 1521. Each of the first and second blocks includes a hierarchy of switches deployed therein. For example, as Fig.15A shown, the first block 1501 includes a first-layer switch (i.e., a layer 1 switch) labeled 1503A and a second-layer switch (i.e., a layer 2 switch) labeled 1503B. The second block 1521 includes a first-layer switch labeled 1523A and a second-layer switch labeled 1523B.

[0204] Similar to Fig. 6A the architecture of, Fig.15A the GPU supercluster architecture of includes third-layer switches (i.e., layer 3 switches) labeled 1503C and 1523C, respectively. It should be appreciated that the multiple switches included in the third-layer switches can be partitioned into multiple groups of third-layer switches (e.g., groups 1503C and 1523C, respectively). Further, note that similar to Fig. 6A the architecture of, the first-layer switches are communicatively coupled to a plurality of GPU clusters at one end and communicatively coupled to the second-layer switches at the other end. In turn, the second-layer switches communicatively couple the first-layer switches to the third-layer switches, where the third-layer switches communicatively couple different blocks included in the GPU supercluster architecture.

[0205] According to some embodiments, one or more switches are selected from the third-layer switches to form a set of target switches. Such target switches are referred to herein as route reflectors. Note that the selection of the set of target switches can (in some implementations) be performed in a random manner. In other implementations, route reflectors can be selected from each of the multiple groups of third-layer switches. For example, the first switch included in each group of the third-layer switches can be designated to perform the function of the route reflector as described below. According to some embodiments, the total number of target switches included in the set of target switches is in the range of 4 to 16.

[0206] According to some embodiments, each switch included in the first - layer switches (1503A, 1523A) forms a peer - to - peer connection (e.g., BGP peer) with each of one or more target switches (i.e., route reflectors) included in the third - layer switches (1503C, 1523C). The route reflectors are configured to reduce the address information (e.g., MAC address) maintained by the layer - 1 switches in the forwarding table, as described below. It should be recognized that when a particular layer - 1 switch receives a data packet from a GPU (i.e., a GPU coupled to the particular layer - 1 switch), the layer - 1 switch can transmit the address information of the GPU to each target switch via the peer - to - peer connection. In this way, each target switch included in the layer - 3 switches receives the address information of each GPU included in multiple GPU clusters.

[0207] After receiving the address information of each GPU included in multiple GPU clusters, each target switch processes the received address information to generate multiple sets of address information. For example, according to one embodiment, a particular target switch can filter the received address information based on certain conditions to generate one or more sets of address information. In one instance, the target switch can filter the received GPU address information based on the VLAN of the customer to which the GPU belongs. In other words, the sets of address information generated by the target switch correspond to grouping together the GPUs associated with a customer. Additionally, each target switch can advertise / transmit the generated one or more sets of address information to each switch included in the layer - 1 switches. A particular layer - 1 switch then stores only a subset of the received one or more sets of address information (in the forwarding table) according to conditions. For example, if a particular layer - 1 switch is associated with a particular customer VLAN, then the particular layer - 1 switch can store only the MAC addresses of the GPUs belonging to the same customer VLAN. In this way, a particular layer - 1 switch can ignore (e.g., discard) the other received sets of address information corresponding to GPUs associated with other customer VLANs. It should be recognized that the target switch can utilize other conditions to generate one or more sets of address information. For example, the target switch can filter the address information of the GPUs based on the type of GPU (i.e., blocks of GPUs operating at different speeds). In this way, the route reflector reduces the MAC - address space in the forwarding table in the control plane.

[0208] According to some embodiments, techniques can be used at the data plane layer to further reduce the number of MAC addresses maintained in the forwarding table. For example, through one technique, a layer 1 switch can be configured to further be configured to purge entries from the address table based on a timer associated with the entry. Additionally, a least recently used mechanism can be employed, where based on the usage of a particular entry (e.g., least used), the entry can be purged from the forwarding table. Yet another approach can include aggregating entries in the IP table (corresponding to addresses of underlying physical network components), thereby providing more storage space for MAC entries in the forwarding table. Thus, through the above techniques, embodiments of the present disclosure limit the size of the forwarding table in order to scale the GPU cluster network without having to compromise in terms of scaling the size of the forwarding table.

[0209] Fig. 15B FIG. 1550 is a flow chart illustrating steps performed by a route reflector in managing the size of an address table stored in a management switch according to certain embodiments. Fig. 15B The processing depicted therein can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of a corresponding system. The software can be stored on a non-transitory storage medium (e.g., a memory device). Fig. 15B The methods presented therein and described below are intended to be illustrative and non-limiting. Although Fig. 15B various processing steps are described as occurring in a particular order or sequence, this is not intended to be limiting. In certain alternative embodiments, these steps can be performed in some different order, or some steps can also be performed in parallel.

[0210] The processing begins at step 1555, where a plurality of GPU clusters are provided in a network architecture, and these clusters are communicatively coupled to each other via a plurality of network devices (e.g., switches) arranged in a hierarchical manner. It should be appreciated that a first GPU cluster among the plurality of GPU clusters operates at a first speed (e.g., 100G), while a second GPU cluster among the plurality of GPU clusters operates at a second speed different from the first speed (e.g., 400G). The hierarchy of switches includes at least a layer 1 switch, a layer 2 switch, and a layer 3 switch.

[0211] In step 1560, the process selects one or more switches from the layer-3 switches to form a set of target switches (i.e., route reflectors). In step 1565, each target switch included in the set of target switches receives address information (e.g., MAC address) of each GPU included in the multiple GPU clusters. As previously described, each layer-1 switch is configured to establish a peer connection (e.g., a BGP connection) with each target switch included in the layer-3 switches. According to some embodiments, the address information of the GPUs can be transmitted to the target switches via the peer connections (e.g., by the layer-1 switches when receiving a data packet from a GPU).

[0212] Thereafter, the process moves to step 1570, where each target switch generates multiple sets of address information. According to certain embodiments, these sets of address information can be generated by filtering / grouping the received GPU address information based on certain conditions. For example, as previously described, the target switch can filter the received GPU address information based on the VLAN of the customer to which the GPU belongs. In other words, the sets of address information generated by the target switch correspond to grouping the GPUs associated with a customer together. In step 1575, the target switch advertises / transmits the multiple sets of address information to each switch included in the layer-1 level switches. After receiving the sets of address information, the switches included in the layer-1 level switches store a subset of the multiple sets of address information according to conditions to save storage space of the forwarding table.

[0213] Examples of cloud infrastructure

[0214] As pointed out above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider can also supply various services to accompany these infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering, etc.). Therefore, since these services may be policy-driven, IaaS users can be able to implement policies to drive load balancing to maintain the availability and performance of applications.

[0215] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and can use the cloud provider's services to install the remaining elements of the application stack. For example, a user can log in to an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as a database, create storage buckets for workloads and backups, and even install enterprise software into that VM. Then, the customer can use the provider's services to perform various functions, including balancing network traffic, troubleshooting applications, monitoring performance, managing disaster recovery, etc.

[0216] In most cases, the cloud computing model will require the involvement of a cloud provider. A cloud provider can be, but is not necessarily, a third-party service that specifically provides (e.g., provisions, rents, sells) IaaS. An entity may also choose to deploy a private cloud and thus become its own infrastructure service provider.

[0217] In some examples, IaaS deployment is the process of placing a new application or a new version of an application onto a prepared application server, etc. It can also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is typically managed by the cloud provider and is below the hypervisor layer (e.g., servers, storage devices, network hardware, and virtualization). Thus, the customer can be responsible for handling the (OS), middleware, and / or application deployment (e.g., on self-service virtual machines, etc. that can be launched on demand).

[0218] In some examples, IaaS provisioning can refer to obtaining computers or virtual hosts for use and even installing the required libraries or services on them. In most cases, deployment does not include provisioning, and provisioning may need to be performed first.

[0219] In some cases, there are two different challenges with IaaS provisioning. First, there is an initial challenge in provisioning the initial set of infrastructure before anything is running. Second, once everything has been provisioned, there is the challenge of evolving the existing infrastructure (e.g., adding new services, changing services, removing services, etc.). In some cases, these two challenges can be addressed by enabling the configuration of the infrastructure to be defined in a declarative manner. In other words, the infrastructure (e.g., which components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which resources and how they work together) can be described in a declarative manner. In some cases, once the topology is defined, a workflow for creating and / or managing the different components described in the configuration files can be generated.

[0220] In some examples, the infrastructure can have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., a potentially on-demand pool of configurable and / or shared computing resources), also referred to as the core network. In some examples, one or more security group rules can also be provisioned to define how the security of the network will be set up and one or more virtual machines (VMs). Other infrastructure elements such as load balancers, databases, etc. can also be provisioned. As more and more infrastructure elements are desired and / or added, the infrastructure can evolve gradually.

[0221] In some cases, continuous deployment techniques can be employed to enable the deployment of infrastructure code across various virtual computing environments. Additionally, the described techniques can enable infrastructure management within these environments. In some examples, a service team can write code that is desired to be deployed to one or more but typically many different production environments (e.g., across various different geographical locations, sometimes spanning the entire world). However, in some examples, the infrastructure on which the code will be deployed must be set up first. In some cases, provisioning can be done manually, resources can be provisioned using provisioning tools, and / or once the infrastructure is provisioned, the code can be deployed using deployment tools.

[0222] Fig.16 FIG. 2000 is a block diagram illustrating an example pattern of an IaaS architecture according to at least one embodiment. A service operator 1602 can be communicatively coupled to a secure host lease 1604 that can include a virtual cloud network (VCN) 1606 and a secure host subnet 1608. In some examples, the service operator 1602 can use one or more client computing devices, which can be portable handheld devices (e.g., cellular phones, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google head-mounted displays), running software such as Microsoft Windows ), and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, etc., and supporting the Internet, email, short message service (SMS), or other communication protocols. Alternatively, the client computing device can be a general-purpose personal computer, including, for example, personal computers and / or laptop computers running various versions of Microsoft Apple and / or Linux operating systems. The client computing device can be running various commercially available A workstation computer running a UNIX-like operating system, including but not limited to any of various GNU / Linux operating systems (such as, for example, Google Chrome OS). Alternatively or additionally, the client computing device can be any other electronic device, such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a gesture input device), and / or a personal messaging device capable of communicating via a network that can access VCN 1606 and / or the Internet.

[0223] VCN 1606 can include a Local Peer Gateway (LPG) 1610, which can be communicatively coupled to a Secure Shell (SSH) VCN 1612 via the LPG 1610 included in the SSH VCN 1612. The SSH VCN 1612 can include an SSH subnet 1614, and the SSH VCN 1612 can be communicatively coupled to a control plane VCN 1616 via the LPG 1610 included in the control plane VCN 1616. Additionally, the SSH VCN 1612 can be communicatively coupled to a data plane VCN 1618 via the LPG 1610. The control plane VCN 1616 and the data plane VCN 1618 can be included in a service lease 1619 that can be owned and / or operated by an IaaS provider.

[0224] The control plane VCN 1616 can include a control plane Demilitarized Zone (DMZ) layer 1620 that acts as a perimeter network (e.g., a portion of a corporate network between an intranet and an external network). Servers based on the DMZ can assume limited liability and help control security vulnerabilities. Additionally, the DMZ layer 1620 can include one or more Load Balancer (LB) subnets 1622, a control plane application layer 1624 that can include one or more application subnets 1626, and a control plane data layer 1628 that can include one or more database (DB) subnets 1630 (e.g., one or more front-end DB subnets and / or one or more back-end DB subnets). The one or more LB subnets 1622 included in the control plane DMZ layer 1620 can be communicatively coupled to the one or more application subnets 1626 included in the control plane application layer 1624 and an Internet gateway 1634 that can be included in the control plane VCN 1616, and the one or more application subnets 1626 can be communicatively coupled to the one or more DB subnets 1630 included in the control plane data layer 1628, as well as a service gateway 1636 and a Network Address Translation (NAT) gateway 1638. The control plane VCN 1616 can include a service gateway 1636 and a NAT gateway 1638.

[0225] The control plane VCN 1616 may include a data plane mirror application layer 1640, which may include one or more application subnets 1626. The one or more application subnets 1626 included in the data plane mirror application layer 1640 may include virtual network interface controllers (VNICs) 1642 that may execute compute instances 1644. The compute instances 1644 may communicatively couple the one or more application subnets 1626 of the data plane mirror application layer 1640 to the one or more application subnets 1626 that may be included in the data plane application layer 1646.

[0226] The data plane VCN 1618 may include a data plane application layer 1646, a data plane DMZ layer 1648, and a data plane data layer 1650. The data plane DMZ layer 1648 may include one or more LB subnets 1622, which may be communicatively coupled to the one or more application subnets 1626 of the data plane application layer 1646 and the Internet gateway 1634 of the data plane VCN 1618. The one or more application subnets 1626 may be communicatively coupled to the service gateway 1636 of the data plane VCN 1618 and the NAT gateway 1638 of the data plane VCN 1618. The data plane data layer 1650 may also include one or more DB subnets 1630 that may be communicatively coupled to the one or more application subnets 1626 of the data plane application layer 1646.

[0227] The Internet gateways 1634 of the control plane VCN 1616 and the data plane VCN 1618 may be communicatively coupled to a metadata management service 1652, which may be communicatively coupled to the public Internet 1654. The public Internet 1654 may be communicatively coupled to the NAT gateways 1638 of the control plane VCN 1616 and the data plane VCN 1618. The service gateways 1636 of the control plane VCN 1616 and the data plane VCN 1618 may be communicatively coupled to cloud services 1656.

[0228] In some examples, the service gateway 1636 of the control plane VCN 1616 or the data plane VCN 1618 may make an application programming interface (API) call to the cloud services 1656 without going through the public Internet 1654. The API call from the service gateway 1636 to the cloud services 1656 may be one-way: the service gateway 1636 may make an API call to the cloud services 1656, and the cloud services 1656 may send the requested data to the service gateway 1636. However, the cloud services 1656 may not initiate an API call to the service gateway 1636.

[0229] In some examples, the secure host lease 1604 can be directly connected to the service lease 1619, which otherwise could be isolated. The secure host subnet 1608 can communicate with the SSH subnet 1614 via the LPG 1610, and the LPG 1610 can enable two-way communication on otherwise isolated systems. Connecting the secure host subnet 1608 to the SSH subnet 1614 can enable the secure host subnet 1608 to access other entities within the service lease 1619.

[0230] The control plane VCN 1616 can allow users of the service lease 1619 to set or otherwise provision desired resources. The desired resources provisioned in the control plane VCN 1616 can be deployed or otherwise used in the data plane VCN 1618. In some examples, the control plane VCN 1616 can be isolated from the data plane VCN 1618, and the data plane mirror application layer 1640 of the control plane VCN 1616 can communicate with the data plane application layer 1646 of the data plane VCN 1618 via the VNIC 1642, and the VNIC 1642 can be included in both the data plane mirror application layer 1640 and the data plane application layer 1646.

[0231] In some examples, a user or customer of the system can make requests, such as create, read, update, or delete (CRUD) operations, via the public Internet 1654, which can transmit the requests to the metadata management service 1652. The metadata management service 1652 can transmit the requests to the control plane VCN 1616 via the Internet gateway 1634. The requests can be received by the (one or more) LB subnets 1622 included in the control plane DMZ layer 1620. The (one or more) LB subnets 1622 can determine that the requests are valid, and in response to that determination, the (one or more) LB subnets 1622 can transmit the requests to the (one or more) application subnets 1626 included in the control plane application layer 1624. If the requests are verified and require a call to the public Internet 1654, then the call to the public Internet 1654 can be transmitted to the NAT gateway 1638, which can make the call to the public Internet 1654. The memory where the requests may expect to be stored can be stored in the (one or more) DB subnets 1630.

[0232] In some examples, the data plane mirroring application layer 1640 can facilitate direct communication between the control plane VCN 1616 and the data plane VCN 1618. For example, it may be desirable to apply configuration changes, updates, or other appropriate modifications to resources contained in the data plane VCN 1618. Via the VNIC 1642, the control plane VCN 1616 can communicate directly with the resources contained in the data plane VCN 1618 and thereby perform configuration changes, updates, or other appropriate modifications.

[0233] In some embodiments, the control plane VCN 1616 and the data plane VCN 1618 can be included in a service lease 1619. In this case, the user or customer of the system may not own or operate the control plane VCN 1616 or the data plane VCN 1618. Instead, the IaaS provider can own or operate the control plane VCN 1616 and the data plane VCN 1618, both of which can be included in the service lease 1619. This embodiment can enable isolation of the network that may prevent a user or customer from interacting with the resources of other users or other customers. Additionally, this embodiment can allow the user or customer of the system to privately store databases without relying on the public Internet 1654, which may not have the desired threat prevention level, for storage.

[0234] In other embodiments, the (one or more) LB subnets 1622 included in the control plane VCN 1616 can be configured to receive signals from the service gateway 1636. In this embodiment, the control plane VCN 1616 and the data plane VCN 1618 can be configured to be invoked by the customers of the IaaS provider without invoking the public Internet 1654. The customers of the IaaS provider may desire this embodiment because the (one or more) databases used by the customers can be controlled by the IaaS provider and can be stored on the service lease 1619, which may be isolated from the public Internet 1654.

[0235] Fig.17 is a block diagram 1700 illustrating another example pattern of an IaaS architecture according to at least one embodiment. A service operator 1702 (e.g., Fig.16 the service operator 1602) can be communicatively coupled to a secure host lease 1704 (e.g., Fig.16 the secure host lease 1604), which can include a virtual cloud network (VCN) 1706 (e.g., Fig.16 the VCN 1606) and a secure host subnet 1708 (e.g., Fig.16 the secure host subnet 1608). The VCN 1706 can include a local peering gateway (LPG) 1710 (e.g., Fig.16 of the LPG 1610), which can be communicatively coupled to a security shell (SSH) VCN 1712 via the LPG 1610 included in the SSH VCN 1712 (e.g., Fig.16 of the SSH VCN 1612). The SSH VCN 1712 can include an SSH subnet 1714 (e.g., Fig.16 of the SSH subnet 1614), and the SSH VCN 1712 can be communicatively coupled to a control plane VCN 1716 via the LPG 1710 included in the control plane VCN 1716 (e.g., Fig.16 of the control plane VCN 1616). The control plane VCN 1716 can be included in a service tenancy 1719 (e.g., Fig.16 of the service tenancy 1619), and a data plane VCN 1718 (e.g., Fig.16 of the data plane VCN 1618) can be included in a customer tenancy 1721 that may be owned or operated by a user or customer of the system.

[0236] The control plane VCN 1716 can include a control plane DMZ layer 1720 that can include one or more LB subnets 1722 (e.g., Fig.16 one or more LB subnets 1622 of), a control plane application layer 1724 that can include one or more application subnets 1726 (e.g., Fig.16 of the control plane DMZ layer 1620), a control plane application layer 1724 that can include one or more application subnets 1726 (e.g., Fig.16 one or more application subnets 1626 of), a control plane data layer 1728 that can include one or more database (DB) subnets 1730 (e.g., similar to Fig.16 one or more DB subnets 1630 of), a control plane data layer 1728 (e.g., Fig.16 of the control plane data layer 1628). The one or more LB subnets 1722 included in the control plane DMZ layer 1720 can be communicatively coupled to the one or more application subnets 1726 included in the control plane application layer 1724 and an Internet gateway 1734 that can be included in the control plane VCN 1716 (e.g., Fig.16 of the Internet gateway 1634), and the one or more application subnets 1726 can be communicatively coupled to the one or more DB subnets 1730 included in the control plane data layer 1728, as well as a service gateway 1736 (e.g., Fig.16 of the service gateway) and a network address translation (NAT) gateway 1738 (e.g., Fig.16 of the), and a network address translation (NAT) gateway 1738 (e.g., Fig.16 The NAT gateway 2438). The control plane VCN 1716 may include a service gateway 1736 and a NAT gateway 1738.

[0237] The control plane VCN 1716 may include a data plane mirror application layer 1740 that may include (one or more) application subnets 1726 (e.g., Fig.16 the data plane mirror application layer 1640). The (one or more) application subnets 1726 included in the data plane mirror application layer 1740 may include virtual network interface controllers (VNICs) 1742 (e.g., the VNICs 1642) that may execute compute instances 1744 (e.g., similar to Fig.16 the compute instances 1644). The compute instances 1744 may facilitate communication between the (one or more) application subnets 1726 of the data plane mirror application layer 1740 and the (one or more) application subnets 1726 that may be included in the data plane application layer 1746 (e.g., Fig.16 the data plane application layer 1646) via the VNICs 1742 included in the data plane mirror application layer 1740 and the VNICs 1742 included in the data plane application layer 1746.

[0238] The Internet gateway 1734 included in the control plane VCN 1716 may be communicatively coupled to a metadata management service 1752 (e.g., Fig.16 the metadata management service 1652), and the metadata management service 1752 may be communicatively coupled to a public Internet 1754 (e.g., Fig.16 the public Internet 1654). The public Internet 1754 may be communicatively coupled to the NAT gateway 1738 included in the control plane VCN 1716. The service gateway 1736 included in the control plane VCN 1716 may be communicatively coupled to a cloud service 1756 (e.g., Fig.16 the cloud service 1656).

[0239] In some examples, the data plane VCN 1718 may be included in a customer lease 1721. In such a case, the IaaS provider may provide a control plane VCN 1716 for each customer, and the IaaS provider may provision a unique compute instance 1744 included in a service lease 1719 for each customer. Each compute instance 1744 may permit communication between the control plane VCN 1716 included in the service lease 1719 and the data plane VCN 1718 included in the customer lease 1721. The compute instances 1744 may permit resources provisioned in the control plane VCN 1716 included in the service lease 1719 to be deployed or otherwise used in the data plane VCN 1718 included in the customer lease 1721.

[0240] In other examples, a customer of an IaaS provider can have a database that resides in customer lease 1721. In this example, the control plane VCN 1716 can include a data plane mirror application layer 1740, which can include one or more application subnets 1726. The data plane mirror application layer 1740 can reside in the data plane VCN 1718, but the data plane mirror application layer 1740 may not be in the data plane VCN 1718. That is, the data plane mirror application layer 1740 can access the customer lease 1721, but the data plane mirror application layer 1740 may not exist in the data plane VCN 1718 or be owned or operated by the customer of the IaaS provider. The data plane mirror application layer 1740 can be configured to make calls to the data plane VCN 1718, but may not be configured to make calls to any entity included in the control plane VCN 1716. The customer may expect to deploy or otherwise use resources provisioned in the control plane VCN 1716 in the data plane VCN 1718, and the data plane mirror application layer 1740 can facilitate the customer's desired deployment or other use of the resources.

[0241] In some embodiments, a customer of an IaaS provider can apply a filter to the data plane VCN 1718. In this embodiment, the customer can determine what the data plane VCN 1718 can access, and the customer can restrict access from the data plane VCN 1718 to the public Internet 1754. The IaaS provider may not be able to apply filters or otherwise control the data plane VCN 1718's access to any external network or database. The customer applying filters and controls to the data plane VCN 1718 included in the customer lease 1721 can help isolate the data plane VCN 1718 from other customers and the public Internet 1754.

[0242] In some embodiments, cloud service 1756 may be invoked by service gateway 1736 to access services that may not be present on public Internet 1754, control plane VCN 1716, or data plane VCN 1718. The connection between cloud service 1756 and control plane VCN 1716 or data plane VCN 1718 may not be real-time or continuous. Cloud service 1756 may exist on a different network owned or operated by an IaaS provider. Cloud service 1756 may be configured to receive calls from service gateway 1736 and may be configured not to receive calls from public Internet 1754. Some cloud services 1756 may be isolated from other cloud services 1756, and control plane VCN 1716 may be isolated from cloud services 1756 that may not be in the same region as control plane VCN 1716. For example, control plane VCN 1716 may be located in "Region 1", and cloud service "Deployment 16" may be located in Region 1 and "Region 2". If service gateway 1736 included in control plane VCN 1716 located in Region 1 makes a call to Deployment 16, then the call may be transmitted to Deployment 16 in Region 1. In this example, control plane VCN 1716 or Deployment 16 in Region 1 may not be communicatively coupled or otherwise communicate with Deployment 16 in Region 2.

[0243] Fig.18 is a block diagram 1800 illustrating another example pattern of an IaaS architecture according to at least one embodiment. Service operator 1802 (e.g., Fig.16 service operator 1602 of Fig.16 may be communicatively coupled to secure host lease 1804 (e.g., Fig.16 secure host lease 1604 of Fig.16 ), which may include virtual cloud network (VCN) 1806 (e.g., Fig.16 VCN 1606 of Fig.16 and secure host subnet 1808 (e.g., Fig.16 secure host subnet 1608 of Fig.16 to the control plane VCN 1616) and is coupled to the data plane VCN 1818 via the LPG 1810 included in the data plane VCN 1818 (e.g., Fig.16 to the data plane 1618). The control plane VCN 1816 and the data plane VCN 1818 can be included in a service tenancy 1819 (e.g., Fig.16 the service tenancy 1619).

[0244] The control plane VCN 1816 can include a control plane DMZ layer 1820 that can include one or more load balancer (LB) subnets 1822 (e.g., Fig.16 one or more LB subnets 1622), a control plane application layer 1824 that can include one or more application subnets 1826 (e.g., similar to Fig.16 one or more application subnets 1626), and a control plane data layer 1828 that can include one or more DB subnets 1830 (e.g., Fig.16 the control plane data layer 1628). The one or more LB subnets 1822 included in the control plane DMZ layer 1820 can be communicatively coupled to the one or more application subnets 1826 included in the control plane application layer 1824 and to an Internet gateway 1834 that can be included in the control plane VCN 1816 (e.g., Fig.16 the Internet gateway 1634), and the one or more application subnets 1826 can be communicatively coupled to the one or more DB subnets 1830 included in the control plane data layer 1828, as well as to a service gateway 1836 (e.g., Fig.16 the service gateway) and a network address translation (NAT) gateway 1838 (e.g., Fig.16 the NAT gateway 2438). The control plane VCN 1816 can include the service gateway 1836 and the NAT gateway 1838. Fig.16 the service gateway) and a network address translation (NAT) gateway 1838 (e.g., Fig.16 the NAT gateway 2438). The control plane VCN 1816 can include the service gateway 1836 and the NAT gateway 1838.

[0245] The data plane VCN 1818 can include a data plane application layer 1846 (e.g., Fig.16 the data plane application layer 1646), a data plane DMZ layer 1848 (e.g., Fig.16 the data plane DMZ layer 1648), and a data plane data layer 1850 (e.g., Fig.16 The data plane data layer 1650). The data plane DMZ layer 1848 may include one or more trusted application subnets 1860 and one or more untrusted application subnets 1862 communicatively coupled to the data plane application layer 1846, and one or more LB subnets 1822 of the Internet gateway 1834 included in the data plane VCN 1818. One or more trusted application subnets 1860 may be communicatively coupled to the service gateway 1836 included in the data plane VCN 1818, the NAT gateway 1838 included in the data plane VCN 1818, and one or more DB subnets 1830 included in the data plane data layer 1850. One or more untrusted application subnets 1862 may be communicatively coupled to the service gateway 1836 included in the data plane VCN 1818 and one or more DB subnets 1830 included in the data plane data layer 1850. The data plane data layer 1850 may include one or more DB subnets 1830 communicatively coupled to the service gateway 1836 included in the data plane VCN 1818.

[0246] One or more untrusted application subnets 1862 may include one or more primary VNICs 1864(1)-(N) communicatively coupled to tenant virtual machines (VMs) 1866(1)-(N). Each tenant VM 1866(1)-(N) may be communicatively coupled to a corresponding application subnet 1867(1)-(N) that may be included in a corresponding container egress VCN 1868(1)-(N), and the corresponding container egress VCNs 1868(1)-(N) may be included in corresponding customer tenancies 1870(1)-(N). Corresponding secondary VNICs 1872(1)-(N) may facilitate communication between one or more untrusted application subnets 1862 included in the data plane VCN 1818 and the application subnets included in the container egress VCNs 1868(1)-(N). Each container egress VCN 1868(1)-(N) may include a NAT gateway 1838 that may be communicatively coupled to the public Internet 1854 (e.g., Fig.16 the public Internet 1654).

[0247] The Internet gateway 1834 included in the control plane VCN 1816 and included in the data plane VCN 1818 may be communicatively coupled to the metadata management service 1852 (e.g., Fig.16 A metadata management system 1652), the metadata management service 1852 can be communicatively coupled to the public Internet 1854. The public Internet 1854 can be communicatively coupled to a NAT gateway 1838 included in the control plane VCN 1816 and included in the data plane VCN 1818. A service gateway 1836 included in the control plane VCN 1816 and included in the data plane VCN 1818 can be communicatively coupled to a cloud service 1856.

[0248] In some embodiments, the data plane VCN 1818 can be integrated with a customer lease 1870. In some cases, such as when it may be desirable to support during code execution, this integration may be useful or desirable for customers of the IaaS provider. The customer may provide code that may be disruptive, may communicate with other customer resources, or may otherwise cause undesirable effects to run. In response to this, the IaaS provider can determine whether to run the code given to the IaaS provider by the customer.

[0249] In some examples, a customer of the IaaS provider can grant the IaaS provider temporary network access and request functionality attached to the data plane layer application 1846. The code running the functionality can be executed in the VMs 1866(1)-(N), and the code can be not configured to run anywhere else on the data plane VCN 1818. Each VM 1866(1)-(N) can be connected to a customer lease 1870. The corresponding containers 1871(1)-(N) included in the VMs 1866(1)-(N) can be configured to run the code. In this case, there can be double isolation (e.g., the containers 1871(1)-(N) run the code, where the containers 1871(1)-(N) may be at least included in the VMs 1866(1)-(N) included in one or more untrusted application subnets 1862), which can help prevent incorrect or otherwise undesirable code from damaging the IaaS provider's network or damaging the networks of different customers. The containers 1871(1)-(N) can be communicatively coupled to the customer lease 1870 and can be configured to transmit or receive data from the customer lease 1870. The containers 1871(1)-(N) can be not configured to transmit or receive data from any other entity in the data plane VCN 1818. After the code execution is completed, the IaaS provider can terminate or otherwise dispose of the containers 1871(1)-(N).

[0250] In some embodiments, the (one or more) trusted application subnets 1860 may run code that may be owned or operated by an IaaS provider. In this embodiment, the (one or more) trusted application subnets 1860 may be communicatively coupled to the (one or more) DB subnets 1830 and configured to perform CRUD operations in the (one or more) DB subnets 1830. The (one or more) untrusted application subnets 1862 may be communicatively coupled to the (one or more) DB subnets 1830, but in this embodiment, the (one or more) untrusted application subnets may be configured to perform read operations in the (one or more) DB subnets 1830. Containers 1871(1)-(N) that may be included in each customer's VMs 1866(1)-(N) and may run code from the customer may not be communicatively coupled to the (one or more) DB subnets 1830.

[0251] In other embodiments, the control plane VCN 1816 and the data plane VCN 1818 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 1816 and the data plane VCN 1818. However, communication may occur indirectly through at least one method. The LPG 1810 may be established by the IaaS provider, which may facilitate communication between the control plane VCN 1816 and the data plane VCN 1818. In another example, the control plane VCN 1816 or the data plane VCN 1818 may invoke a cloud service 1856 via a service gateway 1836. For example, an invocation of the cloud service 1856 from the control plane VCN 1816 may include a request for a service that may communicate with the data plane VCN 1818.

[0252] Fig.19 is a block diagram 1900 illustrating another example pattern of an IaaS architecture according to at least one embodiment. A service operator 1902 (e.g., Fig.16 the service operator 1602) may be communicatively coupled to a secure host lease 1904 (e.g., Fig.16 the secure host lease 1604), which may include a virtual cloud network (VCN) 1906 (e.g., Fig.16 the VCN 1606) and a secure host subnet 1908 (e.g., Fig.16 the secure host subnet 1608). The VCN 1906 may include an LPG 1910 (e.g., Fig.16 the LPG 1610), which may be via an SSH VCN 1912 (e.g., Fig.16 The LPG 1910 in the SSH VCN 1612 is communicatively coupled to the SSH VCN 1912. The SSH VCN 1912 may include an SSH subnet 1914 (e.g., Fig.16 the SSH subnet 1614), and the SSH VCN 1912 may be communicatively coupled to the control plane VCN 1916 via the LPG 1910 included in the control plane VCN 1916 (e.g., Fig.16 the control plane VCN 1616) and coupled to the data plane VCN 1918 via the LPG 1910 included in the data plane VCN 1918 (e.g., Fig.16 the data plane 1618). The control plane VCN 1916 and the data plane VCN 1918 may be included in a service tenancy 1919 (e.g., Fig.16 the service tenancy 1619).

[0253] The control plane VCN 1916 may include a control plane DMZ layer 1920 that may include one or more LB subnets 1922 (e.g., Fig.16 one or more LB subnets 1622), a control plane application layer 1924 that may include one or more application subnets 1926 (e.g., Fig.16 the control plane DMZ layer 1620), a control plane application layer 1924 that may include one or more application subnets 1926 (e.g., Fig.16 one or more application subnets 1626), a control plane data layer 1928 that may include one or more DB subnets 1930 (e.g., Fig.16 the control plane application layer 1624), a control plane data layer 1928 that may include one or more DB subnets 1930 (e.g., Fig.18 one or more DB subnets 1830). One or more LB subnets 1922 included in the control plane DMZ layer 1920 may be communicatively coupled to one or more application subnets 1926 included in the control plane application layer 1924 and an Internet gateway 1934 that may be included in the control plane VCN 1916 (e.g., Fig.16 the control plane data layer 1628), and one or more application subnets 1926 may be communicatively coupled to one or more DB subnets 1930 included in the control plane data layer 1928, a service gateway 1936 (e.g., Fig.16 the Internet gateway 1634), and a network address translation (NAT) gateway 1938 (e.g., Fig.16 the service gateway), and a network address translation (NAT) gateway 1938 (e.g., Fig.16 the NAT gateway 1638). The control plane VCN 1916 may include a service gateway 1936 and a NAT gateway 1938.

[0254] The data plane VCN 1918 may include a data plane application layer 1946 (e.g., Fig.16 the data plane application layer 1646 of Fig.16 ), a data plane DMZ layer 1948 (e.g., Fig.16 the data plane DMZ layer 1648 of Fig.18 ), and a data plane data layer 1950 (e.g., Fig.18 the data plane data layer 1650 of

[0255] . The data plane DMZ layer 1948 may include one or more trusted application subnets 1960 (e.g., Fig.16 one or more trusted application subnets 1860 of

[0256] Fig.18 ) and one or more untrusted application subnets 1962 (e.g., Fig.18 one or more untrusted application subnets 1862 of

[0255] ) that are communicatively coupled to the data plane application layer 1946, and one or more LB subnets 1922 of an Internet gateway 1934 included in the data plane VCN 1918. The one or more trusted application subnets 1960 may be communicatively coupled to a service gateway 1936 included in the data plane VCN 1918, a NAT gateway 1938 included in the data plane VCN 1918, and one or more DB subnets 1930 included in the data plane data layer 1950. The one or more untrusted application subnets 1962 may be communicatively coupled to the service gateway 1936 included in the data plane VCN 1918 and one or more DB subnets 1930 included in the data plane data layer 1950. The data plane data layer 1950 may include one or more DB subnets 1930 that are communicatively coupled to the service gateway 1936 included in the data plane VCN 1918. (One or more) untrusted application subnets 1962 may include one or more primary VNICs 1964(1)-(N) that are communicatively coupled to tenant virtual machines (VMs) 1966(1)-(N) residing within the one or more untrusted application subnets 1962. Each tenant VM 1966(1)-(N) may run code in a respective container 1967(1)-(N) and be communicatively coupled to an application subnet 1926 in the data plane application layer 1946 that may be included in a container egress VCN 1968. Respective secondary VNICs 1972(1)-(N) may facilitate communication between the one or more untrusted application subnets 1962 included in the data plane VCN 1918 and the application subnet included in the container egress VCN 1968. The container egress VCN may include a NAT gateway 1938 that is communicatively coupled to a public Internet 1954 (e.g., Fig.16 the public Internet 1654 of

[0256] The Internet gateway 1934 included in the control plane VCN 1916 and the Internet gateway 1934 included in the data plane VCN 1918 can be communicatively coupled to a metadata management service 1952 (e.g., Fig.16 the metadata management system 1652), and the metadata management service 1952 can be communicatively coupled to the public Internet 1954. The public Internet 1954 can be communicatively coupled to the NAT gateway 1938 included in the control plane VCN 1916 and included in the data plane VCN 1918. The service gateway 1936 included in the control plane VCN 1916 and included in the data plane VCN 1918 can be communicatively coupled to the cloud service 1956.

[0257] In some examples, Fig.19 the pattern shown in the architecture of the block diagram 1900 of Fig.18 can be considered an exception to the pattern shown in the architecture of the block diagram 1800 of

[0258] and may be desired by customers of the IaaS provider if the IaaS provider cannot communicate directly with the customer (e.g., a disconnected region). The customer can access in real time the respective containers 1967(1)-(N) included in each customer's VMs 1966(1)-(N). The containers 1967(1)-(N) can be configured to make calls to the respective secondary VNICs 1972(1)-(N) included in the (one or more) application subnets 1926 of the data plane application layer 1946, and the data plane application layer 1946 can be included in the container egress VCN 1968. The secondary VNICs 1972(1)-(N) can transmit the calls to the NAT gateway 1938, and the NAT gateway 1938 can transmit the calls to the public Internet 1954. In this example, the containers 1967(1)-(N) that can be accessed by the customer in real time can be isolated from the control plane VCN 1916 and can be isolated from other entities included in the data plane VCN 1918. The containers 1967(1)-(N) can also be isolated from resources from other customers.In other examples, a customer can use containers 1967(1)-(N) to invoke cloud service 1956. In this example, the customer can run code in containers 1967(1)-(N) that requests services from cloud service 1956. Containers 1967(1)-(N) can transmit the request to secondary VNICs 1972(1)-(N), which can transmit the request to a NAT gateway that can transmit the request to public Internet 1954. Public Internet 1954 can transmit the request via Internet gateway 1934 to one or more LB subnets 1922 included in control plane VCN 1916. In response to determining that the request is valid, the one or more LB subnets can transmit the request to one or more application subnets 1926, which can transmit the request to cloud service 1956 via service gateway 1936.

[0259] It should be appreciated that the IaaS architectures 1600, 1700, 1800, 1900 depicted in the figures may have other components than those depicted. Additionally, the embodiments shown in the figures are merely some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In some other embodiments, the IaaS system may have more or fewer components than shown in the figures, may combine two or more components, or may have a different configuration or arrangement of components.

[0260] In certain embodiments, the IaaS systems described herein may include application suite, middleware, and database service offerings that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such an IaaS system is the Oracle Cloud Infrastructure (OCI) offered by the present assignee.

[0261] Fig. 20 An example computer system 2000 in which various embodiments may be implemented is illustrated. System 2000 may be used to implement any of the computer systems described above. As shown, computer system 2000 includes a processing unit 2004 that communicates with a plurality of peripheral subsystems via a bus subsystem 2002. These peripheral subsystems may include a processing acceleration unit 2006, an I / O subsystem 2008, a storage subsystem 2018, and a communication subsystem 2024. Storage subsystem 2018 includes a tangible computer-readable storage medium 2022 and system memory 2010.

[0262] The bus subsystem 2002 provides a mechanism for enabling the various components and subsystems of the computer system 2000 to communicate with each other as intended. Although the bus subsystem 2002 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 2002 can be any of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures can include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which can be implemented as a Mezzanine bus manufactured to the IEEE P1786.1 standard.

[0263] The processing unit 2004, which can be implemented as one or more integrated circuits (e.g., a conventional microprocessor or microcontroller), controls the operation of the computer system 2000. One or more processors can be included in the processing unit 2004. These processors can include single-core or multi-core processors. In certain embodiments, the processing unit 2004 can be implemented as one or more independent processing units 2032 and / or 2034, each of which includes a single-core or multi-core processor. In other embodiments, the processing unit 2004 can also be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.

[0264] In various embodiments, the processing unit 2004 can execute various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed can reside in the (one or more) processors 2004 and / or the storage subsystem 2018. Through appropriate programming, the (one or more) processors 2004 can provide the various functions described above. The computer system 2000 can additionally include a processing acceleration unit 2006, which can include a digital signal processor (DSP), a dedicated processor, and so on.

[0265] The I / O subsystem 2008 can include user interface input devices and user interface output devices. User interface input devices can include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keyboard, an audio input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices can include, for example, motion sensing and / or gesture recognition devices, such as Microsoft A motion sensor that enables a user to control and interact with an input device such as a Microsoft 360 game controller through a natural user interface using gestures and voice commands. The user interface input device can also include an eye gesture recognition device, such as one that detects eye activity from the user (e.g., "blinking" when taking a photo and / or making a menu selection) and converts the eye gesture into an input to the input device (e.g., a Google blink detector). Additionally, the user interface input device can include a voice recognition sensing device that enables the user to interact with a voice recognition system (e.g., a navigator) through voice commands. The user interface input device can also include, but is not limited to, a three-dimensional (3D) mouse, joystick or pointing stick, game pad, and graphics tablet, as well as audio / video devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye tracking devices. Additionally, the user interface input device can include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound devices. The user interface input device can also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.

[0266]

[0267] The user interface output device can include a display subsystem, indicator lights, or a non-visual display such as an audio output device, etc. The display subsystem can be a cathode ray tube (CRT), a flat panel device such as one using a liquid crystal display (LCD) or a plasma display, a projection device, a touch screen, etc. In general, the use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from the computer system 2000 to the user or other computers. For example, the user interface output device can include, but is not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, voice output devices, and modems.

[0268] The computer system 2000 can include a storage subsystem 2018 that contains software elements and is shown as currently residing in the system memory 2010. The system memory 2010 can store program instructions that are loadable and executable on the processing unit 2004, as well as data generated during the execution of these programs.

[0269] Depending on the configuration and type of the computer system 2000, the system memory 2010 can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that can be immediately accessed by the processing unit 2004 and / or are currently being operated on and executed by the processing unit 2004. In some implementations, the system memory 2010 can include multiple different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), such as containing basic routines that help transfer information between elements of the computer system 2000 during startup, can typically be stored in the ROM. By way of example, but not limitation, the system memory 2010 is also shown to include application programs 2012 that can include client applications, web browsers, middleware applications, relational database management systems (RDBMS), etc., program data 2014, and an operating system 2016. As an example, the operating system 2016 can include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems such as iOS, Phone, OS, 16OS and OS operating systems.

[0270] The storage subsystem 2018 can also provide a tangible computer-readable storage medium for storing the basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that provides the above functionality when executed by a processor can be stored in the storage subsystem 2018. These software modules or instructions can be executed by the processing unit 2004. The storage subsystem 2018 can also provide a repository for storing data used in accordance with the present disclosure.

[0271] The storage subsystem 2000 can also include a computer-readable storage medium reader 2020 that can be further connected to a computer-readable storage medium 2022. Together with and, optionally, in combination with the system memory 2010, the computer-readable storage medium 2022 can comprehensively represent remote, local, fixed, and / or removable storage devices plus storage media for temporarily and / or more persistently containing, storing, sending, and retrieving computer-readable information.

[0272] The computer-readable storage medium 2022 that contains code or portions of code may also include any suitable medium known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented by any method or technology for the storage and / or transmission of information. This may include tangible computer-readable storage media such as RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or other tangible computer-readable media. This may also include non-tangible computer-readable media such as data signals, data transmissions or any other medium that can be used to transmit the desired information and can be accessed by the computing system 2000.

[0273] For example, the computer-readable storage medium 2022 may include a hard disk drive that reads from or writes to a non-removable non-volatile magnetic medium, a disk drive that reads from or writes to a removable non-volatile disk, and an optical disk drive that reads from or writes to a removable non-volatile optical disk (such as a CD ROM, DVD, and Blu- ray disk or other optical medium). The computer-readable storage medium 2022 may include, but is not limited to, drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD disks, digital audio tapes, and so on. The computer-readable storage medium 2022 may also include solid-state drives (SSDs) based on non-volatile memory (such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc.), SSDs based on volatile memory (such as solid-state RAM, dynamic RAM, static RAM), DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The disk drive and its associated computer-readable medium may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 2000.

[0274] The communication subsystem 2024 provides an interface to other computer systems and networks. The communication subsystem 2024 serves as an interface for receiving data from other systems and sending data from the computer system 2000 to other systems. For example, the communication subsystem 2024 can enable the computer system 2000 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 2024 can include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular phone technologies such as advanced data network technologies like 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), Wi-Fi (IEEE 802.11 series standards), or other mobile communication technologies, or any combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, as an addition or alternative to the wireless interface, the communication subsystem 2024 can provide a wired network connection (e.g., Ethernet).

[0275] In some embodiments, the communication subsystem 2024 can also receive input communications in the form of structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc. on behalf of one or more users who may use the computer system 2000.

[0276] For example, the communication subsystem 2024 can be configured to receive data feeds 2026 from users of social networks and / or other communication services in real time, such as feeds, updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.

[0277] In addition, the communication subsystem 2024 can also be configured to receive data in the form of continuous data streams, which can include event streams 2028 and / or event updates 2030 of real-time events that can be essentially continuous or unbounded without a clear termination. Examples of applications that produce continuous data can include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and so on.

[0278] The communication subsystem 2024 can also be configured to output structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc. to one or more databases, which can communicate with one or more streaming data source computers coupled to the computer system 2000.

[0279] The computer system 2000 can be one of various types, including handheld portable devices (e.g., cellular phones, computing tablets, PDAs), wearable devices (e.g., Glass head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing system.

[0280] Due to the ever-changing nature of computers and networks, the description of the computer system 2000 depicted in the figures is merely to serve as a specific example. Many other configurations with more or fewer components than the systems depicted in the figures are possible. For example, custom hardware can also be used and / or specific elements can be implemented in hardware, firmware, software (including applets), or combinations thereof. Additionally, connections to other computing devices such as network input / output devices can also be employed. Based on the disclosures and teachings provided herein, those of ordinary skill in the art will recognize other ways and / or methods of implementing the various embodiments.

[0281] Although specific embodiments have been described, various modifications, alterations, alternative constructions, and equivalent forms are also included within the scope of the present disclosure. The embodiments are not limited to operating within certain specific data processing environments but can operate freely within multiple data processing environments. Moreover, although the embodiments have been described using a specific series of transactions and steps, those skilled in the art should appreciate that the scope of the present disclosure is not limited to the described series of transactions and steps. The various features and aspects of the above embodiments can be used alone or in combination.

[0282] Furthermore, although the embodiments have been described using a specific combination of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of the present disclosure. The embodiments can be implemented using only hardware, or only software, or using combinations thereof. The various processes described herein can be implemented in any combination on the same processor or on different processors. Accordingly, in cases where a component or module is described as being configured to perform certain operations, such configuration can be accomplished by, for example, designing electronic circuitry to perform the operations, programming a programmable electronic circuit (such as a microprocessor) to perform the operations, or any combination thereof. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes can use different techniques, or the same pair of processes can use different techniques at different times.

[0283] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it is obvious that additions, subtractions, deletions, and other modifications and changes can be made thereto without departing from the broader spirit and scope set forth in the claims. Thus, although specific disclosed embodiments have been described, these are not intended to be limiting. Various modifications and equivalent forms are within the scope of the following claims.

[0284] In the context of describing the disclosed embodiments, particularly in the context of the following claims, the terms "a", "an", "the", and similar references are to be construed to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by context. Unless otherwise noted, the terms "comprising", "having", "including", and "containing" are to be construed as open-ended terms (i.e., meaning "including but not limited to"). The term "connected" shall be construed to mean partly or wholly contained in, attached to, or joined together, even if there is something in between. Unless otherwise indicated herein, the listing of numerical ranges herein is only intended to be a shorthand method for individually referring to each separate value falling within the range, and each separate value is incorporated into the specification as if it were individually recited herein. Unless otherwise indicated herein or clearly contradicted by context, all methods described herein can be performed in any suitable order. The use of any and all examples, or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the disclosure unless otherwise stated. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0285] Disjunctive language such as the phrase "at least one of X, Y, or Z" is, unless otherwise explicitly stated, intended to be understood in the context of generally representing items, terms, etc. and can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to and should not imply that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z each.

[0286] Preferred embodiments of the present disclosure are described herein, including the best mode known for practicing the present disclosure. Variations of those preferred embodiments will become apparent to those of ordinary skill in the art upon reading the above description. Those of ordinary skill in the art should be able to appropriately adopt such variations and practice the present disclosure in a manner different from that specifically described herein. Accordingly, the present disclosure includes all modifications and equivalent forms of the subject matter recited in the appended claims as permitted by applicable law. In addition, unless otherwise indicated herein, the present disclosure includes any combination of the above elements in all possible variations thereof.

[0287] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in full herein.

[0288] In the foregoing specification, aspects of the present disclosure have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present disclosure is not limited thereto. The various features and aspects disclosed above may be used singly or in combination. In addition, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.< / uniqueid> < / realm>

Claims

1. A method, comprising: providing a plurality of graphics processing unit (GPU) clusters, the plurality of GPU clusters being communicatively coupled to each other via a plurality of network devices arranged in a hierarchy, wherein the plurality of GPU clusters at least include a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed; configuring a routing policy for each of the plurality of network devices, wherein the configuring includes establishing a mapping of each incoming port link of the network device to a unique outgoing port link of the network device; and for a data packet transmitted by a GPU of a host machine and received by a first network device, determining the incoming port link of the first network device on which the data packet is received; identifying, based on the configuration, the outgoing port link corresponding to the incoming port link; and forwarding the data packet on the outgoing port link of the network device.

2. The method according to claim 1, wherein the forwarding step further comprises: verifying a condition associated with the outgoing port link of the first network device; and in response to the condition being satisfied, forwarding the data packet on the outgoing port link of the first network device.

3. The method according to claim 2, wherein the condition corresponds to determining whether the outgoing port link of the first network device is active.

4. The method according to claim 2, further comprises: in response to the condition not being satisfied obtaining, by the first network device, flow information associated with the data packet; performing, by the first network device, an equal-cost multi-path algorithm based on the flow information to obtain a new outgoing port link of the first network device; and forwarding, by the first network device, the data packet on the new outgoing port link of the first network device.

5. The method according to any of the preceding claims, further comprises: repeating the determining, the identifying, and the forwarding until the data packet is delivered to a destination host machine.

6. The method according to any of the preceding claims, wherein the data packet belongs to a GPU workload.

7. The method according to any of the preceding claims, wherein the plurality of network devices correspond to a plurality of switches arranged in a hierarchy, the hierarchy including a first-layer switch, a second-layer switch, and a third-layer switch.

8. The method according to claim 7, wherein the first-layer switch communicatively couples the host machine to the second-layer switch, and the second-layer switch communicatively couples the first-layer switch to the third-layer switch, and wherein the third-layer switch communicatively couples a first block including the first one or more racks hosting the first GPU cluster to a second block including the second one or more racks hosting the second GPU cluster.

9. One or more computer-readable non-transitory media storing computer-executable instructions that, when executed by one or more processors, cause: providing a plurality of graphics processing unit (GPU) clusters, the plurality of GPU clusters being communicatively coupled to each other via a plurality of network devices arranged in a hierarchy, wherein the plurality of GPU clusters at least include a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed; Configure a routing policy for each of the multiple network devices, where the configuration includes establishing a mapping of each incoming port link of the network device to a unique outgoing port link of the network device; And For a data packet transmitted by the GPU of the host machine and received by the first network device, Determine the incoming port link of the first network device on which the data packet is received; Identify the outgoing port link corresponding to the incoming port link based on the configuration; And Forward the data packet on the outgoing port link of the network device.

10. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in claim 9, wherein the forwarding further Includes: Verify the conditions associated with the outgoing port link of the first network device; And In response to satisfying the conditions, forward the data packet on the outgoing port link of the first network device.

11. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in claim 10, wherein the conditions correspond to determining whether the outgoing port link of the first network device is active.

12. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in claim 10, further Includes: In response to not satisfying the conditions Obtain flow information associated with the data packet by the first network device; Execute an equal-cost multi-path algorithm by the first network device based on the flow information to obtain a new outgoing port link of the first network device; And Forward the data packet by the first network device on the new outgoing port link of the first network device.

13. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in any one of claims 9 to 12, further Includes: Repeat the determination, the identification, and the forwarding until the data packet is delivered to the destination host machine.

14. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in any one of claims 9 to 13, wherein the data packet belongs to a GPU workload.

15. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in any one of claims 9 to 14, wherein the multiple network devices correspond to multiple switches arranged in a hierarchy, the hierarchy including a first-layer switch, a second-layer switch, and a third-layer switch.

16. The one or more computer-readable non-transitory media storing computer-executable instructions as recited in claim 15, wherein the first-layer switch communicatively couples the host machine to the second-layer switch, and the second-layer switch communicatively couples the first-layer switch to the third-layer switch, and wherein the third-layer switch communicatively couples a first block including the first one or more racks hosting the first GPU cluster to a second block including the second one or more racks hosting the second GPU cluster.

17. A computing device, Includes: One or more processors; And A memory including instructions that, when executed by the one or more processors, cause the computing device to at least: Provided are a plurality of graphics processing unit (GPU) clusters, the plurality of GPU clusters being communicatively coupled to each other via a plurality of network devices arranged in a hierarchy, wherein the plurality of GPU clusters includes at least a first GPU cluster operating at a first speed and a second GPU cluster operating at a second speed different from the first speed; Configure a routing policy for each of the plurality of network devices, wherein the configuration includes establishing a mapping of each incoming port link of the network device to a unique outgoing port link of the network device; And For a data packet transmitted by a GPU of a host machine and received by a first network device, Determine the incoming port link of the first network device on which the data packet is received; Identify the outgoing port link corresponding to the incoming port link based on the configuration; And Forward the data packet on the outgoing port link of the network device.

18. The computing device according to claim 17, wherein the computing device is further configured to: Verify a condition associated with the outgoing port link of the first network device; and In response to the condition being met, forward the data packet on the outgoing port link of the first network device.

19. The computing device according to claim 18, wherein the condition corresponds to determining whether the outgoing port link of the first network device is active.

20. The computing device according to claim 18, wherein the computing device is further configured to: In response to the condition not being met Obtain flow information associated with the data packet by the first network device; Execute an equal-cost multi-path algorithm by the first network device based on the flow information to obtain a new outgoing port link of the first network device; And Forward the data packet by the first network device on the new outgoing port link of the first network device.