Interconnecting global virtual planes
By introducing a global virtual plane into the cloud infrastructure and optimizing packet transmission paths, the low throughput of GPU workloads was addressed, improving network performance and computing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-05-12
AI Technical Summary
The low throughput of GPU workloads in cloud infrastructure is mainly due to the lack of flow entropy and bandwidth contention caused by the failure to consider locality information, which existing technologies cannot effectively solve.
Introducing a global virtual plane in a network environment optimizes packet transmission paths by associating host machines with virtual planes, allowing packets to be transmitted on specific virtual planes and avoiding flow conflicts and bandwidth contention.
It increased the throughput of GPU workloads, optimized network performance, and improved the computing efficiency of cloud infrastructure.
Smart Images

Figure CN122029784A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is a non-provisional application of U.S. Provisional Application No. 63 / 590,269, filed October 13, 2023, and U.S. Provisional Application No. 63 / 611,948, filed December 19, 2023, and claims the benefit of the filing dates thereof, the contents of each of which are incorporated herein by reference in their entirety for all purposes. Technical Field
[0003] This disclosure relates to a network infrastructure for performing artificial intelligence or machine learning workloads, such as graphics processing unit (GPU) workloads. Background Technology
[0004] Organizations are increasingly moving business applications and databases to the cloud to reduce the cost of purchasing, updating, and maintaining on-premises hardware and software. High-performance computing applications consistently consume all available computing power to achieve specific results or outcomes. Such applications require dedicated network performance, fast storage, high computing power, and large amounts of memory—resources that are undersupplied in the virtualized infrastructure that constitutes today's commodity clouds.
[0005] Cloud infrastructure service providers supply newer and faster graphics processing units (GPUs) to meet the requirements of these applications. GPU workloads are typically executed on one or more host machines. Often, such workloads fail to achieve the expected throughput levels. One factor contributing to this problem is the lack of flow entropy, such as Equal Cost Multipath (ECMP) flow entropy. In ECMP, multiple flows (e.g., from different host machines) may be hashed in a way that makes both flows expected to traverse the same outgoing link / port of a switch. Furthermore, the fact that host machines (i.e., hosts) exchange traffic without considering other hosts in their local network neighborhood exacerbates the problem. Other types of workloads are typically executed by randomly (i.e., arbitrarily) selecting one or more host machines from the infrastructure. In other words, workloads are executed without considering locality information (e.g., the physical location of the host machines). Consequently, these workloads have low throughput. This situation often leads to bandwidth contention problems, which are commonly referred to in the literature as flow-collision-based congestion problems. The embodiments discussed in this paper address these and other problems. Summary of the Invention
[0006] This disclosure generally relates to a network infrastructure for performing graphics processing unit (GPU) workloads. Various embodiments are described herein, including methods, systems, non-transitory computer-readable media storing programs, code, or instructions executable by one or more processors. References to these illustrative embodiments are not intended to limit or define this disclosure, but rather to provide examples to aid in understanding it. Additional embodiments are discussed in the detailed description section, and further description is provided therein.
[0007] One embodiment of this disclosure relates to a method comprising: in a network environment comprising a plurality of host machines communicatively coupled to each other via a network architecture comprising a plurality of switches, the plurality of switches comprising a plurality of ports, each of the plurality of host machines comprising one or more GPUs; associating a first subset of the plurality of ports with a first virtual plane, the first virtual plane identifying a first set of resources to be dedicated to transmitting data packets from and to the host machines associated with the first virtual plane; associating a second subset of the plurality of ports with a second virtual plane different from the first virtual plane; associating a first host machine with the first virtual plane and a second host machine with the second virtual plane, the first host machine being directly coupled to a first switch among the plurality of switches and the second host machine being directly coupled to a second switch among the plurality of switches, the second switch being different from the first switch; and for a first data packet originating from a first GPU on the first host machine and destined for a second GPU on the second host machine, transmitting the data packet from the first GPU on the first host machine to the second GPU on the second host machine using ports in the first and second subsets of ports, wherein the first data packet is processed by the first or second switch for transmission on the second virtual plane instead of the first virtual plane.
[0008] One aspect of this disclosure provides a computing device including one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the one or more methods disclosed herein.
[0009] Another aspect of this disclosure provides one or more computer-readable non-transitory media storing computer-executable instructions that, when executed by one or more processors, cause some or all of the methods disclosed herein to be performed.
[0010] The foregoing, together with other features and embodiments, will become more apparent upon reference to the following description, claims and drawings. Attached Figure Description
[0011] The features, embodiments, and advantages of this disclosure will be better understood when reading the following detailed description with reference to the accompanying drawings.
[0012] Figure 1 This is a high-level diagram illustrating a distributed environment of a virtual or overlay cloud network hosted by a cloud service provider infrastructure, according to certain embodiments.
[0013] Figure 2 A simplified architecture diagram of the physical components in the physical network within the CSPI according to certain embodiments is depicted.
[0014] Figure 3 An example arrangement within CSPI according to certain embodiments is shown, in which a host machine is connected to multiple network virtualization devices (NVDs).
[0015] Figure 4 The connectivity between the host machine and the NVD, according to certain embodiments, is described to provide I / O virtualization to support multitenancy.
[0016] Figure 5 A simplified block diagram of a physical network provided by CSPI according to certain embodiments is depicted.
[0017] Figure 6 A simplified block diagram of a cloud infrastructure incorporated into a CLOS network layout according to certain embodiments is depicted.
[0018] Figure 7 An exemplary network architecture is depicted according to the concept of a global virtual plane based on certain embodiments.
[0019] Figure 8 An exemplary network architecture is depicted, illustrating traffic paths established within a global virtual plane according to illustrations of certain embodiments.
[0020] Figure 9 An exemplary flowchart illustrating steps performed when transmitting data packets using network infrastructure is shown according to certain embodiments.
[0021] Figure 10 An exemplary network architecture of a first embodiment of the interconnected global virtual plane is depicted.
[0022] Figure 11 The logical connections of host machines included in a network architecture according to some embodiments are described.
[0023] Figure 12 Another exemplary mechanism for transforming data between different virtual planes according to an embodiment is described.
[0024] Figure 13An exemplary flowchart illustrating steps performed when transmitting data packets using a first embodiment of an interconnected global virtual plane, according to certain embodiments, is shown.
[0025] Figure 14 An exemplary network architecture is depicted in the second embodiment of the interconnected global virtual plane.
[0026] Figure 15 An exemplary flowchart illustrating steps performed when transmitting data packets using a second embodiment of an interconnected global virtual plane, according to certain embodiments, is shown.
[0027] Figure 16 The illustration depicts a virtual partition of a Layer 1 switch according to certain embodiments.
[0028] Figure 17A An exemplary network architecture is depicted in a third embodiment of the interconnected global virtual plane.
[0029] Figure 17B An exemplary flowchart illustrating steps performed when transmitting data packets using a third embodiment of an interconnected global virtual plane, according to certain embodiments, is shown.
[0030] Figure 18 This is a block diagram illustrating a pattern for implementing a cloud infrastructure-as-a-service system according to at least one embodiment.
[0031] Figure 19 This is a block diagram illustrating another pattern for implementing a cloud infrastructure-as-a-service system according to at least one embodiment.
[0032] Figure 20 This is a block diagram illustrating another pattern for implementing a cloud infrastructure-as-a-service system according to at least one embodiment.
[0033] Figure 21 This is a block diagram illustrating another pattern for implementing a cloud infrastructure-as-a-service system according to at least one embodiment.
[0034] Figure 22 This is a block diagram illustrating an example computer system according to at least one embodiment. Detailed Implementation
[0035] In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of certain embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not intended to be limiting. The word “exemplary” is used herein to mean “serves as an example, instance, or illustration.” Any embodiment or design described herein as “exemplary” is not necessarily to be construed as preferred or superior to other embodiments or designs.
[0036] Example architecture of cloud infrastructure
[0037] The term cloud service generally refers to services provided by a cloud service provider (CSP) to users or customers on demand (e.g., via a subscription model) using systems and infrastructure (cloud infrastructure) provided by the CSP. Typically, the servers and systems that make up the CSP's infrastructure are separate from the customer's own on-premises servers and systems. Therefore, customers can utilize cloud services provided by the CSP without having to purchase separate hardware and software resources for the service. Cloud services are designed to provide subscribers with simple, scalable access to applications and computing resources without requiring customers to invest in the infrastructure used to provide the service.
[0038] Several cloud service providers offer various types of cloud services. There are various different types or models of cloud services, including Software as a Service (SaaS), Platform as a Service (PaaS), Infrastructure as a Service (IaaS), etc.
[0039] A customer can subscribe to one or more cloud services provided by a CSP. A customer can be any entity, such as an individual, organization, or enterprise. When a customer subscribes to or registers for a service provided by a CSP, a lease or account is created for that customer. The customer can then access one or more subscribed cloud resources associated with that account.
[0040] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing service. In the IaaS model, a CSP provides customers with infrastructure (called Cloud Service Provider Infrastructure or CSPI) that they can use to build their own customizable networks and deploy customer resources. Therefore, the customer's resources and network are hosted in a distributed environment by the infrastructure provided by the CSP. This differs from traditional computing, where the customer's resources and network are hosted by the infrastructure provided by the customer.
[0041] CSPI can include interconnected high-performance computing resources forming a physical network, including various host machines, memory resources, and network resources. This physical network is also known as the base network or underlying network. Resources in the CSPI can be distributed across one or more data centers, which may be geographically distributed across one or more geographic regions. Virtualization software can be executed by these physical resources to provide a virtualized distributed environment. Virtualization creates overlay networks (also known as software-based networks, software-defined networks, or virtual networks) on the physical network. The CSPI physical network provides the underlying foundation for creating one or more overlay or virtual networks on top of the physical network. The physical network (or base network or underlying network) includes physical network devices such as physical switches, routers, computers, and host machines. An overlay network is a logical (or virtual) network that runs on top of the physical base network. A given physical network can support one or more overlay networks. Overlay networks typically use encapsulation techniques to distinguish traffic belonging to different overlay networks. Virtual or overlay networks are also known as Virtual Cloud Networks (VCNs). Virtual networks are created using software virtualization technologies (e.g., hypervisors, virtualization functions implemented by network virtualization devices (NVDs) (e.g., smartNICs), top-of-rack (TOR) switches, intelligent TORs that implement one or more functions performed by NVDs, and other mechanisms) to create a layer of network abstraction that can run on top of the physical network. Virtual networks can take many forms, including peer-to-peer networks, IP networks, etc. Virtual networks are typically Layer 3 IP networks or Layer 2 VLANs. This method of virtual or overlay networking is often referred to as virtual or overlay Layer 3 networking. Examples of protocols developed for virtual networks include IP-in-IP (or Generic Routing Encapsulation (GRE)), Virtual Extensible LAN (VXLAN—IETF RFC 7348), Virtual Private Networks (VPNs) (e.g., MPLS Layer 3 Virtual Private Network (RFC 4364)), VMware's NSX, GENEVE (Generic Network Virtualization Encapsulation), etc.
[0042] For IaaS, the infrastructure provided by a CSP (Center for Service Providers) can be configured to offer virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, cloud service providers can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, IaaS providers can also provision various services accompanying those infrastructure components (e.g., billing, monitoring, logging, security, load balancing, and clustering, etc.). Therefore, since these services can be policy-driven, IaaS users can implement policies to drive load balancing to maintain application availability and performance. CSPI provides a collection of infrastructure and complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available, hosted, distributed environment. CSPI provides high-performance computing resources and capabilities, as well as storage capacity, in a flexible virtual network securely accessible from various networked locations, such as the customer's on-premises network. When a customer subscribes to or enrolls in an IaaS service provided by a CSP, the lease created for that customer is a secure and isolated partition within the CSP, where the customer can create, organize, and manage their cloud resources.
[0043] Customers can build their own virtual networks using the compute, storage, and networking resources provided by CSPI. One or more customer resources or workloads, such as compute instances, can be deployed on these virtual networks. For example, a customer can use resources provided by CSPI to build one or more customizable and private virtual networks, called Virtual Cloud Networks (VCNs). A customer can deploy one or more customer resources, such as compute instances, on a customer VCN. Compute instances can take the form of virtual machines, bare metal instances, etc. Therefore, CSPI provides a collection of infrastructure and complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available, virtually managed environment. Customers do not manage or control the underlying physical resources provided by CSPI, but they have control over the operating system, storage, and deployed applications; and may also have limited control over selected networking components, such as firewalls.
[0044] The CSP can provide a console that enables customers and network administrators to configure, access, and manage resources deployed in the cloud using CSPI resources. In some embodiments, the console provides a web-based user interface that can be used to access and manage the CSPI. In other embodiments, the console is a web-based application provided by the CSP.
[0045] CSPI can support single-tenant or multi-tenant architectures. In a single-tenant architecture, software (e.g., applications, databases) or hardware components (e.g., host machines or servers) serve a single customer or tenant. In a multi-tenant architecture, the software or hardware components serve multiple customers or tenants. Therefore, in a multi-tenant architecture, CSPI resources are shared among multiple customers or tenants. In a multi-tenant scenario, precautions and safeguards are implemented in CSPI to ensure that each tenant's data is isolated and invisible to other tenants.
[0046] In a physical network, a network endpoint is a computing device or system that connects to and communicates with the physical network. Network endpoints in a physical network can connect to a Local Area Network (LAN), a Wide Area Network (WAN), or other types of physical networks. Examples of traditional endpoints in a physical network include modems, hubs, bridges, switches, routers and other networking devices, physical computers (or host machines), etc. Each physical device in a physical network has a fixed network address that can be used to communicate with the device. This fixed network address can be a Layer 2 address (e.g., a MAC address), a fixed Layer 3 address (e.g., an IP address), etc. In a virtualized environment or virtual network, endpoints can include various virtual endpoints, such as virtual machines hosted by components of the physical network (e.g., hosted by physical host machines). These endpoints in a virtual network are addressed using overlay addresses, such as overlay Layer 2 addresses (e.g., overlay MAC addresses) and overlay Layer 3 addresses (e.g., overlay IP addresses). Network overlays provide flexibility by allowing network administrators to move around the overlay addresses associated with network endpoints using software management (e.g., via software implementing a control plane for virtual networks). Accordingly, unlike physical networks, in virtual networks, overlay addresses (e.g., overlay IP addresses) can be moved from one endpoint to another using network management software. Since virtual networks are built on top of physical networks, communication between components within a virtual network involves both the virtual network and the underlying physical network. To facilitate this communication, CSPI components are configured to learn and store mappings that map overlay addresses in the virtual network to actual physical addresses in the base network, and vice versa. These mappings are then used to facilitate communication. Client traffic is encapsulated to facilitate routing within the virtual network.
[0047] Accordingly, physical addresses (e.g., physical IP addresses) are associated with components in a physical network, and overlay addresses (e.g., overlay IP addresses) are associated with entities in a virtual or overlay network. A physical IP address is an IP address associated with a physical device (e.g., a network device) in the underlying or physical network. For example, each NVD has an associated physical IP address. An overlay IP address is an overlay address associated with an entity in an overlay network, such as an overlay address associated with a compute instance in a customer's Virtual Cloud Network (VCN). Two different customers or tenants (each with its own private VCN) can potentially use the same overlay IP address in their VCN without knowing about each other. Both physical IP addresses and overlay IP addresses are types of real IP addresses. These addresses are separate from virtual IP addresses. A virtual IP address is typically a single IP address that represents or maps to multiple real IP addresses. A virtual IP address provides a one-to-many mapping between virtual IP addresses and multiple real IP addresses. For example, a load balancer can use a VIP to map or represent multiple servers, each with its own real IP address.
[0048] Cloud infrastructure, or CSPI, is physically hosted in one or more data centers in one or more regions of the world. CSPI may include components in the physical or underlying network and virtualized components (e.g., virtual networks, compute instances, virtual machines, etc.) in virtual networks built on top of the physical network components. In some embodiments, CSPI is organized and hosted in domains, regions, and availability domains. A region is typically a localized geographical area containing one or more data centers. Regions are generally independent of each other and can be geographically distant, for example, spanning countries or even continents. For example, one region might be in Australia, another in Japan, another in India, and so on. CSPI resources are partitioned between regions, such that each region has its own independent subset of CSPI resources. Each region can provide a set of core infrastructure services and resources, such as compute resources (e.g., bare metal servers, virtual machines, containers, and related infrastructure); storage resources (e.g., block volume storage, file storage, object storage, archive storage); networking resources (e.g., virtual cloud networks (VCNs), load balancing resources, connections to on-premises networks), database resources; edge networking resources (e.g., DNS); and access management and monitoring resources, etc. Each region typically has multiple paths connecting it to other regions within the domain.
[0049] Generally, applications are deployed in the areas where they are used most frequently (i.e., on the infrastructure associated with that area) because using nearby resources is faster than using resources far away. Applications may also be deployed in different areas for various reasons, such as redundancy to mitigate the risk of events within a region (such as large weather systems or earthquakes), or to meet different requirements such as legal jurisdiction, tax domain, and other business or social standards.
[0050] Data centers within a region can be further organized and subdivided into Availability Domains (ADs). An Availability Domain can correspond to one or more data centers located within the region. A region can consist of one or more Availability Domains. In this distributed environment, CSPI resources are either region-specific, such as Virtual Cloud Networks (VCNs), or Availability Domain-specific, such as compute instances.
[0051] Availability Zones (ADs) within a region are isolated from each other, fault-tolerant, and configured to make simultaneous failures highly unlikely. This is achieved by ensuring that ADs do not share critical infrastructure resources (such as networking, physical cabling, cable paths, cable entry points, etc.), making a failure at one AD within a region unlikely to affect the availability of other ADs in the same region. ADs within the same region can be interconnected via low-latency, high-bandwidth networks, enabling highly available connectivity to other networks (e.g., the internet, customer on-premises networks, etc.) and allowing for replication systems across multiple ADs to achieve both high availability and disaster recovery. Cloud services use multiple ADs to ensure high availability and prevent resource failures. As the infrastructure provided by the IaaS provider grows, more regions and ADs, along with additional capacity, can be added. Traffic between availability domains is typically encrypted.
[0052] In some embodiments, regions are grouped into domains. A domain is a logical collection of regions. Domains are isolated from each other and do not share any data. Regions within the same domain can communicate with each other, but regions in different domains cannot. A customer's lease or account with the CSP resides in a single domain and can be distributed across one or more regions belonging to that domain. Typically, when a customer subscribes to an IaaS service, a lease or account is created for that customer in a region within the domain that the customer designates (called the "primary" region). A customer can extend their lease to one or more other regions within the domain. A customer cannot access regions that are not within the domain where their lease resides.
[0053] IaaS providers can offer multiple domains, each catering to the needs of a specific set of customers or users. For example, a business domain can be offered for business customers. As another example, a domain can be offered for customers within a specific country. As yet another example, a government domain can be offered for governments, and so on. For instance, a government domain can meet the needs of a specific government and may have a higher level of security than a business domain. For example, Oracle Cloud Infrastructure (OCI) currently offers domains for its business region and two domains (e.g., FedRAMP licensed and IL5 licensed) for its government cloud region.
[0054] In some embodiments, an Active Directory (AD) can be subdivided into one or more fault domains. A fault domain is a grouping of infrastructure resources within an AD to provide anti-affinity. Fault domains allow for the distribution of compute instances such that these instances do not reside on the same physical hardware within a single AD. This is called anti-affinity. A fault domain refers to a collection of hardware components (computers, switches, etc.) that share a single point of failure. Compute pools are logically divided into fault domains. Therefore, a hardware failure or compute hardware maintenance event affecting one fault domain does not affect instances in other fault domains. Depending on the embodiment, the number of fault domains per AD can vary. For example, in some embodiments, each AD contains three fault domains. Fault domains act as logical data centers within the AD.
[0055] When a customer subscribes to IaaS services, resources from CSPI are provisioned to the customer and associated with the customer's lease. Customers can use these provisioned resources to build private networks and deploy resources on those networks. Customer networks hosted in the cloud by CSPI are called Virtual Cloud Networks (VCNs). Customers can use the CSPI resources allocated to them to set up one or more VCNs. A VCN is a virtual or software-defined private network. Customer resources deployed in a customer's VCN can include compute instances (e.g., virtual machines, bare metal instances) and other resources. These compute instances can represent various customer workloads, such as applications, load balancers, databases, etc. Compute instances deployed on a VCN can communicate with publicly accessible endpoints (“public endpoints”) via public networks (such as the Internet), with other instances in the same VCN or other VCNs (e.g., other VCNs belonging to the customer or VCNs not belonging to the customer), with the customer's on-premises data center or network, and with service endpoints and other types of endpoints.
[0056] CSPs can use CSPIs to provide various services. In some cases, CSPI clients can act as service providers themselves and use CSPI resources to provide services. Service providers can expose service endpoints characterized by identification information such as IP addresses, DNS names, and ports. Client resources (e.g., compute instances) can access a specific service by visiting the service endpoints exposed by the service for that specific service. These service endpoints are generally publicly accessible to users via public communication networks such as the Internet using the public IP addresses associated with the endpoints. Publicly accessible network endpoints are sometimes also called public endpoints.
[0057] In some embodiments, a service provider may expose the service via an endpoint used for the service (sometimes referred to as a service endpoint). Customers of the service can then use this service endpoint to access the service. In some implementations, the service endpoint provided for the service can be accessed by multiple customers intending to consume the service. In other implementations, a dedicated service endpoint can be provided to a customer, so that only that customer can use that dedicated service endpoint to access the service.
[0058] In some embodiments, when a VCN is created, it is associated with a Private Overlay Classless Inter-Domain Routing (CIDR) address space, which is a set of private overlay IP addresses (e.g., 10.0 / 16) assigned to the VCN. A VCN includes associated subnets, routing tables, and gateways. A VCN resides within a single area but can span one or more of the availability domains within that area. A gateway is a virtual interface configured for the VCN and enables traffic to and from the VCN to one or more endpoints outside the VCN. One or more different types of gateways can be configured for a VCN to enable communication to and from different types of endpoints.
[0059] A VCN can be subdivided into one or more subnets, such as one or more subnets. Therefore, a subnet is a configurable unit or subdivision that can be created within a VCN. A VCN can have one or more subnets. Each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24), which do not overlap with other subnets within that VCN and represent a subset of the VCN's address space.
[0060] Each compute instance is associated with a Virtual Network Interface Card (VNIC), which enables the compute instance to participate in subnets within a VCN. A VNIC is the logical representation of a physical network interface card (NIC). Generally, a VNIC is the interface between an entity (e.g., a compute instance, a service) and a virtual network. A VNIC exists within a subnet and has one or more associated IP addresses, along with associated security rules or policies. A VNIC is equivalent to a Layer 2 port on a switch. A VNIC is attached to both the compute instance and the subnet within the VCN. The VNIC associated with a compute instance enables the compute instance to be part of a VCN's subnet and allows the compute instance to communicate (e.g., send and receive packets) with endpoints on the same subnet as the compute instance, with endpoints in different subnets within the VCN, or with endpoints outside the VCN. Therefore, the VNIC associated with a compute instance determines how the compute instance connects to endpoints inside and outside the VCN. When a compute instance is created and added to a subnet within a VCN, a VNIC is created for the compute instance and associated with that compute instance. For a subnet that includes a set of compute instances, the subnet contains a VNIC corresponding to that set of compute instances, with each VNIC attached to a compute instance within that set of compute instances.
[0061] A private overlay IP address is assigned to each compute instance via the VNIC associated with it. This private overlay network IP address is assigned to the VNIC associated with the compute instance when the compute instance is created and is used to route traffic to and from the compute instance. All VNICs within a given subnet use the same routing table, security list, and DHCP options. As described above, each subnet within a VCN is associated with a contiguous range of overlay IP addresses (e.g., 10.0.0.0 / 24 and 10.0.1.0 / 24) that do not overlap with other subnets within that VCN and represent a subset of the address space within the VCN's address space. For a VNIC on a specific subnet of a VCN, the private overlay IP address assigned to that VNIC is an address from the contiguous range of overlay IP addresses allocated to the subnet.
[0062] In some embodiments, in addition to a private overlay IP address, a compute instance may optionally be assigned additional overlay IP addresses, such as one or more public IP addresses, for example, if in a public subnet. These multiple addresses are assigned either on the same VNIC or on multiple VNICs associated with the compute instance. However, each instance has a primary VNIC, which is created during instance startup and associated with the overlay private IP address assigned to that instance—this primary VNIC cannot be deleted. Additional VNICs, called secondary VNICs, can be added to an existing instance in the same availability domain as the primary VNIC. All VNICs are in the same availability domain as the instance. The secondary VNIC can be located in a subnet within the same VCN as the primary VNIC, or in a different subnet within the same VCN or different VCNs.
[0063] If a compute instance is in a public subnet, it can optionally be assigned a public IP address. When creating a subnet, it can be specified as either a public or private subnet. A private subnet means that resources within the subnet (such as compute instances) and associated VNICs cannot have public overriding IP addresses. A public subnet means that resources within the subnet and associated VNICs can have public IP addresses. Customers can specify that a subnet exists within a single availability domain or across multiple availability domains in a region or domain.
[0064] As described above, a VCN can be subdivided into one or more subnets. In some embodiments, a virtual router (VR) configured for a VCN (referred to as a VCN VR or simply VR) enables communication between the subnets of the VCN. For a subnet within a VCN, the VR represents a logical gateway for that subnet, enabling that subnet (i.e., compute instances on that subnet) to communicate with endpoints on other subnets within the VCN as well as with other endpoints outside the VCN. The VCN VR is a logical entity configured to route traffic between the VNIC within the VCN and the virtual gateway (“gateway”) associated with the VCN. The following section discusses… Figure 1Further description of the gateway. A VCN VR is a Layer 3 / IP layer concept. In one embodiment, there exists a VCN VR for a VCN, where the VCN VR potentially has an unlimited number of ports addressable via IP addresses, one port for each subnet of the VCN. In this way, the VCN VR has a different IP address for each subnet within the VCN to which the VCN VR is attached. The VR also connects to various gateways configured for the VCN. In some embodiments, a specific overlay IP address within a range of overlay IP addresses for a subnet is reserved for the port of the VCN VR for that subnet. For example, consider a VCN with two subnets, associated with address ranges 10.0 / 16 and 10.1 / 16. For the first subnet in the VCN with an address range of 10.0 / 16, addresses within this range are reserved for the port of the VCN VR for that subnet. In some cases, the first IP address within the range can be reserved for the VCN VR. For example, for a subnet with an overlay IP address range of 10.0 / 16, the IP address 10.0.0.1 could be reserved for the port of the VCN VR for that subnet. For a second subnet within the same VCN with an address range of 10.1 / 16, the VCN VR can have a port with IP address 10.1.0.1 for the second subnet. The VCN VR has a different IP address for each subnet within the VCN.
[0065] In some other embodiments, each subnet within a VCN may have its own associated VR, which can be addressed by the subnet using a reserved or default IP address associated with the VR. For example, the reserved or default IP address may be the first IP address in a range of IP addresses associated with the subnet. The VNIC in the subnet can use this default or reserved IP address to communicate with the VR associated with the subnet (e.g., send and receive data packets). In this embodiment, the VR is the ingress / egress point of the subnet. The VR associated with a subnet within the VCN can communicate with other VRs associated with other subnets within the VCN. The VR can also communicate with the gateway associated with the VCN. The VR functionality of the subnet runs on or is performed by one or more NVDs that perform the VNIC functionality of the VNICs within the subnet.
[0066] You can configure routing tables, security rules, and DHCP options for a VCN. A routing table is a virtual routing table used by the VCN and contains rules that route traffic from subnets within the VCN to destinations outside the VCN via gateways or specially configured instances. You can customize the VCN's routing table to control how packets are forwarded / routed to and from the VCN. DHCP options refer to configuration information automatically provided to the instance when it starts up.
[0067] Security rules configured for a VCN represent overlay firewall rules used by the VCN. Security rules can include ingress and egress rules, specifying the types of traffic allowed to enter and exit instances within the VCN (e.g., based on protocol and port). Clients can choose whether a given rule is stateful or stateless. For example, a client can set up a stateful ingress rule with source CIDR 0.0.0.0 / 0 and destination TCP port 22 to allow a set of incoming SSH traffic from anywhere to an instance. Security rules can be implemented using network security groups or security lists. A network security group consists of a set of security rules that apply only to resources within that group. A security list, on the other hand, includes rules applicable to all resources in any subnet using that security list. A default security list with default security rules can be provided to the VCN. DHCP options configured for the VCN provide configuration information that is automatically provided to instances within the VCN at instance startup.
[0068] In some embodiments, configuration information for a VCN is determined and stored by the VCN control plane. For example, the configuration information for a VCN may include information about: the address range associated with the VCN, subnets and associated information within the VCN, one or more VRs associated with the VCN, compute instances and associated VNICs within the VCN, NVDs (e.g., VNICs, VRs, gateways) that perform various virtualization network functions associated with the VCN, status information for the VCN, and other VCN-related information. In some embodiments, a VCN distribution service publishes the configuration information or portions thereof stored by the VCN control plane to the NVD. The distributed information can be used to update information stored by the NVD and used to forward data packets to and from compute instances within the VCN (e.g., forwarding tables, routing tables, etc.).
[0069] In some embodiments, the creation of VCNs and subnets is handled by the VCN control plane (CP), and the initiation of compute instances is handled by the compute control plane. The compute control plane is responsible for allocating physical resources to the compute instances and then invoking the VCN control plane to create VNICs and attach them to the compute instances. The VCN CP also sends VCN data maps to the VCN data plane, which is configured to perform packet forwarding and routing functions. In some embodiments, the VCN CP provides a distribution service responsible for providing updates to the VCN data plane. Examples of VCN control planes are also available in... Figure 6 , Figure 7 , Figure 8 and Figure 9 The figures are depicted in (see reference numerals 616, 716, 816 and 916) and described below.
[0070] Customers can create one or more VCNs using resources hosted by CSPI. Compute instances deployed on a customer's VCN can communicate with different endpoints. These endpoints can include endpoints hosted by CSPI and endpoints outside of CSPI.
[0071] Various architectures are used to implement cloud-based services using CSPI. Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 18-22 It is described in the text and further described below. Figure 1 This is a high-level diagram of a distributed environment 100, illustrating an overlay or client VCN hosted by CSPI according to certain embodiments. Figure 1 The distributed environment described includes multiple components in the overlay network. Figure 1 The distributed environment 100 depicted herein is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Many variations, substitutions, and modifications are possible. For example, in some implementations, Figure 1 The distributed environment described in the text can have more than Figure 1 The more or fewer systems or components shown can be combined into two or more systems, or can have different system configurations or arrangements.
[0072] like Figure 1 As illustrated in the example, distributed environment 100 includes a CSPI 101 that provides services and resources that customers can subscribe to and use to build their Virtual Cloud Network (VCN). In some embodiments, CSPI 101 provides IaaS services to subscribing customers. Data centers within CSPI 101 can be organized into one or more regions. Figure 1 The example shown is Region US 102. The customer has already configured a customer VCN c / o Oracle International for Region 102. The customer can deploy various compute instances on VCN 104, which can include virtual machines or bare metal instances. Examples of instances include applications, databases, load balancers, etc.
[0073] exist Figure 1 In the embodiment depicted, customer VCN 104 includes two subnets, namely "Subnet-1" and "Subnet-2", each with its own CIDR IP address range. Figure 1In this configuration, subnet-1 covers the IP address range of 10.0 / 16, and subnet-2 covers the address range of 10.1 / 16. VCN Virtual Router 105 represents a logical gateway for the VCN, enabling communication between subnets within VCN 104 and with other endpoints outside the VCN. VCN VR 105 is configured to route traffic between the VNICs within VCN 104 and the gateway associated with VCN 104. VCN VR 105 provides a port for each subnet of VCN 104. For example, VR 105 could provide a port with IP address 10.0.0.1 for subnet-1 and a port with IP address 10.1.0.1 for subnet-2.
[0074] Multiple compute instances can be deployed on each subnet, where compute instances can be virtual machine instances and / or bare metal instances. Compute instances within a subnet can be hosted by one or more host machines within CSPI 101. Compute instances participate in the subnet via a VNIC associated with the compute instance. For example, as... Figure 1 As shown, compute instance C1 becomes part of subnet-1 via the VNIC associated with it. Similarly, compute instance C2 becomes part of subnet-1 via the VNIC associated with it. In a similar manner, multiple compute instances (which can be virtual machine instances or bare metal instances) can be part of subnet-1. Each compute instance is assigned a private overlay IP address and MAC address via its associated VNIC. For example, in... Figure 1 In this context, compute instance C1 has an overlay IP address of 10.0.0.2 and a MAC address of M1, while compute instance C2 has a private overlay IP address of 10.0.0.3 and a MAC address of M2. Each compute instance in subnet-1 (including compute instances C1 and C2) has a default route to VCN VR 105 using IP address 10.0.0.1, which is the IP address of the port used by VCN VR 105 in subnet-1.
[0075] Multiple compute instances, including virtual machine instances and / or bare metal instances, can be deployed on subnet-2. For example, such as Figure 1 As shown, compute instances D1 and D2 become part of subnet-2 via the VNIC associated with the respective compute instance. Figure 1 In the illustrated embodiment, compute instance D1 has an overlay IP address of 10.1.0.2 and a MAC address of MM1, while compute instance D2 has a private overlay IP address of 10.1.0.3 and a MAC address of MM2. Each compute instance in subnet-2 (including compute instances D1 and D2) has a default route to VCN VR 105 using IP address 10.1.0.1, which is the IP address of the port of VCN VR105 in subnet-2.
[0076] VCN A 104 may also include one or more load balancers. For example, a load balancer can be provided for a subnet, and the load balancer can be configured to load balance traffic across multiple compute instances on the subnet. Load balancers can also be provided to load balance traffic across subnets within a VCN.
[0077] A specific compute instance deployed on VCN 104 can communicate with a variety of different endpoints. These endpoints can include endpoints hosted by CSPI 200 and endpoints outside of CSPI 200. Endpoints hosted by CSPI 101 can include: endpoints on the same subnet as the specific compute instance (e.g., communication between two compute instances in subnet-1); endpoints on different subnets but within the same VCN (e.g., communication between a compute instance in subnet-1 and a compute instance in subnet-2); endpoints in different VCNs within the same region (e.g., communication between a compute instance in subnet-1 and an endpoint in a VCN within the same region 106 or 110, or communication between a compute instance in subnet-1 and an endpoint in service point 110 within the same region); or endpoints in VCNs in different regions (e.g., communication between a compute instance in subnet-1 and an endpoint in a VCN within a different region 108). Compute instances in subnets hosted by CSPI 101 can also communicate with endpoints not hosted by CSPI 101 (i.e., outside of CSPI 101). These external endpoints include endpoints in the customer’s on-premises network 116, endpoints in other remote cloud-hosted networks 118, public endpoints 114 that are accessible via public networks such as the Internet, and other endpoints.
[0078] VNICs associated with source and destination compute instances facilitate communication between compute instances on the same subnet. For example, compute instance C1 in subnet-1 might want to send a packet to compute instance C2 in subnet-1. For a packet originating from the source compute instance and destined for another compute instance in the same subnet, the packet is first processed by the VNIC associated with the source compute instance. The processing performed by the VNIC associated with the source compute instance may include determining the packet's destination information from the packet header, identifying any policies (e.g., security lists) configured for the VNIC associated with the source compute instance, determining the packet's next hop, performing any packet encapsulation / decapsulation functions as needed, and then forwarding / routing the packet to the next hop to facilitate communication to its intended destination. When the destination compute instance is in the same subnet as the source compute instance, the VNIC associated with the source compute instance is configured to identify the VNIC associated with the destination compute instance and forward the packet to that VNIC for processing. The VNIC associated with the destination compute instance then performs the forwarding and forwarding of the packet to the destination compute instance.
[0079] For data packets to be transmitted from a compute instance in a subnet to an endpoint in a different subnet within the same VCN, communication is facilitated by the VNIC associated with the source and destination compute instances, as well as the VCN VR. For example, if Figure 1 If compute instance C1 in subnet-1 wants to send a data packet to compute instance D1 in subnet-2, the packet is first processed by the VNIC associated with compute instance C1. The VNIC associated with compute instance C1 is configured to route the packet to VCN VR 105 using the default route or port 10.0.0.1 of the VCN VR. VCN VR 105 is configured to route the packet to subnet-2 using port 10.1.0.1. Then, the VNIC associated with D1 receives and processes the packet and forwards it to compute instance D1.
[0080] For data packets to be transmitted from a compute instance in VCN 104 to an endpoint outside VCN 104, communication is facilitated by the VNIC associated with the source compute instance, VCN VR 105, and the gateway associated with VCN 104. One or more types of gateways can be associated with VCN 104. A gateway is an interface between a VCN and another endpoint located outside the VCN. A gateway is a Layer 3 / IP concept and enables a VCN to communicate with endpoints outside the VCN. Therefore, a gateway facilitates traffic flow between a VCN and other VCNs or networks. Various different types of gateways can be configured for a VCN to facilitate different types of communication with different types of endpoints. Depending on the gateway, communication can be conducted over a public network (e.g., the Internet) or over a private network. Various communication protocols can be used for these communications.
[0081] For example, compute instance C1 might want to communicate with an endpoint outside of VCN 104. The data packet can first be processed by the VNIC associated with the source compute instance C1. The VNIC processing determines that the packet's destination is outside C1's subnet-1. The VNIC associated with C1 can then forward the packet to VCN VR 105 for VCN 104. VCN VR 105 then processes the packet and, as part of the processing, determines the specific gateway associated with VCN 104 as the next hop for the packet based on its destination. VCN VR 105 can then forward the packet to the identified specific gateway. For example, if the destination is an endpoint within a customer's on-premises network, the packet can be forwarded by VCN VR 105 to the Dynamic Routing Gateway (DRG) gateway 122 configured for VCN 104. The packet can then be forwarded from the gateway to the next hop to facilitate delivery to its final intended destination.
[0082] Various types of gateways can be configured for a VCN. Examples of gateways that can be configured for a VCN are available in [link to example]. Figure 1 It is depicted in [the text] and described below. Examples of gateways associated with VCNs are also [described in the text]. Figure 18 , Figure 19 , Figure 20 and Figure 21 The gateways described herein (e.g., referenced by reference numerals 1834, 1836, 1838, 1934, 1936, 1938, 2034, 2036, 2038, 2134, 2136, and 2138) are as follows. Figure 1As depicted in the embodiments, a Dynamic Routing Gateway (DRG) 122 can be added to or associated with a customer VCN 104 and provides a path for private network traffic communication between the customer VCN 104 and another endpoint, which can be the customer's on-premises network 116, a VCN 108 in a different region of CSPI 101, or another remote cloud network 118 not hosted by CSPI 101. The customer's on-premises network 116 can be a customer network or customer data center built using the customer's resources. Access to the customer's on-premises network 116 is generally very restricted. For customers who have both an on-premises network 116 and one or more VCNs 104 deployed or hosted in the cloud by CSPI 101, the customer may want their on-premises network 116 and their cloud-based VCN 104 to be able to communicate with each other. This allows customers to build extended hybrid environments encompassing the customer's VCN 104 hosted by CSPI 101 and their on-premises network 116. The DRG 122 enables this communication. To enable this type of communication, a communication channel 124 is established, with one endpoint located in the customer's on-premises network 116 and the other endpoint located in CSPI 101 and connected to the customer's VCN 104. Communication channel 124 can be over a public communication network (such as the Internet) or a private communication network. Various communication protocols can be used, such as IPsec VPN technology over a public communication network (such as the Internet), Oracle's FastConnect technology using a private network instead of a public network, etc. The device or equipment forming one endpoint of communication channel 124 in the customer's on-premises network 116 is referred to as customer premises equipment (CPE), such as... Figure 1 The CPE126 is depicted in the diagram. On the CSPI 101 side, the endpoint can be a host machine executing DRG 122.
[0083] In some embodiments, a remote peering connection (RPC) can be added to the DRG, which allows a customer to peer one VCN with another VCN in a different region. Using such an RPC, customer VCN 104 can use DRG 122 to connect to VCN 108 in another region. DRG 122 can also be used to communicate with other remote cloud networks 118 not hosted by CSPI 101, such as Microsoft Azure Cloud, Amazon AWS Cloud, etc.
[0084] like Figure 1As shown, an Internet Gateway (IGW) 120 can be configured for customer VCN 104, enabling compute instances on VCN 104 to communicate with a public endpoint 114 accessible via a public network, such as the Internet. IGW 120 is a gateway connecting the VCN to a public network, such as the Internet. IGW 120 enables public subnets within the VCN (such as VCN 104), where resources have publicly overridden IP addresses, to directly access a public endpoint 112 on the public network 114 (such as the Internet). Using IGW 120, connections can be initiated from subnets within VCN 104 or from the Internet.
[0085] Network Address Translation (NAT) gateway 128 can be configured for a customer's VCN 104, enabling cloud resources within the customer's VCN that do not have dedicated public overlay IP addresses to access the Internet, and it does so without exposing those resources to directly incoming Internet connections (e.g., L4-L7 connections). This allows private subnets within the VCN (such as private subnet-1 in VCN 104) to privately access public endpoints on the Internet. In a NAT gateway, connections can only be initiated from a private subnet to the public Internet, and not from the Internet to a private subnet.
[0086] In some embodiments, a Service Gateway (SGW) 126 may be configured for a customer's VCN 104 and provide a path for private network traffic between VCN 104 and service endpoints supported in Service Network 110. In some embodiments, Service Network 110 may be provided by a CSP and may offer a variety of services. An example of such a service network is Oracle's service network, which provides a variety of services available to customers. For example, compute instances (e.g., database systems) in a private subnet of customer VCN 104 may back up data to service endpoints (e.g., object storage) without requiring a public IP address or access to the Internet. In some embodiments, a VCN may have only one SGW, and connections may only originate from subnets within the VCN, not from Service Network 110. If a VCN is peered to another, resources in the other VCN typically cannot access the SGW. Resources in an on-premises network connected to a VCN using FastConnect or VPN Connect may also use the Service Gateway configured for that VCN.
[0087] In some implementations, the SGW 126 uses the concept of a Service Classless Inter-Domain Routing (CIDR) label, which is a string representing the range of all regional public IP addresses used for the service or group of services of interest. Customers use the Service CIDR label when configuring the SGW and associated routing rules to control traffic to the service. Customers can optionally use it when configuring security rules without needing to adjust the security rules if the public IP addresses of the service change in the future.
[0088] Local peering gateway (LPG) 132 is a gateway that can be added to customer VCN 104 and enable VCN 104 to peer with another VCN in the same area. Peering refers to VCNs communicating using private IP addresses without traffic traversing public networks (such as the Internet) or routing traffic through the customer's on-premises network 116. In a preferred embodiment, a VCN has a separate LPG for each peering it establishes. Local peering, or VCN peering, is a common practice for establishing network connectivity between different applications or infrastructure management functions.
[0089] Service providers (such as service providers in service network 110) can offer access to services using different access models. Under the public access model, a service can be exposed as a public endpoint accessible to compute instances within a customer's VCN via a public network (such as the Internet), and / or privately accessible via SGW 126. Under a specific private access model, a service can be accessed as a private IP endpoint within a private subnet of the customer's VCN. This is called Private Endpoint (PE) access and enables service providers to expose their services as instances within the customer's private network. A private endpoint resource represents a service within the customer's VCN. Each PE is represented as a VNIC (called a PE-VNIC, with one or more private IPs) within the customer's VCN in a subnet chosen by the customer. Thus, the PE provides a way to present services within a private customer VCN subnet using a VNIC. Because the endpoint is exposed as a VNIC, all characteristics associated with the VNIC (such as routing rules, security lists, etc.) are now available for the PE VNIC.
[0090] Service providers can register their services to enable access via PE. Providers can associate policies with services, which restricts the visibility of services to customer leases. Providers can register multiple services under a single Virtual IP address (VIP), especially for multi-tenant services. Multiple such private endpoints (across multiple VCNs) can represent the same service.
[0091] Compute instances in a private subnet can then access the service using the private IP address or service DNS name of the PE VNIC. Compute instances in a customer VCN can access the service by sending traffic to the private IP address of the PE in the customer VCN. A Private Access Gateway (PAGW) 130 is a gateway resource that can be attached to a service provider VCN (e.g., a VCN in service network 110), which acts as the ingress / egress point for all traffic originating from / to the private endpoint of the customer subnet. PAGW 130 enables providers to scale the number of PE connections without utilizing their internal IP address resources. A provider only needs to configure one PAGW for any number of services registered in a single VCN. A provider can represent a service as a private endpoint in multiple VCNs of one or more customers. From the customer's perspective, the PE VNIC is not an instance attached to the customer, but rather appears to be attached to the service the customer wishes to interact with. Traffic to the private endpoint is routed to the service via PAGW 130. These are referred to as customer-to-service private connections (C2S connections).
[0092] The PE concept can also be used to extend private access to services to the customer's on-premises network and data center by allowing traffic to flow through FastConnect / IPsec links and private endpoints within the customer's VCN. Private access to services can also be extended to the customer's peering VCN by allowing traffic to flow between the LPG 132 and the PE within the customer's VCN.
[0093] Customers can control routing within a VCN at the subnet level, allowing them to specify which subnets within a customer's VCN (such as VCN104) use each gateway. The VCN's routing table is used to determine whether traffic is allowed to leave the VCN via a specific gateway. For example, in a given instance, the routing table for public subnets within customer VCN 104 might send non-local traffic via IGW 120. The routing table for private subnets within the same customer VCN 104 might send traffic destined for CSP services via SGW 126. All remaining traffic might be sent via NAT gateway 128. The routing table only controls traffic leaving the VCN.
[0094] Security lists associated with a VCN are used to control traffic entering the VCN via a gateway through inbound connections. All resources within a subnet use the same routing tables and security lists. Security lists can be used to control specific types of traffic allowed to enter or leave instances within a subnet of the VCN. Security list rules can include inbound and outbound rules. For example, inbound rules can specify allowed source address ranges, while outbound rules can specify allowed destination address ranges. Security rules can specify specific protocols (e.g., TCP, ICMP), specific ports (e.g., port 22 for SSH, port 3389 for Windows RDP), etc. In some implementations, the instance's operating system can enforce its own firewall rules that conform to the security list rules. Rules can be stateful (e.g., tracking connections and automatically allowing responses without explicit security list rules for response traffic) or stateless.
[0095] Access from a customer's VCN (i.e., through resources or compute instances deployed on VCN 104) can be categorized as public access, private access, or dedicated access. Public access refers to an access model that uses a public IP address or NAT to access a public endpoint. Private access enables customer workloads with private IP addresses within VCN 104 (e.g., resources in a private subnet) to access services without traversing a public network such as the Internet. In some embodiments, CSPI 101 enables customer VCN workloads with private IP addresses to access the service's (public service endpoint) using a service gateway. Thus, the service gateway provides a private access model by establishing a virtual link between the customer's VCN and the public endpoint of the service residing outside the customer's private network.
[0096] Furthermore, CSPI can provide private public access using technologies such as FastConnect public peering, where on-premises instances can access one or more services within a customer's VCN using FastConnect connections without traversing public networks such as the internet. CSPI can also provide private private access using FastConnect private peering, where on-premises instances with private IP addresses can access a customer's VCN workloads using FastConnect connections. FastConnect is a network connectivity alternative to using the public internet to connect a customer's on-premises network to CSPI and its services. Compared to internet-based connections, FastConnect offers a simple, flexible, and cost-effective way to create private and private connections with higher bandwidth options and a more reliable and consistent network experience.
[0097] Figure 1The accompanying description above describes the various virtualized components in the example virtual network. As mentioned above, the virtual network is built on the underlying physical or base network. Figure 2 A simplified architecture diagram of the physical components within the physical network of the CSPI 200, which provides the underlying layer for virtual networks according to certain embodiments, is depicted. As shown, the CSPI 200 provides a distributed environment including components and resources (e.g., compute, memory, and network resources) provided by a cloud service provider (CSP). These components and resources are used to provide cloud services (e.g., IaaS services) to subscribed customers (i.e., customers who have subscribed to one or more services provided by the CSP). Based on the services subscribed to by the customer, a subset of the resources of the CSPI 200 (e.g., compute, memory, and network resources) is provisioned to the customer. The customer can then use the physical compute, memory, and networking resources provided by the CSPI 200 to build their own cloud-based (i.e., CSPI-hosted) customizable and private virtual networks. As indicated above, these customer networks are referred to as Virtual Cloud Networks (VCNs). Customers can deploy one or more customer resources, such as compute instances, on these customer VCNs. Compute instances can take the form of virtual machines, bare metal instances, etc. The CSPI 200 provides a collection of infrastructure and complementary cloud services that enable customers to build and run a wide range of applications and services in a highly available managed environment.
[0098] exist Figure 2 In the example embodiment depicted, the physical components of CSPI 200 include one or more physical host machines or physical servers (e.g., 202, 206, 208), network virtualization devices (NVDs) (e.g., 210, 212), top-of-rack (TOR) switches (e.g., 214, 216), and a physical network (e.g., 218), as well as switches within physical network 218. The physical host machines or servers can host and execute various compute instances participating in one or more subnets of the VCN. Compute instances can include virtual machine instances and bare metal instances. For example, Figure 1 The various computational examples described in the text can be derived from... Figure 2 The physical host machine described in the diagram is used for hosting virtual machine compute instances in a VCN. Virtual machine compute instances in a VCN can be executed by one host machine or multiple different host machines. A physical host machine can also host virtual host machines, container-based hosts, or functions, etc. Figure 1 The VNIC and VCN VR described in the text can be generated by Figure 2 The NVD execution described in the text. Figure 1 The gateway described herein can be a host machine and / or Figure 2 The NVD execution described in [the document / document].
[0099] A host machine or server can execute a hypervisor (also known as a virtual machine monitor or VMM) that creates and enables virtualized environments on the host machine. Virtualized or virtualized environments facilitate cloud-based computing. One or more compute instances can be created, executed, and managed on the host machine by a hypervisor on that host machine. The hypervisor on the host machine enables the host machine's physical computing resources (e.g., compute, storage, and network resources) to be shared among various compute instances executed by the host machine.
[0100] For example, such as Figure 2 As depicted, host machines 202 and 208 execute hypervisors 260 and 266, respectively. These hypervisors can be implemented using software, firmware, or hardware, or a combination thereof. Typically, a hypervisor is a process or software layer that sits above the host machine's operating system (OS), which in turn executes on the host machine's hardware processor. Hypervisors provide a virtualized environment by enabling the host machine's physical computing resources (e.g., processing resources such as processors / cores, memory resources, network resources) to be shared among various virtual machine computing instances executed by the host machine. For example, in... Figure 2 In this configuration, the hypervisor 260 can reside on top of the operating system of the host machine 202 and enable the computing resources (e.g., processing, memory, and network resources) of the host machine 202 to be shared among computing instances (e.g., virtual machines) executed by the host machine 202. Virtual machines can have their own operating systems (called guest operating systems), which can be the same as or different from the host machine's operating system. The operating system of a virtual machine executed by the host machine can be the same as or different from the operating system of another virtual machine executed by the same host machine. Therefore, the hypervisor enables multiple operating systems to be executed simultaneously, while sharing the same computing resources of the host machine. Figure 2 The host machines described in the text may have the same or different types of management programs.
[0101] A compute instance can be a virtual machine instance or a bare metal instance. Figure 2 In the diagram, compute instance 268 on host machine 202 and compute instance 274 on host machine 208 are examples of virtual machine instances. Host machine 206 is an example of a bare metal instance provided to a customer.
[0102] In some cases, an entire host machine can be provisioned to a single customer, and one or more compute instances (or virtual machines or bare metal instances) hosted by that host machine all belong to the same customer. In other cases, the host machine can be shared among multiple customers (i.e., multiple tenants). In this multi-tenancy scenario, the host machine can host virtual machine compute instances belonging to different customers. These compute instances can be members of different VCNs for different customers. In some embodiments, bare metal compute instances are hosted by bare metal servers without a hypervisor. When provisioning bare metal compute instances, a single customer or tenant maintains control over the physical CPU, memory, and network interfaces of the host machine hosting the bare metal instance, and the host machine is not shared with other customers or tenants.
[0103] As previously described, each compute instance, as part of a VCN, is associated with a VNIC that enables that compute instance to become a member of a subnet of the VCN. The VNIC associated with a compute instance facilitates communication of packets or frames to and from the compute instance. The VNIC is associated with the compute instance when it is created. In some embodiments, for compute instances executed by a host machine, the VNIC associated with that compute instance is executed by an NVD connected to the host machine. For example, in Figure 2 In this example, host machine 202 executes a virtual machine compute instance 268 associated with VNIC 276, and VNIC 276 is executed by NVD 210 connected to host machine 202. As another example, a bare metal instance 272 hosted by host machine 206 is associated with VNIC 280 executed by NVD 212 connected to host machine 206. As yet another example, VNIC 284 is associated with compute instance 274 executed by host machine 208, and VNIC 284 is executed by NVD 212 connected to host machine 208.
[0104] For compute instances hosted by a host machine, NVDs connected to that host machine also execute VCN VR corresponding to the VCN where the compute instance is a member. For example, in Figure 2 In the embodiment depicted, NVD 210 executes VCN VR 277 corresponding to the VCN of compute instance 268, which is a member of NVD 212. NVD 212 may also execute one or more VCN VR 283 corresponding to the VCNs of compute instances hosted by host machines 206 and 208.
[0105] The host machine may include one or more network interface cards (NICs) that enable the host machine to connect to other devices. The NIC on the host machine may provide one or more ports (or interfaces) that allow the host machine to communicatively connect to another device. For example, the host machine may use one or more ports (or interfaces) provided on the host machine and the NVD to connect to the NVD. The host machine may also connect to other devices (such as another host machine).
[0106] For example, in Figure 2 In this configuration, host machine 202 is connected to NVD 210 via link 220, which extends between port 234 provided by NIC 232 of host machine 202 and port 236 of NVD 210. Host machine 206 is connected to NVD 212 via link 224, which extends between port 246 provided by NIC 244 of host machine 206 and port 248 of NVD 212. Host machine 208 is connected to NVD 212 via link 226, which extends between port 252 provided by NIC 250 of host machine 208 and port 254 of NVD 212.
[0107] The NVD is then connected to top-of-rack (TOR) switches via communication links, which are connected to physical network 218 (also known as a switch architecture). In some embodiments, the links between the host machine and the NVD, and between the NVD and the TOR switches, are Ethernet links. For example, in Figure 2 In this configuration, NVDs 210 and 212 are connected to TOR switches 214 and 216 via links 228 and 230, respectively. In some embodiments, links 220, 224, 226, 228, and 230 are Ethernet links. The collection of host machines and NVDs connected to the TOR is sometimes referred to as a rack.
[0108] Physical network 218 provides a communication architecture that enables TOR switches to communicate with each other. Physical network 218 can be a multi-layer network. In some implementations, physical network 218 is a multi-layer Clos network of switches, where TOR switches 214 and 216 represent leaf-level nodes of the multi-layer and multi-node physical switching network 218. Different Clos network configurations are possible, including but not limited to Layer 2 networks, Layer 3 networks, Layer 4 networks, Layer 5 networks, and general "n"-layer networks. Examples of Clos networks are provided in... Figure 5 It is depicted in the middle and described below.
[0109] Various connection configurations can exist between the host machine and the NVD, such as one-to-one, many-to-one, and one-to-many configurations. In a one-to-one configuration, each host machine connects to its own individual NVD. For example, in... Figure 2 In this configuration, host machine 202 connects to NVD 210 via its NIC 232. In a many-to-one configuration, multiple host machines connect to a single NVD. For example, in... Figure 2 In this configuration, host machines 206 and 208 are connected to the same NVD 212 via NICs 244 and 250, respectively.
[0110] In a one-to-many configuration, a host machine connects to multiple NVDs. Figure 3 An example within the CSPI 300 is shown, where a host machine is connected to multiple NVDs. (Example follows) Figure 3 As shown, host machine 302 includes a network interface card (NIC) 304, which includes multiple ports 306 and 308. Host machine 300 is connected to a first NVD 310 via port 306 and link 320, and to a second NVD 312 via port 308 and link 322. Ports 306 and 308 may be Ethernet ports, and links 320 and 322 between host machine 302 and NVDs 310 and 312 may be Ethernet links. NVD 310 is further connected to a first TOR switch 314, and NVD 312 is connected to a second TOR switch 316. Links between NVDs 310 and 312 and TOR switches 314 and 316 may be Ethernet links. TOR switches 314 and 316 represent Layer 0 switching devices in a multi-layer physical network 318.
[0111] Figure 3 The layout depicted provides two separate physical network paths to and from physical switch network 318 to host machine 302: the first path traverses TOR switch 314 to NVD 310 and then to host machine 302, and the second path traverses TOR switch 316 to NVD 312 and then to host machine 302. These separate paths provide enhanced availability (referred to as high availability) for host machine 302. If one of the paths (e.g., a link in one of the paths breaks) or a device (e.g., a particular NVD is not running) experiences a problem, the other path can be used for communication to / from host machine 302.
[0112] exist Figure 3 In the configuration depicted, the host machine connects to two different NVDs using two different ports provided by the host machine's NIC. In other embodiments, the host machine may include multiple NICs that enable the host machine to connect to multiple NVDs.
[0113] Go back to reference Figure 2An NVD is a physical device or component that performs one or more network and / or storage virtualization functions. An NVD can be any device with one or more processing units (e.g., CPU, Network Processing Unit (NPU), FPGA, packet processing pipeline, etc.), cached memory, and ports. Various virtualization functions can be executed by software / firmware performed by one or more processing units of the NVD.
[0114] NVDs can be implemented in various different forms. For example, in some embodiments, an NVD is implemented as an interface card called a smartNIC or a smart NIC with an onboard embedded processor. A smartNIC is a device separate from the NIC on the host machine. Figure 2 In this context, NVD 210 and 212 can be implemented as smartNICs connected to host machine 202 and host machines 206 and 208, respectively.
[0115] However, smartNIC is just one example of an NVD implementation. Various other implementations are possible. For example, in some other implementations, the NVD, or one or more functions performed by the NVD, may be integrated into or performed by one or more host machines, one or more TOR switches, and other components of the CSPI 200. For instance, the NVD may be implemented within a host machine, where the functions performed by the NVD are performed by the host machine. As another example, the NVD may be part of a TOR switch, or the TOR switch may be configured to perform functions performed by the NVD, enabling the TOR switch to perform various complex packet transformations for public clouds. A TOR performing the functions of the NVD is sometimes referred to as a smart TOR. In other implementations that serve virtual machine (VM) instances rather than bare metal (BM) instances to customers, the functions performed by the NVD may be implemented within the hypervisor of the host machine. In some other implementations, some of the functions of the NVD may be offloaded to a centralized service running on a set of host machines.
[0116] In some embodiments, such as when implemented as Figure 2 As shown in the smartNIC diagram, the NVD can include multiple physical ports that enable it to connect to one or more host machines and one or more TOR switches. Ports on the NVD can be classified as host-facing ports (also known as "south ports") or network-facing or TOR-facing ports (also known as "north ports"). The host-facing ports of the NVD are those used to connect the NVD to the host machine. Figure 2 Examples of host-facing ports include port 236 on the NVD 210 and ports 248 and 254 on the NVD 212. Network-facing ports on the NVD are used to connect the NVD to a TOR switch. Figure 2 Examples of network-facing ports include port 256 on the NVD 210 and port 258 on the NVD 212. Figure 2 As shown, NVD 210 is connected to TOR switch 214 via link 228, which extends from port 256 of NVD 210 to TOR switch 214. Similarly, NVD 212 is connected to TOR switch 216 via link 230, which extends from port 258 of NVD 212 to TOR switch 216.
[0117] The NVD receives packets and frames from the host machine (e.g., packets and frames generated by compute instances hosted on the host machine) via its host-facing port, and after performing the necessary packet processing, can forward the packets and frames to the TOR switch via its network-facing port. The NVD can also receive packets and frames from the TOR switch via its network-facing port, and after performing the necessary packet processing, can forward the packets and frames to the host machine via its host-facing port.
[0118] In some embodiments, there can be multiple ports and associated links between the NVD and TOR switches. These ports and links can be aggregated to form a link aggregation group (called a LAG) of multiple ports or links. Link aggregation allows multiple physical links between two endpoints (e.g., between the NVD and TOR switches) to be treated as a single logical link. All physical links in a given LAG can operate at the same speed in full-duplex mode. LAGs help increase the bandwidth and reliability of the connection between two endpoints. If one of the physical links in the LAG fails, traffic will be dynamically and transparently reassigned to one of the other physical links in the LAG. Aggregated physical links deliver higher bandwidth than each individual link. Multiple ports associated with an LAG are treated as a single logical port. Traffic can be load balanced across multiple physical links in the LAG. One or more LAGs can be configured between two endpoints. These endpoints can be located between the NVD and TOR switches, between a host machine and the NVD, etc.
[0119] The NVD implements or performs network virtualization functions. These functions are performed by software / firmware executed by the NVD. Examples of network virtualization functions include, but are not limited to: packet encapsulation and decapsulation functions; functions for creating VCN networks; functions for implementing network policies, such as VCN security list (firewall) functionality; functions for facilitating the routing and forwarding of packets to and from compute instances in the VCN; and so on. In some embodiments, upon receiving a packet, the NVD is configured to execute a packet processing pipeline for processing the packet and determining how to forward or route the packet. As part of this packet processing pipeline, the NVD may perform one or more virtual functions associated with the overlay network, such as executing a VNIC associated with a compute instance in the VCN, executing a virtual router (VR) associated with the VCN, packet encapsulation and decapsulation to facilitate forwarding or routing in the virtual network, execution of certain gateways (e.g., local peer gateways), implementation of security lists, network security groups, Network Address Translation (NAT) functionality (e.g., per-host translation of public IPs to private IPs), throttling functions, and other functions.
[0120] In some embodiments, the packet processing data path in the NVD may include multiple packet pipelines, each consisting of a series of packet transformation stages. In some implementations, upon receiving a packet, the packet is parsed and classified into a single pipeline. The packet is then processed linearly, stage by stage, until the packet is dropped or sent out through the NVD's interface. These stages provide basic functional packet processing building blocks (e.g., header verification, throttling, insertion of new Layer 2 headers, L4 firewall enforcement, VCN encapsulation / decapsulation, etc.) so that new pipelines can be built by combining existing stages, and new functionality can be added by creating new stages and inserting them into existing pipelines.
[0121] NVD can perform both the control plane and data plane functions corresponding to those of VCN. An example of the VCN control plane is also available in... Figure 18 , Figure 19 , Figure 20 and Figure 21 The VCN data plane is depicted in (see reference numerals 1816, 1916, 2016, and 2116) and described below. An example of the VCN data plane is... Figure 18 , Figure 19 , Figure 20 and Figure 21The following describes the control plane functionality (see reference numerals 1818, 1918, 2018, and 2118). Control plane functionality includes features for configuring how control data is forwarded on the network (e.g., setting routes and routing tables, configuring VNICs, etc.). In some embodiments, a VCN control plane is provided that centrally computes all overlay mappings to the base layer and publishes them to the NVD and virtual network edge devices (such as various gateways, such as DRGs, SGWs, IGWs, etc.). Firewall rules can also be published using the same mechanism. In some embodiments, the NVD only receives mappings associated with that NVD. Data plane functionality includes the ability to actually route / forward packets based on the configuration established using the control plane. The VCN data plane is implemented by encapsulating client network packets before they traverse the base network. Encapsulation / decapsulation functionality is implemented on the NVD. In some embodiments, the NVD is configured to intercept all network packets entering and leaving the host machine and perform network virtualization functions.
[0122] As indicated above, NVD performs various virtualization functions, including VNIC and VCN VR. NVD can execute VNICs associated with compute instances hosted on one or more host machines connected to the VNIC. For example, as... Figure 2 As depicted, NVD 210 performs the functionality of VNIC 276 associated with compute instance 268 hosted by host machine 202 connected to NVD 210. As another example, NVD 212 performs VNIC 280 associated with bare-metal compute instance 272 hosted by host machine 206, and VNIC 284 associated with compute instance 274 hosted by host machine 208. Host machines can host compute instances belonging to different VCNs (which belong to different customers), and NVDs connected to host machines can perform VNICs corresponding to compute instances (i.e., perform VNIC-related functionality).
[0123] NVD also executes a VCN virtual router corresponding to the VCN of the compute instance. For example, in Figure 2 In the embodiments depicted, NVD 210 executes VCN VR 277 corresponding to the VCN to which compute instance 268 belongs. NVD 212 executes one or more VCN VR 283 corresponding to one or more VCNs to which compute instances hosted by host machines 206 and 208 belong. In some embodiments, the VCN VR corresponding to a VCN is executed by all NVDs connected to host machines hosting at least one compute instance belonging to that VCN. If a host machine hosts compute instances belonging to different VCNs, then NVDs connected to that host machine can execute VCN VR corresponding to those different VCNs.
[0124] In addition to VNIC and VCN VR, NVD can also execute various software (e.g., daemons) and include components, one or more of which facilitate various network virtualization functions performed by NVD. For simplicity, these various components are grouped together as... Figure 2 The “packet processing component” shown is illustrated. For example, NVD 210 includes packet processing component 286 and NVD 212 includes packet processing component 288. For example, a packet processing component for an NVD may include a packet processor configured to interact with the NVD’s ports and hardware interface to monitor all packets received by and transmitted using the NVD and to store network information. Network information may include, for example, network flow information identifying different network flows handled by the NVD and per-flow information (e.g., per-flow statistics). In some embodiments, network flow information may be stored on a per-VNIC basis. The packet processor may perform per-packet manipulation and implement stateful NAT and L4 firewall (FW). As another example, a packet processing component may include a replication agent configured to copy information stored by the NVD to one or more different replication target repositories. As yet another example, a packet processing component may include a logging agent configured to perform logging functions of the NVD. The packet processing component may also include software for monitoring the performance and health of the NVD and may also monitor the status and health of other components connected to the NVD.
[0125] Figure 1 The components of an example virtual or overlay network are shown, including a VCN, subnets within the VCN, compute instances deployed on the subnets, VNICs associated with the compute instances, VRs for the VCN, and a collection of gateways configured for the VCN. Figure 1 The overlay component described in the text can be made by Figure 2 One or more executions or hosts are described in the physical components. For example, a compute instance in a VCN can be executed or managed by... Figure 2 The VNIC described herein is executed or hosted by one or more host machines. For a compute instance hosted by a host machine, the VNIC associated with that compute instance is typically executed by an NVD connected to that host machine (i.e., VNIC functionality is provided by an NVD connected to that host machine). The VCN VR functionality for a VCN is executed by all NVDs connected to the host machine hosting or executing a compute instance as part of that VCN. The gateway associated with a VCN can be executed by one or more different types of NVDs. For example, some gateways can be executed by smartNICs, while others can be executed by one or more host machines or other implementations of NVDs.
[0126] As described above, compute instances in a client VCN can communicate with various endpoints, which may be in the same subnet as the source compute instance, in a different subnet but within the same VCN as the source compute instance, or outside the source compute instance's VCN. These communications are facilitated using VNICs, VCN VRs, and gateways associated with the VCNs.
[0127] For communication between two compute instances on the same subnet within a VCN, a VNIC associated with both the source and destination compute instances facilitates the communication. The source and destination compute instances can be hosted by the same host machine or different host machines. Packets originating from the source compute instance can be forwarded from the host machine hosting the source compute instance to an NVD connected to that host machine. On the NVD, packets are processed using a packet processing pipeline, which may include the execution of the VNIC associated with the source compute instance. Because the destination endpoint of the packet is within the same subnet, the execution of the VNIC associated with the source compute instance results in the packet being forwarded to the NVD executing the VNIC associated with the destination compute instance, which then processes the packet and forwards it to the destination compute instance. The VNIC associated with the source and destination compute instances can execute on the same NVD (e.g., when the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., when the source and destination compute instances are hosted by different host machines connected to different NVDs). The VNIC can use a routing / forwarding table stored by the NVD to determine the next hop for the packet.
[0128] For packets destined for endpoints in different subnets within the same VCN, the packets originating from the source compute instance are routed from the host machine hosting the source compute instance to the NVD connected to that host machine. On the NVD, the packets are processed using a packet processing pipeline, which may include the execution of one or more VNICs and the VR associated with the VCN. For example, as part of the packet processing pipeline, the NVD executes or invokes functionality associated with the VNIC associated with the source compute instance (also known as executing the VNIC). Functionality executed by the VNIC may include viewing VLAN tags on the packets. Since the packet's destination is outside the subnet, the VCN VR functionality is then invoked and executed by the NVD. The VCN VR then routes the packets to the NVD executing the VNIC associated with the destination compute instance. The VNIC associated with the destination compute instance then processes the packets and forwards them to the destination compute instance. The VNICs associated with the source and destination compute instances may execute on the same NVD (e.g., when the source and destination compute instances are hosted by the same host machine) or on different NVDs (e.g., when the source and destination compute instances are hosted by different host machines connected to different NVDs).
[0129] If the destination of a data packet is outside the VCN of the source compute instance, the packet originating from the source compute instance is forwarded from the host machine hosting the source compute instance to the NVD connected to that host machine. The NVD executes the VNIC associated with the source compute instance. Since the destination endpoint of the packet is outside the VCN, the packet is subsequently processed by the VCN VR used by that VCN. The NVD invokes the VCN VR functionality, which may result in the packet being forwarded to the NVD executing the appropriate gateway associated with the VCN. For example, if the destination is an endpoint within a customer's on-premises network, the packet may be forwarded by the VCN VR to the NVD executing the DRG gateway configured for the VCN. The VCN VR may execute on the same NVD as the NVD executing the VNIC associated with the source compute instance, or it may be executed by a different NVD. The gateway may be executed by the NVD, which may be a smartNIC, a host machine, or another NVD implementation. The packet is then processed by the gateway and forwarded to the next hop, which facilitates the packet's delivery to its intended destination endpoint. For example, in Figure 2In the embodiment depicted, data packets originating from compute instance 268 can be transmitted from host machine 202 to NVD 210 via link 220 (using NIC 232). On NVD 210, VNIC 276 is invoked because it is the VNIC associated with the source compute instance 268. VNIC 276 is configured to examine the information encapsulated in the data packets and determine the next hop for forwarding the data packets, with the aim of facilitating the transmission of the data packets to their intended destination endpoint, and then forwarding the data packets to the determined next hop.
[0130] Compute instances deployed on a VCN can communicate with a variety of endpoints. These endpoints can include endpoints hosted by CSPI 200 and endpoints outside of CSPI 200. Endpoints hosted by CSPI 200 can include instances within the same VCN or other VCNs, which can be the customer's VCN or a VCN not belonging to the customer. Communication between endpoints hosted by CSPI 200 can be performed via physical network 218. Compute instances can also communicate with endpoints not hosted by CSPI 200 or outside of CSPI 200. Examples of these endpoints include endpoints within the customer's on-premises network or data center, or public endpoints accessible via public networks such as the Internet. Communication with endpoints outside of CSPI 200 can use various communication protocols over public networks (e.g., the Internet). Figure 2 (not shown in the image) or a dedicated network ( Figure 2 (Not shown in the image) to execute.
[0131] Figure 2 The architecture of the CSPI 200 depicted herein is merely an example and is not intended to be limiting. Variations, alternatives, and modifications are possible in alternative embodiments. For example, in some implementations, the CSPI 200 may have more advanced features than... Figure 2 The systems or components shown may include more or fewer systems or components, and may combine two or more systems, or may have different system configurations or arrangements. Figure 2 The systems, subsystems, and other components described herein may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective system, using hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., a memory device).
[0132] Figure 4 The connectivity between the host machine and the NVD, according to certain embodiments, is described for providing I / O virtualization to support multi-tenancy. For example... Figure 4As depicted, host machine 402 executes a hypervisor 404 that provides a virtualized environment. Host machine 402 executes two virtual machine instances, VM1 406 belonging to customer / tenant #1 and VM2 408 belonging to customer / tenant #2. Host machine 402 includes a physical NIC 410 connected to NVD 412 via link 414. Each compute instance is attached to a VNIC executed by NVD 412. Figure 4 In the embodiment, VM1 406 is attached to VNIC-VM1 420 and VM2 408 is attached to VNIC-VM2 422.
[0133] like Figure 4 As shown, NIC 410 includes two logical NICs, logical NIC A 416 and logical NIC B 418. Each virtual machine is attached to its own logical NIC and configured to work with its own logical NIC. For example, VM1 406 is attached to logical NIC A 416 and VM2 408 is attached to logical NIC B 418. Although host machine 402 includes only one physical NIC 410 shared by multiple tenants, each tenant's virtual machines believe they have their own host machine and network interface card due to the logical NICs.
[0134] In some embodiments, each logical NIC is assigned its own VLAN ID. Thus, a specific VLAN ID is assigned to logical NIC A 416 for tenant #1, and a different VLAN ID is assigned to logical NIC B 418 for tenant #2. When a packet is transmitted from VM1 406, the hypervisor appends a tag assigned to tenant #1 to the packet, and the packet is then transmitted from host machine 402 to NVD 412 via link 414. Similarly, when a packet is transmitted from VM2 408, the hypervisor appends a tag assigned to tenant #2 to the packet, and the packet is then transmitted from host machine 402 to NVD 412 via link 414. Accordingly, the packet 424 transmitted from host machine 402 to NVD 412 has an associated tag 426 identifying the specific tenant and the associated VM. On the NVD, for a data packet 424 received from the host machine 402, the tag 426 associated with the data packet is used to determine whether the data packet is processed by VNIC-VM1 420 or VNIC-VM2 422. The data packet is then processed by the corresponding VNIC. Figure 4 The configuration described in [the document] enables each tenant's compute instance to believe that it owns its own host machine and NIC. Figure 4 The setup described in [the document] provides I / O virtualization to support multi-tenancy.
[0135] Figure 5A simplified block diagram of a physical network 500 according to certain embodiments is depicted. Figure 5 The embodiments depicted are structured as Clos networks. Clos networks are a specific type of network topology designed to provide connectivity redundancy while maintaining high bandwidth and maximum resource utilization. Clos networks are non-blocking, multi-stage or multi-layer switching networks, where the number of stages or layers can be two, three, four, five, etc. Figure 5 The embodiment depicted is a Layer 3 network, including Layer 1, Layer 2, and Layer 3. TOR switch 504 represents a Layer 0 switch in a Clos network. One or more NVDs are connected to the TOR switch. Layer 0 switches are also referred to as edge devices of the physical network. Layer 0 switches are connected to Layer 1 switches, also known as leaf switches. Figure 5 In the embodiments depicted, a set of "n" Layer 0 TOR switches is connected to a set of "n" Layer 1 switches, forming a pod. Each Layer 0 switch in the pod is interconnected to all Layer 1 switches in that pod, but there is no switch connectivity between pods. In some implementations, two pods are referred to as blocks. Each block is served by or connected to a set of "n" Layer 2 switches (sometimes called backbone switches). There can be several blocks in the physical network topology. The Layer 2 switches are then connected to "n" Layer 3 switches (sometimes called super backbone switches). Communication of packets on the physical network 500 is typically performed using one or more Layer 3 communication protocols. Typically, all layers of the physical network (except the TOR layer) are n-way redundant, thus allowing high availability. Policies can be specified for pods and blocks to control the visibility of switches to each other in the physical network, thereby enabling scaling of the physical network.
[0136] A key characteristic of Clos networks is that the maximum number of hops from one Layer 0 switch to another (or from an NVD connected to a Layer 0 switch to another NVD connected to a Layer 0 switch) is fixed. For example, in a Layer 3 Clos network, a packet takes a maximum of seven hops to reach another NVD, where the source and destination NVDs are connected to the leaf layers of the Clos network. Similarly, in a Layer 4 Clos network, a packet takes a maximum of nine hops to reach another NVD, where the source and destination NVDs are connected to the leaf layers of the Clos network. Therefore, the Clos network architecture maintains consistent latency throughout the network, which is important for communication within and between data centers. Clos topologies are horizontally scalable and cost-effective. Network bandwidth / throughput capacity can be easily increased by adding more switches at each layer (e.g., more leaf switches and backbone switches) and by increasing the number of links between switches in adjacent layers.
[0137] In some implementations, each resource within CSPI is assigned a unique identifier called a Cloud Identifier (CID). This identifier is included as part of the resource's information and can be used to manage the resource, for example, via a console or API. An example syntax for a CID is:
[0138] ocid1.<RESOURCE TYPE> . <realm>[REGION][FUTURE USE].<UNIQUE ID>
[0139] in,
[0140] ocid1: A text string indicating the version of the CID;
[0141] resource type: The type of resource (e.g., instance, volume, VCN, subnet, user, group, etc.);
[0142] realm: The realm where the resource resides. Example values are "c1" for the commercial realm, "c2" for the government cloud realm, or "c3" for the federal government cloud realm, etc. Each realm can have its own domain name;
[0143] region: The region where the resource is located. This section may be empty if the region is not applicable to the resource.
[0144] future use: to be reserved for future use.
[0145] Unique ID: The unique part of the ID. The format may vary depending on the type of resource or service.
[0146] Global Virtual Plane
[0147] Cloud infrastructure service providers supply newer and faster graphics processing units (GPUs) to address the growing demands (e.g., bandwidth requirements) of high-performance computing applications. GPU workloads typically run on one or more host machines. Often, such workloads fail to achieve the expected throughput levels. One factor contributing to this problem is the lack of flow entropy, such as Equal Equivalent Multipath (ECMP) flow entropy. In ECMP, multiple flows (e.g., from different host machines) may be hashed in a way that makes both flows expected to traverse the same outgoing link / port of a switch. Furthermore, the fact that host machines exchange traffic without considering which other hosts are in their local network neighborhood exacerbates the problem. This situation often leads to bandwidth contention issues, commonly referred to in the literature as flow-collision-based congestion problems.
[0148] The following is for reference. Figure 6 This describes in detail the problems caused by ECMP flow entropy. To address the problems caused by ECMP flow entropy, according to some embodiments of this disclosure, a novel concept (referred to herein as a "global virtual plane") is provided that eliminates ECMP traffic load balancing decisions on switches, thereby providing a significant throughput improvement. In other words, this disclosure provides a scalable networking scheme for performing GPU workloads in hierarchical network configurations (e.g., Layer 2 or Layer 3 CLOS network configurations). This scheme avoids (i.e., eliminates) ECMP-based traffic load balancing decisions, thereby avoiding flow hash collisions that lead to network congestion. Therefore, the workload can achieve theoretically maximum network performance. References are provided below. Figure 7-9 Describe the details regarding scalable networking solutions.
[0149] Figure 6 A block diagram of a cloud infrastructure 600 arranged in a CLOS network according to certain embodiments is depicted. The cloud infrastructure 600 includes multiple racks (e.g., rack 1 610 and rack 2 620). Each rack includes multiple host machines (also referred to herein as hosts). Rack 1 610 is depicted as including two host machines, namely, host 1-A 612 and host 1-B 614, and rack 2 620 is depicted as including two host machines, namely, host 2-A 622 and host 2-B 624. It should be appreciated that... Figure 6 The illustrations in the diagram (i.e., each rack includes two host machines) are intended to be illustrative and non-limiting. For example, a cloud infrastructure may include more than two racks, where each rack may include more than two host machines. Furthermore, it is important to note that each rack is not limited to having the same number of hosts. Rather, one rack may have more or fewer host machines compared to the number of host machines contained in another rack.
[0150] Each host machine includes multiple graphics processing units (GPUs). For example, host machine 1-A 612 includes N GPUs, such as GPU 1, 613. Furthermore, it should be recognized that... Figure 6 The illustration of each host machine including the same number of GPUs (i.e., N GPUs) is intended to be illustrative and non-limiting; that is, each host machine may include a different number of GPUs. Each rack includes a top-of-rack (TOR) switch that is communicatively coupled to the GPUs hosted on the host machines within that rack. For example, rack 1 610 includes a TOR switch (i.e., TOR 1) 616 communicatively coupled to host machines host 1-A, 612 and host 1-B, 614, while rack 2 620 includes a TOR switch (i.e., TOR 2) 626 communicatively coupled to host machines host 2-A, 622 and host 2-B, 624. It should be recognized that... Figure 6 The TOR switches depicted (i.e., TOR 1 616 and TOR 2 626) each include N ports, which are used to communicatively couple the TOR switches to the N GPUs hosted on each host machine included in the rack. Figure 6 The depiction of the coupling between the TOR switch and the GPU is intended to be illustrative and non-limiting. For example, in some embodiments, the TOR switch may have multiple ports, each corresponding to a GPU on a host machine; that is, the GPU on the host machine may be connected to a single port of the TOR via a communication link.
[0151] The TOR switches from each rack are communicatively coupled to multiple backbone switches (also referred to herein as upper-layer switches), such as backbone switch 1,630 and backbone switch P,640. For example, as... Figure 6 As shown, TOR 1, 616 is connected to backbone switch 1, 630 via two links and to backbone switch P, 640 via two additional links. Information transmitted from a specific TOR switch to the backbone switch is referred to herein as communication via an uplink, while information transmitted from the backbone switch to a TOR switch is referred to herein as communication via a downlink. According to some embodiments, the TOR switches and backbone switches are connected in a CLOS network configuration (e.g., a multi-stage switching network), where each TOR switch forms a "leaf" node in the CLOS network.
[0152] In some embodiments, GPUs included in the host machine perform machine learning-related tasks. In this setup, a single task can be executed / distributed across a large number of GPUs (e.g., 64 GPUs), which may be distributed across multiple host machines and racks. Since all these GPUs are processing the same task (i.e., the workload), they all need to communicate with each other in a time-synchronized manner. Furthermore, at any given time, the GPUs are either in compute mode or communication mode; that is, the GPUs communicate with each other approximately at the same time. The speed of the workload is determined by the speed of the slowest GPU.
[0153] Typically, Equal-Cost Multipath (ECMP) routing is used to route packets from a source GPU to a destination GPU. In ECMP routing, when multiple equal-cost paths exist to route traffic from the sender to the receiver, a selection technique is used to choose a specific path. Accordingly, at the network device receiving the traffic (e.g., a TOR switch or backbone switch), a selection algorithm is used to select the outgoing link to forward the traffic from the network device to subsequent devices. This outgoing link selection occurs at each network device along the path from the sender to the receiver. Hash-based selection is a widely used ECMP selection technique, where the hash can be based on, for example, a 4-tuple of the packet (e.g., source port, destination port, source IP, destination IP).
[0154] ECMP routing is a flow-aware routing technique where each flow (i.e., a sequence of packets) is hashed to the same path for the duration of the flow. Therefore, packets in a flow are forwarded from network devices using specific outgoing ports / links. This is typically done to ensure that packets in a flow arrive in order, i.e., without needing to reorder them. However, ECMP routing is bandwidth (or throughput) insensitive. In other words, TOR and backbone switches perform statistical flow-aware (throughput-insensitive) ECMP load balancing on parallel links.
[0155] In standard ECMP routing (i.e., flow-aware routing only), a problem arises when flows received by a network device via two separate inbound links might be hashed onto the same outbound link, leading to flow collisions. For example, consider a scenario where two flows arrive via two separate 100G inbound links, and each flow is hashed onto the same 100G outbound link. This situation causes congestion (i.e., flow collisions) and results in packet drops because the inbound bandwidth is 200G while the outbound bandwidth is only 100G. Figure 6 As shown, there are two flows: flow 1 641, which flows from the first GPU on host machine 1-A, 612 to TOR switch 616; and flow 2 643, which flows from another GPU on host machine 614 to TOR switch 616. Note that these two flows point to the TOR switch on separate links. Assume... Figure 6 All links depicted have a capacity (i.e., bandwidth) of 100G. When the TOR switch 616 executes the ECMP routing algorithm, these two flows may be hashed to the same outgoing link using TOR, for example, link 650 connecting the TOR switch 616 to the backbone switch 630. In this case, a collision exists between the two flows (indicated by an "X" symbol), causing packets to be dropped.
[0156] Regardless of the protocol used, this congestion scenario is generally problematic for all types of traffic. For example, TCP is intelligent because when a packet is dropped and the sender does not receive an acknowledgment of the dropped packet, it will be retransmitted. However, the situation is more severe for Remote Direct Memory Access (RDMA) traffic. There are several reasons why RDMA networks do not use TCP (e.g., TCP has complex logic that is detrimental to low latency and high performance). RDMA networks use protocols such as RDMA over Infiniband or RDMA over Converged Ethernet (RoCE). In RoCE, there is a congestion control algorithm where the sender slows down packet transmission when it identifies congestion or dropped packets. For dropped packets, not only is the dropped packet retransmitted, but several packets surrounding it are also retransmitted, further consuming available bandwidth and causing performance degradation.
[0157] Therefore, flow collisions are a critical issue for workloads such as GPU workloads due to stringent time synchronization requirements. This paper describes a scalable networking scheme for executing GPU workloads in Layer 2 or Layer 3 CLOS network configurations. This scheme avoids (i.e., eliminates) the ECMP-based traffic load balancing decisions of ToR switches. In other words, this scheme eliminates the flow hash collision problem that leads to network congestion. Therefore, the workload can achieve theoretically maximum network performance.
[0158] Now go to Figure 7 This document depicts an exemplary network architecture 700 based on the concept of a global virtual plane according to certain embodiments. Network architecture 700 includes multiple host machines (labeled nodes) communicatively coupled to each other via multiple switches. Each host machine includes one or more GPUs. The one or more GPUs of each host machine can be configured to perform artificial intelligence or machine learning workloads. The multiple switches are arranged in a hierarchical structure, such as a CLOS network architecture like a Layer 2 or Layer 3 CLOS network. In one embodiment, network architecture 700 includes a three-layer structure of switches, including a Layer 1 switch (T0), a Layer 2 switch (T1), and a Layer 3 switch (T2).
[0159] In one implementation, the network architecture is provided as multiple blocks. For example, such as Figure 7 As shown, network architecture 700 comprises K blocks, namely blocks 1, 705 to K, 745. Each block includes multiple host machines, multiple Layer 1 switches, and multiple Layer 2 switches. For example, block 705 includes "N" Layer 1 switches (T0) labeled 717 and 727, respectively. Furthermore, block 705 includes "M" Layer 2 switches (T1) labeled 715A, 715B, 715C, and 715D. Figure 7 As shown, multiple host machines are directly coupled to each switch in the Layer 1 switch. For example, host machines 706 and 707 are directly coupled to Layer 1 switch 717, and host machines 725 and 726 are coupled to Layer 1 switch 727.
[0160] Note that each host machine (i.e., node) can be represented by a tuple containing three identifiers (x, y, z), where identifier "x" corresponds to the ID of the block containing the host machine, identifier "y" corresponds to the ID of the Layer 1 (T0) switch directly coupled to the host machine, and identifier "z" corresponds to the host machine's identifier. Each host machine may include one or more GPUs, for example, a GPU 706A contained in the host machine labeled 1-1-1. Furthermore, as... Figure 7 As shown, N Layer 1 switches (i.e., 717, 727) communicatively couple multiple host machines to M Layer 2 switches, such as switches 715A, 715B, 715C and 715D.
[0161] Block K 745 has a configuration similar to that of block 705. For example, block 745 includes "N" Layer 1 switches (T0) labeled 737 and 747 respectively. Furthermore, block 745 includes "M" Layer 2 switches (T1) labeled 735A, 735B, 735C, and 735D. Figure 7 As shown, multiple host machines are directly coupled to each switch in the Layer 1 switches. For example, host machines 734 and 736 are directly coupled to Layer 1 switch 737, and host machines 738 and 740 are coupled to Layer 1 switch 747. Furthermore, as... Figure 7 As shown, N Layer 1 switches (i.e., 734 and 747) communicatively couple multiple host machines to M Layer 2 switches, such as switches 735A, 735B, 735C, and 735D. Note that, although Figure 7 The illustration depicts two host machines directly coupled to each switch included in the Layer 1 switch, but this is not intended to be limiting. Rather, each switch included in the Layer 1 switch can have a greater number of host machines coupled to that switch.
[0162] Furthermore, the network architecture includes multiple groups of upper-layer switches, such as Upper Layer 1 of switch 701 and Upper Layer K / 2 of switch 731. Each group in the upper layers of the switches includes multiple Layer 3 switches (T2). For example, Upper Layer 1 of switch 701 includes "M" Layer 3 switches labeled 713A, 713B, 713C, and 713D. Similarly, Upper Layer K / 2 of switch includes M Layer 3 switches labeled 733A, 733B, 733C, and 733D. The Layer 3 switches are communicatively coupled to different blocks included in the network architecture.
[0163] As previously described, this disclosure provides a novel concept called a "global virtual plane". The goal of implementing a global virtual plane (also referred to herein as a "virtual plane") is to eliminate ECMP traffic load balancing decisions on switches, thereby providing a significant throughput improvement. According to some embodiments, a virtual plane is established as follows: the resources of the entire network architecture 700 are partitioned into multiple parts. Each part is assigned a virtual plane / associated with a virtual plane. In other words, the resources of the network architecture may correspond to multiple switches included in a hierarchical structure of switches, i.e., T0, T1, and T2 layer switches. Each switch includes multiple ports. Therefore, the network architecture as a whole can be interpreted as including multiple switches with multiple ports. In one implementation, a first subset of ports among the multiple ports is associated with a virtual plane (e.g., a first virtual plane). Thus, the first virtual plane identifies / corresponds to a first set of resources that will be dedicated to transmitting packets from / to the host machine associated with the first virtual plane. The resources associated with the first virtual plane may include: (i) a first subset of ports included in each of the first layer switches; (ii) a first subset of switches included in the second layer switches; and (iii) a first subset of switches included in the third layer switches.
[0164] Similarly, a second subset of ports among the multiple ports included in the network architecture is associated with another virtual plane (e.g., a second virtual plane). Therefore, the second virtual plane identifies / corresponds to a second set of resources specifically used for transmitting packets from / to the host machine associated with the second virtual plane. Resources associated with the second virtual plane may include: (i) a second subset of ports included in each of the first-layer switches; (ii) a second subset of switches included in the second-layer switches; and (iii) a second subset of switches included in the third-layer switches. Note that each of the three resources associated with the first virtual plane is different from the three resources associated with the second virtual plane.
[0165] As an example, consider Layer 1 switch (T0) 717. The ports of this switch are partitioned into two parts: port 717A and port 717B. Similarly, the ports of switch 727 are partitioned into two parts: port 727A and port 727B, while the ports of switches 737 and 747 (included in block K) are partitioned into ports 737A and 737B, and ports 747A and 747B, respectively. Ports labeled 717A and 727A (of the switches included in block 1) and ports labeled 737A and 747A (of the switches included in block K) are associated with the first global virtual plane. In contrast, ports labeled 717B, 727B, 737B, and 747B (of the Layer 1 switches) are associated with the second virtual plane.
[0166] Furthermore, among the "M" Layer 2 switches (i.e., 715A, 715B, 715C, and 715D) contained in block 1 705, a first subset of switches (denoted as 750A), such as switches 1 715A to switch B 715B, can be associated with a first virtual plane. Similarly, among the "M" Layer 2 switches (i.e., 735A, 735B, 735C, and 735D) contained in block K 745, a first subset of switches (denoted as 750D), such as switches 1 735A to switch B 735B, is associated with a first virtual plane. In a similar manner, a subset of switches contained in the upper layer of the switches, i.e., the Layer 3 switches, can be assigned to the first virtual plane, such as the switch group denoted as 750B and 750C. Similarly, the Layer 2 switches 715C, 715D, 735C, and 735D, as well as the upper-layer switches 713C, 713D, 733C, and 733D, are associated with the second virtual plane. It should be recognized that when a particular switch in a Layer 2 (or Layer 3) switch is associated with a particular virtual plane, it means that all ports of that particular switch—i.e., uplink and downlink ports—are associated with that particular virtual plane.
[0167] Furthermore, each Layer 1 switch has multiple host machines directly coupled to that switch. In one implementation, each host machine is assigned to a different virtual plane. For example, as... Figure 7 As shown, host machines 706, 725, 737, and 747 are associated with the first virtual plane. The host machines and resources associated with the first virtual plane are... Figure 7 Solid lines are used to draw them. Similarly, host machines and resources associated with another virtual plane (e.g., a second virtual plane) are also represented in this way. Figure 7 The virtual planes are drawn using dashed lines. According to some embodiments, the number of virtual planes a network architecture can support corresponds to the number of host machines directly coupled to a switch contained within a Layer 1 switch. For example, as... Figure 7 As shown, each switch in the Layer 1 switches includes two host machines coupled to that switch. Therefore, Figure 7 A total of two global virtual planes are depicted. Note that, Figure 7 The two global virtual planes depicted are for illustrative purposes only. Network architectures can support a much larger number of global virtual planes.
[0168] Therefore, within the framework described above, it's important to note that each virtual plane is associated with a unique set of resources (i.e., the ports of the switches contained in the Layer 1 switches) and a unique set of switches in the Layer 1 and Layer 2 switches (i.e., the uplink and downlink ports of the switches). Furthermore, port-to-port traffic forwarding is implemented at each layer within the network architecture. This results in traffic on a particular virtual plane remaining on the same virtual plane in an end-to-end manner. Therefore, if a specific GPU on a particular host machine expects to communicate with another GPU on another host machine (assuming both host machines are associated with the same virtual plane), then an end-to-end path (i.e., referred to herein as a "traffic path" within the virtual plane, and referenced later) is pre-established. Figure 8 (Description to be provided). Furthermore, it should be recognized that within the global virtual plane, port-to-port traffic forwarding can be accomplished by pre-determining the traffic path. This can be achieved through policy-based routing mechanisms, static routing mechanisms, and / or hardware forwarding rules similar to OpenFlow.
[0169] In this way, a host machine assigned to a specific virtual plane is restricted to using only the resources associated with that virtual plane when transmitting data packets to other host machines associated with that virtual plane. In other words, a first host machine associated with a first virtual plane is not allowed to communicate with a second host machine associated with a second virtual plane. In this way, traffic isolation between different clients can be achieved by assigning host machines associated with different virtual planes to different clients. Note that in... Figure 7 In this network architecture, ECMP traffic load balancing decisions are not required on T0, T1, or T2 layer switches. In other words, for traffic between different GPUs, this network design eliminates flow hash collisions that cause congestion. In this way, workloads can achieve theoretically maximum network performance.
[0170] Figure 8 An exemplary network architecture 800 is depicted, illustrating traffic paths established within a global virtual plane according to certain embodiments. For example, a first traffic path (shown by bold lines) represents a traffic path (which can be predetermined using a port-to-port traffic forwarding mechanism) for transmitting data packets originating from host machine 1-1-1 (706), i.e., the source host machine, and terminating at host machine K-1-1 (734), i.e., the destination host machine. Similarly, host machine 1-N-1 (725) can, for example, communicate with node KN-1 (738), both nodes being associated with the first virtual plane using a different set of resources than those used by the first traffic path. In other words, the different traffic paths contained within the first virtual plane use non-overlapping (or different) resources to transmit traffic. Furthermore, Figure 8 Another traffic path (e.g., a second traffic path) established from node 1-N-2 (726) to node KN-2 (740) is depicted, drawn by dashed lines. It should be recognized that the second traffic path does not utilize any resources allocated to the first virtual plane in the network architecture.
[0171] Figure 9 The illustration depicts a scenario according to certain embodiments in use. Figure 7 An exemplary flowchart 900 shows the steps performed when a network infrastructure transmits data packets. Figure 9 The processing described herein can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of various systems. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 9 The method 900 presented and described below is intended to be illustrative and non-limiting. Although Figure 9 Various processing steps that occur in a specific sequence or order are described, but this is not intended to be limiting. In some alternative embodiments, these steps may be performed in a different order, or some steps may be performed in parallel. It should be recognized that the control plane of the network infrastructure can be configured to perform... Figure 9 The steps described in the text.
[0172] Figure 9 The steps described can be implemented in a network environment comprising multiple host machines, each communicatively coupled to other host machines via a network architecture. Each host machine includes one or more GPUs. In one implementation, the network architecture may include multiple switches (including multiple ports) arranged in a hierarchical structure such as a Layer 2 CLOS network or a Layer 3 CLOS network. For illustration, consider a Layer 3 CLOS network; the network architecture may include Layer 1 switches, Layer 2 switches, and Layer 3 switches. In one implementation, multiple host machines are directly coupled to a switch contained within a Layer 1 switch. The Layer 2 switch then communicatively couples the Layer 1 switch to the Layer 3 switch.
[0173] In this setting, and as Figure 9 As shown, the process begins in step 905, where a first subset of ports from a plurality of ports is associated with a first virtual plane. The first virtual plane identifies a first set of resources specifically for transmitting data packets from and to the host machine associated with the first virtual plane. The process then proceeds to step 910, where a second subset of ports from a plurality of ports is associated with a second virtual plane. Note that the second virtual plane is different from the first virtual plane.
[0174] Subsequently, in step 915, the first host machine and the second host machine among the plurality of host machines are each associated with the first virtual plane. In step 920, for a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine, the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using only ports in the first subset of ports. For example, refer to Figure 8 Consider the transmission of data packets from the GPU on the first host machine (e.g., node 1-1-1 (706)) to the GPU on the second host machine (e.g., node K-1-1 (734)). Note that each of host machines 706 and 734 is associated with the first virtual plane. It should be recognized that the transmission of data packets specifically uses the resources associated with the first virtual plane, namely, one of the ports 717A associated with the first-layer switch 717, one of the ports of switch 715A (i.e., the second-layer switch), one of the ports of switch 713A (i.e., the third-layer switch), one of the ports of switch 735A (i.e., another second-layer switch), and one of the ports 737A associated with another first-layer switch 734.
[0175] Interconnect Global Virtual Plane
[0176] Previous reference Figure 7-9 The concept of a "global virtual plane" is described. As mentioned earlier, the advantage of a global virtual plane is that it avoids (i.e., eliminates) top-of-rack (ToR) switches (e.g., Figure 7 The architecture described herein (including switches in layers T0, T1, and T2) uses ECMP-based traffic load balancing decisions. In other words, the global virtual plane eliminates flow hash collisions that cause network congestion.
[0177] However, a limitation of global virtual plane design is that GPUs contained in a first host machine associated with a specific plane (e.g., the first virtual plane) cannot communicate with another GPU associated with a different virtual plane (e.g., the second virtual plane) (on a different host machine). There are situations where it is necessary to build a larger topology by connecting such GPUs. For example, host machines associated with different virtual planes may be assigned to specific clients. In such cases, interconnecting host machines associated with different virtual planes is clearly necessary or required. However, doing so carelessly would mean that different virtual planes might end up overlapping on a single link. This would imply traffic over-subscription and hash collisions. In the case of assigning host machines associated with different virtual planes to specific clients, it is worth noting that a naive approach to avoiding interconnecting virtual planes is to break down the client's overall workload into multiple sub-workloads. Furthermore, different sub-workloads can be assigned to different virtual planes. However, this approach may be not scalable, and therefore there is a requirement to interconnect different virtual planes. A novel technique for interconnecting global virtual planes is described below.
[0178] Example 1: In this example, the virtual planes are connected only on the ToR switch. Specifically, the virtual planes are connected on a Tier 0 switch. This is because doing so keeps traffic between virtual planes local to the ToR, and therefore avoids hash collisions. Note that in this example of connected virtual planes, only nodes (i.e., host machines) located behind the same ToR (i.e., Tier 0 switch) are allowed to switch packet transmission from one virtual plane to another.
[0179] refer to Figure 10 The illustration depicts an exemplary network architecture of a first embodiment of an interconnected global virtual plane. Figure 10 The network architecture 1000 describes the previous reference Figure 7 This describes a portion of the overall network architecture. Specifically, network architecture 1000 includes a first switch block 1005, which is communicatively coupled to the upper layer 1 of switch (1001) and the upper layer K / 2 of switch (1031). Note that the first switch block 1005 corresponds to... Figure 7 Block 1 705, and the upper layer 1 of switch (1001) and the upper layer K / 2 of switch (1031) respectively correspond to Figure 7 The upper layer 1705 and the upper layer K / 2731 of the switch.
[0180] For the sake of explanation, in Figure 10 In the diagram, the first switch block 1005 is depicted as including two switches 1017 and 1027 contained in the first-layer switches (i.e., the T0 layer), and multiple switches contained in the second-layer switches, namely, switches 715A, 715B, 715C, and 715D. Furthermore, the first-layer switches 1017 and 1027 correspond to... Figure 7 The first-layer switches 717 and 727 shown are included, respectively, ports labeled 1017A and 1017B (in switch 1017) and ports labeled 1027A and 1027B (in switch 1027). Compared to ( Figure 7 (The) 717 / 727 switch, ( Figure 10 The difference between the 1017 / 1027 and the VxLAN 1017 / 1027 is that the 1017 / 1027 is a switch (or router) capable of routing traffic between overlay networks. An overlay network can be one of VxLAN, Geneve, MPLS, or a similar overlay network. Therefore, the 1017 / 1027 is a switch, such as a VxLAN router / switch, capable of converting data transmitted on one overlay network (e.g., a first virtual plane) into data compatible with other data transmitted on another overlay network (e.g., a second virtual plane).
[0181] According to some embodiments, the data transformation process by switch 1017 / 1027 (e.g., a VxLAN router) includes: (i) decapsulating packets received on a first overlay network (e.g., a first virtual plane) to extract information corresponding to a first header associated with the first virtual plane, and (ii) encapsulating the first packet with information corresponding to a second header associated with another overlay network (e.g., a second virtual plane). Note that in this case, the traffic is Layer 3 routed, and the data transmission switching process can be performed internally within the switch (e.g., via software components such as programming the node's chip).
[0182] According to some embodiments, as referenced above Figure 10 The flexibility described in performing traffic routing between different overlay networks is provided on each switch in the Layer 1 (T0) switch of the network architecture, for example, Figure 7 The switches are 717, 727, 737, and 747. However, it should be noted that if a customer needs to switch between different virtual planes, then the customer must be assigned a node located behind the same ToR (i.e., a T0-level switch) (belonging to different virtual planes). Figure 11 Depicting Figure 7 The logical connections of the nodes (i.e., host machines) depicted in the diagram (i.e., node 1-1-1 (706), node 1-1-2 (707), node 1-N-1 (725), node 1-N-2 (726), node K-1-1 (734), node K-1-2 (736), node KN-1 (738), and node KN-1 (740)). Figure 11 As shown in the exemplary embodiment, a first conversion may occur between data packets transmitted on the second virtual plane and data packets sent on the first virtual plane, i.e., between node K-1-2 and node K-1-1 (labeled 1110). Note that these node / host machines are located on the same T0 switch, i.e. Figure 7 Behind switch 737. Similarly, a second conversion may occur between packets transmitted on the first virtual plane and packets sent on the second virtual plane, i.e., between node 1-N-1 and node 1-N-2 (labeled 1120). Note that these node / host machines are located on the same T0 switch, i.e. Figure 7 The back of the 727 switch.
[0183] Now go to Figure 12 This describes another exemplary mechanism for transforming data between different virtual planes according to an embodiment. Specifically, Figure 12 The embodiments depicted correspond to bridging mechanisms between different global planes. Figure 10 similar, Figure 12 The network architecture 1200 describes the previous reference Figure 7 This describes a portion of the overall network architecture. Specifically, network architecture 1200 includes a first switch block 1205, which is communicatively coupled to the upper layer 1 of switch (1201) and the upper layer K / 2 of switch (1231). Note that the first switch block 1205 corresponds to... Figure 7 Block 1 705, and the upper layer 1 of switch (1201) and the upper layer K / 2 of switch (1231) respectively correspond to Figure 7 The upper layer 1 705 and the upper layer K / 2 of the switch 731.
[0184] For the sake of explanation, in Figure 12 In the diagram, the first switch block 1205 is depicted as including two switches 1217 and 1227 contained in the first-layer switches (i.e., T0 layer), and multiple switches contained in the second-layer switches (T1 layer), namely, switches 715A, 715B, 715C, and 715D. Furthermore, the first-layer switches 1217 and 1227 correspond to... Figure 7 The first-layer switches 717 and 727 shown include ports labeled 1217A and 1217B (in switch 1217) and ports labeled 1227A and 1227B (in switch 1027), respectively. Compared to Figure 7 717 / 727 switches (and Figure 10 (1017 / 1027 switches) Figure 12 The difference between the 1217 and 1227 switches is that the 1217 / 1227 is a switch that includes the physical cabling between the ports of the 1210 switch. For example, as Figure 12 As shown, port 1217A of switch 1217 (associated with the first virtual plane) is physically connected to port 1217B of the same switch (associated with the second virtual plane). Similarly, port 1227A of switch 1227 (associated with the first virtual plane) is physically connected to port 1227B of the same switch (associated with the second virtual plane). Specifically, in the bridging mechanism, traffic is mixed at Layer-2. For example, traffic received on a particular VxLAN (e.g., the first virtual plane) is converted into VLAN traffic (by known means) and sent on another virtual plane (e.g., the second virtual plane) via physical cables interconnecting the ports associated with the first and second virtual planes. Therefore, Figure 12 The embodiments presented provide layer 2 traffic mixing and can be used in scenarios where T0 switches do not have layer 3 routing capabilities (e.g., not VxLAN routers).
[0185] Figure 13 An exemplary flowchart illustrating steps performed when transmitting data packets using a first embodiment of an interconnected global virtual plane, according to certain embodiments, is shown. Figure 13 The illustration depicts a scenario according to certain embodiments in use. Figure 10 An exemplary flowchart 1300 shows the steps performed when the network infrastructure transmits data packets. Figure 13 The processing described herein can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of a corresponding system. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 13 The method 1300 presented and described below is intended to be illustrative and non-limiting. Although Figure 13 Various processing steps that occur in a specific sequence or order are described, but this is not intended to be limiting. In some alternative embodiments, these steps may be performed in a different order, or some steps may be performed in parallel. It should be recognized that the control plane of the network infrastructure can be configured to perform... Figure 13 The steps described in the text.
[0186] Figure 13 The steps described can be implemented in a network environment comprising multiple host machines, each communicatively coupled to other host machines via a network architecture. Each host machine includes one or more GPUs. In one implementation, the network architecture may include multiple switches (including multiple ports) arranged in a hierarchical structure such as a Layer 2 CLOS network or a Layer 3 CLOS network. For illustration, consider a Layer 3 CLOS network; the network architecture may include Layer 1 switches, Layer 2 switches, and Layer 3 switches. In one implementation, multiple host machines are directly coupled to a switch contained within a Layer 1 switch. The Layer 2 switch then communicatively couples the Layer 1 switch to the Layer 3 switch.
[0187] In this setting, and as Figure 13 As shown, the process begins in step 1305, where a first subset of ports from a plurality of ports is associated with a first virtual plane. The first virtual plane identifies a first set of resources to be dedicated to transmitting packets from and to the host machine associated with the first virtual plane. The process then proceeds to step 1310, where a second subset of ports from a plurality of ports is associated with a second virtual plane. Note that the second virtual plane is different from the first virtual plane.
[0188] Subsequently, in step 1315, a first host machine among the plurality of host machines is associated with a first virtual plane, and a second host machine among the plurality of host machines is associated with a second virtual plane. The first host machine is directly coupled to a first switch among the plurality of switches, and the second host machine is directly coupled to a second switch among the plurality of switches. In one embodiment, the second switch is different from the first switch, and both the first switch and the second switch are connected to a T0 layer switch (which may be different T0 switches).
[0189] In step 1320, for a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine (e.g., a first data packet), the data packet is transferred from the first GPU on the first host machine to the second GPU on the second host machine using ports from the first port subset and the second port subset. Note that in the transmission of the first data packet, the first data packet is processed by either the first switch or the second switch to be transmitted on the second virtual plane instead of the first virtual plane. An example of transmitting a data packet (step 1320) can be found in [reference needed]. Figure 7 Describe and assume Figure 7 All Layer 1 switches (e.g., 717, 727, etc.) have the aforementioned reference. Figure 10 The data transformation described is provided. In one case, assume that node 1-1-1 (associated with the first virtual plane) corresponds to the first host machine, and node 1-N-2 (associated with the second virtual plane) corresponds to the second host machine. Assume that the GPU contained in node 1-1-1 wants to send data packets to another GPU on node 1-N-2. One way to achieve this communication is as follows: node 1-1-1 (associated with the first virtual plane 1) can transmit data packets to a third host machine (e.g., node 1-N-1) that is also associated with the first virtual plane. Note that the third host machine and the second host machine (1-N-2) are directly coupled to a second switch (i.e., switch 727). In this case, the data packets can be processed by the second switch, for example via... Figure 10 The routing mechanism transmits data to the second host machine on the second virtual plane. It should be recognized that both the third and second host machines are coupled to the same T0 layer switch, namely switch 727.
[0190] In another implementation, assume that node 1-1-1 (associated with the first virtual plane) corresponds to the first host machine, node 1-N-2 (associated with the second virtual plane) corresponds to the second host machine, and node 1-1-2 (associated with the second virtual plane) corresponds to the third host machine. Assume that the GPU contained in node 1-1-1 wants to send a data packet to another GPU on node 1-N-2. In this implementation, the T0 switch, i.e., switch 717 (coupled to the first and third host machines), can perform processing of data packets that will be transmitted by the third host machine (associated with the second virtual plane), i.e., node 1-1-2, on the second virtual plane. Note that in this case, the conversion also occurs between host machines coupled to the same T0 layer switch. Furthermore, the third host machine can send the first data packet to the second host machine on the second virtual plane. Note that in the implementation discussed above, it is assumed that a particular client (i.e., whose data packet is being transmitted) is assigned to a host machine coupled to the same T0 layer switch.
[0191] Example 2: In this example, virtual planes are interconnected with each other via exchange blocks. Figure 14 An exemplary network architecture 1400, illustrating a second embodiment of an interconnected global virtual plane, is depicted. (As shown...) Figure 14 As shown, the network architecture comprises two parts: a first part (labeled 1405) and a second part (labeled 1410). Note that the first part 1405 of the network architecture 1400 is similar to that in the previous reference. Figure 7 The network architecture described is 700. Therefore, for the sake of brevity, this article will not repeat the description of Part 1 1405. Instead, we describe Part 2 1410 of network architecture 1400 here.
[0192] The second part of network architecture 1400 is referred to herein as switching block 1410. Switching block 1410 includes multiple switches arranged in a hierarchical structure, which includes first-layer switches (T0 layer), such as switches labeled 1411, 1412, etc., and second-layer switches (T1 layer), such as switches labeled 1415, 1417, etc. Switching block 1410 is coupled to the first part 1405 of network architecture via upper-layer switches, namely T2 layer switches (e.g., switches in blocks 701 and 731, respectively). Therefore, network architecture 1400 as a whole can be conceived as a network architecture including multiple switches, comprising a first group of switches (included in the first part 1405) and a second group of switches (included in switching block 1410).
[0193] According to one embodiment, the architecture of switching block 1210 (i.e., the connectivity between the two layer switches (i.e., the T0 and T1 layer switches) in the switching block, and the interconnection between switching block 1410 and the first portion 1405 of network architecture 1400) is similar to that of the previously referenced Figure 10 The architecture of the first switch block 1005 is described. For example, Figure 14 Switches 1411 and 1412 (T0 layer switches) correspond to Figure 10 The switches 1017 and 1027, and Figure 14 Switches 1415 and 1417 (T1 layer switches) correspond to Figure 10 The switches are 715A and 715D.
[0194] In one implementation, only the switches in the first-layer switches included in the switching block 1410 of network architecture 1400 (e.g., switches 1411, 1412) are configured to transform data packets transmitted on a first virtual plane into data packets to be transmitted on a second virtual plane, and vice versa. The mechanism by which the switches included in the first-layer switches in the switching block perform packet transformation can be as previously referenced. Figure 10 The described example is one of the overlay routing technologies for VxLAN routing, or the switching mechanism may correspond to, as previously referenced. Figure 12 The virtual plane bridging technology described.
[0195] According to one embodiment, multiple host machines (706, 707, etc.) included in a first portion 1405 of network architecture 1400 are communicatively coupled to a switching block via a first set of switches (i.e., T0, T1, and T2 layer switches) included in the first portion 1405. It is noteworthy that, in one implementation, the switches in the first layer switches (T0) included in the first portion 1405 of network architecture are not configured to transform data packets transmitted on a first virtual plane for transmission on a second virtual plane, and vice versa. Therefore, if it is desired to transmit a data packet originating from a first GPU on a first host machine (associated with the first virtual plane) to a second GPU on a second host machine (associated with the second virtual plane), then the data packet must be transmitted to the second GPU on the second host machine via the switching block (i.e., for transforming the data packet for transmission on the second virtual plane). It should be understood that only the components included in switching block 1410 are switches arranged in a two-layer manner (1411, 1412, 1415, 1417, etc.). Specifically, the switching block does not include any host machines. Thus, the entire computing / processing power of the switches contained in the first-layer switches of the switching block is dedicated to the sole purpose of switching data transmission between different virtual planes.
[0196] Now go to Figure 15 An exemplary flowchart 1500, illustrating steps performed when transmitting data packets using a second embodiment of an interconnected global virtual plane according to certain embodiments, is shown. Specifically, Figure 15 The illustration depicts a scenario according to certain embodiments in use. Figure 14 An exemplary flowchart 1500 shows the steps performed when a network infrastructure transmits data packets. Figure 15 The processing described herein can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of a corresponding system. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 15 The method 1500 presented and described below is intended to be illustrative and non-limiting. Although Figure 15 Various processing steps that occur in a specific sequence or order are described, but this is not intended to be limiting. In some alternative embodiments, these steps may be performed in a different order, or some steps may be performed in parallel. It should be recognized that the control plane of the network infrastructure can be configured to perform... Figure 15 The steps described in the text.
[0197] like Figure 15 As shown, the process begins in step 1505, where a first subset of the plurality of ports is associated with a first virtual plane. The first virtual plane identifies a first set of resources to be dedicated to transmitting data packets from and to the host machine associated with the first virtual plane. The process then proceeds to step 1510, where a second subset of the plurality of ports is associated with a second virtual plane. Note that the second virtual plane is different from the first virtual plane.
[0198] Subsequently, in step 1315, a first host machine among the plurality of host machines is associated with a first virtual plane, and a second host machine among the plurality of host machines is associated with a second virtual plane. The first host machine can be directly coupled to a first group of switches (i.e., included in...). Figure 14 In the first part 1405 of the network architecture 1400, the first switch (e.g., a T0 layer switch) is directly coupled to the second switch (e.g., another T0 layer switch) in the first group of switches.
[0199] At step 1520, a switching block is provided in the network architecture. For example, a switching block can be provided as follows: Figure 14 The switching block 1410 shown is coupled to multiple host machines via a first set of switches. In step 1520, for a data packet (e.g., a first data packet) originating from a first GPU on a first host machine (associated with a first virtual plane) and destined for a second GPU on a second host machine (associated with a second virtual plane), the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine via the switching block (i.e., for transforming / converting the first data packet to make it compatible with transmission on the second virtual plane).
[0200] Example 3: In this example, virtual planes are interconnected via coupling of switches in the T2 layer of a switch hierarchy (i.e., switches contained in the upper layer). The interconnection of virtual planes can be described via two phases: (a) a first phase, referred to herein as the virtual partitioning phase, which involves virtual partitioning of the switches contained in the first layer; and (b) a coupling phase, also referred herein as the T2 cross-plane interconnection phase.
[0201] Figure 16 The illustration 1600 shows a schematic diagram depicting a virtual partition of a Layer 1 (T0) switch according to one embodiment. Figure 16 A T0 layer switch 1617 is depicted, communicatively coupled to multiple T1 layer switches (e.g., M switches labeled 1615A, 1615B, 1615C, and 1615D). Furthermore, Figure 16 Two host machines, 1606 (node 1-1-1-) and 1607 (node 1-1-2), are depicted respectively. Similar to previous references. Figure 7 The described framework associates each host machine in a group of host machines coupled to the same T0 layer switch with a unique virtual plane. For example, as Figure 16 As shown, host machine 1606 is associated with a first virtual plane, and host machine 1607 is associated with a second virtual plane (different from the first virtual plane). Each host machine includes one or more GPUs, such as GPU 1606A included in host machine 1606. Note that... Figure 16 Only two virtual planes are depicted. However, it should be noted that this is not intended to limit the scope of this disclosure and is for illustrative purposes only. Rather, the Layer 1 switch (T0) can be virtually partitioned into any number of parts, i.e., corresponding to the number of virtual planes.
[0202] like Figure 16 As shown, two Virtual Tunnel Endpoints (VTEPs) are created in the Layer 1 (T0) switch 1617. A first VTEP 1617A is created for the first virtual plane, and a second VTEP 1617B is created for the second virtual plane. The GPU of a specific host machine is coupled to the corresponding VTEP. For example, as... Figure 16 As shown, the GPU of host machine 1606, for example, GPU 1606A, is coupled to the first VTEP 1617A, while the GPU contained in the second host machine 1607 is coupled to the second VTEP 1617B.
[0203] According to one embodiment, the first VTEP 1617A, i.e., the VTEP associated with the first virtual plane, is advertised only to other switches (e.g., switches 1615A and 1615B) in the network architecture associated with the first virtual plane, via links that are associated with the first virtual plane and interconnect the first VTEP to these switches. Examples of such links for the first virtual plane are shown in […]. Figure 16 The second VTEP 1617B, i.e., the VTEP associated with the second virtual plane, is advertised only to other switches in the network architecture associated with the second virtual plane (e.g., switches 1615C and 1615D) via links that are associated with the second virtual plane and interconnect the second VTEP to those switches. Examples of such links for the second virtual plane are shown in [reference to a specific example]. Figure 16 The lines are depicted as dashed lines. In this way, switches in the network architecture learn which specific VTEP to use in order to reach a specific interface on a host machine. Furthermore, it is worth noting that each VTEP (i.e., the first VTEP 1617A and the second VTEP 1617B) is associated with a unique Autonomous System Number (ASN) (e.g., an identifier).
[0204] Figure 17A An exemplary network architecture 1700 infrastructure is depicted, illustrating a third embodiment of an interconnected global virtual plane. Figure 17A The network architecture infrastructure of the 1700 is similar to (previous reference) Figure 7 The network architecture infrastructure 700 described includes multiple host machines communicatively coupled to each other via multiple switches. These switches are arranged in a hierarchical structure comprising Layer 1 switches, Layer 2 switches, and Layer 3 switches. The host machines are directly coupled to switches included in the Layer 1 switches, and the Layer 2 switches communicatively couple the Layer 1 switches to the Layer 3 switches. Figure 7 Compared to the network architecture of 700) Figure 17A The network architecture 1700 is different in that it includes multiple interconnections in the upper-layer switch (i.e., the T2-layer switch).
[0205] As Figure 17A shown, the upper layer 1701 of the switch includes multiple switches, for example, M switches labeled 713A, 713B, 713C, and 713D. Similar to Figure 7 this, a first subset of these switches, for example, switches 1713A - switch B 713B (depicted as being surrounded by ellipse 750B), is associated with a first virtual plane, while a second subset, for example, switch B+1 (713C) - switch M (713D), is associated with a second virtual plane. In one embodiment, the interconnection between the first virtual plane and the second virtual plane is achieved through switches in the upper layer of the interconnected switches.
[0206] For example, referring to the upper layer 1 (701) of the switch and assuming a value of B = M / 2, it results in half of the switches in the upper layer 1 (switches 1 to switch M / 2) being associated with the first virtual plane, and the other half (switches M / 2 +1 to switch M) being associated with the second virtual plane. The interconnection is obtained by directly coupling a pair of switches, where one switch is selected from the subset of switches associated with the first virtual plane, and the other switch is selected from the subset of switches associated with the second virtual plane. For example, as Figure 17A shown, switch 1 (713A) is directly coupled to switch B+1 (713C), as shown by link 1710, while switch B (713B) is directly coupled to switch M (713D), as shown by link 1720. Similar interconnections can be performed between the upper-layer switches included in block 731 (shown by links 1730 and 1740 respectively).
[0207] In one embodiment, a BGP session can be instantiated on each of the links 1710 - 1740 such that network address information (e.g., IP address) of the switches associated with the second virtual plane and information about the VTEP associated with the second virtual plane can be exchanged with the switches associated with the first virtual plane (included in the upper layer), and vice versa. In this way, the switches associated with each virtual plane (e.g., the first virtual plane) obtain / learn network information about the switches associated with another virtual plane (e.g., the second virtual plane). According to one embodiment, each link / cable (e.g., links 1710, 1720, 1730, and 1740) can use multiple interfaces, for example, sub-lines, where the number of interfaces corresponds to the total number of blocks below the upper-layer switch, for example, as Figure 17A The diagram shows K blocks. This approach provides non-blocking performance for traffic across planes in all blocks, thus maintaining good performance for cross-plane traffic while still ensuring workload isolation.
[0208] Therefore, according to the third embodiment of the interconnected virtual planes (e.g., a first virtual plane and a second virtual plane), it should be recognized that for a data packet originating from a first GPU (e.g., GPU 706A) on a first host machine (e.g., host machine 1-1-1 (706)) and destined for a second GPU (e.g., GPU 726A) on a second host machine (e.g., host machine 1-N-2 (726)), the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using ports in a first subset of ports (associated with the first virtual plane) and a second subset of ports (associated with the second virtual plane). Essentially, in the transmission of such data packets (i.e., data packets traversing different virtual planes), according to the third embodiment, at least one switch (e.g., a port of the switch) in the upper layer of the switch associated with the first virtual plane, and at least one switch (e.g., a port of the switch) in the upper layer of the switch associated with the second virtual plane, are utilized. For the example considered above, i.e., the transmission of data packets from GPU 706A to GPU 726A, the path taken by the data packet is in Figure 17A The diagram describes the process by numbering the corresponding links (i.e., link A to link G). Furthermore, it should be recognized that the transmission of information (e.g., data packets) can be done using methods previously described... Figure 12 The bridging mechanism described herein transitions from transmission on the first virtual plane to transmission on the second virtual plane (and vice versa).
[0209] Figure 17B An exemplary flowchart illustrating steps performed when transmitting data packets using a third embodiment of an interconnected global virtual plane, according to certain embodiments, is shown.
[0210] Figure 17B The processing described herein can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of a corresponding system. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 17B The method 1750 presented and described below is intended to be illustrative and non-limiting. Although Figure 17B Various processing steps that occur in a specific sequence or order are described, but this is not intended to be limiting. In some alternative embodiments, these steps may be performed in a different order, or some steps may be performed in parallel. It should be recognized that the control plane of the network infrastructure can be configured to perform... Figure 17B The steps described in the text.
[0211] Figure 17B The steps described can be implemented in a network environment comprising multiple host machines, each communicatively coupled to other host machines via a network architecture. Each host machine includes one or more GPUs. In one implementation, the network architecture may include multiple switches (including multiple ports) arranged in a hierarchical structure such as a Layer 2 CLOS network or a Layer 3 CLOS network. For illustration, consider a Layer 3 CLOS network; the network architecture may include Layer 1 switches, Layer 2 switches, and Layer 3 switches. In one implementation, multiple host machines are directly coupled to switches included in a Layer 1 switch. A Layer 2 switch then communicatively couples the Layer 1 switch to a Layer 3 switch. Furthermore, switches included in a Layer 3 switch (i.e., Layer T2) are coupled in pairs, as previously referenced. Figure 17A Described.
[0212] In this setting, and as Figure 17B As shown, the process begins in step 1755, where a first subset of ports from a plurality of ports is associated with a first virtual plane. The first virtual plane identifies a first set of resources to be dedicated to transmitting packets from and to the host machine associated with the first virtual plane. The process then proceeds to step 1760, where a second subset of ports from a plurality of ports is associated with a second virtual plane. Note that the second virtual plane is different from the first virtual plane.
[0213] Subsequently, in step 1765, the first host machine is associated with the first virtual plane, and the second host machine is associated with the second virtual plane. In step 1770, for a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine, the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using ports in the first subset of ports and ports in the second subset of ports. Specifically, for transmitting the data packet, at least one port of the first switch included in the T2 layer (and associated with the first plane) and at least one port of the second switch included in the T2 layer (and associated with the second virtual plane) are utilized, wherein the first switch is directly coupled to the second switch.
[0214] Example cloud infrastructure implementation
[0215] As noted above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, cloud providers can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, IaaS providers can also provision various services to accompany these infrastructure components (example services include billing software, monitoring software, logging software, load balancing software, and clustering software, etc.). Therefore, because these services may be policy-driven, IaaS users can implement policies to drive load balancing to maintain application availability and performance.
[0216] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the internet and can use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log in to the IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create buckets for workloads and backups, and even install enterprise software into that VM. The customer can then use the provider's services to perform various functions, including balancing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0217] In most cases, cloud computing models will require the involvement of cloud providers. Cloud providers can, but are not necessarily, third-party providers specializing in (e.g., provisioning, renting, selling) IaaS services. Entities may also choose to deploy private clouds, thus becoming their own infrastructure service providers.
[0218] In some examples, IaaS deployment is the process of placing a new application or a new version of an application onto a prepared application server, etc. It may also include the processing of server preparation (e.g., installation libraries, daemons, etc.). This is typically managed by the cloud provider, below the hypervisor layer (e.g., servers, storage devices, network hardware, and virtualization). Therefore, the customer can be responsible for processing (OS), middleware, and / or application deployment (e.g., on self-service virtual machines, etc., which can be started on demand).
[0219] In some examples, IaaS provisioning can refer to acquiring computers or virtual hosts for use, or even installing necessary libraries or services on them. In most cases, deployment does not include provisioning, and provisioning may need to be performed first.
[0220] In some cases, IaaS provisioning presents two distinct challenges. First, there's the initial challenge of provisioning the initial infrastructure set before anything is operational. Second, once everything is provisioned, there's the challenge of evolving the existing infrastructure (e.g., adding new services, changing services, removing services, etc.). In some cases, both challenges can be addressed by enabling configuration that declaratively defines the infrastructure. In other words, the infrastructure (e.g., which components are needed and how they interact) can be defined by one or more configuration files. Therefore, the overall topology of the infrastructure (e.g., which resources depend on which resources and how they work together) can be described declaratively. In some cases, once the topology is defined, workflows for creating and / or managing the different components described in the configuration files can be generated.
[0221] In some examples, the infrastructure can have many interconnected components. For example, there may be one or more Virtual Private Clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as the core network. In some examples, one or more inbound / outbound traffic group rules may also be provided to define how inbound / outbound traffic to the network and one or more virtual machines (VMs). Other infrastructure elements, such as load balancers, databases, etc., may also be provided. The infrastructure can evolve incrementally as more and / or more infrastructure elements are expected and added.
[0222] In some cases, continuous deployment techniques can be used to enable the deployment of infrastructure code across various virtual computing environments. Furthermore, the described techniques enable infrastructure management within these environments. In some examples, service teams may write code that they expect to deploy to one or more, but often many, different production environments (e.g., across various geographical locations, sometimes spanning the entire world). However, in some examples, the infrastructure on which the code will be deployed must first be set up. In some cases, provisioning can be done manually, resources can be provisioned using provisioning tools, and / or once the infrastructure is provisioned, the code can be deployed using deployment tools.
[0223] Figure 18 This is a block diagram 1800 illustrating an example pattern of an IaaS architecture according to at least one embodiment. Service operator 1802 may communicatively couple to secure host lease 1804, which may include a virtual cloud network (VCN) 1806 and a secure host subnet 1808. In some examples, service operator 1802 may use one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, cellular phone, iPad®, computing tablet, personal digital assistant (PDA)) or wearable devices (e.g., Google Glass® head-mounted display), running software (such as Microsoft Windows Mobile®) and / or various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, etc.), and supporting the Internet, email, short message service (SMS), Blackberry®, or other communication protocols. Alternatively, client computing devices may be general-purpose personal computers, including, for example, personal computers and / or laptops running various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems. Client computing devices can be workstation computers running a variety of commercially available UNIX® or UNIX-like operating systems, including but not limited to any of the various GNU / Linux operating systems (such as, for example, Google Chrome OS). Alternatively or additionally, client computing devices can be any other electronic device, such as thin client computers, internet-enabled gaming systems (e.g., Microsoft Xbox game consoles with or without Kinect® gesture input devices), and / or personal messaging devices capable of communicating over a network that can access VCN 1806 and / or the internet.
[0224] VCN 1806 may include a local peering gateway (LPG) 1810, which may be communicatively coupled to a secure shell (SSH) VCN 1812 via LPG 1810 included in SSH VCN 1812. SSH VCN 1812 may include an SSH subnet 1814, and SSH VCN 1812 may be communicatively coupled to a control plane VCN 1816 via LPG 1810 included in control plane VCN 1816. Furthermore, SSH VCN 1812 may be communicatively coupled to a data plane VCN 1818 via LPG 1810. Control plane VCN 1816 and data plane VCN 1818 may be contained within a service lease 1819 that may be owned and / or operated by an IaaS provider.
[0225] The control plane VCN 1816 may include a control plane demilitarized zone (DMZ) layer 1820 that acts as a peripheral network (e.g., part of a corporate network between a corporate intranet and an external network). DMZ-based servers can assume limited liability and help control vulnerabilities. Furthermore, the DMZ layer 1820 may include one or more load balancer (LB) subnets 1822, a control plane application layer 1824 that may include one or more application subnets 1826, and a control plane data layer 1828 that may include one or more database (DB) subnets 1830 (e.g., one or more front-end DB subnets and / or one or more back-end DB subnets). One or more LB subnets 1822 contained in the control plane DMZ layer 1820 may be communicatively coupled to one or more application subnets 1826 contained in the control plane application layer 1824 and an Internet gateway 1834 that may be contained in the control plane VCN 1816. The application subnets 1826 may be communicatively coupled to one or more DB subnets 1830 contained in the control plane data layer 1828, as well as a service gateway 1836 and a Network Address Translation (NAT) gateway 1838. The control plane VCN 1816 may include the service gateway 1836 and the NAT gateway 1838.
[0226] The control plane VCN 1816 may include a data plane mirror application layer 1840, which may include one or more application subnets 1826. The one or more application subnets 1826 included in the data plane mirror application layer 1840 may include a virtual network interface controller (VNIC) 1842 capable of executing a compute instance 1844. The compute instance 1844 may communicatively couple the one or more application subnets 1826 of the data plane mirror application layer 1840 to the one or more application subnets 1826 that may be included in the data plane application layer 1846.
[0227] Data plane VCN 1818 may include data plane application layer 1846, data plane DMZ layer 1848, and data plane data layer 1850. Data plane DMZ layer 1848 may include one or more LB subnets 1822 communicatively coupled to one or more application subnets 1826 of data plane application layer 1846 and Internet gateway 1834 of data plane VCN 1818. One or more application subnets 1826 may be communicatively coupled to service gateway 1836 and NAT gateway 1838 of data plane VCN 1818. Data plane data layer 1850 may also include one or more DB subnets 1830 communicatively coupled to one or more application subnets 1826 of data plane application layer 1846.
[0228] The Internet gateway 1834 of the control plane VCN 1816 and data plane VCN 1818 can be communicatively coupled to the metadata management service 1852, which can be communicatively coupled to the public Internet 1854. The public Internet 1854 can be communicatively coupled to the NAT gateway 1838 of the control plane VCN 1816 and data plane VCN 1818. The service gateway 1836 of the control plane VCN 1816 and data plane VCN 1818 can be communicatively coupled to the cloud service 1856.
[0229] In some examples, the service gateway 1836 of the control plane VCN 1816 or data plane VCN 1818 can make application programming interface (API) calls to the cloud service 1856 without traversing the public internet 1854. API calls from the service gateway 1836 to the cloud service 1856 can be unidirectional: the service gateway 1836 can make API calls to the cloud service 1856, and the cloud service 1856 can send requested data to the service gateway 1836. However, the cloud service 1856 may not initiate API calls to the service gateway 1836.
[0230] In some examples, secure host lease 1804 can be directly connected to service lease 1819, which would otherwise be isolated. Secure host subnet 1808 can communicate with SSH subnet 1814 via LPG 1810, which enables bidirectional communication between otherwise isolated systems. Connecting secure host subnet 1808 to SSH subnet 1814 allows secure host subnet 1808 to access other entities within service lease 1819.
[0231] Control plane VCN 1816 allows users of service lease 1819 to configure or otherwise provision desired resources. Desired resources provisioned in control plane VCN 1816 can be deployed or otherwise used in data plane VCN 1818. In some examples, control plane VCN 1816 can be isolated from data plane VCN 1818, and the data plane mirror application layer 1840 of control plane VCN 1816 can communicate with the data plane application layer 1846 of data plane VCN 1818 via VNIC 1842, which can be included in both the data plane mirror application layer 1840 and the data plane application layer 1846.
[0232] In some examples, users or clients of the system can make requests, such as create, read, update, or delete (CRUD) operations, via the public internet 1854, which can transmit requests to the metadata management service 1852. The metadata management service 1852 can transmit requests to the control plane VCN 1816 via internet gateway 1834. Requests can be received by one or more LB subnets 1822 contained in the control plane DMZ layer 1820. The LB subnets 1822 can determine that the request is valid, and in response to this determination, they can transmit the request to one or more application subnets 1826 contained in the control plane application layer 1824. If the request is validated and requires a call to the public internet 1854, the call to the public internet 1854 can be transmitted to a NAT gateway 1838 that can make calls to the public internet 1854. The request may expect stored metadata to be stored in one or more DB subnets 1830.
[0233] In some examples, the data plane mirroring application layer 1840 can facilitate direct communication between the control plane VCN 1816 and the data plane VCN 1818. For example, it might be desirable to apply configuration changes, updates, or other appropriate modifications to resources contained in the data plane VCN 1818. Through VNIC 1842, the control plane VCN 1816 can communicate directly with the resources contained in the data plane VCN 1818, and thus can perform configuration changes, updates, or other appropriate modifications.
[0234] In some embodiments, the control plane VCN 1816 and data plane VCN 1818 may be included in service lease 1819. In this case, the system's users or customers may not own or operate the control plane VCN 1816 or data plane VCN 1818. Alternatively, the IaaS provider may own or operate both the control plane VCN 1816 and data plane VCN 1818, and both planes may be included in service lease 1819. This embodiment can enable the isolation of networks that might prevent users or customers from interacting with resources of other users or customers. Furthermore, this embodiment can allow users or customers of the system to privately store databases without relying on the public Internet 1854, which may not have the desired level of threat protection.
[0235] In other embodiments, one or more LB subnets 1822 included in the control plane VCN 1816 may be configured to receive signals from the service gateway 1836. In this embodiment, the control plane VCN 1816 and the data plane VCN 1818 may be configured to be invoked by the IaaS provider's customers without invoking the public internet 1854. The IaaS provider's customers may expect this embodiment because the database(s) used by the customer can be controlled by the IaaS provider and can be stored on a service lease 1819, which may be isolated from the public internet 1854.
[0236] Figure 19 This is a block diagram 1900 illustrating another example pattern of an IaaS architecture according to at least one embodiment. Service operator 1902 (e.g., Figure 18 Service providers (1802) can communicatively couple to secure host leases (1904, e.g., Figure 18 Secure hosting lease 1804), the secure hosting lease 1904 may include a virtual cloud network (VCN) 1906 (e.g., Figure 18 VCN 1806) and Secure Host Subnet 1908 (e.g., Figure 18 The secure host subnet 1808). VCN 1906 may include a local peering gateway (LPG) 1910 (e.g., Figure 18 The LPG 1810, which can be communicatively coupled to the Secure Shell (SSH) VCN 1912 (e.g., via the LPG 1810 included in the SSH VCN 1912) Figure 18 SSH VCN 1812). SSH VCN 1912 can include SSH subnet 1914 (e.g., Figure 18 SSH subnet 1814), and SSH VCN 1912 can be communicatively coupled to control plane VCN 1916 via LPG 1910 contained in control plane VCN 1916 (e.g., Figure 18 Control plane VCN 1816). Control plane VCN 1916 may be included in service lease 1919 (e.g., Figure 18 In the service lease 1819), and the data plane VCN 1918 (e.g., Figure 18 The data plane (VCN 1818) may be included in a customer lease 1921 that may be owned or operated by the system's users or customers.
[0237] The control plane VCN 1916 may include one or more LB subnets 1922 (e.g., Figure 18 The control plane DMZ layer 1920 of (one or more) LB subnets 1822 (e.g., Figure 18 The control plane DMZ layer 1820 can contain one or more application subnets 1926 (e.g., Figure 18 The control plane application layer 1924 of (one or more) application subnets 1826 (e.g., Figure 18 The control plane application layer 1824) may contain one or more database (DB) subnets 1930 (e.g., similar to...). Figure 18 The control plane data layer 1928 of (one or more) DB subnets 1830 (e.g., Figure 18 The control plane data layer 1828). One or more LB subnets 1922 contained in the control plane DMZ layer 1920 can be communicatively coupled to one or more application subnets 1926 contained in the control plane application layer 1924 and an Internet gateway 1934 that can be contained in the control plane VCN 1916 (e.g., Figure 18 Internet gateway 1834), and application subnet(s) 1926 can communicatively couple to DB subnet(s) 1930 contained in control plane data layer 1928 and service gateway 1936 (e.g., Figure 18 Service gateway 1836) and Network Address Translation (NAT) gateway 1938 (e.g., Figure 18 (NAT gateway 1838). The control plane VCN 1916 may include the service gateway 1936 and the NAT gateway 1938.
[0238] The control plane VCN 1916 may include a data plane mirror of the application layer 1940, which may contain one or more application subnets 1926 (e.g., Figure 18 The data plane mirror application layer 1840). One or more application subnets 1926 contained in the data plane mirror application layer 1940 may include computational instances 1944 (e.g., similar to...). Figure 18 The virtual network interface controller (VNIC) 1942 (e.g., the VNIC of 1842) of the computing instance 1844. The computing instance 1944 may facilitate the mirroring of the application subnet(s) 1926 of the application layer 1940 in the data plane and may be included in the application layer 1946 in the data plane (e.g., Figure 18 Communication between one or more application subnets 1926 in the data plane application layer 1846 via VNIC 1942 contained in the data plane mirror application layer 1940 and VNIC 1942 contained in the data plane application layer 1946.
[0239] The Internet gateway 1934, included in the control plane VCN 1916, can be communicatively coupled to the metadata management service 1952 (e.g., Figure 18 Metadata management service 1852), which can communicatively couple to the public Internet 1954 (e.g., Figure 18 The public internet 1954 can communicatively couple to a NAT gateway 1938 contained in a control plane VCN 1916. A service gateway 1936 contained in a control plane VCN 1916 can communicatively couple to a cloud service 1956 (e.g., ...). Figure 18 Cloud services (1856).
[0240] In some examples, data plane VCN 1918 may be included in customer lease 1921. In this case, the IaaS provider may provide control plane VCN 1916 for each customer, and the IaaS provider may set up a unique compute instance 1944 for each customer, included in service lease 1919. Each compute instance 1944 may allow communication between control plane VCN 1916 included in service lease 1919 and data plane VCN 1918 included in customer lease 1921. Compute instance 1944 may allow resources provisioned in control plane VCN 1916 included in service lease 1919 to be deployed or otherwise used in data plane VCN 1918 included in customer lease 1921.
[0241] In other examples, an IaaS provider's customer may have a database residing in customer lease 1921. In this example, control plane VCN 1916 may include data plane mirror application layer 1940, which may include one or more application subnets 1926. Data plane mirror application layer 1940 may reside in data plane VCN 1918, but may not reside in data plane VCN 1918. That is, data plane mirror application layer 1940 may have access to customer lease 1921, but may not reside in data plane VCN 1918 or be owned or operated by an IaaS provider's customer. Data plane mirror application layer 1940 may be configured to invoke data plane VCN 1918, but may not be configured to invoke any entity contained in control plane VCN 1916. Customers may expect to deploy or otherwise use resources provided in the control plane VCN 1916 in the data plane VCN 1918, and the data plane mirroring application layer 1940 can facilitate customers' desired deployments or other uses of resources.
[0242] In some embodiments, an IaaS provider's customer can apply filters to data plane VCN 1918. In this embodiment, the customer can determine what data plane VCN 1918 can access, and the customer can restrict access from data plane VCN 1918 to the public Internet 1954. The IaaS provider may not be able to apply filters or otherwise control data plane VCN 1918's access to any external networks or databases. Applying filters and controls to data plane VCN 1918, which is included in customer lease 1921, can help isolate data plane VCN 1918 from other customers and the public Internet 1954.
[0243] In some embodiments, cloud service 1956 may be invoked by service gateway 1936 to access services that may not exist on public internet 1954, control plane VCN 1916, or data plane VCN 1918. The connection between cloud service 1956 and control plane VCN 1916 or data plane VCN 1918 may not be real-time or continuous. Cloud service 1956 may reside on different networks owned or operated by an IaaS provider. Cloud service 1956 may be configured to receive calls from service gateway 1936 and may be configured not to receive calls from public internet 1954. Some cloud services 1956 may be isolated from other cloud services 1956, and control plane VCN 1916 may be isolated from cloud services 1956 that may not be in the same region as control plane VCN 1916. For example, control plane VCN 1916 may be located in "Region 1," and cloud service "Deployment 18" may be located in both "Region 1" and "Region 2." If the service gateway 1936, contained in the control plane VCN 1916 located in region 1, makes a call to deployment 18, then that call can be transmitted to deployment 18 in region 1. In this example, the control plane VCN 1916 or deployment 18 in region 1 may not be communicatively coupled to or otherwise communicate with deployment 18 in region 2.
[0244] Figure 20 This is a block diagram 2000 illustrating another example pattern of an IaaS architecture according to at least one embodiment. Service operator 2002 (e.g., Figure 18 Service providers (1802) can communicatively couple to secure host rental (2004) (e.g., Figure 18 Secure hosting lease 1804), the secure hosting lease 2004 may include Virtual Cloud Network (VCN) 2006 (e.g., Figure 18 VCN 1806) and Secure Host Subnet 2008 (e.g., Figure 18 The secure host subnet 1808). VCN 2006 can include LPG 2010 (e.g., Figure 18 The LPG 1810), which can be communicatively coupled to SSH VCN 2012 via the LPG 2010 included in SSH VCN 2012 (e.g., Figure 18 SSH VCN 1812). SSH VCN 2012 can include SSH subnets 2014 (e.g., Figure 18 SSH subnet 1814), and SSH VCN 2012 can be communicatively coupled to control plane VCN 2016 via LPG 2010 included in control plane VCN 2016 (e.g., Figure 18 The control plane VCN 1816) and coupled to the data plane VCN 2018 via the LPG 2010 contained in the data plane VCN 2018 (e.g., Figure 18 Data plane 1818). Control plane VCN 2016 and data plane VCN 2018 can be included in service lease 2019 (e.g., Figure 18 In the service rental (1819).
[0245] The control plane VCN 2016 may include one or more load balancer (LB) subnets 2022 (e.g., Figure 18 The control plane DMZ layer of (one or more) LB subnets 1822) 2020 (e.g., Figure 18 The control plane DMZ layer 1820 may include one or more application subnets 2026 (e.g., similar to...). Figure 18 The control plane application layer 2024 of (one or more) application subnets 1826 (e.g., Figure 18 The control plane application layer 1824), which may include (one or more) DB subnets 2030, and the control plane data layer 2028 (e.g., Figure 18 The control plane data layer 1828). One or more LB subnets 2022 contained in the control plane DMZ layer 2020 can be communicatively coupled to one or more application subnets 2026 contained in the control plane application layer 2024 and an Internet gateway 2034 that can be contained in the control plane VCN 2016 (e.g., Figure 18 Internet gateway 1834), and application subnet(s) 2026 can communicatively couple to DB subnet(s) 2030 contained in control plane data layer 2028 and service gateway 2036 (e.g., Figure 18 The service gateway) and Network Address Translation (NAT) gateway 2038 (e.g., Figure 18 (NAT gateway 1838). The control plane VCN 2016 may include service gateway 2036 and NAT gateway 2038.
[0246] Data plane VCN 2018 may include data plane application layer 2046 (e.g., Figure 18 Data plane application layer 1846), data plane DMZ layer 2048 (e.g., Figure 18 Data plane DMZ layer 1848), and data plane data layer 2050 (e.g., Figure 18 The data plane data layer 1850). The data plane DMZ layer 2048 may include one or more trusted application subnets 2060 and one or more untrusted application subnets 2062 that are communicatively coupled to the data plane application layer 2046, and one or more LB subnets 2022 that are included in the Internet gateway 2034 in the data plane VCN 2018. One or more trusted application subnets 2060 may be communicatively coupled to the service gateway 2036, the NAT gateway 2038, and the DB subnets 2030 included in the data plane VCN 2018. One or more untrusted application subnets 2062 may be communicatively coupled to the service gateway 2036 and the DB subnets 2030 included in the data plane VCN 2018. The data plane data layer 2050 may include one or more DB subnets 2030 that can be communicatively coupled to the service gateway 2036 contained in the data plane VCN 2018.
[0247] One or more untrusted application subnets 2062 may include one or more primary VNICs 2064(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 2066(1)-(N). Each tenant VM 2066(1)-(N) may be communicatively coupled to a corresponding application subnet 2067(1)-(N) that may be contained in a corresponding container egress VCN 2068(1)-(N), which may be contained in a corresponding customer lease 2070(1)-(N). A corresponding secondary VNIC 2072(1)-(N) may facilitate communication between one or more untrusted application subnets 2062 contained in the data plane VCN 2018 and the application subnets contained in the container egress VCN 2068(1)-(N). Each container egress VCN 2068(1)-(N) may include a NAT gateway 2038, which may communicatively couple to the public Internet 2054 (e.g., Figure 18 The public internet (1854).
[0248] The Internet gateway 2034, contained in the control plane VCN 2016 and the data plane VCN 2018, can communicatively couple to the metadata management service 2052 (e.g., Figure 18 A metadata management system 1852 is provided, which can communicatively couple to the public internet 2054. The public internet 2054 can communicatively couple to a NAT gateway 2038 contained in a control plane VCN 2016 and a data plane VCN 2018. A service gateway 2036 contained in a control plane VCN 2016 and a data plane VCN 2018 can communicatively couple to a cloud service 2056.
[0249] In some embodiments, the data plane VCN 2018 can be integrated with customer leases 2070. Such integration may be useful or desired by the IaaS provider's customers in certain situations, such as when support might be expected during code execution. Customers may provide code that could be destructive, might communicate with other customer resources, or might otherwise cause undesirable effects. In response, the IaaS provider can determine whether to run the code provided by the customer.
[0250] In some examples, an IaaS provider's customer can grant temporary network access to the IaaS provider and request functionality attached to the data plane application layer 2046. The code running this functionality can execute in VMs 2066(1)-(N), and this code may not be configured to run anywhere else on the data plane VCN 2018. Each VM 2066(1)-(N) can be connected to a customer lease 2070. The corresponding container 2071(1)-(N) contained in VM 2066(1)-(N) can be configured to run the code. In this case, there can be dual isolation (e.g., container 2071(1)-(N) runs the code, where container 2071(1)-(N) may be contained in at least one or more untrusted application subnets 2062 containing VM 2066(1)-(N)), which can help prevent incorrect or otherwise unintended code from corrupting the IaaS provider's network or the networks of different customers. Containers 2071(1)-(N) may be communicatively coupled to customer lease 2070 and may be configured to transmit or receive data from customer lease 2070. Containers 2071(1)-(N) may not be configured to transmit or receive data from any other entity in the data plane VCN 2018. After the code execution is complete, the IaaS provider may terminate or otherwise dispose of containers 2071(1)-(N).
[0251] In some embodiments, one or more trusted application subnets 2060 may run code that can be owned or operated by an IaaS provider. In this embodiment, one or more trusted application subnets 2060 may be communicatively coupled to one or more database subnets 2030 and configured to perform CRUD operations in one or more database subnets 2030. One or more untrusted application subnets 2062 may be communicatively coupled to one or more database subnets 2030, but in this embodiment, one or more untrusted application subnets may be configured to perform read operations in one or more database subnets 2030. Containers 2071(1)-(N) that may be contained in each customer's VM 2066(1)-(N) and may run code from the customer may not be communicatively coupled to one or more database subnets 2030.
[0252] In other embodiments, the control plane VCN 2016 and the data plane VCN 2018 may be coupled without direct communication. In this embodiment, there may be no direct communication between the control plane VCN 2016 and the data plane VCN 2018. However, communication can occur indirectly through at least one method. LPG 2010 may be established by an IaaS provider, which can facilitate communication between the control plane VCN 2016 and the data plane VCN 2018. In another example, the control plane VCN 2016 or the data plane VCN 2018 may invoke cloud service 2056 via service gateway 2036. For example, an invocation of cloud service 2056 from the control plane VCN 2016 may include a request for a service that can communicate with the data plane VCN 2018.
[0253] Figure 21 This is a block diagram 2100 illustrating another example pattern of an IaaS architecture according to at least one embodiment. Service operator 2102 (e.g., Figure 18 The service provider 1802) can communicatively couple to the secure host lease 2104 (e.g., Figure 18 Secure hosting lease 1804), the secure hosting lease 2104 may include a virtual cloud network (VCN) 2106 (e.g., Figure 18 VCN 1806) and Secure Host Subnet 2108 (e.g., Figure 18 The secure host subnet 1808). VCN 2106 may include LPG 2110 (e.g., Figure 18 LPG 1810), the LPG 2110 can be contained in SSH VCN 2112 (e.g., Figure 18 LPG 2110 in SSH VCN 2112 is communicatively coupled to SSH VCN 2112. SSH VCN 2112 may include SSH subnet 2114 (e.g., Figure 18 SSH subnet 1814), and SSH VCN 2112 can be communicatively coupled to control plane VCN 2116 via LPG 2110 included in control plane VCN 2116 (e.g., Figure 18 The control plane VCN 1816) and coupled to the data plane VCN 2118 via the LPG 2110 contained in the data plane VCN 2118 (e.g., Figure 18 Data plane 1818). Control plane VCN 2116 and data plane VCN 2118 may be contained in service lease 2119 (e.g., Figure 18 In the service rental (1819).
[0254] The control plane VCN 2116 may include one or more LB subnets 2122 (e.g., Figure 18 The control plane DMZ layer 2120 of (one or more) LB subnets 1822) (e.g., Figure 18 The control plane DMZ layer 1820 may include (one or more) application subnets 2126 (e.g., Figure 18 The control plane application layer 2124 of (one or more) application subnets 1826 (e.g., Figure 18 The control plane application layer 1824) may include one or more DB subnets 2130 (e.g., Figure 20 The control plane data layer 2128 of (one or more) DB subnets 2030 (e.g., Figure 18 The control plane data layer 1828). One or more LB subnets 2122 contained in the control plane DMZ layer 2120 can be communicatively coupled to one or more application subnets 2126 contained in the control plane application layer 2124 and an Internet gateway 2134 that can be contained in the control plane VCN 2116 (e.g., Figure 18 Internet gateway 1834), and application subnet(s) 2126 can communicatively couple to DB subnet(s) 2130 contained in control plane data layer 2128 and service gateway 2136 (e.g., Figure 18 The service gateway) and the Network Address Translation (NAT) gateway 2138 (e.g., Figure 18 (NAT gateway 1838). The control plane VCN 2116 may include the service gateway 2136 and the NAT gateway 2138.
[0255] Data plane VCN 2118 may include data plane application layer 2146 (e.g., Figure 18 Data plane application layer 1846), data plane DMZ layer 2148 (e.g., Figure 18 Data plane DMZ layer 1848), and data plane data layer 2150 (e.g., Figure 18 The data plane data layer 1850). The data plane DMZ layer 2148 may include one or more trusted application subnets 2160 that can be communicatively coupled to the data plane application layer 2146 (e.g., Figure 20 (one or more) trusted application subnets 2060) and (one or more) untrusted application subnets 2162 (e.g., Figure 20 The data plane VCN 2118 may include one or more untrusted application subnets 2062 and one or more LB subnets 2122 of Internet gateway 2134. One or more trusted application subnets 2160 may communicatively couple to service gateway 2136, NAT gateway 2138, and DB subnets 2130 in data plane VCN 2118. One or more untrusted application subnets 2162 may communicatively couple to service gateway 2136 and DB subnets 2130 in data plane VCN 2118. Data plane VCN 2150 may include one or more DB subnets 2130 that may communicatively couple to service gateway 2136 in data plane VCN 2118.
[0256] One or more untrusted application subnets 2162 may include a primary VNIC 2164(1)-(N) communicatively coupled to tenant virtual machines (VMs) 2166(1)-(N) residing within one or more untrusted application subnets 2162. Each tenant VM 2166(1)-(N) may run code in a corresponding container 2167(1)-(N) and is communicatively coupled to an application subnet 2126 that may be contained in a data plane application layer 2146 contained in a container egress VCN 2168. A corresponding secondary VNIC 2172(1)-(N) may facilitate communication between one or more untrusted application subnets 2162 contained in a data plane VCN 2118 and the application subnets contained in a container egress VCN 2168. The container egress VCN may include a public internet 2154 (e.g., Figure 18 The public internet (1854) uses NAT gateway 2138.
[0257] Internet gateway 2134, contained in control plane VCN 2116 and data plane VCN 2118, can be communicatively coupled to metadata management service 2152 (e.g., Figure 18 The metadata management system 1852 can communicatively couple to the public internet 2154. The public internet 2154 can communicatively couple to a NAT gateway 2138 contained in a control plane VCN 2116 and a data plane VCN 2118. The service gateway 2136 contained in a control plane VCN 2116 and a data plane VCN 2118 can communicatively couple to a cloud service 2156.
[0258] In some examples, Figure 21 The architecture shown in block diagram 2100 can be considered as... Figure 20 This is an exception to the pattern shown in the architecture diagram 2000, and this pattern may be what the IaaS provider's customers would expect if the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected region). The customer can access in real time the corresponding container 2167(1)-(N) contained in each customer's VM 2166(1)-(N). Container 2167(1)-(N) can be configured to invoke the corresponding auxiliary VNIC 2172(1)-(N) contained in one or more application subnets 2126 of the data plane application layer 2146, which may be contained in the container egress VCN 2168. The auxiliary VNIC 2172(1)-(N) can transmit the call to the NAT gateway 2138, which can then transmit the call to the public internet 2154. In this example, containers 2167(1)-(N), which can be accessed by clients in real time, can be isolated from the control plane VCN 2116 and from other entities contained in the data plane VCN 2118. Containers 2167(1)-(N) can also be isolated from resources from other clients.
[0259] In other examples, a client can use container 2167(1)-(N) to invoke cloud service 2156. In this example, the client can run code within container 2167(1)-(N) requesting a service from cloud service 2156. Container 2167(1)-(N) can then transmit the request to auxiliary VNIC 2172(1)-(N), which can then transmit the request to a NAT gateway, which can then transmit the request to the public internet 2154. The public internet 2154 can then transmit the request via internet gateway 2134 to one or more LB subnets 2122 contained in control plane VCN 2116. In response to determining that the request is valid, one or more LB subnets can then transmit the request to one or more application subnets 2126, which can then transmit the request to cloud service 2156 via service gateway 2136.
[0260] It should be recognized that the IaaS architectures 1800, 1900, 2000, and 2100 depicted in the figures may have other components besides those depicted. Furthermore, the embodiments shown in the figures are merely some examples of cloud infrastructure systems that can be incorporated into embodiments of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than shown in the figures, may combine two or more components, or may have different configurations or component arrangements.
[0261] In some embodiments, the IaaS system described herein may include application suites, middleware, and database service offerings delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such an IaaS system is the Oracle Cloud Infrastructure (OCI) provided by this assignee.
[0262] Figure 22 An example computer system 2200, in which various embodiments can be implemented, is illustrated. System 2200 can be used to implement any of the computer systems described above. As shown, computer system 2200 includes a processing unit 2204 that communicates with a plurality of peripheral subsystems via a bus subsystem 2202. These peripheral subsystems may include a processing acceleration unit 2206, an I / O subsystem 2208, a storage subsystem 2218, and a communication subsystem 2224. Storage subsystem 2218 includes a tangible computer-readable storage medium 2222 and system memory 2210.
[0263] Bus subsystem 2202 provides a mechanism for allowing various components and subsystems of computer system 2200 to communicate with each other as intended. While bus subsystem 2202 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2202 can be any of several types of bus architectures, including memory buses or memory controllers, peripheral buses, and local buses using any of the various bus architectures. For example, such architectures may include Industry Standard Architecture (ISA) buses, Microchannel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses, which may be implemented as Mezzanine buses manufactured according to the IEEE P1386.1 standard.
[0264] A processing unit 2204, which may be implemented as one or more integrated circuits (e.g., a conventional microprocessor or microcontroller), controls the operation of the computer system 2200. One or more processors may be included in the processing unit 2204. These processors may include single-core or multi-core processors. In some embodiments, the processing unit 2204 may be implemented as one or more independent processing units 2232 and / or 2234, each including a single-core or multi-core processor. In other embodiments, the processing unit 2204 may also be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0265] In various embodiments, processing unit 2204 can execute various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed can reside in processor(s) 2204 and / or storage subsystem 2218. With appropriate programming, processor(s) 2204 can provide the various functions described above. Computer system 2200 may additionally include processing acceleration unit 2206, which may include digital signal processor (DSP), dedicated processor, etc.
[0266] I / O subsystem 2208 may include user interface input devices and user interface output devices. User interface input devices may include keyboards, pointing devices such as mice or trackballs, touchpads or touchscreens integrated into a display, scroll wheels, click wheels, dials, buttons, switches, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may include, for example, motion sensing and / or gesture recognition devices, such as the motion sensor of Microsoft Kinect®, which enables users to control and interact with input devices such as game controllers for the Microsoft Xbox® 360 via a natural user interface using gestures and voice commands. User interface input devices may also include eye posture recognition devices, such as the Google Glass® blink detector, which detects eye activity from the user (e.g., "blinking" when taking a photo and / or making menu selections) and translates the eye posture into input in an input device (e.g., Google Glass®). Furthermore, user interface input devices may include voice recognition sensing devices that enable users to interact with a voice recognition system (e.g., the Siri® navigator) via voice commands.
[0267] User interface input devices may also include, but are not limited to, 3D mice, joysticks or pointing sticks, game panels and drawing tablets, as well as audio / video devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Furthermore, user interface input devices may include, for example, medical imaging input devices such as computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), and medical ultrasound equipment. User interface input devices may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.
[0268] User interface output devices may include display subsystems, indicator lights, or non-visual displays such as audio output devices, etc. Display subsystems may be cathode ray tubes (CRTs), flat panel devices such as those using liquid crystal displays (LCDs) or plasma displays, projection devices, touchscreens, etc. Generally, the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 2200 to the user or other computers. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, voice output devices, and modems.
[0269] Computer system 2200 may include storage subsystem 2218, which provides a tangible, non-transitory, computer-readable storage medium for storing software and data constructs that provide the functionality of the embodiments described in this disclosure. The software may include programs, code modules, instructions, scripts, etc., which, when executed by one or more cores or processors of processing unit 2204, provide the aforementioned functionality. Storage subsystem 2218 may also provide a repository for storing data used according to this disclosure.
[0270] like Figure 22 As illustrated in the example, storage subsystem 2218 may include various components, including system memory 2210, computer-readable storage medium 2222, and computer-readable storage medium reader 2220. System memory 2210 may store program instructions that can be loaded and executed by processing unit 2204. System memory 2210 may also store data used during instruction execution and / or data generated during program instruction execution. Various types of programs may be loaded into system memory 2210, including but not limited to client applications, web browsers, middleware applications, relational database management systems (RDBMS), virtual machines, containers, etc.
[0271] System memory 2210 may also store operating system 2216. Examples of operating system 2216 may include various versions of Microsoft Windows®, Apple Macintosh® and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.) and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS. In some implementations of computer system 2200 that execute one or more virtual machines, the virtual machine, along with the guest operating system (GOS), may be loaded into system memory 2210 and executed by one or more processors or cores of processing unit 2204.
[0272] System memory 2210 can be configured differently depending on the type of computer system 2200. For example, system memory 2210 can be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). Different types of RAM configurations can be provided, including static random access memory (SRAM), dynamic random access memory (DRAM), etc. In some embodiments, system memory 2210 may include a basic input / output system (BIOS), which contains basic routines such as those that facilitate the transfer of information between components within computer system 2200 during startup.
[0273] Computer-readable storage medium 2222 may represent remote, local, fixed and / or removable storage devices and storage media for temporarily and / or more permanently containing and storing computer-readable information for use by computer system 2200, including instructions executable by processing unit 2204 of computer system 2200.
[0274] Computer-readable storage medium 2222 may include any suitable medium known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage or other magnetic storage devices, or other tangible computer-readable media.
[0275] As an example, computer-readable storage medium 2222 may include hard disk drives that read from or write to non-removable non-volatile magnetic media, disk drives that read from or write to removable non-volatile magnetic disks, and optical disc drives that read from or write to removable non-volatile optical discs (such as CD ROMs, DVDs, and Blu-ray® discs or other optical media). Computer-readable storage medium 2222 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash memory drives, Secure Digital (SD) cards, DVD discs, digital audio tapes, and so on. Computer-readable storage medium 2222 may also include solid-state drives (SSDs) based on non-volatile memory (such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc.), volatile memory-based SSDs (such as solid-state RAM, dynamic RAM, static RAM), DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs using a combination of DRAM-based and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide non-volatile storage for computer-readable instructions, data structures, program modules and other data for computer system 2200.
[0276] Machine-readable instructions executable by one or more processors or cores of processing unit 2204 may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include physically tangible memory or storage devices, including volatile memory storage devices and / or non-volatile memory devices. Examples of non-transitory computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard disk drives, floppy disk drives, removable memory drives (e.g., USB drives), or other types of storage devices.
[0277] The communication subsystem 2224 provides an interface to other computer systems and networks. The communication subsystem 2224 serves as an interface for receiving data from other systems and sending data from computer system 2200 to other systems. For example, the communication subsystem 2224 enables computer system 2200 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 2224 may include radio frequency (RF) transceiver components (e.g., advanced data network technologies using cellular telephone technologies, such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.11 series standards), or other mobile communication technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components for accessing wireless voice and / or data networks. In some embodiments, as an addition to or alternative to the wireless interface, the communication subsystem 2224 may provide a wired network connection (e.g., Ethernet).
[0278] In some embodiments, the communication subsystem 2224 may also represent one or more users who may use the computer system 2200 to receive input communications in the form of structured and / or unstructured data feeds 2226, event streams 2228, event updates 2230, etc.
[0279] As an example, the communication subsystem 2224 can be configured to receive data feeds 2226 in real time from users of social networks and / or other communication services, such as Twitter® feeds, Facebook® updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0280] Furthermore, the communication subsystem 2224 can also be configured to receive data in the form of a continuous data stream, which may include event streams 2228 and / or event updates 2230 that are essentially continuous or unbounded real-time events without a clearly defined termination. Examples of applications that generate continuous data may include, for example, sensor data applications, financial quote machines, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, vehicle traffic monitoring, and so on.
[0281] The communication subsystem 2224 can also be configured to output structured and / or unstructured data feeds 2226, event streams 2228, event updates 2230, etc. to one or more databases, which can communicate with one or more streaming data source computers coupled to the computer system 2200.
[0282] The computer system 2200 can be one of a variety of types, including handheld portable devices (e.g., iPhone® cellular phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google® Glass head-mounted displays), PCs, workstations, mainframes, information stations, server racks, or any other data processing systems.
[0283] Due to the ever-evolving nature of computers and networks, the description of the computer system 2200 depicted in the figures is merely a concrete example. Many other configurations with more or fewer components than the system depicted in the figures are possible. For example, custom hardware may be used and / or specific elements may be implemented using hardware, firmware, software (including applets), or a combination thereof. Additionally, connections to other computing devices, such as network input / output devices, may also be employed. Based on the disclosure and teachings provided herein, those skilled in the art will recognize other ways and / or methods for implementing the various embodiments.
[0284] While specific embodiments have been described, various modifications, alterations, alternative constructions, and equivalents are also included within the scope of this disclosure. The embodiments are not limited to operation within certain specific data processing environments, but can be freely operated within multiple data processing environments. Furthermore, although the embodiments have been described using a specific series of transactions and steps, those skilled in the art will understand that the scope of this disclosure is not limited to the described series of transactions and steps. Various features and aspects of the above embodiments can be used individually or in combination.
[0285] Furthermore, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments may be implemented using only hardware, or only software, or a combination thereof. The various processes described herein can be implemented in any combination on the same processor or on different processors. Accordingly, where a component or service is described as being configured to perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits to perform operations, by programming programmable electronic circuits (such as microprocessors), or any combination thereof. Processes may communicate using a variety of technologies, including but not limited to conventional technologies for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.
[0286] Accordingly, the specification and drawings are to be considered illustrative rather than restrictive. However, it will be apparent that additions, omissions, deletions, and other modifications and changes may be made therein without departing from the broader spirit and scope set forth in the claims. Therefore, while specific disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.
[0287] In the context of describing the disclosed embodiments (particularly in the context of the following claims), the terms "a," "an," and "the," and similar designations, are to be interpreted as covering both singular and plural, unless otherwise indicated herein or obviously contradicted by the context. Unless otherwise stated, the terms "comprising," "having," "including," and "containing" are to be interpreted as open-ended terms (i.e., meaning "including but not limited to"). The term "connected" should be interpreted as partially or wholly contained in, attached to, or joined together, even if something exists in between. Unless otherwise indicated herein, the enumeration of value ranges herein is intended only as a shorthand method for individually referencing each individual value falling within that range, and each individual value is incorporated into the specification as if it were individually enumerated herein. Unless otherwise indicated herein or obviously contradicted by the context, all methods described herein can be performed in any suitable order. The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the embodiments and does not constitute a limitation on the scope of this disclosure, unless otherwise stated. Nothing in the specification should be construed as indicating that any unclaimed element is essential to the practice of this disclosure.
[0288] Exclusion language, such as the phrase "at least one of X, Y, or Z", is intended to be understood in the context generally used to represent items, terms, etc., unless otherwise explicitly stated, and may be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Therefore, such exclusion language is generally not intended to, and should not, imply that some embodiments require the presence of at least one of X, at least one of Y, or at least one of Z.
[0289] This document describes preferred embodiments of the present disclosure, including the best modes known for carrying out the present disclosure. Variations of those preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Those skilled in the art should be able to suitably employ such variations and may practice the present disclosure in ways other than those specifically described herein. Accordingly, the present disclosure includes all modifications and equivalents to the subject matter recited in the appended claims, where permitted by applicable law. Furthermore, unless otherwise indicated herein, the present disclosure includes any combination of the foregoing elements in all its possible variations.
[0290] All references cited in this article, including publications, patent applications and patents, are incorporated into this article by reference to the same extent as if each reference individually and specifically indicated to be incorporated by reference and elaborated in full in this article.
[0291] In the foregoing specification, various aspects of this disclosure have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that this disclosure is not limited thereto. The various features and aspects of the foregoing disclosure may be used individually or in combination. Furthermore, embodiments may be used in any number of settings and applications other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the accompanying drawings should be considered illustrative rather than restrictive.< / realm>
Claims
1. A method comprising: In a network environment comprising multiple host machines, the multiple host machines are communicatively coupled to each other via a network architecture comprising multiple switches comprising multiple ports, each of the multiple host machines comprising one or more GPUs, a first subset of the multiple ports is associated with a first virtual plane, the first virtual plane identifying a first set of resources to be dedicated to transmitting data packets from and to the host machines associated with the first virtual plane; Associate a subset of the second ports among the plurality of ports with a second virtual plane that is different from the first virtual plane; Associate the first host machine with the first virtual plane and associate the second host machine with the second virtual plane; as well as For a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine, the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using ports in the first port subset and the second port subset.
2. The method of claim 1, wherein the plurality of switches are arranged in a hierarchical structure including a first-layer switch, a second-layer switch and a third-layer switch, wherein the plurality of host machines are directly coupled to a switch included in the first-layer switch, and wherein the second-layer switch communicatively couples the first-layer switch to the third-layer switch.
3. The method of claim 2, wherein at least one port in the first port subset belongs to a first switch contained in a third-layer switch and associated with a first virtual plane.
4. The method of claim 3, wherein at least one port in the second port subset belongs to a second switch contained in a third-layer switch and associated with a second virtual plane.
5. The method of claim 4, wherein the first switch in the third-layer switch and the second switch in the third-layer switch are directly coupled to each other via a cable.
6. The method of claim 2, further comprising: For each of the plurality of switches included in the first-layer switches, a first virtual tunnel endpoint (VTEP) is created associated with a first virtual plane and a second VTEP is created associated with a second virtual plane. The first VTEP has a first autonomous system number (ASN) and the second VTEP has a second ASN that is different from the first ASN.
7. The method of claim 2, wherein a subset of host machines in the plurality of host machines is directly coupled to a first switch contained in a first-layer switch, each host machine in the subset of host machines being associated with a different virtual plane.
8. The method of claim 7, wherein the number of virtual planes supported by the network architecture corresponds to the number of host machines included in the subset of host machines that are directly coupled to a first switch included in a first-layer switch.
9. The method of claim 7, wherein each host machine is associated with a unique virtual tunnel endpoint (VTEP) created in a first switch included in a first-layer switch, and wherein one or more GPUs of the host machine communicate with one or more other GPUs of other host machines in the network environment via the unique VTEP associated with the host machine.
10. One or more computer-readable non-transitory media storing computer-executable instructions, said instructions, when executed by one or more processors, such that: In a network environment comprising multiple host machines, the multiple host machines are communicatively coupled to each other via a network architecture comprising multiple switches comprising multiple ports, each of the multiple host machines comprising one or more GPUs, a first subset of the multiple ports is associated with a first virtual plane, the first virtual plane identifying a first set of resources to be dedicated to transmitting data packets from and to the host machines associated with the first virtual plane; Associate a subset of the second ports among the plurality of ports with a second virtual plane that is different from the first virtual plane; Associate the first host machine with the first virtual plane and associate the second host machine with the second virtual plane; as well as For a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine, the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using ports in the first port subset and the second port subset.
11. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 10, wherein the plurality of switches are arranged in a hierarchical structure including a first-layer switch, a second-layer switch and a third-layer switch, wherein the plurality of host machines are directly coupled to a switch included in the first-layer switch, and wherein the second-layer switch communicatively couples the first-layer switch to the third-layer switch.
12. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 11, wherein at least one port in the first port subset belongs to a first switch contained in a third-layer switch and associated with a first virtual plane.
13. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 12, wherein at least one port in the second port subset belongs to a second switch contained in a third-layer switch and associated with a second virtual plane.
14. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 13, wherein a first switch included in a Layer 3 switch and a second switch included in a Layer 3 switch are directly coupled to each other via a cable.
15. The computer-readable non-transitory medium storing one or more computer-executable instructions as described in claim 11, further comprising: For each of the plurality of switches included in the first-layer switches, a first virtual tunnel endpoint (VTEP) is created associated with a first virtual plane and a second VTEP is created associated with a second virtual plane. The first VTEP has a first autonomous system number (ASN) and the second VTEP has a second ASN that is different from the first ASN.
16. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 11, wherein a subset of host machines in the plurality of host machines is directly coupled to a first switch included in a first-layer switch, each host machine in the subset of host machines being associated with a different virtual plane.
17. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 16, wherein the number of virtual planes supported by the network architecture corresponds to the number of host machines included in the subset of host machines that are directly coupled to a first switch included in a first-layer switch.
18. The computer-readable non-transitory medium for storing computer-executable instructions as claimed in claim 16, wherein each host machine is associated with a unique virtual tunnel endpoint (VTEP) created in a first switch included in a first-layer switch, and wherein one or more GPUs of the host machine communicate with one or more other GPUs of other host machines in the network environment via the unique VTEP associated with the host machine.
19. A computing device, comprising: One or more processors; as well as The memory includes instructions that, when executed by the one or more processors, cause the computing device to at least: In a network environment comprising multiple host machines, the multiple host machines are communicatively coupled to each other via a network architecture comprising multiple switches comprising multiple ports, each of the multiple host machines comprising one or more GPUs, a first subset of the multiple ports is associated with a first virtual plane, the first virtual plane identifying a first set of resources to be dedicated to transmitting data packets from and to the host machines associated with the first virtual plane; Associate a subset of the second ports among the plurality of ports with a second virtual plane that is different from the first virtual plane; Associate the first host machine with the first virtual plane and associate the second host machine with the second virtual plane; as well as For a data packet originating from the first GPU on the first host machine and destined for the second GPU on the second host machine, the data packet is transmitted from the first GPU on the first host machine to the second GPU on the second host machine using ports in the first port subset and the second port subset.
20. The computing device of claim 19, wherein the plurality of switches are arranged in a hierarchical structure including a first-layer switch, a second-layer switch and a third-layer switch, wherein the plurality of host machines are directly coupled to a switch included in the first-layer switch, and wherein the second-layer switch communicatively couples the first-layer switch to the third-layer switch.