System and method for providing bandwidth congestion control in a private fabric in a high-performance computing environment
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-08-07
Smart Images

Figure 0007902234000001 
Figure 0007902234000002 
Figure 0007902234000003
Abstract
Description
[Technical Field]
[0001] Copyright Notice Some of the disclosures in this patent document are protected by copyright. The copyright holder will not object to any reproduction of this patent document or patent disclosure by any party, provided that it is in the patent files or records of the Patent and Trademark Office; otherwise, all copyrights are reserved in all cases.
[0002] Priority claims and cross-references to related applications: This application claims priority to U.S. Provisional Patent Application No. 62 / 937,594, filed on 19 November 2019, entitled "SYSTEM AND METHOD FOR PROVIDING QUALITY-OF-SERVICE AND SERVICE- LEVEL AGREEMENTS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT," which is incorporated herein by reference in its entirety.
[0003] This application also claims priority to the following patent application, which is incorporated herein by reference in its entirety: "SYSTEM AND METHOD FOR SUPPORTING" filed on 11 May 2020. RDMA BANDWIDTH RESTRICTIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING U.S. Patent Application No. 16 / 872,035, titled "ENVIRONMENT (System and Method for Supporting RDMA Bandwidth Limiting in a Private Fabric in a High-Performance Computing Environment)," filed May 11, 2020, "SYSTEM U.S. Patent Application No. 16 / 872,038, filed on May 11, 2020, entitled "SYSTEM AND METHOD FOR SUPPORTING TARGET GROUPS FOR CONGESTION CONTROL IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT" (System and Method for Providing Bandwidth Congestion Control in a Private Fabric in a High-Performance Computing Environment); U.S. Patent Application No. 16 / 872,039, entitled "System and Method for Supporting Target Groups for Congestion Control in Private Fabrics in a Cutting Environment"; and "SYSTEM AND" filed on 11 May 2020. METHOD FOR SUPPORTING USE OF FORWARD AND BACKWARD CONGESTION NOTIFICATIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT U.S. Patent Application No. 16 / 872,043, entitled "System and Method for Supporting the Use of Forward and Reverse Congestion Notifications in Private Fabrics in a Pleasant Environment."
[0004] field This instruction concerns systems and methods for implementing quality of service (QoS) and service level agreements (SLAs) in private high-performance interconnect fabrics such as InfiniBand (IB) and RoCE (RDMA (Remote Direct Memory Access) over Converged Ethernet®). [Background technology]
[0005] background As larger cloud computing architectures are deployed, traditional network and storage-related performance and management bottlenecks are becoming a significant problem. Interest has been growing in using high-performance lossless interconnects such as InfiniBand (IB) technology as the foundation for cloud computing fabrics. This is a common area that embodiments of this teaching are intended to address. [Overview of the Initiative] [Means for solving the problem]
[0006] overview: Specific aspects are described in the independent claims. Various optional embodiments are described in the dependent claims.
[0007] This specification describes a system and method for providing bandwidth congestion control in a private fabric in a high-performance computing environment. An exemplary method may provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches and a plurality of host channel adapters, each host channel adapter comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes. The method may provide an end node ingress bandwidth allocation associated with an end node attached to the host channel adapter. The method may allow an end node of the host channel adapter to receive ingress bandwidth, the ingress bandwidth exceeding the end node's ingress bandwidth allocation. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example of an InfiniBand environment according to one embodiment. [Figure 2] This figure shows an example of a partitioned cluster environment according to one embodiment. [Figure 3] This figure shows an example of a tree topology in a network environment according to one embodiment. [Figure 4] This figure shows an exemplary shared port architecture according to one embodiment. [Figure 5] This figure shows an exemplary vSwitch architecture according to one embodiment. [Figure 6] This figure shows an exemplary vPort architecture according to one embodiment. [Figure 7] This figure shows an exemplary vSwitch architecture in which LIDs according to one embodiment are prepopulated. [Figure 8] This figure shows an exemplary vSwitch architecture with dynamic LID assignment according to one embodiment. [Figure 9] FIG. [Figure 9] is a diagram showing an exemplary vSwitch architecture in which dynamic LID assignment is performed on a vSwitch and LIDs are pre-populated. [Figure 10] FIG. is a diagram showing an exemplary multi-subnet InfiniBand fabric according to one embodiment. [Figure 11] FIG. [Figure 10] is a diagram showing the interconnection between two subnets in a high-performance computing environment according to one embodiment. [Figure 12] FIG. is a diagram showing the interconnection between two subnets via a dual-port virtual router configuration in a high-performance computing environment according to one embodiment. [Figure 13] FIG. [Figure 11] is a flowchart showing a method for supporting a dual-port virtual router in a high-performance computing environment according to one embodiment. [Figure 14] FIG. shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to one embodiment. [Figure 15] FIG. [Figure 12] shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to one embodiment. [Figure 16] FIG. shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to one embodiment. [Figure 17] FIG. [Figure 13] shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment. [Figure 18] FIG. shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment. [Figure 19] FIG. [Figure 14] shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment. [Figure 20]This is a flowchart of a method for providing RDMA (Remote Direct Memory Access) read requests as a restricted feature in a high-performance computing environment. [Figure 21] This document describes a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment. [Figure 22] This document describes a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment. [Figure 23] This document describes a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment. [Figure 24] This document describes a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment. [Figure 25] This document describes a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment. [Figure 26] This is a flowchart of a method for combining multiple shared bandwidth segments in a high-performance computing environment, according to one embodiment. [Figure 27] This invention presents a system for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 28] This invention presents a system for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 29] This invention presents a system for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 30]This invention presents a system for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 31] This is a flowchart of a method for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 32] This invention presents a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment, according to one embodiment. [Figure 33] This invention presents a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment, according to one embodiment. [Figure 34] This invention presents a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment, according to one embodiment. [Figure 35] This is a flowchart illustrating a method for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment, according to one embodiment. [Figure 36] This describes a system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment. [Figure 37] This describes a system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment. [Figure 38] This describes a system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment. [Figure 39] This is a flowchart illustrating a method for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment. [Modes for carrying out the invention]
[0009] Detailed explanation: This teaching is illustrated, not limited, in the drawings of the accompanying drawings where similar reference numbers refer to similar elements. References to “some” or “one” or “several” embodiments in this disclosure do not necessarily refer to the same embodiment, but rather mean at least one. Specific implementations are described, but it will be understood that these specific implementations are provided for illustrative purposes only. Those skilled in the art will recognize that other components and configurations may be used without departing from the spirit and scope of this disclosure.
[0010] Common reference numbers may be used to indicate similar elements throughout the drawings and detailed descriptions. Therefore, a reference number used in one drawing may or may not be referenced in the detailed description specific to that drawing if the element is described elsewhere.
[0011] This specification describes systems and methods for providing quality of service (QOS) and service level agreements (SLAs) in private fabrics within high-performance computing environments.
[0012] According to one embodiment, the following description in this teaching applies to InfiniBand as an example of a high-performance network. TM Use the (IB) network. Throughout the following description, InfiniBand TMThis document may refer to the specifications (which are referred to in various ways, such as the InfiniBand specification, IB specification, or legacy IB specification). Such references are understood to be to the InfiniBand Trade Association Architecture Specification, Volume 1, Version 1.3, published in March 2015 and available at http: / / www.inifinibandta.org, which is incorporated herein by reference in its entirety. It will be apparent to those skilled in the art that other types of high-performance networks may be used without limitation. The following description also uses a fat tree topology as an example of a fabric topology. It will be apparent to those skilled in the art that other types of fabric topologies may be used without limitation.
[0013] According to the embodiment, the following description uses RoCE (RDMA over Converged Ethernet (Remote Direct Memory Access)). RDMA over Converged Ethernet (RoCE) is a standard protocol that enables efficient data transfer of RDMA over Ethernet networks, allowing offload transport and superior performance through hardware RDMA engine implementations. RoCE is a standard protocol defined in the InfiniBand Trade Association (IBTA) standard. RoCE utilizes UDP (User Datagram Protocol) encapsulation to enable it to cross Layer 3 networks. RDMA is a key capability natively used by InfiniBand interconnect technology. Both InfiniBand and Ethernet RoCE share a common user API but have different physical and link layers.
[0014] According to one embodiment, various parts of this specification include references to InfiniBand fabric when describing various implementations, but those skilled in the art will understand the manner described herein. It will be easy to see that various embodiments can also be realized with RoCE fabric.
[0015] To meet the demands of the cloud in today's era (for example, the exascale era), virtual machines must be able to utilize low-overhead network communication paradigms such as Remote Direct Memory Access (RDMA). Desirable. Because RDMA bypasses the OS stack and communicates directly with the hardware, pass-through technologies such as Single Root I / O Virtualization (SR-IOV) network adapters can be used. According to one embodiment, a virtual switch (vSwitch) SR-IOV architecture can be applied to a high-performance lossless interconnect network. Since network reconfiguration time is critical to making live migration a realistic option, in addition to the network architecture, a scalable and topology-independent dynamic reconfiguration mechanism can be provided.
[0016] According to one embodiment, a routing strategy for a virtualized environment using vSwitch can be provided, and an efficient routing algorithm can be provided for a network topology (e.g., a fat tree topology). By further adjusting the dynamic reconfiguration mechanism, the overhead imposed on the fat tree can be minimized.
[0017] According to one embodiment of this teaching, virtualization can be beneficial for efficient resource utilization and flexible resource allocation in cloud computing. Live migration enables optimized resource utilization by moving virtual machines (VMs) between physical servers in a way that is transparent to applications. Thus, virtualization, through live migration, enables consolidation, on-demand resource provisioning, and flexibility.
[0018] InfiniBand TM InfiniBand TM (IB) is InfiniBand TM Trade Association (InfiniBand) TM It is an open standard lossless networking technology developed by the Trade Association. This technology is based on serial point-to-point full-duplex interconnect, which provides high-throughput and low-latency communication, particularly for high-performance computing (HPC) applications and data centers.
[0019] InfiniBand TM • The architecture (InfiniBand Architecture: IBA) is 2 It supports layered topology partitioning. At the lower layer, IB networks are called subnets, and one subnet may contain a set of hosts interconnected using switches and point-to-point links. At higher levels, one IB fabric consists of one or more subnets that can be interconnected using routers.
[0020] Within a subnet, hosts may be connected using switches and point-to-point links. In addition, there may be a single master management entity, or subnet manager (SM), residing on a designated device within the subnet. The subnet manager is responsible for configuring, starting, and maintaining IB subnets. Furthermore, the subnet manager (SM) may perform routing table calculations within the IB fabric. Here, for example, routing in an IB network aims for proper load balancing between all source-destination pairs within the local subnet.
[0021] Through the subnet management interface, the subnet manager exchanges control packets called subnet management packets (SMPs) with the subnet management agent (SMA). The agent resides on all IB subnet devices. By using SMP, the subnet manager can discover the fabric, configure end nodes and switches, and receive notifications from the SMP.
[0022] According to one embodiment, routing within a subnet in an IB network is performed using a linear forwarding table (LF) stored in the switch. Based on T), the LFT is calculated by the SM according to the routing mechanism in use. In a subnet, the Host Channel adapter on the end node Adapter (HCA) ports and switches are addressed using local identifiers (LIDs). Each entry in the LFT is a destination LID (DLI). It consists of D) and the output port. Only one entry is supported for each LID in the table. When a packet arrives at a switch, its output port is determined by looking up the DLID in the switch's forwarding table. Routing is deterministic because packets travel the same path in a network between a given source-destination pair (LID pair).
[0023] Generally, all subnet managers except the master subnet manager operate in standby mode for fault tolerance. However, in the event of a master subnet manager failure, the standby subnet managers will arrange for a new master subnet manager. The master subnet manager also performs periodic sweeps of subnets to detect any topology changes and adjusts the network accordingly. Reconstruct the code.
[0024] Furthermore, hosts and switches within a subnet can be addressed using local identifiers (LIDs), and a single subnet may be limited to 49,151 unicast LIDs. In addition to LIDs, which are valid local addresses within a subnet, each IB device may have a 64-bit global unique identifier (GUID). GUIDs can be used to form global identifiers (GIDs), which are IB Layer 3 (L3) addresses.
[0025] The SM may compute a routing table (i.e., connections / routes between each pair of nodes in a subnet) during network initialization time. Furthermore, whenever the topology changes, the routing table may be updated to ensure connectivity and optimal performance. During normal operation, the SM may perform periodic light sweeps of the network to check for topology changes. If a change is detected during a light sweep, Alternatively, if an SM receives a message (trap) that signals a change in the network, the SM may reconfigure the network according to the discovered change.
[0026] For example, a Smart Module (SM) can reconfigure the network when the network topology changes, such as when a link goes down, a device is added, or a link is removed. The reconfiguration step may include steps performed during network initialization. Furthermore, the reconfiguration may have a local scope, limited to the subnet where the network change occurred. Additionally, segmenting a large fabric using routers may limit the scope of the reconfiguration.
[0027] Figure 1 shows an example of an InfiniBand environment 100 according to one embodiment, and an example of an InfiniBand fabric. In the example shown in Figure 1, nodes A101 to E105 are InfiniBand. A niband fabric 120 is used for communication via the respective host channel adapters 111-115. According to one embodiment, various nodes (e.g., nodes A101-E105) can be represented by various physical devices. According to one embodiment, various nodes (e.g., nodes A101-E105) can be represented by various virtual devices such as virtual machines.
[0028] Partitioning in InfiniBand According to one embodiment, an IB network may support partitioning as a security mechanism for isolating logical groups of systems sharing a network fabric. Each HCA port on a node in the fabric may be a member of one or more partitions. Partition membership is managed by a centralized partition manager, which may be part of the SM. The SM can organize partition membership information for each port as a table of 16-bit partition keys (P_Key). The SM also manages these Switch ports and router ports can be configured using partition enforcement tables that contain P_Key information associated with end nodes sending or receiving data traffic through the port. In addition, in typical cases, the partition membership of a switch port may represent the set of all memberships indirectly associated with LIDs routed through the port toward the outbound direction (towards the link).
[0029] In one embodiment, a partition is a logical group of ports, and members of a group can only communicate with other members of the same logical group. In host channel adapters (HCAs) and switches, isolation can be implemented by filtering packets using partition membership information. Packets with invalid partitioning information can be dropped immediately upon reaching the entry port. In a partitioned IB system, partitions can be used to create tenant clusters. When partitioning is implemented appropriately, nodes cannot communicate with other nodes belonging to different tenant clusters. In this way, the security of the system can be guaranteed even if there are faulty or malicious tenant nodes.
[0030] In one embodiment, for communication between nodes, queue pairs (QP) and end-to-end contexts (EEC), excluding management queue pairs (QP0 and QP1), can be assigned to specific partitions. Next, P_Key information can be added to all transmitted IB transport packets. When a packet arrives at an HCA port or switch, its P_Key value can be checked against a table configured by the SM. If an invalid P_Key value is found, the packet is immediately discarded. In this way, communication is permitted only between ports sharing the same partition.
[0031] Figure 2, which illustrates an example of a partitioned cluster environment according to one embodiment, shows an example of an IB partition. In the example shown in Figure 2, nodes A101 to E105 communicate using the InfiniBand fabric 120 via their respective host channel adapters 111 to 115. Nodes A to E are located in partitions, namely partition 1 130, partition 2 140, and partition 3 150. Partition 1 includes node A 101 and node D 104. Partition 2 includes node A 101, node B 102, and node C 103. Partition 3 includes node C 103 and node E 105. This arrangement of partitions means that nodes D 104 and node E 105 share one partition. Since it is not available, communication is not possible. On the other hand, for example, node A 101 and node C 103 are both members of partition 2 140, so they can communicate with each other.
[0032] Virtual Machines in InfiniBand Over the past decade, CPU overhead has been virtually eliminated by hardware virtualization support, memory overhead has been significantly reduced by virtualizing the memory management unit, storage overhead has been reduced by the use of high-speed SAN storage or distributed network file systems, and device passthrough technologies such as Single Root Input / Output Virtualization (SR-IOV) have been introduced. As network I / O overhead has been reduced by using [this technology], the future prospects for virtualized High Performance Computing (HPC) environments have improved significantly. Currently, the cloud supports virtual HPC (vHPC) clusters using high-performance interconnect solutions, and meets the necessary requirements. It can provide the ability.
[0033] However, when connected to lossless networks such as InfiniBand (IB), some cloud functions, such as live migration of virtual machines (VMs), still present challenges due to the complex addressing and routing schemes used in these solutions. IB is an interconnect network technology that provides high bandwidth and low latency, making it very well suited for HPC and other communication-intensive workloads.
[0034] The traditional approach to connecting IB devices to VMs involves using directly assigned SR-IOVs. However, achieving live migration of VMs assigned to IB host channel adapters (HCAs) using SR-IOVs has proven difficult. Each node to which an IB is connected has three different addresses (i.e., LID, GUID, and GID). When live migration occurs, one or more of these addresses change. Other nodes communicating with the VM in migration (VM-in-migration) may lose connectivity. When this occurs, the IB Subnet Manager (SM) should send a Subnet Administration (SA) routing query to reconnect. By determining the virtual machine's new address, we can attempt to restore the lost connection.
[0035] IB uses three different types of addresses. The first type of address is a 16-bit Local Identifier (LID). At least one unique LID is assigned by the SM to each HCA port and each switch. LIDs are used to route traffic within the subnet. Since LIDs are 16 bits long, 65,536 unique address combinations can be formed, of which only 49,151 (0×0001-0×BFFF) can be used as unicast addresses. As a result, the number of available unicast addresses defines the maximum size of the IB subnet. The second type of address is a 64-bit Globally Unique Identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM may assign additional subnet-specific GUIDs to HCA ports, which are useful when SR-IOV is used. The third type of address is a 128-bit Global Identifier (GID). A GID is a valid IPv6 unicast address, and at least one is assigned to each HCA port. The GID is formed by combining a globally unique 64-bit prefix assigned by the Fabric Administrator with the GUID address of each HCA port.
[0036] Fat Tree (FTree) topology and routing In one embodiment, some IB-based HPC systems employ a fat tree topology to leverage the beneficial properties that fat trees offer. These properties include full bisection bandwidth and inherent fault tolerance, resulting from the availability of multiple paths between each source-destination pair. The initial concept behind fat trees was to employ fatter links between nodes using more available bandwidth as the tree approaches the root of the topology. These fatter links can help avoid congestion at higher-level switches while maintaining bisection bandwidth.
[0037] Figure 3 shows an example of a tree topology in a network environment according to one embodiment. As shown in Figure 3, one or more end nodes 201-204 may be connected in the network fabric 200. The network fabric 200 may be based on a fat tree topology including multiple leaf switches 211-214 and multiple spine switches or root switches 231-234. In addition, the network fabric 200 may include one or more intermediate switches such as switches 221-224.
[0038] Furthermore, as shown in Figure 3, each of the end nodes 201 to 204 may be a multi-homed node, that is, a single node connected to two or more parts of the network fabric 200 via multiple ports. For example, node 201 may include ports H1 and H2, node 202 may include ports H3 and H4, node 203 may include ports H5 and H6, and node 204 may include ports H7 and H8.
[0039] In addition, each switch may have multiple switch ports. For example, root switch 231 may have switch ports 1-2, root switch 232 may have switch ports 3-4, root switch 233 may have switch ports 5-6, and root switch 234 may have switch ports 7-8.
[0040] According to the embodiment, the fat tree routing mechanism is one of the most popular routing algorithms for IB-based fat tree topologies. The fat tree routing mechanism is also used in OFED (Open Fabric Enterprise Distribution: standard software for building and deploying IB-based applications). This is implemented in OpenSM, specifically through a subnet manager (a shared stack).
[0041] The purpose of the fat tree routing mechanism is to generate LFTs that uniformly spread the shortest path routes across links in the network fabric. This mechanism traverses the fabric in an indexed order, assigning the target LID of each end node, and consequently the corresponding route, to each switch port. For end nodes connected to the same leaf switch, the indexed order may depend on the switch port to which the end node is connected (i.e., the port numbering sequence). For each port, the mechanism can maintain a port usage counter, and each time a new route is added, it can use the port usage counter to select the least used port.
[0042] In one embodiment, in a partitioned subnet, nodes that are not members of a common partition are not allowed to communicate. In practice, this means that some of the routes allocated by the fat tree routing algorithm are not used for user traffic. Problems arise if the fat tree routing mechanism generates LFTs for those routes in the same way as other functional routes. This behavior can degrade balancing on the link because nodes are routed in indexed order. Because routing is performed without noticing this, subnets routed via fat trees generally result in poor isolation between partitions.
[0043] According to one embodiment, a fat tree is a hierarchical network topology that can scale with available network resources. Further, a fat tree is easily constructed using commodity switches arranged in different levels of hierarchy. Further, various variants of fat trees, including k-ary-n-tree, Extended Generalized Fat-Tree (XGFT), Parallel Ports Generalized Fat-Tree (PGFT), and Real Life Fat-Tree (RLFT), are generally available.
[0044] Also, a k-ary-n-tree is an n-level fat tree having n end nodes and n·k n-1 switches, each having 2k ports. Each switch has the same number of connections in the up and down directions in the tree. The XGFT fat tree extends the k-ary-n-tree by allowing different numbers of up and down connections for switches and different numbers of connections at each level in the tree. The PGFT definition further extends the XGFT topology to allow multiple connections between switches. A wide variety of topologies can be defined using XGFT and PGFT. However, for practical use, a restricted version of PGFT, RLFT, has been introduced to define the fat tree commonly found in modern HPC clusters. RLFT uses the same port count switches at all levels in the fat tree.
[0045] Input / Output (I / O) virtualization In one embodiment, I / O virtualization (IOV) can make I / O available by allowing virtual machines (VMs) to access the underlying physical resources. The combination of storage traffic and inter-server communication can place an unbearably high load on a single server's I / O resources, potentially resulting in backlogs and processor idle time while data is waiting. As the number of I / O requests increases, IOV can provide availability and improve the performance, scalability, and flexibility of (virtualized) I / O resources to match the performance levels seen in modern CPU virtualization.
[0046] In one embodiment, an IOV is desired that can enable the sharing of I / O resources and protect access to those resources from the VM. The IOV decouples the logical units exposed to the VM from their physical implementation. Currently, emulation, paravirtualization, direct assignment (DA), and single-root I / O are available. Various types of IOV technologies can exist, such as virtualization (SR-IOV).
[0047] In one embodiment, software emulation is a type of IOV technology. Software emulation can enable a separate front-end / back-end software architecture. The front-end can be a device driver deployed in the VM and can communicate with the back-end, which is implemented by the hypervisor to provide I / O access. The physical device sharing ratio is high, and live migration of VMs can be achieved with only a few milliseconds of network downtime. However, software emulation introduces further undesirable computational overhead.
[0048] According to one embodiment, another type of IOV technology is direct device assignment. In direct device assignment, the I / O device needs to be linked to the VM, but the device Chairs are not shared between VMs. Direct allocation or device passthrough provides nearly unique performance with minimal overhead. Physical devices bypass the hypervisor and are attached directly to the VM. However, the drawback of such direct device allocation is that scalability is limited because there is no sharing between virtual machines, such as one physical network card being connected to one VM.
[0049] According to one embodiment, a Single Root IOV (SR-IOV) is, Hardware virtualization can allow a physical device to appear as multiple independent, lightweight instances of the same device. These instances can be assigned to VMs as pass-through devices and accessed as virtual functions (VFs). The hypervisor accesses the devices through unique, fully functional physical functions (PFs) (for each device). SR-I OV mitigates scalability issues when allocating purely directly. However, the problem presented by SR-IOV is that it can impair VM migration. Among these IOV technologies, SR-IOV can extend the PCI Express (PCIe) standard by providing a means to allow multiple VMs to directly access a single physical device while maintaining nearly inherent performance. This allows SR-IOV to offer superior performance and scalability.
[0050] SR-IOV allows a PCIe device to expose multiple virtual devices that can be shared among multiple guests by assigning one virtual device to each guest. Each SR-IOV device has at least one physical function (PF) and one or more associated virtual functions (VFs). The PF is controlled by a virtual machine monitor (VMM) or hypervisor. While PCIe is a standard PCIe function, VF is a lightweight PCIe function. Each VF has its own base address (BAR) and is assigned a unique requester ID. The unique requester ID is assigned to the I / O memory management unit (I / O memory). The management unit (IOMMU) enables the differentiation of traffic streams to and from various VFs. The IOMMU also applies memory to interrupt the translation between PFs and VFs.
[0051] Unfortunately, direct device allocation technology presents a barrier for cloud providers in situations where transparent live migration of virtual machines is desired for data center optimization. The essence of live migration is that the memory contents of the VM are copied to the remote hypervisor. Furthermore, the VM is suspended on the source hypervisor and its operation is resumed at the destination. When using software emulation methods, network interfaces are virtual so that their internal state is stored in memory and then copied. As a result, downtime can be reduced to a few milliseconds.
[0052] However, migration becomes more difficult when direct device assignment technologies such as SR-IOV are used. In such situations, the entire internal state of the network interface cannot be copied because it is tied to the hardware. Instead, the SR-IOV VF assigned to the VM is isolated, a live migration is performed, and a new VF is assigned at the destination. In the case of InfiniBand and SR-IOV, this process can result in downtime on the order of seconds. Furthermore, in the SR-IOV shared port model, the VM's address changes after migration, which adds overhead to the SM and negatively impacts the performance of the underlying network fabric. This will happen.
[0053] InfiniBand SR-IOV Architecture - Shared Ports There can be various types of SR-IOV models (e.g., shared port model, virtual switch model, and virtual port model).
[0054] Figure 4 shows an exemplary shared port architecture according to one embodiment. As shown in the figure, host 300 (e.g., host channel adapter) can interact with hypervisor 310. Hypervisor 310 can assign various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions can be handled by hypervisor 310.
[0055] According to one embodiment, when using a shared port architecture as shown in Figure 4, the host (e.g., HCA) appears as a single port in a network where there is a single shared LID and shared queue pair (QP) space between the physical function 320 and the virtual functions 330, 350, 350. However, each function (i.e., the physical function and the virtual function) may have its own GID.
[0056] As shown in Figure 4, according to one embodiment, various GIDs can be assigned to virtual and physical functions, and a special queue pair QP0 and QP1 (i.e., InfiniBand) can be used. TM A dedicated queue pair (QP) used for management packets is owned by the physical function. These QPs are also exposed to the VF, but the VF is not permitted to use QP0 (all incoming SMPs from the VF to QP0 are discarded), and QP1 can act as a proxy for the actual QP1 owned by the PF.
[0057] According to one embodiment, a shared port architecture can enable a highly scalable data center that is not limited by the number of VMs (which are associated with the network by being assigned to virtual functions), because only the physical machines and switches in the network consume LID space.
[0058] However, a drawback of the shared port architecture is its inability to provide transparent live migration, which hinders the possibility of flexible VM placement. Since each LID is associated with a specific hypervisor and shared among all VMs residing on that hypervisor, migrating VMs (i.e., virtual machines migrating to a destination hypervisor) must change their LID to that of the destination hypervisor. Furthermore, as a result of restricted QP0 access, the subnet manager cannot be run inside a VM.
[0059] InfiniBand SR-IOV Architecture Model - Virtual Switch (vSwitch) Figure 5 shows an exemplary vSwitch architecture according to one embodiment. As shown in the figure, host 400 (e.g., host channel adapter) can interact with hypervisor 410, which can assign various virtual functions 430, 440, and 450 to several virtual machines. Similarly, physical functions can be handled by hypervisor 410. Virtual switch 415 can also be handled by hypervisor 401.
[0060] According to one embodiment, in the vSwitch architecture, each virtual function 430, 440, and 450 is a complete virtual host channel adapter. er:vHCA) means that, in hardware, the VMs assigned to the VF are assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. For the rest of the network and SM, HCA400 appears as a switch with additional nodes connected via virtual switch 415. Hypervisor 410 can use PF420, and VMs (assigned to virtual functions) use VF.
[0061] In one embodiment, the vSwitch architecture provides transparent virtualization. However, since each virtual function is assigned a unique LID, the available LIDs are consumed rapidly. Similarly, if many LID addresses are in use (i.e., one for each physical function and each virtual function), more communication paths must be computed by the SM, and more subnet management packets (SMPs) must be sent to the switch to update those LFTs. For example, calculating communication paths can take several minutes in large networks. Since the LID space is limited to 49,151 unicast LIDs, and each VM (via VF) occupies one LID each for the physical node and switch, the number of active VMs is limited by the number of physical nodes and switches in the network, and vice versa.
[0062] InfiniBand SR-IOV Architecture Model - Virtual Port (vPort) Figure 6 shows an exemplary vPort concept according to one embodiment. As shown in the figure, host 300 (e.g., host channel adapter) can interact with hypervisor 410, which can assign various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions can be handled by hypervisor 310.
[0063] In one embodiment, the vPort concept is loosely defined to give vendors implementation freedom (for example, the definition does not stipulate that implementations should be SRIOV-specific), and the purpose of vPort is to standardize how VMs are handled within a subnet. The vPort concept can define both SR-IOV shared port-like architectures and vSwitch-like architectures, or a combination of these architectures, which can be more scalable in both the spatial and performance domains. In addition, vPort supports optional LIDs, and unlike shared ports, the SM can recognize all vPorts available in the subnet, even if the vPort does not use a dedicated LID.
[0064] InfiniBand SR-IOV Architecture Model - LID pre-populated vSwitch According to one embodiment, the disclosure provides a system and method for providing a vSwitch architecture with prepopulated LIDs.
[0065] Figure 7 shows an exemplary vSwitch architecture with pre-populated LIDs according to one embodiment. As shown in the figure, several switches 501-504 are in InfiniBand within a network switching environment 600 (e.g., an IB subnet). TM Communication can be established between members of a fabric, such as a fabric. The fabric may include several hardware devices such as host channel adapters 510, 520, and 530. Furthermore, host channel adapters 510, 520, and 530 can interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, in conjunction with the host channel adapters, can also interact with and configure several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. It can be assigned to several virtual machines. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. In addition, hypervisor 511 can assign virtual machine 2 551 to virtual function 2 515 and virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapters via physical functions 513, 523, and 533, each of which has sufficient functionality on each host channel adapter.
[0066] According to one embodiment, each of switches 501 to 504 may include several ports (not shown). Some of these ports are used to configure a linear forwarding table to direct traffic within the network switching environment 600.
[0067] In one embodiment, virtual switches 512, 522, and 532 can be handled by their respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), meaning that in hardware, the VMs assigned to the VF are assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. For the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches to which additional nodes are connected via virtual switches.
[0068] In one embodiment, the disclosure provides a system and method for providing a vSwitch architecture with prepopulated LIDs. Referring to Figure 7, LIDs are prepopulated for various physical functions 513, 523, and 533, as well as for virtual functions 514-516, 524-526, and 534-536 (even virtual functions not associated with any currently active virtual machines). For example, physical function 513 is prepopulated with LID1, and virtual function 534 is prepopulated with LID10. When the network is booted, LIDs are prepopulated in the SR-IOV vSwitch-enabled subnet. Even if not all VFs are occupied by VMs in the network, the populated VFs are assigned LIDs as shown in Figure 7.
[0069] According to one embodiment, many similar physical host channel adapters may have two or more ports (two ports shared for redundancy), and a virtual HCA may also be represented by two ports and connected to an external IB subnet via one or more virtual switches.
[0070] In one embodiment, in a vSwitch architecture where LIDs are prepopulated, each hypervisor consumes one LID for itself via the PF, and can consume one or more LIDs for each additional VF. The sum of all available VFs across all hypervisors in an IB subnet gives the maximum number of VMs that can run in the subnet. For example, in an IB subnet with 16 virtual functions per hypervisor, each hypervisor consumes 17 LIDs in the subnet (one LID for each of the 16 virtual functions and one LID for the physical function). In such an IB subnet, the theoretical hypervisor limit for a single subnet is defined by the number of available unicast LIDs, which is 2891 (obtained by dividing 49151 available LIDs by 17 LIDs per hypervisor), and the total number of VMs (i.e., the limit) is 46256 (obtained by multiplying 2891 hypervisors by 16 VFs per hypervisor) (effectively, this is the limit for each switch, router, etc. in the IB subnet). (Alternatively, dedicated SM nodes consume LIDs, so these numbers are actually smaller.) Note that vSwitches can share LIDs with PFs, so they do not need to occupy additional LIDs.
[0071] In one embodiment, in a vSwitch architecture where LIDs are prepopulated, once the network is booted, communication paths are calculated for all LIDs. If a new VM needs to be started, the system does not need to add a new LID in the subnet. Otherwise, operations that could completely reconfigure the network, including recalculating paths, would be the most time-consuming. Instead, available ports for VMs reside in one of the hypervisors (i.e., available virtual functions), and virtual machines are assigned to available virtual functions.
[0072] According to one embodiment, a vSwitch architecture with pre-populated LIDs also enables the ability to compute and use different paths to reach different VMs hosted by the same hypervisor. Essentially, this is achieved by using LID-Mask-Control-like (LMC-like) features in such subnets and networks to provide alternative paths to a single physical machine without being constrained by the LMC constraints that require LIDs to be contiguous. This makes it possible. The ability to freely use discontinuous LIDs is particularly useful when migrating a VM and needing to deliver its associated LID to a destination.
[0073] In one embodiment, along with the aforementioned advantages of a vSwitch architecture with pre-populated LIDs, several considerations can be taken into account. For example, because LIDs are pre-populated in SR-IOV vSwitch-enabled subnets when the network is booting, the initial routing calculation (e.g., at startup) may take longer than if LIDs were not pre-populated.
[0074] InfiniBand SR-IOV Architecture Model - vSwitch with Dynamic LID Assignment According to one embodiment, the disclosure provides a system and method for providing a vSwitch architecture with dynamic LID assignment.
[0075] Figure 8 shows an exemplary vSwitch architecture with dynamic LID assignment according to one embodiment. As shown in the figure, several switches 501-504 are in InfiniBand within a network switching environment 700 (e.g., IB subnet). TMCommunication can be established between members of a fabric, such as a fabric. The fabric may include several hardware devices, such as host channel adapters 510, 520, and 530. Host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor can further interact with, configure, and assign to several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536, together with the host channel adapters. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can also assign virtual machine 2 551 to virtual function 2 515 and virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapters via physical functions 513, 523, and 533, which are fully functional on each of the host channel adapters.
[0076] According to one embodiment, each of switches 501 to 504 may include several ports (not shown). Some of these ports are used to configure a linear forwarding table to direct traffic within the network switching environment 700.
[0077] In one embodiment, virtual switches 512, 522, and 532 can be handled by their respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), meaning that in hardware, the VMs assigned to the VF are assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. For the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches to which additional nodes are connected via the virtual switches.
[0078] In one embodiment, the disclosure provides a system and method for providing a vSwitch architecture with dynamic LID assignment. Referring to Figure 8, LIDs are dynamically assigned to various physical functions 513, 523, and 533, with physical function 513 receiving LID1, physical function 523 receiving LID2, and physical function 533 receiving LID3. Those virtual functions associated with active virtual machines can also receive dynamically assigned LIDs. For example, since virtual machine 1 550 is active and associated with virtual function 1 514, virtual function 514 may be assigned LID5. Similarly, virtual functions 2 515, 3 516, and 1 534 are each associated with an active virtual function. Therefore, LIDs are assigned to these virtual functions, with LID7 assigned to virtual function 2 515, LID11 assigned to virtual function 3 516, and LID9 assigned to virtual function 1 534. Unlike vSwitches where LIDs are prepopulated, virtual functions that are not currently associated with an active virtual machine do not receive an LID assignment.
[0079] According to one embodiment, if dynamic LID assignment is performed, the initial routing calculation can be substantially reduced. When the network is booting for the first time and no VMs exist, a relatively small number of LIDs can be used for the initial routing and LFT distribution.
[0080] According to one embodiment, many similar physical host channel adapters may have two or more ports (two ports shared for redundancy), and a virtual HCA may also be represented by two ports and connected to an external IB subnet via one or more virtual switches.
[0081] According to one embodiment, when a new VM is created in a system using a vSwitch with dynamic LID assignment, a free VM slot is discovered, and a unique unused unicast LID is also discovered, in order to determine which hypervisor the newly added VM should boot on. However, there is no known route in the switch's LFT and network to handle the newly added LID. Calculating a new set of routes to handle the newly added VM is undesirable in a dynamic environment where several VMs may be booted every minute. In large IB subnets, calculating a new set of routes can take several minutes, and this procedure would have to be repeated every time a new VM is booted.
[0082] Advantageously, according to one embodiment, since all VFs in the hypervisor share the same uplink as the PF, there is no need to compute routes for a new set. The LFT is repeated for all physical switches in the network and forwarded from the LID entries belonging to the PF of the hypervisor (where the VMs are created) to the newly added LIDs. It is sufficient to simply copy the ports and send a single SMP to update the corresponding LFT blocks on a specific switch. This eliminates the need for the system and method to calculate the routes for the new set.
[0083] According to one embodiment, the assigned LIDs in a vSwitch with a dynamic LID assignment architecture do not need to be contiguous. Comparing the LIDs assigned on VMs on each hypervisor between a vSwitch with pre-populated LIDs and a vSwitch with dynamic LID assignment, it can be seen that the assigned LIDs in the dynamic LID assignment architecture are discontinuous, while the pre-populated LIDs are essentially contiguous. Furthermore, in the vSwitch dynamic LID assignment architecture, when a new VM is created, the next available LID is used throughout the VM's lifetime. Conversely, in a vSwitch with pre-populated LIDs, each VM inherits the LID already assigned to its corresponding VF, and in a network without live migration, VMs continuously assigned to a given VF receive the same LID.
[0084] In one embodiment, a vSwitch with a dynamic LID assignment architecture can overcome the shortcomings of a vSwitch with a pre-popular LID architecture model, at the expense of some additional network and runtime SM overhead. Each time a VM is created, the LFT of the physical switch in the subnet is updated with the newly added LID associated with the created VM. This operation requires sending one subnet management packet (SMP) per switch. Features such as LMC are also unavailable because each VM uses the same path as its host hypervisor. However, there is no limit to the total number of VFs present across all hypervisors, and the number of VFs may exceed the unicast LID limit. In such cases, it is naturally not possible for all VFs to be assigned simultaneously on an active VM, and having more spare hypervisors and VFs adds flexibility to recover from and optimize fragmented network failures when operating near the unicast LID limit.
[0085] InfiniBand SR-IOV Architecture Model - A vSwitch with Dynamically Assigned and Pre-populated LIDs Figure 9 shows an exemplary vSwitch architecture with a vSwitch that has been dynamically assigned and has prepopulated LIDs, according to one embodiment. As shown in the figure, several switches 501-504 are in InfiniBand within a network switching environment 800 (e.g., an IB subnet). TM Communication can be established between members of a fabric, such as a fabric. The fabric may include several hardware devices, such as host channel adapters 510, 520, and 530. Host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, together with the host channel adapters, can further interact with, configure, and assign to several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536 to several virtual machines. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can also assign virtual machine 2 551 to virtual function 2 515. Hypervisor 521 can assign virtual machine 3 552 to virtual function 3 526. The hypervisor 531 can further assign virtual machine 4 553 to virtual function 2 535. The hypervisor can access the host channel adapters via physical functions 513, 523, and 533, each of which are fully functional on the host channel adapter.
[0086] According to one embodiment, each of switches 501 to 504 may include several ports (not shown). These ports are used to configure a linear forwarding table to direct traffic within the network switching environment 800.
[0087] In one embodiment, virtual switches 512, 522, and 532 can be handled by their respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), meaning that in hardware, the VMs assigned to the VF are assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. For the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches to which additional nodes are connected via the virtual switches.
[0088] According to one embodiment, the present disclosure provides a system and method for providing a hybrid vSwitch architecture in which dynamic LID assignment is performed and LIDs are prepopulated. Referring to Figure 9, a hypervisor 511 may be configured with a vSwitch having a prepopulated LID architecture, a hypervisor 521 may be configured with a vSwitch in which LIDs are prepopulated and dynamic LID assignment is performed, and a hypervisor 531 may be configured with a vSwitch in which dynamic LID assignment is performed. Therefore, the physical functions 513 and virtual functions 514-516 have their LIDs prepopulated (i.e., LIDs are assigned even to virtual functions that are not assigned to an active virtual machine). Physical function 523 and virtual function 1 524 may have their LIDs prepopulated, while virtual function 2 525 and virtual function 3 526 have their LIDs dynamically assigned (i.e., virtual function 2 525 is available for dynamic LID assignment, and virtual function 3 526 is dynamically assigned the LID 11 because virtual machine 3 552 is assigned to it). Finally, functions (physical and virtual functions) associated with hypervisor 3 531 can have their LIDs dynamically assigned. As a result, virtual function 1 534 and virtual function 3 536 become available for dynamic LID assignment, and virtual function 2 535 is dynamically assigned the LID 9 because virtual machine 4 553 is assigned to it.
[0089] According to one embodiment, as shown in Figure 9, in which both vSwitches with pre-populated LIDs and vSwitches with dynamic LID assignments are used (independently or in combination within any given hypervisor), the number of pre-populated LIDs per host channel adapter can be defined by the fabric administrator and can be within the range of 0 <= pre-populated VF <= total VF (per host channel adapter). The VFs available for dynamic LID assignment can be found (per host channel adapter) by subtracting the number of pre-populated VFs from the total number of VFs.
[0090] According to one embodiment, many similar physical host channel adapters may have two or more ports (two ports shared for redundancy), and a virtual HCA may also be represented by two ports and connected to an external IB subnet via one or more virtual switches.
[0091] InfiniBand - Inter-subnet communication (Fabric Manager) According to one embodiment, in addition to providing an InfiniBand fabric within a single subnet, embodiments of the present disclosure provide an InfiniBand fabric spanning two or more subnets. We can also provide the fabric.
[0092] Figure 10 shows an exemplary multi-subnet InfiniBand fabric according to one embodiment. As shown in this figure, a number of switches 1001-1004 within subnet A 1000 can provide communication between members of a fabric, such as an InfiniBand fabric, within subnet A 1000 (e.g., an IB subnet). This fabric may include a number of hardware devices, such as channel adapters 1010. The host channel adapter 1010 can interact with the hypervisor 1011. The hypervisor can set up a number of virtual functions 1014 together with the host channel adapter with which it is interacting. In addition, the hypervisor can assign virtual machines to each virtual function. For example, virtual machine 1 1015 is assigned to virtual function 1 1014. The hypervisor can access its associated host channel adapters through fully functional physical functions, such as physical functions 1013 on each host channel adapter. Multiple switches 1021-1024 can provide communication between members of a fabric, such as an InfiniBand fabric, within subnet B 1040 (e.g., an IB subnet). This fabric may include multiple hardware devices, such as host channel adapters 1030. Host channel adapter 1030 can interact with hypervisor 1031. The hypervisor can set up multiple virtual functions 1034 together with the host channel adapter with which it interacts. In addition, the hypervisor can assign virtual machines to each virtual function. For example, virtual machine 2 1035 is assigned to virtual function 2 1034. The hypervisor can access its associated host channel adapters through fully functional physical functions, such as physical functions 1033 on each host channel adapter. Note that although only one host channel adapter is shown for each subnet (i.e., subnet A and subnet B), it should be understood that each subnet may contain multiple host channel adapters and their corresponding components.
[0093] In one embodiment, each host channel adapter may further be associated with virtual switches such as virtual switch 1012 and virtual switch 1032, and each HCA may be set up with a different architectural model as described above. Although both subnets in Figure 10 are shown using the architecture model of a vSwitch with prepopulated LIDs, this is not intended to suggest that all such subnet configurations may follow the same architectural model.
[0094] In one embodiment, at least one switch within each subnet may be associated with a router. For example, switch 1002 in subnet A 1000 is associated with router 1005, and switch 1021 in subnet B 1040 is associated with router 1006.
[0095] In one embodiment, at least one device (e.g., a switch, node, etc.) can be associated with a fabric manager (not shown). The fabric manager can be used, for example, to discover inter-subnet fabric topologies, create fabric profiles (e.g., virtual machine fabric profiles), and build virtual machine-related database objects that form the basis for building virtual machine fabric profiles. In addition, the fabric manager can define legal inter-subnet connectivity, specifying which subnets are allowed to communicate through which router ports and using which partition numbers.
[0096] According to one embodiment, a tracking source such as virtual machine 1 in subnet A If the fix is destined for a different subnet, such as virtual machine 2 in subnet B, the traffic should be directed to a router in subnet A, namely router 1005, which can then forward this traffic to subnet B via the link with router 1006.
[0097] Virtual Dual-Port Router According to one embodiment, a dual-port router abstraction is derived from the GRH (Global Route Header) to the LR. This provides a simple way to define inter-subnet router functionality based on a switch hardware implementation that has the capability to perform conversion to H (local route header) in addition to normal LRH-based switching.
[0098] In one embodiment, a virtual dual-port router can be logically connected outside of the corresponding switch port. This virtual dual-port router can provide an InfiniBand-compliant view to standard management entities such as subnet managers.
[0099] According to one embodiment, a dual-port router model has different subnets, with each subnet handling packet forwarding and address mapping in the ingress path to the subnet. This demonstrates that connections can be made in a way that provides complete control and does not affect routing or logical connectivity within any of the incorrectly connected subnets.
[0100] According to one embodiment, in situations involving a misconnected fabric, a virtual dual-port router abstraction can be used to ensure that management entities such as subnet managers and IB diagnostic software function correctly in the presence of unintended physical connections to remote subnets.
[0101] Figure 11 shows the interconnection between two subnets in a high-performance computing environment according to one embodiment. Before configuring with a virtual dual-port router, subnet A Switch 1120 within 1101 can be connected to switch 1130 in subnet B 1102 via switch port 1121 of switch 1120, and then via physical connection 1110, through switch port 1131 of switch 1130. In this embodiment, each of the switch ports 1121 and 1131 can function as both a switch port and a router port.
[0102] According to one embodiment, the problem with this configuration is that a management entity, such as a subnet manager, within an InfiniBand subnet cannot distinguish between a physical port that is both a switch port and a router port. In such a situation, the SM can treat a switch port as having a router port connected to it. However, if the switch port is connected to another subnet with a different subnet manager, for example via a physical link, the subnet manager can send discovery messages to the physical link. However, such discovery messages are not permitted in the other subnet.
[0103] Figure 12 shows an interconnection between two subnets via a dual-port virtual router configuration in a high-performance computing environment according to one embodiment.
[0104] According to one embodiment, after configuration, the dual-port virtual router configuration is configured such that the appropriate end node, which indicates the end of the subnet, is responsible for the subnet manager. It can be provided in a way that is easy to understand.
[0105] According to one embodiment, a switch port on switch 1220 in subnet A 1201 can be connected (i.e., logically connected) to router port 1211 in virtual router 1210 via virtual link 1223. The virtual router 1210 (e.g., a dual-port virtual router) is shown in the embodiment as being outside switch 1220, but can be logically included within switch 1220 and may also include a second router port, router port II 1212. According to one embodiment, a physical link 1203, which may have two ends, can connect subnet A 1201 to subnet B 1202 via the first end of the physical link, via the second end of the physical link, via router port II 1212, and router port II included in virtual router 1230 in subnet B 1202. A connection can be made via 1232. The virtual router 1230 may further include a router port 1231 that can be connected (i.e., logically connected) to switch port 1241 on switch 1240 via virtual link 1233.
[0106] According to one embodiment, a subnet manager (not shown) on subnet A can discover router port 1211 on virtual router 1210 as the endpoint of the subnet controlled by the subnet manager. The dual-port virtual router abstraction allows the subnet manager on subnet A to treat subnet A in the usual way (for example, as defined in the InfiniBand standard). At the subnet management agent level, the dual-port virtual router An abstraction can be provided to make the SM aware of a regular switch port, and then, at the SMA level, that abstraction can be provided to indicate that there is another port connected to this switch port and that this port becomes a router port on a dual-port virtual router. The local SM can continue to use the traditional fabric topology (in which the SM considers the port to be a standard switch port), and therefore the SM considers the router port to be an end port. Physical connectivity can be established between two switch ports that are also configured as router ports in two different subnets.
[0107] According to one embodiment, a dual-port virtual router can also solve the problem that a physical link may be mistakenly connected to any other switch port within the same subnet, or to a switch port not intended to provide connectivity to a different subnet. Thus, the methods and systems described herein also represent those outside of a subnet.
[0108] According to one embodiment, a local SM within a subnet such as subnet A determines a switch port, and then determines the router port connected to this switch port (for example, router port 1211 connected to switch port 1221 via virtual link 1223). Since the SM considers router port 1211 to be the edge of the subnet managed by the SM, the SM cannot send discovery and / or management messages beyond this point (for example, to router port II 1212).
[0109] According to one embodiment, the dual-port virtual router offers the advantage that the dual-port virtual router abstraction is fully managed by a management entity (e.g., SM or SMA) within the subnet to which the dual-port virtual router belongs. By making management local only, the system does not need to provide an external, independent management entity. In other words, each side of the inter-subnet connection plays a role in configuring its own dual-port virtual router.
[0110] According to one embodiment, if a packet such as an SMP packet destined for a remote destination (i.e., outside the local subnet) arrives at a local target port that is not configured via the dual-port virtual router, the local port can return a message indicating that it is not a router port.
[0111] Many of the features of this disclosure can be implemented using or with the support of hardware, software, firmware, or a combination thereof. Therefore, the features of this disclosure can be implemented using a processing system (for example, including one or more processors).
[0112] Figure 13 shows a method for supporting a dual-port virtual router in a high-performance computing environment according to one embodiment. In step 1310, this method can provide a first subnet on one or more computers, each containing one or more microprocessors. The first subnet includes a plurality of switches, each of which includes at least a leaf switch, and each of which includes a plurality of switch ports. The first subnet further includes a plurality of host channel adapters, each containing at least one host channel adapter port, a plurality of end nodes, each associated with at least one host channel adapter among the plurality of host channel adapters, and a subnet manager, the subnet manager running on one of the plurality of switches and the plurality of host channel adapters.
[0113] In step 1320, this method allows one of several switch ports on one of several switches to be configured as a router port.
[0114] In switch 1330, this method allows a switch port configured as a router port to be logically connected to a virtual router, which includes at least two virtual router ports.
[0115] Service Quality and Service Level Agreement in Private Fabrics In one embodiment, a high-performance computing environment in the cloud, as well as in larger clouds at customer and on-premises locations, such as a switched network operating on InfiniBand or RoCE, is capable of deploying virtual machine (VM)-based workloads, and a specific requirement is the ability to define and control quality of service (QoS) for different types of communication flows. In addition, workloads belonging to different tenants must run within the boundaries of the relevant service level agreements (SLAs), while minimizing interference between such workloads and maintaining QoS assumptions for different communication types.
[0116] RDMA readout as a restricted feature (ORA20Q246-US-NP-1) According to one embodiment, when defining bandwidth limits in a system using conventional network interfaces (NICs), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.
[0117] However, in RDMA-based networking, according to one embodiment, different nodes can generate RDMA read requests (i.e., egress bandwidth), which may represent a small amount of egress bandwidth. However, such RDMA read requests may have the potential to represent a very large amount of ingress RDMA traffic in response to such requests. In such a situation, limiting the egress bandwidth of all nodes / VMs to control the overall traffic generation in the system is no longer sufficient. isn't it.
[0118] According to one embodiment, it is possible to limit total bandwidth utilization while limiting only the transmit (egress) bandwidth for untrusted nodes / VMs by restricting RDMA read operations and allowing such read requests only for nodes / VMs that are trusted not to generate excessive RDMA read-based ingress bandwidth.
[0119] Figure 14 shows a system according to one embodiment for providing RDMA read requests as restricted features in a high-performance computing environment.
[0120] More specifically, according to one embodiment, Figure 14 shows a host channel adapter 1401 including a hypervisor 1411. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 1414-1416, and physical functions (PFs) 1413. The host channel adapter may further support or have several ports, such as ports 1402 and 1403, used to connect the host channel adapter to a network, such as network 1400. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 1401 can be connected to several other nodes, such as switches, additional separate HCAs.
[0121] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 1450, VM2 1451, and VM3 1452.
[0122] According to one embodiment, the host channel adapter 1401 can further support a virtual switch 1412 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0123] According to one embodiment, the host channel adapter can implement a trusted RDMA read restriction 1460, which can be configured to block any of the virtual machines (e.g., VM1, VM2, and / or VM3) from sending any RDMA read requests to the network (e.g., via port 1402 or 1403).
[0124] According to one embodiment, the trusted RDMA read restriction 1460 can achieve blocking at the host channel adapter level from certain types of packets from certain endpoints, such as virtual machines, or from other physical nodes that utilize the HCA 1401 to connect to the network, from generating (i.e., egressing) RDMA read request packets. This configurable restriction component 1460 can, for example, allow only trusted nodes (e.g., VMs or physical end nodes) to generate such types of packets.
[0125] According to one embodiment, the trusted RDMA read restriction component can be configured based on instructions received, for example, by a host channel adapter, or it can be configured directly by, for example, a subnet manager (not shown).
[0126] Figure 15 shows a system according to one embodiment for providing RDMA read requests as restricted features in a high-performance computing environment.
[0127] More specifically, according to one embodiment, Figure 15 shows a host channel adapter 1501 including a hypervisor 1511. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 1514-1516, and physical functions (PFs) 1513. In addition, the host channel adapter may support or have several ports, such as ports 1502 and 1503, which are used to connect the host channel adapter to a network, such as network 1500. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 1501 can be connected to several other nodes, such as switches, additional separate HCAs.
[0128] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 1550, VM2 1551, and VM3 1552.
[0129] According to one embodiment, the host channel adapter 1501 can further support a virtual switch 1512 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0130] According to one embodiment, the host channel adapter can implement a trusted RDMA read restriction 1560, which can be configured to block any of the virtual machines (e.g., VM1, VM2, and / or VM3) from sending any RDMA read requests to the network (e.g., via port 1502 or 1503).
[0131] According to one embodiment, the trusted RDMA read restriction 1560 can achieve blocking at the host channel adapter level from certain types of packets from certain endpoints, such as virtual machines, or from other physical nodes that utilize the HCA 1501 to connect to the network, from generating (i.e., egressing) RDMA read request packets. This configurable restriction component 1560 can, for example, allow only trusted nodes (e.g., VMs or physical end nodes) to generate such types of packets.
[0132] According to one embodiment, the trusted RDMA read restriction component can be configured based on instructions received, for example, by a host channel adapter, or it can be configured directly by, for example, a subnet manager (not shown).
[0133] In one embodiment, for example, the trusted RDMA read restriction 1560 can be configured to trust VM1 1550 but not VM2 1551. Thus, an RDMA read request 1554 originating from VM1 may be permitted, while an RDMA read request 1555 originating from VM2 may be blocked before it exits and reaches the host channel adapter 1501 (shown outside the HCA in Figure 15, but this is simply for illustrative purposes).
[0134] Figure 16 shows a system according to one embodiment for providing RDMA read requests as restricted features in a high-performance computing environment.
[0135] According to one embodiment, a switched network or high subnet such as subnet 1600 Within the performance computing environment, several end nodes 1601 and 1602 can support several virtual machines VM1-VM4 1650-1653 interconnected via several switches such as leaf switches 1611 and 1612, switches 1621 and 1622, and root switches 1631 and 1632.
[0136] According to one embodiment, various host channel adapters providing functionality for connecting nodes 1601 and 1602, as well as virtual machines to be connected to the subnet, are not shown. Such embodiments are discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0137] In one embodiment, a typical system limits the RDMA egress bandwidth from any one virtual machine on an end node to prevent any single virtual machine from monopolizing the bandwidth of any link connecting an end node to a subnet. However, while such egress bandwidth limitation is effective in the general case, it does not prevent a virtual machine from issuing RDMA read requests such as RDMA read requests 1654 and 1655. This is because such RDMA read requests are generally small packets and do not utilize much egress bandwidth.
[0138] However, according to one embodiment, such an RDMA read request may result in the generation of a large amount of return traffic to issuing entities such as VM1 and VM3. In such a situation, the RDMA read request may then lead to link congestion and degraded network performance, for example, when read request 1654 results in a large amount of data traffic flowing back to VM1 as a result of the execution of the read request at the destination.
[0139] According to one embodiment, this can lead to a loss of subnet performance, particularly in situations where multiple tenants share subnet 1600.
[0140] In one embodiment, each node (or host channel adapter) can be configured with RDMA read restrictions 1660 and 1661 that block any untrusted VM from issuing RDMA read requests to that VM. Such RDMA read restrictions can vary from always blocking the issuance of RDMA read requests to a limit that places a time frame when a virtual machine configured with RDMA read request restrictions can issue RDMA read requests (for example, during periods of slow network traffic). In addition, RDMA read restrictions 1660 and 1661 may further allow trusted VMs to issue RDMA read requests.
[0141] In one embodiment, a scenario is conceivable in which multiple VMs / tenants share a “new” HCA, i.e., an HCA that has support for the relevant new features, but are making RDMA requests to a remote “old” HCA that does not have such support. In such a scenario, it would be meaningful to have a way to limit the ingress bandwidth that such VMs can generate with respect to RDMA read responses without relying on a static rate configuration on the “old” RDMA read response-side HCA. There is no simple way to do this, as long as VMs are allowed to generate “arbitrary” RDMA read sizes. Also, since multiple RDMA read requests generated over a period of time may all receive response data simultaneously, it is impossible to guarantee that the ingress bandwidth will not exceed the maximum bandwidth beyond a very limited time unless there is both a limit on the RDMA read size that can be generated in a single request and a limit on the total number of pending RDMA read requests from the same vHCA port.
[0142] Therefore, according to one embodiment, if a maximum read size is defined for the vHCA, bandwidth control may be based on an allocation to the sum of all outstanding read sizes, or, in a simpler scheme, simply limit the maximum number of outstanding RDMA reads based on the “worst-case” read size. Thus, in either case, there is no limit on peak bandwidth within short intervals (except for the maximum link bandwidth of the HCA port), but the duration of such peak bandwidth “windows” will be limited. However, in addition, assuming that responses with data are received at the same rate, the transmission rate of RDMA read requests must also be throttled so that the transmission rate of requests does not exceed the maximum allowable ingress rate. In other words, the maximum outstanding request limit defines the worst-case short-interval bandwidth, and the request transmission rate limit will ensure that new requests cannot be generated immediately after a response is received, but only after an associated delay representing the allowable average ingress bandwidth for RDMA read responses. Thus, in the worst case, a permitted number of requests are sent without any responses, and then all of these responses are received at the “same time.” At this point, the next request can be sent immediately upon arrival of the first response, but the next request must be delayed for a specified delay period. Therefore, over time, the average ingress bandwidth cannot exceed what is defined by the request rate. However, the smaller the maximum number of pending requests, the lower the possible "variability".
[0143] Explicit use of RDMA read bandwidth limiting (ORA200246-US-NP-1) According to one embodiment, when defining bandwidth limits in a system using conventional network interfaces (NICs), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.
[0144] However, in RDMA-based networking, according to one embodiment, different nodes can generate RDMA read requests that represent small request messages but potentially very large response messages, and limiting the egress bandwidth of all nodes / VMs to control the total traffic generation in the system is no longer sufficient.
[0145] According to one embodiment, it is possible to control the total traffic generation in the system without relying on restricting the use of RDMA reads by untrusted nodes / VMs by defining an explicit allocation of how much RDMA read ingress bandwidth a node / VM is allowed to generate, independently of any transmit / egress bandwidth limits.
[0146] According to one embodiment, the system and method can support, in addition to supporting average ingress bandwidth utilization resulting from locally generated RDMA read requests, the duration / length of the worst-case maximum link bandwidth burst (i.e., as a result of RDMA read responses "stacking up").
[0147] Figure 17 shows a system, according to one embodiment, for providing explicit RDMA read bandwidth limiting in a high-performance computing environment.
[0148] More specifically, according to one embodiment, Figure 17 shows a host channel adapter 1701 including a hypervisor 1711. The hypervisor can host or associate with several virtual functions (VFs), such as VFs 1714-1716, and physical functions (PFs) 1713. The host channel adapter has port 17 used to connect the host channel adapter to a network such as network 1700. It may further support or have several ports such as 02 and 1703. The network may have a switched network such as an InfiniBand network or a RoCE network, where the HCA1701 can be connected to several other nodes such as switches, additional separate HCAs, etc.
[0149] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 1750, VM2 1751, and VM3 1752.
[0150] According to one embodiment, the host channel adapter 1701 can further support a virtual switch 1712 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0151] According to one embodiment, the host channel adapter can implement an RDMA read limit 1760, which can be configured to impose an allocation on the amount of ingress bandwidth that any VM (of the HCA 1701) can generate with respect to responses to RDMA read requests sent by a particular VM. Such ingress bandwidth limiting is performed locally in the host channel adapter.
[0152] According to one embodiment, the RDMA read restriction component can be configured based on instructions received, for example, by a host channel adapter, or it can be configured directly by, for example, a subnet manager (not shown).
[0153] Figure 18 shows a system, according to one embodiment, for providing explicit RDMA read bandwidth limiting in a high-performance computing environment.
[0154] More specifically, according to one embodiment, Figure 18 shows a host channel adapter 1801 including a hypervisor 1811. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 1814-1816, and physical functions (PFs) 1813. The host channel adapter may further support or have several ports, such as ports 1802 and 1803, used to connect the host channel adapter to a network, such as network 1800. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 1801 can be connected to several other nodes, such as switches, additional separate HCAs.
[0155] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 1850, VM2 1851, and VM3 1852.
[0156] According to one embodiment, the host channel adapter 1801 can further support a virtual switch 1812 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0157] According to one embodiment, the host channel adapter achieves an RDMA read limit of 1860. This allows the read limit 1860 to be configured to impose an allocation on the amount of ingress bandwidth that any VM (of HCA1701) can generate with respect to responses to RDMA read requests sent by a particular VM. Limiting such ingress bandwidth is performed locally in the host channel adapter.
[0158] According to one embodiment, the RDMA read restriction component can be configured based on instructions received, for example, by a host channel adapter, or it can be configured directly by, for example, a subnet manager (not shown).
[0159] According to one embodiment, for example, VM1 may have previously sent at least two RDMA read requests, requesting that the read operation be performed on the connected node. In response, VM1 may be in the process of receiving multiple responses to the RDMA read requests, shown in the figure as RDMA read responses 1855 and 1854. Since these RDMA read responses can be very large, especially compared to the initial RDMA read request sent by VM1, these read responses 1854 and 1855 may be subject to the RDMA read limit 1860, and the ingress bandwidth may be limited or throttled. This throttling may be based on an explicit ingress bandwidth limit, or on the QoS and / or SLA of VM1 set within the RDMA limit 1860.
[0160] Figure 19 shows a system for providing explicit RDMA read bandwidth limiting in a high-performance computing environment according to one embodiment.
[0161] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 1900, several end nodes 1901 and 1902 can support several virtual machines VM1 to VM4 1950 to 1953 interconnected via several switches such as leaf switches 1911 and 1912, switches 1921 and 1922, and root switches 1931 and 1932.
[0162] According to one embodiment, various host channel adapters providing functionality for connecting nodes 1901 and 1902, as well as virtual machines to be connected to the subnet, are not shown. Such embodiments are discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0163] In one embodiment, a typical system limits the RDMA egress bandwidth from any single virtual machine on an end node to prevent any single virtual machine from monopolizing the bandwidth of any link connecting an end node to a subnet. However, while such egress bandwidth limiting is effective in the general case, it cannot prevent an influx of RDMA read responses from monopolizing the link between the requesting VM and the network.
[0164] In other words, according to one embodiment, if VM1 sends out several RDMA read requests, VM1 cannot control when responses to such read requests are returned to VM1. This can result in a back-up / stack of responses to RDMA read requests, each attempting to use the same link to return the requested information to VM1 (via RDMA read response 1954). This results in network congestion and backlogs of traffic.
[0165] According to one embodiment, RDMA limits 1960 and 1961 can impose an allocation on the amount of ingress bandwidth that a given VM can generate with respect to responses to RDMA read requests sent by a particular VM. Limiting such ingress bandwidth is performed locally.
[0166] In one embodiment, if a maximum read size is defined for a vHCA, bandwidth control may be based on an allocation to the sum of all outstanding read sizes, or, in a simpler scheme, simply limiting the maximum number of outstanding RDMA reads based on the “worst-case” read size. Thus, in either case, there is no limit on peak bandwidth within short intervals (except for the maximum link bandwidth of the HCA port), but the duration of such peak bandwidth “windows” will be limited. However, in addition, assuming that responses with data are received at the same rate, the transmission rate of RDMA read requests must also be throttled so that the transmission rate of requests does not exceed the maximum allowable ingress rate. In other words, the maximum outstanding request limit defines the worst-case short-interval bandwidth, and the request transmission rate limit will ensure that new requests cannot be generated immediately upon receipt of a response, but only after an associated delay representing the allowable average ingress bandwidth for RDMA read responses. Thus, in the worst case, a permitted number of requests are sent without any responses, and then all of these responses are received at the “same time.” At this point, the next request can be sent immediately upon arrival of the first response, but the next request must be delayed for a specified delay period. Therefore, over time, the average ingress bandwidth cannot exceed what is defined by the request rate. However, the smaller the maximum number of pending requests, the lower the possible "variability".
[0167] Figure 20 is a flowchart of a method for providing RDMA (Remote Direct Memory Access) read requests as restricted features in a high-performance computing environment, according to one embodiment.
[0168] According to one embodiment, in step 2010, the method can provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches and a plurality of host channel adapters, each host channel adapter comprising at least one host channel adapter port, and the plurality of host channel adapters being interconnected via the plurality of switches.
[0169] According to one embodiment, in step 2020, the method can provide multiple end nodes, each containing multiple virtual machines.
[0170] According to one embodiment, in step 2030, the method can associate a host channel adapter with a selective RDMA restriction.
[0171] According to one embodiment, in step 2040, the method can host one of several virtual machines in a host channel adapter that includes selective RDMA restriction.
[0172] Combining multiple shared bandwidth segments (ORA20Q246-US-NP-3) According to one embodiment, conventional bandwidth / rate limiting schemes for network interfaces are typically limited to a combination of the overall aggregated transmit rate and, in some cases, the maximum rate for individual destinations. However, in many cases, there is a shared bottleneck in the intermediate network / fabric topology, and the target This means that the total bandwidth available to the set is limited by this shared bottleneck. Therefore, if such a shared bottleneck is not taken into consideration when determining what rates various data flows can be sent, the shared bottleneck is likely to become overloaded, even if rate limits for each target are adhered to.
[0173] According to one embodiment, the system and method herein can introduce an object “target group” that can associate multiple individual flows, and this target group can represent rate limits on individual (potentially shared) links or other bottlenecks in the network / fabric path used by the flows. Furthermore, the system and method can enable each flow to be associated with a hierarchy of such target groups so that it can represent all link segments and any other (shared) bottlenecks in the path between the source and target for each individual flow.
[0174] According to one embodiment, in order to limit egress bandwidth, the system and method can establish groups of destinations that share bandwidth allocations to reduce the possibility of congestion on a shared ISL (Inter-Switch Link). This requires a destination / route-related lookup mechanism that can be managed regarding which destination / route maps to which group at a logical level. This means that the hyper-privileged communications infrastructure must be aware of the actual location of peer nodes in the fabric topology, as well as the relevant routing and capacity information that can be mapped to “target groups” (i.e., HCA-level object types) within the local HCA with associated bandwidth allocations. However, it is not practical to have the hardware perform a direct lookup of WQE (Work Queue Entry) / packet address information to map to the relevant target groups. Instead, the HCA implementation can provide an association between RC (Trusted Connected) QP (Queue Pair) and address handles that represent the transmission context for outgoing traffic and the relevant target groups. In this way, this association is transparent at the verbs level. It may be set up at the hyper-privileged software level and then implemented at the HCA HW (and firmware) level. A significant additional complexity associated with this scheme is that live VM migrations, where the relevant VM or vHCA port address information is maintained across migrations, may still mean that there are changes in target groups for different communication peers. However, the target group associations do not need to be updated synchronously, as long as the system and method allow for some transient period during which the relevant bandwidth allocation is not 100% accurate. Thus, logical connectivity and the ability to communicate may not change due to VM migration, but the target groups associated with the RC connections and address handles in both the migrated VM and its communication peer VMs may be "completely wrong" after the migration. This may mean both that less bandwidth is used than is available (e.g., when a VM is moved from a remote location to the same "leaf group" as its peer) and that excess bandwidth is generated (e.g., when a VM is moved from the same "leaf group" as its peer to a remote location that implies a shared ISL with limited bandwidth).
[0175] According to one embodiment, in order to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric that the target group represents, the target group-specific bandwidth allocation may, in principle, also be divided into allocations for specific priorities ("QoS classes").
[0176] According to one embodiment, a target group can send objects to a specific destination address. By decoupling, the system and method gain the ability to represent intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits, in addition to the target limits.
[0177] In one embodiment, this system and method can consider a hierarchy of target groups (bandwidth allocation) that reflect bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if the target limit is 30 Gb / s and the intermediate uplink limit is 50 Gb / s, the maximum rate toward the target can never exceed 30 Gb / s. On the other hand, if multiple 30 Gb / s targets share the same 50 Gb / s intermediate limit, using the relevant target rate limit for flows toward these targets may mean an overrun of the intermediate rate limit. Therefore, to ensure the best possible utilization and throughput within the relevant limits, all target groups in the relevant hierarchy can be considered in their relevant strict order. This means that packets can only be sent toward their relevant destinations if each target group in the hierarchy represents available bandwidth. Thus, in the above example, if a single flow is active toward one of the targets, this flow will be allowed to operate at 30 Gb / s. However, if another flow becomes active (via a shared intermediate target group) toward another target, each flow will be limited to 25 Gb / s. In the next round, if an additional flow toward one of the two targets becomes active, the two flows toward the same target will each be operating at 12.5 Gb / s (i.e., on average, and unless they have any additional bandwidth allocation / limitations).
[0178] In one embodiment, when multiple tenants share a server / HCA, both the initial egress bandwidth and the actual target bandwidth may be shared in addition to any sharing of intermediate ISL bandwidth. On the other hand, in a scenario where each tenant has a dedicated server / HCA, the intermediate ISL bandwidth represents the only possible “inter-tenant” bandwidth sharing.
[0179] In one embodiment, target groups should typically be global for HCA ports, and VF / tenant assignments at the HCA level would represent the maximum local traffic a tenant can generate for any combination of targets, either globally or for specific priorities. Furthermore, it would also be possible to use target groups specific to certain tenants alongside the "global" target groups within the same hierarchy.
[0180] According to one embodiment, there are several possible ways to implement target groups and represent target group associations (hierarchies) for a particular QP or address handle. However, a 16-bit target group ID space can be provided, as well as support for up to four or eight target group associations for each QP and address handle. Each target group ID value would then represent some hardware state that reflects the associated IPD (inter-packet delay) value for the associated rate, as well as timer information defining when the next packet associated with this target group may be transmitted.
[0181] According to one embodiment, different flows / paths can use different "QoS IDs" (i.e., service levels, priorities, etc.) on the same shared link segment, so it is also possible to associate different target groups on the same link segment so that different target groups represent bandwidth allocations for such different QoS IDs. It is also possible to represent both a target group specific to a QoS ID and a single target group representing a physical link in the same link segment.
[0182] In one embodiment, the system and method can also implement different “sub-assignments” to mediate between different such flow types by distinguishing between different flow types defined by an explicit flow type packet header parameter and / or by taking into account the operation type (e.g., RDMA read / write / transmit). In particular, this may be useful in distinguishing flows that represent responding mode bandwidth to requesting mode traffic initially initiated by the local node itself (i.e., typically RDMA read response traffic).
[0183] According to one embodiment, it is possible, in principle, to avoid "any" congestion by strictly using target groups and rate limiting all relevant transmitting HCAs up to a total maximum rate that does not exceed the capacity of any target or shared ISL segment. However, this may mean strict limits on both sustained bandwidth for different flows and low average utilization of available link bandwidth. Thus, various rate limits may be set to allow different HCAs to use more optimistic maximum rates. In this case, the aggregated sum may be greater than the sustainable maximum and therefore may lead to congestion.
[0184] Figure 21 shows a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment.
[0185] More specifically, according to one embodiment, Figure 21 shows a host channel adapter 2101 including a hypervisor 2111. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2114-2116, and physical functions (PFs) 2113. The host channel adapter may further support or have several ports, such as ports 2102 and 2103, used to connect the host channel adapter to a network, such as network 2100. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2101 can be connected to several other nodes, such as switches, additional separate HCAs, etc.
[0186] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 2150, VM2 2151, and VM3 2152.
[0187] According to one embodiment, the host channel adapter 2101 can further support a virtual switch 2112 via the hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0188] According to one embodiment, the network 2100 may comprise several switches, such as switches 2140, 2141, 2142, and 2143, which are interconnected and can be connected to a host channel adapter 2101 via, for example, leaf switches 2140 and 2141, as shown in the figure.
[0189] According to one embodiment, switches 2140-2143 can be interconnected, and furthermore, other switches and other end nodes (e.g., other HCAs) not shown in the figure can be interconnected. It can be connected to ).
[0190] According to one embodiment, target groups such as target groups 2170 and 2171 can be defined along inter-switch links (ISLs), such as ISLs between leaf switch 2140 and switch 2142, and between leaf switch 2141 and switch 2143. These target groups 2170 and 2171 can represent bandwidth allocations as HCA objects, stored in a target group repository 2161 associated with the HCA, which is accessible, for example, by rate limiting component 2160.
[0191] According to one embodiment, target groups 2170 and 2171 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0192] In one embodiment, target groups 2170 and 2171 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2151 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2170 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0193] In one embodiment, target groups can be inherently hierarchical, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2170 represents a higher bandwidth limit than target group 2171, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2171 is the control bandwidth limiting factor.
[0194] In one embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided depending on the QoS and SLA associated with each flow. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2170, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2170 would represent a 5 Gb / s limit for each flow. This sharing or division of target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0195] Figure 22 shows a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment.
[0196] More specifically, according to one embodiment, Figure 22 shows a host channel adapter 2201 including a hypervisor 2211. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2214-2216, and physical functions (PFs) 2213. The host channel adapter may further support or have several ports, such as ports 2202 and 2203, used to connect the host channel adapter to a network, such as network 2200. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2201 can be connected to several other nodes, such as switches, additional separate HCAs, etc.
[0197] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 2250, VM2 2251, and VM3 2252.
[0198] According to one embodiment, the host channel adapter 2201 can further support a virtual switch 2212 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0199] According to one embodiment, the network 2200 may comprise several switches, such as switches 2240, 2241, 2242, and 2243, which are interconnected and can be connected to a host channel adapter 2201 via, for example, leaf switches 2240 and 2241, as shown in the figure.
[0200] According to one embodiment, switches 2240-2243 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0201] According to one embodiment, target groups such as target groups 2270 and 2271 can be defined, for example, on a switch port. As shown in the figure, target groups 2270 and 2271 are defined on the switch ports of switches 2242 and 2243, respectively. These target groups 2270 and 2271 can represent bandwidth allocations as HCA objects, stored in a target group repository 2261 associated with the HCA, which is accessible, for example, by the rate limiting component 2260.
[0202] According to one embodiment, target groups 2270 and 2271 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0203] According to one embodiment, target groups 2270 and 2271 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2251 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2270 which sets a lower bandwidth limit, then the egress bandwidth from VM2 will be less than the default / original egress limit imposed on VM2. It may be limited to a level lower than the limit. The HCA can be responsible for adjusting such throttling / egress bandwidth limits depending on the target group involved in routing packets from VM2, for example.
[0204] In one embodiment, target groups can be inherently hierarchical, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2270 represents a higher bandwidth limit than target group 2271, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2271 is the control bandwidth limiting factor.
[0205] In one embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided depending on the QoS and SLA associated with each flow. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2270, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2270 would represent a 5 Gb / s limit for each flow. This sharing or division of target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0206] According to one embodiment, Figures 21 and 22 show target groups defined at an inter-switch link and a switch port, respectively. Those skilled in the art will readily understand that target groups can be defined at various locations within a subnet, and that no given subnet is limited to having target groups defined only at ISLs and switch ports, but generally such target groups can be defined at both ISLs and switch ports within any given subnet.
[0207] Figure 23 shows a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment.
[0208] More specifically, according to one embodiment, Figure 23 shows a host channel adapter 2301 including a hypervisor 2311. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2314-2316, and physical functions (PFs) 2313. The host channel adapter may further support or have several ports, such as ports 2302 and 2303, used to connect the host channel adapter to a network, such as network 2300. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2301 can be connected to several other nodes, such as switches, additional separate HCAs.
[0209] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 2350, VM2 2351, and VM3 2352.
[0210] According to one embodiment, the host channel adapter 2301 can further support the virtual switch 2312 via the hypervisor. This is a vSwitch adapter This is due to the circumstances under which the architecture is realized. Although not shown in the figures, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0211] According to one embodiment, the network 2300 may comprise several switches, such as switches 2340, 2341, 2342, and 2343, which are interconnected and can be connected to a host channel adapter 2301 via, for example, leaf switches 2340 and 2341, as shown in the figure.
[0212] According to one embodiment, switches 2340-2343 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0213] According to one embodiment, target groups such as target groups 2370 and 2371 can be defined along inter-switch links (ISLs), such as ISLs between leaf switch 2340 and switch 2342, and between leaf switch 2341 and switch 2343. These target groups 2370 and 2371 can represent bandwidth allocations as HCA objects, stored in a target group repository 2361 associated with the HCA, which is accessible, for example, by rate limiting component 2360.
[0214] According to one embodiment, target groups 2370 and 2371 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0215] According to one embodiment, target groups 2370 and 2371 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2351 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2370 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0216] In one embodiment, target groups can be inherently hierarchical, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2370 represents a higher bandwidth limit than target group 2371, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2371 is the control bandwidth limiting factor.
[0217] According to one embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with each flow, the target Bandwidth allocations represented by groups can be divided. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2370, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2370 would represent a 5 Gb / s limit for each flow. This sharing or division of target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0218] In one embodiment, the target group repository may query the target group 2370 2375 to determine, for example, the bandwidth allocation of the target group. Once the bandwidth allocation of the target group is determined, the target group repository may store the allocation value associated with the target group. This allocation may then be used by the rate limiting component as follows: a) determine, based on QoS or SLA, whether the bandwidth allocation of the target group is lower than that of the VM's bandwidth allocation, and b) in such determination, update the VM's bandwidth allocation 2376 based on the path across the target group 2370.
[0219] Figure 24 shows a system for combining multiple shared bandwidth segments in a high-performance computing environment according to one embodiment.
[0220] More specifically, according to one embodiment, Figure 24 shows a host channel adapter 2401 including a hypervisor 2411. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2414-2416, and physical functions (PFs) 2413. The host channel adapter may further support or have several ports, such as ports 2402 and 2403, used to connect the host channel adapter to a network, such as network 2400. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2401 can be connected to several other nodes, such as switches, additional separate HCAs, etc.
[0221] According to one embodiment, as described above, each of the virtual functions is VM1 2450, VM2 It can host virtual machines (VMs) such as 2451, VM3, and 2453.
[0222] According to one embodiment, the host channel adapter 2401 can further support a virtual switch 2412 via the hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0223] According to one embodiment, the network 2400 may comprise several switches, such as switches 2440, 2441, 2442, and 2443, which are interconnected and can be connected to a host channel adapter 2401 via, for example, leaf switches 2440 and 2441, as shown in the figure.
[0224] According to one embodiment, switches 2440-2443 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0225] According to one embodiment, target groups such as target groups 2470 and 2471 can be defined, for example, on a switch port. As shown in the figure. Target groups 2470 and 2471 are defined on the switch ports of switches 2442 and 2443, respectively. These target groups 2470 and 2471 can represent bandwidth allocations as HCA objects, stored in a target group repository 2461 associated with the HCA, which is accessible, for example, by the rate limiting component 2460.
[0226] According to one embodiment, target groups 2470 and 2471 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0227] According to one embodiment, target groups 2470 and 2471 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2451 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2470 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0228] In one embodiment, target groups can be hierarchical in nature, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2470 represents a higher bandwidth limit than target group 2471, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2471 is the control bandwidth limiting factor.
[0229] In one embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided depending on the QoS and SLA associated with each flow. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2470, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2470 would represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0230] According to one embodiment, the target group repository can perform a query 2475 on the target group 2470, for example, to determine the bandwidth allocation of the target group. Once the bandwidth allocation of the target group is determined, the target group repository may store the allocation value associated with the target group. This allocation can then be used by the rate limiting component as follows: a) Determine whether the bandwidth allocation of the target group is lower than that of the VM's bandwidth allocation based on QoS or SLA, and b) In such a determination, update 2476 the VM's bandwidth allocation based on the path across the target group 2470. Based on this, update the bandwidth allocation of the VM 2476.
[0231] FIG. 25 shows a system for combining a plurality of shared bandwidth segments in a high-performance computing environment according to an embodiment.
[0232] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 2500, some end nodes 2501 and 2502 can support some virtual machines VM1 to VM4 2550 to 2553 interconnected via some switches such as leaf switches 2511 and 2512, switches 2521 and 2522, and root switches 2531 and 2532.
[0233] According to one embodiment, various host channel adapters providing functions for connecting nodes 2501 and 2502, and the virtual machines to be connected to the subnet are not shown. The discussion of such embodiments has been described above with respect to SR-IOV, and each virtual machine can be associated with the virtual function of the hypervisor on the host channel adapter.
[0234] In one embodiment, as discussed above, a concept unique to such a switched fabric is that each end node or VM may have its own egress / ingress bandwidth limits that traffic entering and leaving it must adhere to, while within a subnet, there may also be links or ports that represent bottlenecks for traffic entering them. Therefore, when determining what rates traffic should enter and leave such end nodes, such as VM1, VM2, VM3, or VM4, rate limiting components 2560 and 2561 can query various target groups, such as 2550 and 2551, to determine whether such target groups represent bottlenecks for traffic flow. In response to such determinations, rate limiting components 2560 and 2561 can then set different or new bandwidth limits on the endpoints controlled by the rate limiting components.
[0235] According to one embodiment, if the target groups utilize both target groups 2550 and 2551 for traffic from VM1 to VM3, rate limit 2560 can be queried in a nested / hierarchical manner so that it can take into account the limits from both such target groups when determining the bandwidth limit from VM1 to VM3.
[0236] Figure 26 is a flowchart illustrating a method for supporting target groups for congestion control in a private fabric in a high-performance computing environment, according to one embodiment.
[0237] According to one embodiment, in step 2610, the method can provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches, each of the plurality of switches comprising at least a leaf switch, each of the plurality of switches comprising a plurality of switch ports, the first subnet further comprising a plurality of host channel adapters, each of the host channel adapters comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes comprising a plurality of virtual machines.
[0238] According to one embodiment, in step 2620, the method may define a target group on an inter-switch link between two switches of a plurality of switches or on at least one port of a switch of a plurality of switches, wherein the target group defines a bandwidth limit on an inter-switch link between two switches of a plurality of switches or on at least one port of a switch of a plurality of switches.
[0239] According to one embodiment, in step 2630, the method can provide a target group repository stored in the memory of the host channel adapter.
[0240] According to one embodiment, in step 2640, the method can record the defined target group in the target group repository.
[0241] Target-specific transmit / RDMA write and RDMA read bandwidth limit combinations (ORA200246-US-NP-2) In one embodiment, a node / VM can target incoming data traffic, which is the result of both transmit and RDMA write operations initiated by peer nodes / VMs and RDMA read operations initiated by the local node / VM itself. In such a situation, the issue becomes ensuring that the local node / VM's maximum or average ingress bandwidth is within the required boundaries, unless all of these flows are coordinated with respect to rate limiting.
[0242] In one embodiment, the systems and methods described herein can achieve target-specific egress rate control in a manner that enables all flows representing data fetching from local memory and transmission of data to associated remote targets to be subject to the same shared rate limiting and associated flow scheduling and arbitration. Furthermore, different flow types may be given different priorities and / or different shares of the available bandwidth.
[0243] According to one embodiment, the association of a target group with respect to flows from a “producer / source” node means bandwidth throttling for all outgoing data packets, including UD (Untrusted Datagram) transmissions, RDMA writes, RDMA transmissions, and RDMA reads (i.e., RDMA read responses with data). This is independent of whether the VM owning the target vHCA port is generating an “excessive” amount of RDMA read requests to multiple peer nodes.
[0244] According to one embodiment, coupling a target group to both flow-specific and "unclaimed" BECN signaling means that the ingress bandwidth per vHCA port can be dynamically throttled for any number of remote peers.
[0245] According to one embodiment, the “unsolicited BECN” message can also be used to communicate specific rate values in addition to pure CE flagging / unflaggling for different stage numbers. In this way, it is possible to have a scheme in which an initial incoming packet from a new peer (e.g., a communications management (CM) packet) can trigger the generation of one or more “unsolicited BECN” messages to both the HCA (i.e., the associated firmware / hyper-privileged software) from which the incoming packet originated and the current communications peer.
[0246] According to one embodiment, there is a case where both ports on the HCA are used simultaneously (i.e.) In an active-active scheme, sharing target groups across local HCA ports can make sense if concurrent flows may share several ISLs or even target the same destination port.
[0247] According to one embodiment, another reason for sharing target groups between HCA ports is whether the HCA local memory bandwidth can maintain full-speed link speed for both (all) HCA ports. In this case, the target group can be configured such that the aggregated total link bandwidth never exceeds the local memory bandwidth, regardless of whether each port is involved with the source HCA or the destination HCA.
[0248] In one embodiment, for a fixed route destined for a specific destination, any intermediate target group would typically represent only a single ISL at a particular stage in the path. However, if dynamic forwarding is active, both target groups and ECN processing must take this into account. If dynamic forwarding decisions arise only to balance traffic between parallel ISLs between a pair of switches (e.g., an uplink from a single leaf switch to a single spine switch), then all processing is, in principle, very similar to the case where only a single ISL is used. ECN notifications would presumably be based on the state of all ports in the group in question, and signaling could be “aggressive” in the sense that it signals based on congestion indications from any of the ports, or it could be more conservative and based on the size of the shared egress queue for all ports in the group. The target group configuration would typically represent the aggregated bandwidth for all links in the group, insofar as it allows forwarding to select the best egress port for any packet at that moment. However, if there is a concept of strict packet ordering per flow, evaluating bandwidth allocation becomes more complex because some flows “must” use the same ISL at some point. If such a flow sequencing scheme is based on well-defined header fields, it may be best to represent each port within a group as an independent target group. In this case, the selection of target groups at the source HCA must be such that the evaluation of the header fields associated with the RC QP connection or address handle is the same as that performed by the switch at runtime for all packets.
[0249] According to one embodiment, by default, the initial target group rate for a new remote target may be set conservatively low. In this way, there is inherent throttling until the target has an opportunity to update the associated rate. Thus, all such rate control is independent of the VMs involved, although a VM could request the hypervisor to update different remote peer assignments for both ingress and egress traffic, but this would only be permitted within the total constraints defined for both local and remote vHCA ports.
[0250] Figure 27 shows a system, according to one embodiment, for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment.
[0251] More specifically, according to one embodiment, Figure 27 shows a host channel adapter 2701 including a hypervisor 2711. The hypervisor can host or associate with several virtual functions (VFs), such as VFs 2714-2716, and physical functions (PFs) 2713. The host channel adapter has port 27 used to connect the host channel adapter to a network such as network 2700. It may further support or have several ports such as 02 and 2703. The network may have a switched network such as an InfiniBand network or a RoCE network, where the HCA2701 can be connected to several other nodes such as switches, additional separate HCAs, etc.
[0252] According to one embodiment, as described above, each of the virtual functions can host virtual machines (VMs) such as VM1 2750, VM2 2751, and VM3 2752.
[0253] According to an embodiment, the host channel adapter 2701 can further support a virtual switch 2712 via a hypervisor. This is due to the situation where the vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0254] According to an embodiment, the network 2700 can include several switches interconnected as shown, such as switches 2740, 2741, 2742, and 2743, and can be connected to the host channel adapter 2701 via, for example, leaf switches 2740 and 2741.
[0255] According to an embodiment, switches 2740 - 2743 can be interconnected and can further be connected to other switches and other end nodes (such as other HCAs) not shown in the figure.
[0256] According to an embodiment, target groups such as target groups 2770 and 2771 can be defined in an inter - switch link (ISL) such as the ISL between leaf switch 2740 and switch 2742 and between leaf switch 2741 and switch 2743. These target groups 2770 and 2771 can represent bandwidth allocations as HCA objects stored in a target group repository 2761 associated with the HCA, which is accessible, for example, by a rate - limiting component 2760.
[0257] According to an embodiment, target groups 2770 and 2771 can represent specific (and different) bandwidth allocations. These bandwidth allocations can be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for different priorities within the relevant paths in the fabric represented by the target groups.
[0258] According to one embodiment, target groups 2770 and 2771 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2751 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2770 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0259] According to one embodiment, the target group can also be hierarchical in nature, Therefore, this system and method can consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2770 represents a higher bandwidth limit than target group 2771, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2771 is the control bandwidth limiting factor.
[0260] In one embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided depending on the QoS and SLA associated with each flow. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2770, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2770 would represent a 5 Gb / s limit for each flow. This sharing or division of target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0261] In one embodiment, bandwidth allocation and performance issues may arise when a VM, for example VM1 2750, receives excessive ingress bandwidth 2790 from multiple sources. This may occur, for example, when VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, for example, a target group 2770 on an inter-switch link, may be updated, for example, via query 2775, to reflect a lower bandwidth allocation than would typically be allowed.
[0262] According to one embodiment, the HCA's rate limiting component 2760 may further include VM-specific rate limits 2762 that can be negotiated with other peer HCAs to coordinate, for example, an ingress bandwidth limit for VM1 with an egress bandwidth limit for the node responsible for generating ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.
[0263] Figure 28 shows a system, according to one embodiment, for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment.
[0264] More specifically, according to one embodiment, Figure 28 shows a host channel adapter 2801 including a hypervisor 2811. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2814-2816, and physical functions (PFs) 2813. The host channel adapter may further support or have several ports, such as ports 2802 and 2803, used to connect the host channel adapter to a network, such as network 2800. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2801 can be connected to several other nodes, such as switches, additional separate HCAs.
[0265] According to one embodiment, as described above, each of the virtual functions is VM1 2850, VM It can host virtual machines (VMs) such as VM2 2851 and VM3 2852.
[0266] According to one embodiment, the host channel adapter 2801 can further support a virtual switch 2812 via the hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0267] According to one embodiment, the network 2800 may comprise several switches, such as switches 2840, 2841, 2842, and 2843, which are interconnected and can be connected to a host channel adapter 2801 via, for example, leaf switches 2840 and 2841, as shown in the figure.
[0268] According to one embodiment, switches 2840-2843 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0269] According to one embodiment, target groups such as target groups 2870 and 2871 can be defined, for example, on a switch port. As shown in the figure, target groups 2870 and 2871 are defined on the switch ports of switches 2842 and 2843, respectively. These target groups 2870 and 2871 can represent bandwidth allocations as HCA objects, stored in a target group repository 2861 associated with the HCA, which is accessible, for example, by the rate limiting component 2860.
[0270] According to one embodiment, target groups 2870 and 2871 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0271] According to one embodiment, target groups 2870 and 2871 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2851 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2870 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0272] In one embodiment, target groups can be inherently hierarchical, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, target group 2870 represents a higher bandwidth limit than target group 2871, and packets between both switches represented by the two target groups. When addressed via a link, the bandwidth limit for target group 2871 is the control bandwidth limiting factor.
[0273] In one embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided depending on the QoS and SLA associated with each flow. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2870, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2870 would represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0274] In one embodiment, bandwidth allocation and performance issues may arise when a VM, for example VM1 2850, receives excessive ingress bandwidth 2890 from multiple sources. This may occur, for example, when VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, for example, a target group 2870 on an inter-switch link, may be updated, for example, via query 2875, to reflect a lower bandwidth allocation than would typically be allowed.
[0275] According to one embodiment, the HCA's rate limiting component 2860 may further include VM-specific rate limits 2862 that can be negotiated with other peer HCAs to coordinate, for example, an ingress bandwidth limit for VM1 with an egress bandwidth limit for the node responsible for generating ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.
[0276] Figure 29 shows a system, according to one embodiment, for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment.
[0277] More specifically, according to one embodiment, Figure 29 shows a host channel adapter 2901 including a hypervisor 2911. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 2914-2916, and physical functions (PFs) 2913. The host channel adapter may further support or have several ports, such as ports 2902 and 2903, used to connect the host channel adapter to a network, such as network 2900. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 2901 can be connected to several other nodes, such as switches, additional separate HCAs.
[0278] According to one embodiment, as described above, each of the virtual functions can host virtual machines (VMs) such as VM1 2950, VM2 2951, and VM3 2952.
[0279] According to one embodiment, the host channel adapter 2901 can further support a virtual switch 2912 via the hypervisor. This is due to the circumstances under which the vSwitch architecture is realized. Although not shown, embodiments of this disclosure are As mentioned above, this further supports the virtual port (vPort) architecture.
[0280] According to one embodiment, the network 2900 may comprise several switches, such as switches 2940, 2941, 2942, and 2943, which are interconnected and can be connected to a host channel adapter 2901 via, for example, leaf switches 2940 and 2941, as shown in the figure.
[0281] According to one embodiment, switches 2940-2943 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0282] According to one embodiment, a target group such as target group 2971 can be defined along an inter-switch link (ISL), such as an ISL between leaf switch 2941 and switch 2943. Other target groups can be defined, for example, at switch ports. As shown in the figure, target group 2970 is defined at a switch port of switch 2952. These target groups 2970 and 2971 can represent bandwidth allocations as HCA objects, stored in a target group repository 2961 associated with the HCA, which is accessible, for example, by rate limiting component 2960.
[0283] According to one embodiment, target groups 2970 and 2971 may represent specific (and different) bandwidth allocations. These bandwidth allocations may be divided into allocations for specific priorities ("QoS classes") to reflect the expected bandwidth usage for various priorities within the relevant paths in the fabric represented by the target groups.
[0284] According to one embodiment, target groups 2970 and 2971 decouple objects from specific destination addresses, and the system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2951 is set to one threshold, but the destination of packets sent from VM2 will pass through target group 2970 which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustments depending on the target group involved in routing packets from VM2, for example.
[0285] In one embodiment, target groups can be inherently hierarchical, thereby allowing the system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if target group 2970 represents a higher bandwidth limit than target group 2971, and packets are addressed across both inter-switch links represented by the two target groups, then the bandwidth limit of target group 2971 is the control bandwidth limiting factor.
[0286] According to one embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with each flow, the target Bandwidth allocations represented by groups can be divided. For example, if both VM1 and VM2 simultaneously send flows that would involve target group 2970, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has the same QoS and SLA associated with it, then target group 2970 would represent a 5 Gb / s limit for each flow. This sharing or division of target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.
[0287] In one embodiment, bandwidth allocation and performance issues may arise when a VM, for example VM1 2950, receives excessive ingress bandwidth 2990 from multiple sources. This may occur, for example, when VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, for example, a target group 2970 on an inter-switch link, may be updated, for example, via query 2975, to reflect a lower bandwidth allocation than would typically be allowed.
[0288] According to one embodiment, the HCA's rate limiting component 2960 may further include VM-specific rate limits 2962 that can be negotiated with other peer HCAs to coordinate, for example, an ingress bandwidth limit for VM1 with an egress bandwidth limit for the node responsible for generating ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.
[0289] Figure 30 shows a system, according to one embodiment, for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment.
[0290] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3000, several end nodes 3001 and 3002 can support several virtual machines VM1 to VM4 3050 to 3053 interconnected via several switches such as leaf switches 3011 and 3012, switches 3021 and 3022, and root switches 3031 and 3032.
[0291] According to one embodiment, various host channel adapters providing functionality for connecting nodes 3001 and 3002, as well as virtual machines to be connected to the subnet, are not shown. Such embodiments are discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0292] According to one embodiment, a node such as VM3 3052 may enter a bandwidth limit (for example, from rate limiting 3061) while simultaneously processing an RDMA read response 3050 and an RDMA write request 3051 (incoming bandwidth).
[0293] According to one embodiment, rate limits 3060 and 3061 can be configured to ensure that ingress bandwidth allocation is not violated by, for example, coordinating RDMA requests (i.e., messages sent by VM3 to VM4 requesting an RDMA read and resulting in an RDMA read response 3050) and RDMA write operations (e.g., an RDMA write from VM2 to VM3).
[0294] For each individual node, the system and method can have a chain of such target groups so that the flow is always coordinated with all other flows that share link bandwidth in different parts of the fabric represented in the target group.
[0295] Figure 31 is a flowchart of a method for combining target-specific RDMA write bandwidth limits and RDMA read bandwidth limits in a high-performance computing environment, according to one embodiment.
[0296] According to one embodiment, in step 3110, the method can provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches, each of the plurality of switches comprising at least a leaf switch, each of the plurality of switches comprising a plurality of switch ports, the first subnet further comprising a plurality of host channel adapters, each of the host channel adapters comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes comprising a plurality of virtual machines.
[0297] According to one embodiment, in step 3120, the method may define a target group on an inter-switch link between two switches of a plurality of switches or on at least one port of a switch of a plurality of switches, the target group defining a bandwidth limit on an inter-switch link between two switches of a plurality of switches or on at least one port of a switch of a plurality of switches.
[0298] According to one embodiment, in step 3130, the method can provide a target group repository stored in the memory of the host channel adapter.
[0299] According to one embodiment, in step 3140, the method can record the defined target group in the target group repository.
[0300] According to one embodiment, in step 3150, the method allows the host channel adapter end node to receive ingress bandwidth from at least two remote end nodes, and the ingress bandwidth exceeds the end node's ingress bandwidth limit.
[0301] According to one embodiment, in 3160, in response to the reception of ingress bandwidth from at least two sources, the method can update the bandwidth allocation of the target group.
[0302] Combining ingress bandwidth arbitration with congestion feedback (ORA200246-US-NP-2) In one embodiment, when each and / or all of multiple sending nodes / VMs are sending to a single receiving node / VM, achieving an optimal balance of fairness among senders to avoid congestion, while simultaneously limiting the ingress bandwidth usage consumed by the receiving node / VM to a maximum that is (well) below the maximum physical link bandwidth that the relevant network interface can provide for ingress traffic, is not straightforward. Furthermore, the equation becomes even more complex when different senders are assumed to be allocated different bandwidth due to different SLA levels.
[0303] According to one embodiment, the systems and methods of this specification can extend legacy schemes for end-to-end congestion feedback to include both initial negotiation of bandwidth allocation, dynamic adjustment of such bandwidth allocation (e.g., to adapt to changes in the number of transmitting nodes sharing available bandwidth or changes in SLAs), and dynamic congestion feedback to indicate that the transmitter needs to temporarily slow down the relevant egress data rate even though the overall bandwidth allocation remains the same. Relevant information is transmitted from the target node to the transmitting node using both explicit unsolicited messages and "piggyback" information in data packets.
[0304] Figure 32 shows a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to one embodiment.
[0305] More specifically, according to one embodiment, Figure 32 shows a host channel adapter 3201 including a hypervisor 3211. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 3214-3216, and physical functions (PFs) 3213. The host channel adapter may further support or have several ports, such as ports 3202 and 3203, used to connect the host channel adapter to a network, such as network 3200. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 3201 can be connected to several other nodes, such as switches, additional separate HCAs.
[0306] According to one embodiment, as described above, each of the virtual functions can host virtual machines (VMs) such as VM1 3250, VM2 3251, and VM3 3252.
[0307] According to one embodiment, the host channel adapter 3201 can further support a virtual switch 3212 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0308] According to one embodiment, the network 3200 may comprise several switches, such as switches 3240, 3241, 3242, and 3243, which are interconnected and can be connected to a host channel adapter 3201 via, for example, leaf switches 3240 and 3241, as shown in the figure.
[0309] According to one embodiment, switches 3240-3243 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0310] According to one embodiment, bandwidth allocation and performance issues may arise when a VM, for example VM1 3250, receives excessive ingress bandwidth 3290 from multiple sources. This may occur, for example, when VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from an attached VM and one RDMA write request from another attached VM).
[0311] According to one embodiment, the rate limiting component 3260 of the HCA is also For example, the ingress bandwidth limit for VM1 may further include VM-specific rate limits 3261 that can be negotiated with other peer HCAs to coordinate with the egress bandwidth limit for the node responsible for generating ingress bandwidth on VM1. Such initial negotiations may be performed to adapt, for example, to changes in the number of transmitting nodes sharing available bandwidth or changes in the SLA. These other HCAs / nodes are not shown in the diagram.
[0312] According to one embodiment, the above negotiation can be updated based on, for example, an explicit and unsolicited feedback message 3291 generated as a result of ingress bandwidth. Such feedback messages 3291 may be sent, for example, to multiple remote nodes responsible for generating ingress bandwidth 3290 on VM1. Upon receiving such feedback messages, the sending nodes (senders of the bandwidth responsible for the ingress bandwidth on VM1) can update their associated egress bandwidth limits to the sending nodes so as not to overload, for example, the links connected to VM1, while attempting to maintain QoS and SLA.
[0313] Figure 33 shows a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to one embodiment.
[0314] More specifically, according to one embodiment, Figure 33 shows a host channel adapter 3301 including a hypervisor 3311. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 3314-3316, and physical functions (PFs) 3313. The host channel adapter may further support or have several ports, such as ports 3302 and 3303, used to connect the host channel adapter to a network, such as network 3300. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 3301 can be connected to several other nodes, such as switches, additional separate HCAs.
[0315] According to one embodiment, as described above, each of the virtual functions can host virtual machines (VMs) such as VM1 3350, VM2 3351, and VM3 3352.
[0316] According to one embodiment, the host channel adapter 3301 can further support a virtual switch 3312 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0317] According to one embodiment, the network 3300 may comprise several switches, such as switches 3340, 3341, 3342, and 3343, which are interconnected and can be connected to a host channel adapter 3301 via, for example, leaf switches 3340 and 3341, as shown in the figure.
[0318] According to one embodiment, switches 3340-3343 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0319] According to one embodiment, bandwidth allocation and performance issues may arise when a VM, for example VM1 3350, receives excessive ingress bandwidth 3390 from multiple sources. This can occur, for example, when VM1 is performing one or more RDMA write operations simultaneously with one or more RDMA write operations. This can occur when a DMA read response is received and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from an attached VM and one RDMA write request from another attached VM).
[0320] According to one embodiment, the rate limiting component 3360 of the HCA may further include VM-specific rate limits 3361 that can be negotiated with other peer HCAs to coordinate, for example, an ingress bandwidth limit for VM1 with an egress bandwidth limit for the node responsible for generating ingress bandwidth on VM1. Such initial negotiations may be performed to adapt, for example, to changes in the number of transmitting nodes sharing available bandwidth or changes in the SLA. These other HCAs / nodes are not shown in the figure.
[0321] In one embodiment, the above negotiation can be updated based, for example, on a piggyback message 3391 (a message that resides on top of normal data or other communication packets transmitted between end nodes) generated as a result of ingress bandwidth. Such a piggyback message 3391 may be sent, for example, to multiple remote nodes responsible for generating ingress bandwidth 3390 on VM1. Upon receiving such a feedback message, the sending node (the sender of the bandwidth responsible for the ingress bandwidth on VM1) can update its associated egress bandwidth limits to the sending node so as not to overload, for example, the links connected to VM1, while attempting to maintain QoS and SLA.
[0322] Figure 34 shows a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to one embodiment.
[0323] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3400, several end nodes 3401 and 3402 can support several virtual machines VM1 to VM4 3450 to 3453 interconnected via several switches such as leaf switches 3411 and 3412, switches 3421 and 3422, and root switches 3431 and 3432.
[0324] According to one embodiment, various host channel adapters providing functionality for connecting nodes 3401 and 3402, as well as virtual machines to be connected to the subnet, are not shown. Such embodiments are discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0325] According to one embodiment, a node such as VM3 3452 may enter an ingress bandwidth limit (e.g., from rate limit 3461) upon receiving multiple RDMA ingress bandwidth packets (e.g., multiple RDMA writes), such as 3451 and 3452. This can occur, for example, when there is no communication between various sending nodes to adjust the bandwidth limit.
[0326] According to one embodiment, the System and Method of this Specification includes both initial negotiation of bandwidth allocation (i.e., VM3, or the bandwidth limit associated with VM3, negotiates with all transmitting nodes targeting VM3 in ingress traffic), dynamic adjustment of such bandwidth allocation (e.g., to adapt to changes in the number of transmitting nodes sharing available bandwidth or changes in SLAs), and dynamic congestion feedback to indicate that a transmitter needs to temporarily slow down the associated egress data rate even though the overall bandwidth allocation remains the same, thus providing end-to-end congestion feedback. The scheme can be extended. Such dynamic congestion feedback may occur in various return messages to different sending nodes (e.g., feedback message 3470) instructing each sending node about updated bandwidth limits to be used when sending traffic to VM3. Such feedback messages 3460 can take the form of explicit unsolicited messages and "piggyback" information in data packets to convey relevant information from the target node (i.e., VM3 in the shown embodiment) to the sending node.
[0327] Figure 35 is a flowchart illustrating a method for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to one embodiment.
[0328] According to one embodiment, in step 3510, the method can provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches, each of the plurality of switches comprising at least a leaf switch, each of the plurality of switches comprising a plurality of switch ports, the first subnet further comprising a plurality of host channel adapters, each of the host channel adapters comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes comprising a plurality of virtual machines.
[0329] According to one embodiment, in step 3520, the method can provide an end node ingress bandwidth allocation associated with an end node attached to the host channel adapter in a host channel adapter.
[0330] According to one embodiment, in step 3530, the method can negotiate bandwidth allocation between an end node attached to a host channel adapter and a remote end node.
[0331] According to one embodiment, in step 3540, the method enables an end node attached to a host channel adapter to receive ingress bandwidth from a remote end node, and the ingress bandwidth exceeds the ingress bandwidth limit of the end node.
[0332] According to one embodiment, in 3550, in response to receiving ingress bandwidth from a remote end node, the method can send a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress bandwidth allocation of the end node attached to the host channel adapter has been exceeded.
[0333] Use of multiple CE (Congestion Experience) flags in both FECN (Forward Explicit Congestion Notification) signaling and BECN (Backward Explicit Congestion Notification) signaling (ORA200246-US-NP-4) According to one embodiment, conventional congestion notifications are based on data packets that encounter congestion at a certain point (for example, some link segment between several node / switch pairs along the path from the sender to the target through the network / fabric topology), which are marked with a "congested" status flag (also called the CE flag), and this status is then reflected in the response packets sent back from the target to the sender.
[0334] According to one embodiment, the problem with this scheme is that the sending node cannot distinguish between congested flows on the same link segment, even though they represent different targets. Also, multiple paths between the sending node and the target When available between the pair of network-side nodes, any information about congestion on different alternative routes requires that some flow is active for the relevant target via the relevant route.
[0335] In one embodiment, the system and method described herein extend a congestion marking scheme to facilitate multiple CE flags in the same packet and configure switch ports to represent stage numbers that define which CE flag index should be updated. A specific path between a particular sender and a particular target, through an ordered sequence of switch ports, represents a specific ordered list of unique stage numbers, which in turn also represents CE flag index numbers.
[0336] In one embodiment, a transmitting node receiving congestion feedback with multiple CE flags set can map the various CE flags to different “target group” contexts, which will represent the associated congestion condition states and associated dynamic rate reductions. Furthermore, different flows to different targets will share congestion information and dynamic rate reduction states associated with shared link segments represented by shared “target groups” at the transmitting node.
[0337] According to one embodiment, when congestion occurs, the key issue is that congestion feedback should ideally be associated with all relevant target groups in the hierarchy associated with the flow receiving the congestion feedback. The affected target groups should then dynamically adjust their maximum rates accordingly. Therefore, the hardware state of each target group must also include any current congestion status and associated "throttling information".
[0338] In one embodiment, a key aspect here is that the FECN signaling should have the ability to include multiple “Congestion Experience” (CE) flags so that a switch detecting congestion can mark a flag corresponding to its stage in the topology. – In a typical fat tree, each switch has a unique (maximum) stage number upwards and another unique (maximum) stage number downwards. Thus, a flow using a particular path will then be associated with a specific sequence of stage numbers that will include all or only a subset of the entire set of stage numbers in the complete fabric. However, for that particular flow, the various stage numbers associated with its path can then be mapped to one or more target groups associated with that flow. In this way, the received BECN for a flow can mean that the target groups associated with each CE-flagped stage in the BECN will be updated to indicate congestion, and the dynamic maximum rates for these target groups can then be adjusted accordingly.
[0339] According to one embodiment, while inherently suited to fat tree topologies, the concept of a switch's "stage number" can be generalized to represent almost any topology to which such a number can be assigned to a switch. However, in this general case, the stage number is not simply a function of the output ports, but a function of each input / output port number tuple. The required number of stage numbers and the route-specific mapping to target groups are also more complex in the general case. Therefore, in this context, the reasoning assumes only fat tree topologies.
[0340] According to one embodiment, multiple CE flags in a single packet are not currently a supported feature with respect to the standard protocol header. Therefore, this may be supported based on extensions to the standard header and / or additional to the flow. This might be supported by inserting a separate FECN packet. Conceptually, generating an additional packet in a flow is similar to using an encapsulation scheme within a switch, the effect being that packets received at wire speed cannot be forwarded at the same wire speed because more "overhead bytes" must be sent downstream. Inserting an additional packet typically results in more overhead than encapsulation, but this overhead is likely to be acceptable as long as it is amortized across multiple data packets (it is not necessary to send such additional notice for every data packet).
[0341] According to one embodiment, it is also possible to have a scheme in which the switch firmware monitors congestion conditions within the switch and, as a result, can send an "unsolicited BECN" to the relevant transmitting node. However, this means that the switch firmware must have more state information regarding the mapping between ports, priorities, and associated transmitting nodes, which may also include dynamic information regarding the relevant transmitting nodes and which addresses are involved in the packets experiencing congestion.
[0342] According to one embodiment, in the case of RC QP, the mapping "from CE flag to target group" is typically part of the QP context, and any BECN information received in an ACK / response packet is then handled in a straightforward manner with respect to the relevant QP context and associated target group. However, in the case of "unclaimed BECN" (e.g., as a result of datagram traffic with only application-level responses / ACKs, or as a result of a "congestion warning" being broadcast to multiple potential senders), the reverse mapping is not straightforward—at least in that it is handled automatically by the hardware. A better approach would therefore be to have a scheme in which FECN can lead to an automatically hardware-generated BECN in the case of connected (RC) flows, but both FECN events with hardware-automatic BECN generation and FECN events without hardware-generated BECN can be handled by firmware and / or hyper-privileged software associated with the HCA receiving the FECN. In this way, there may be a FW / SW-generated "unclaimed BECN" sent to one or more potential senders affected by the observed congestion. Upon receiving these "unclaimed BECNs," the FW / SW can then perform a mapping to the relevant local target group based on the payload data within the received "BECN message," and subsequently trigger local hardware to update the target group state, similar to what happens in the full hardware control processing of RC-related BECNs.
[0343] According to one embodiment, an RC ACK / response packet without BECN notification, or a subset of stage numbers with the CE flag set that differs from (fewer than) a previously recorded state, may lead to a corresponding update of the relevant target group within the local HCA. Similarly, an "unsolicited BECN" may be sent by the responding HCA (i.e., the relevant sw / fw) to indicate that the previously signaled congestion no longer exists.
[0344] According to one embodiment, as described above, the target group concept combined with dynamic congestion feedback at either the hardware level or the firewall / software level provides flexible control of the egress bandwidth generated by the HCA, as well as by tenants sharing individual vHCAs and physical HCAs.
[0345] According to one embodiment, the target group is identified completely independently of the associated remote address and routing information at the VM level, so the use of the target group and There is no dependency between the extent to which communication from the VM is based on overlay or other virtual networking schemes. The only requirement is that the hyper-privileged software controlling the HCA resources can define the relevant mappings. It would also be possible to use a scheme at the VM / vHCA level that has a "logical target group ID" which is mapped by the HCA to the actual target group. However, it is not clear that this is useful except for hiding the actual target group ID from the tenant - if the underlying path has changed and it is necessary to change which target groups are associated with a particular destination, this may not involve other destinations. Therefore, in general, updating target groups should involve updating all involved QPs and address handles, rather than simply updating the logical-physical target group ID mapping.
[0346] In one embodiment, for a virtualized target HCA, it is possible to represent individual vHCA ports, rather than physical HCA ports, as the final destination target group. In this way, the target group hierarchy for the remote peer node can include both a target group representing the destination physical HCA ports and an additional target group representing the final destination with respect to the vHCA ports. Thus, the system and method have the ability to limit the ingress bandwidth of individual vHCA ports (VFs) while meaning that the bandwidth per physical HCA port and the associated transmitting target group do not need to be kept below the physical HCA port bandwidth (or associated bandwidth allocation).
[0347] According to one embodiment, within the transmitting HCA, target groups can be used to represent the sharing of physical HCA ports in the egress direction by assigning different target groups to different tenants. Furthermore, to facilitate multiple VMs from the same tenant sharing tenant-level target groups for physical HCA ports, different target groups can be assigned to different such VMs. Such target groups are then set up as initial target groups for all egress communications from those VMs.
[0348] Figure 36 shows a system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment.
[0349] More specifically, according to one embodiment, Figure 36 shows a host channel adapter 3601 including a hypervisor 3611. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 3614-3616, and physical functions (PFs) 3613. The host channel adapter may further support or have several ports, such as ports 3602 and 3603, which are used to connect the host channel adapter to a network, such as network 3600. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 3601 can be connected to several other nodes, such as switches, additional separate HCAs.
[0350] According to one embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 3650, VM2 3651, and VM3 3652.
[0351] According to one embodiment, the host channel adapter 3601 can further support the virtual switch 3612 via the hypervisor. This is due to the circumstances under which the vSwitch architecture is realized. Although not shown, embodiments of this disclosure are As mentioned above, this further supports the virtual port (vPort) architecture.
[0352] According to one embodiment, the network 3600 may comprise several switches, such as switches 3640, 3641, 3642, and 3643, which are interconnected and can be connected to a host channel adapter 3601 via, for example, leaf switches 3640 and 3641, as shown in the figure.
[0353] According to one embodiment, switches 3640-3643 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0354] According to one embodiment, an ingress packet 3690 may experience congestion at any stage of its path while traversing the network, and the packet may be marked by a switch when it detects such congestion at any stage. In addition to marking the packet as having experienced congestion, the switch performing the marking may further indicate the stage at which the packet experienced congestion. Upon arrival at a destination node, e.g., VM1 3650, VM1 may (e.g., automatically) send a response packet via an explicit feedback message 3691 that can indicate to the sending node that the packet experienced congestion and at which stage the packet experienced congestion.
[0355] According to one embodiment, an ingress packet may include a bit field that is updated to indicate where the packet experienced congestion, and an explicit feedback message may mirror / represent this bit field when notifying the sending node of such congestion.
[0356] In one embodiment, each switch port represents a stage within the entire subnet. Thus, each packet transmitted within the subnet can traverse a maximum number of stages. To identify where congestion is detected (which can be in multiple locations), congestion marking (e.g., the CE flag) is extended from a simple binary flag (congestion experienced) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a fat tree consisting of three stages, the maximum number of stages would be three. If the system has a route from A to B and the routing is known, each end node can determine which switch ports a packet traversed through at any given stage of the route. By doing so, each end node can determine which separate switch ports the packet experienced congestion at by correlating the routing with the received congestion message.
[0357] According to one embodiment, the system may provide a return congestion feedback indicating at which stage congestion is detected, and then, if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer-grained information regarding congestion.
[0358] According to one embodiment, by providing such finer granularity, the end node can then use alternative routes when routing future packets. Or, for example, if an end node has multiple flows, all going to different destinations, but congestion is detected at a common stage of the route, rerouting may be triggered. The system and method have 10 different congestion notices, rather than associating them with each other. It provides an immediate response to throttling. This is a more efficient way to handle congestion notifications.
[0359] Figure 37 shows a system, according to one embodiment, for using multiple CE flags in both FECN and BECN in a high-performance computing environment.
[0360] More specifically, according to one embodiment, Figure 37 shows a host channel adapter 3701 including a hypervisor 3711. The hypervisor may host or associate with several virtual functions (VFs), such as VFs 3714-3716, and physical functions (PFs) 3713. The host channel adapter may further support or have several ports, such as ports 3702 and 3703, used to connect the host channel adapter to a network, such as network 3700. The network may comprise a switched network, such as an InfiniBand network or a RoCE network, to which the HCA 3701 can be connected to several other nodes, such as switches, additional separate HCAs.
[0361] According to one embodiment, as described above, each of the virtual functions can host virtual machines (VMs) such as VM1 3750, VM2 3751, and VM3 3752.
[0362] According to one embodiment, the host channel adapter 3701 can further support a virtual switch 3712 via a hypervisor. This is due to the circumstances under which a vSwitch architecture is realized. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture as described above.
[0363] According to one embodiment, the network 3700 may comprise several switches, such as switches 3740, 3741, 3742, and 3743, which are interconnected and can be connected to a host channel adapter 3701 via, for example, leaf switches 3740 and 3741, as shown in the figure.
[0364] According to one embodiment, switches 3740-3743 can be interconnected and further connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0365] According to one embodiment, an ingress packet 3790 may experience congestion at any stage of its path while traversing the network, and the packet may be marked by a switch when it detects such congestion at any stage. In addition to marking the packet as having experienced congestion, the switch performing the marking may further indicate the stage at which the packet experienced congestion. Upon arrival at a destination node, e.g., VM1 3750, VM1 may (e.g., automatically) send a response packet via a piggyback message (a message on another message / packet sent from the receiving node to the sending node) 3791 that indicates to the sending node that the packet experienced congestion and at which stage the packet experienced congestion.
[0366] According to one embodiment, an ingress packet may include a bit field that is updated to indicate where the packet experienced congestion, and an explicit feedback message may mirror / represent this bit field when notifying the sending node of such congestion.
[0367] In one embodiment, each switch port represents a stage within the entire subnet. Thus, each packet transmitted within the subnet can traverse a maximum number of stages. To identify where congestion is detected (which can be in multiple locations), congestion marking (e.g., the CE flag) is extended from a simple binary flag (congestion experienced) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a fat tree consisting of three stages, the maximum number of stages would be three. If the system has a route from A to B and the routing is known, each end node can determine which switch ports a packet traversed through at any given stage of the route. By doing so, each end node can determine which separate switch ports the packet experienced congestion at by correlating the routing with the received congestion message.
[0368] According to one embodiment, the system may provide a return congestion feedback indicating at which stage congestion is detected, and then, if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer-grained information regarding congestion.
[0369] According to one embodiment, by providing such finer granularity, the end node can then use alternative routes when routing future packets. Alternatively, for example, if an end node has multiple flows, all going to different destinations, but congestion is detected at a common stage of the route, rerouting may be triggered. The system and method provide an immediate response to the throttling associated with it, rather than having 10 different congestion notices. This is a more efficient handling of congestion notices.
[0370] Figure 38 shows a system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment.
[0371] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3800, several end nodes 3801 and 3802 can support several virtual machines VM1 to VM4 3850 to 3853 interconnected via several switches such as leaf switches 3811 and 3812, switches 3821 and 3822, and root switches 3831 and 3832.
[0372] According to one embodiment, various host channel adapters providing functionality for connecting nodes 3801 and 3802, as well as virtual machines to be connected to the subnet, are not shown. Such embodiments are discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0373] According to one embodiment, a packet 3851 sent from node VM3 3852 to VM1 3850 can traverse subnet 3800 via several links or stages such as stages 1 through 6, as shown in the figure. While traversing the subnet, packet 3851 may experience congestion at any of these stages, and if such congestion is detected at any of the stages, it may be marked by the switch. In addition to marking the packet as having experienced congestion, the switch performing the marking may further indicate the stage at which the packet experienced congestion. Upon reaching the destination node VM1, VM1 can (for example, automatically) send a response packet via a feedback message 3870 that indicates to VM3 3852 that the packet experienced congestion and at what stage the packet experienced congestion.
[0374] In one embodiment, each switch port represents a stage within the entire subnet. Thus, each packet transmitted within the subnet can traverse a maximum number of stages. To identify where congestion is detected (which can be in multiple locations), congestion marking (e.g., the CE flag) is extended from a simple binary flag (congestion experienced) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a fat tree consisting of three stages, the maximum number of stages would be three. If the system has a route from A to B and the routing is known, each end node can determine which switch ports a packet traversed through at any given stage of the route. By doing so, each end node can determine which separate switch ports the packet experienced congestion at by correlating the routing with the received congestion message.
[0375] According to one embodiment, the system may provide a return congestion feedback indicating at which stage congestion is detected, and then, if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer-grained information regarding congestion.
[0376] According to one embodiment, by providing such finer granularity, the end node can then use alternative routes when routing future packets. Alternatively, for example, if an end node has multiple flows, all going to different destinations, but congestion is detected at a common stage of the route, rerouting may be triggered. The system and method provide an immediate response to the throttling associated with it, rather than having 10 different congestion notices. This is a more efficient handling of congestion notices.
[0377] Figure 39 is a flowchart illustrating a method for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to one embodiment.
[0378] According to one embodiment, in step 3910, the method can provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches, each of the plurality of switches comprising at least a leaf switch, each of the plurality of switches comprising a plurality of switch ports, the first subnet further comprising a plurality of host channel adapters, each of the host channel adapters comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes comprising a plurality of virtual machines.
[0379] According to one embodiment, in step 3920, the method can receive ingress packets from a remote end node at an end node attached to a host channel adapter, the ingress packets traversing at least a portion of a first subnet before being received at the end node, and the ingress packets include markings indicating that the ingress packets experienced congestion while traversing at least a portion of the first subnet.
[0380] According to one embodiment, upon receiving an ingress packet, in step 3930, the method may have an end node send a response message from the end node attached to the host channel adapter to a remote end node, the response message indicating that the ingress packet experienced congestion while traversing at least a portion of the first subnet, and the response message includes a bit field.
[0381] QOS and SLA in switched fabrics such as private fabrics According to one embodiment, a private network fabric in the cloud and larger clouds in customer and on-premises locations (e.g., a private fabric used to build dedicated distributed appliances or general-purpose high-performance computing resources) desires the ability to deploy VM-based workloads, and a specific requirement is the ability to define and control quality of service (QoS) for different types of communication flows. In addition, workloads belonging to different tenants must run within the boundaries of the relevant service level agreements (SLAs) while minimizing interference between such workloads and maintaining QoS assumptions for different types of communication.
[0382] In one embodiment, the following sections discuss relevant problem scenarios, objectives, and potential solutions.
[0383] In one embodiment, the initial scheme for provisioning fabric resources to a cloud customer (also known as a "tenant") is that a tenant can be allocated a dedicated portion of a rack (e.g., an allocated rack) or one or more full racks. This granularity means that each tenant is guaranteed to have a communications SLA that is always met, as long as the allocated resources are fully operational. This is also true when a single rack is divided into multiple parts, because the granularity is always a full physical server with a HCA. Connectivity between different such servers within a single rack can, in principle, always be via a single full crossbar switch. In this case, there are no resources shared in a way that could lead to circuit competition or congestion between flows belonging to different tenants as a result of communication traffic between sets of servers belonging to the same tenant.
[0384] However, according to the embodiment, since redundant switches are shared, it is crucial that traffic generated by a workload on one server cannot target servers belonging to another tenant. Such traffic does not facilitate any inter-tenant communication or data leakage / observation, but the result could be serious interference with communication flows belonging to other tenants or even a DoS (Dynamic Service)-like effect.
[0385] In one embodiment, despite the fact that a full crossbar-leaf switch inherently means that all communication between local servers can take place only through local switches, there are some cases where this may not be possible or cannot be achieved due to other practical issues. According to one embodiment, for example, if host bus (PCIe) generation can only sustain the bandwidth of one fabric link at a time, it is important that only one HCA port is used for data traffic at any given time. Therefore, if not all servers agree on which local switch to use for data traffic, some traffic will have to traverse inter-switch links (ISLs) between local leaf switches. According to one embodiment, if one or more servers lose connectivity to one of the switches, all communication must be conducted through the other switch. Again, if all pairs of servers cannot agree on using the same single switch, some data traffic will have to go through the ISL. According to one embodiment, if a server can use both HCA ports (and therefore both leaf switches), but it is impossible to implement that connections are established only through HCA ports connected to the same switch, some data traffic may pass through the ISL. One reason this scheme ends up being problematic is that the lack of sockets / port numbers in the fabric host stack means that a process can only establish one socket to accept incoming connections. In that case, this socket can only be associated with one HCA port at a time. As long as the same single socket is also used when establishing outgoing connections, a number of processes will end up with a certain number of connections that require an ISL, even though those single sockets are evenly distributed among the local HCA ports.
[0386] According to one embodiment, in addition to the single-rack scenario of the special case imposing the ISL usage / sharing outlined above, when the provisioning granularity is extended to a multi-rack configuration in which leaf switches within each rack are interconnected by spine switches, then communication SLAs for different tenants become highly dependent on which servers are allocated to which tenant and how different communication flows are mapped over different switch-switch links by the fabric-level routing scheme. A key issue in this scenario is that the two optimization aspects are somewhat contradictory. According to one embodiment, on the one hand, in order to provide the best possible performance, all concurrent flows targeting different destination ports should use as many different paths through the fabric (i.e., different switch-switch link-ISL) as possible. • According to the embodiment, on the other hand, in order to provide predictable QoS and SLAs to different tenants, it is important that flows belonging to different tenants do not simultaneously compete for bandwidth on the same ISL. In general, this means that there must be restrictions on which routes different tenants can use.
[0387] However, according to certain embodiments, in some situations, depending on the size of the system, the number of tenants, and how servers are provisioned for different tenants, it may not be possible to avoid bandwidth competition for flows belonging to different tenants on the same ISL. In this situation, there are two main approaches that can be used from a fabric perspective to address the problem and reduce the likelihood of circuit competition. According to one embodiment, to ensure that flows from different tenants have forward progress independently of other tenants, even though they are competing for the same ISL bandwidth, the system restricts which switch buffer resources can be occupied by different tenants (or groups of tenants). According to one embodiment, a "permission control" mechanism is implemented that limits the maximum bandwidth that can be consumed by one tenant at the expense of other tenants.
[0388] In one embodiment, one challenge with respect to the physical fabric configuration is that the binary bandwidth should be as high as possible, ideally non-blocking or even over-provisioned. However, even with non-blocking binary bandwidth, there may be scenarios where achieving the desired SLA for one or more tenants is difficult given the current allocation of servers to those different tenants. In such situations, the best approach would generally be to re-provision at least some of the servers for the different tenants to reduce the need for independent ISLs and binary bandwidth.
[0389] In one embodiment, several multi-rack systems have a blocking fat tree topology, the assumption being that workloads will be provisioned so that the relevant communication servers are located to a considerable extent within the same rack, meaning that a significant portion of bandwidth utilization resides only between ports on local leaf switches. Also, in traditional workloads, the majority of data traffic is from one set of fixed nodes to another. However, with next-generation servers, which include non-volatile memory and newer communication and storage middleware, communication workloads become even more demanding and less predictable, as different servers may simultaneously provide multiple functions, according to one embodiment.
[0390] In one embodiment, one objective is to provide per-tenant provisioning granularity at the VM level, as opposed to the physical server level. Another objective is that up to tens of VMs can be deployed on the same physical server, different sets of VMs on the same physical server may belong to different tenants, and various tenants may each represent multiple workloads with different characteristics.
[0391] In addition to the current fabric deployments using different types of service (TOS) associations to provide basic QoS (traffic isolation) for different flow types (for example, to prevent lock messages from being “stalled” after large bulk data transfers), it is also desirable to provide communication SLAs for different tenants. These SLAs are expected to ensure that tenants experience expected workload throughput and response times, even if their workloads are provisioned on physical infrastructure shared by other tenants. Relevant SLAs for tenants are expected to be met independently of concurrent activity by workloads belonging to other tenants.
[0392] In one embodiment, a workload may provision a fixed (minimum) set of CPU cores / threads and physical memory on a fixed (minimum and / or maximum) set of physical servers, but provisioning fixed / guaranteed networking resources is generally not so straightforward, insofar as deployment implies sharing of HCAs / NICs on servers. HCA sharing also inherently implies that at least the ingress and egress links to the fabric are shared by different tenants. Thus, different CPU cores / threads can operate truly in parallel, but there is no way to divide the capacity of a single fabric link except for some kind of bandwidth multiplexing or “time-sharing”. This fundamental bandwidth sharing may or may not be combined with the use of different “QoS IDs” (e.g., service level, priority, DSCP, traffic class, etc.) that are considered when implementing buffer selection / allocation and bandwidth arbitration within the fabric.
[0393] In one embodiment, the overall server memory bandwidth should be very high compared to the typical memory bandwidth needs of any individual CPU core to prevent memory-intensive workloads on some cores from imposing latency on other cores. Similarly, in the ideal case, the available fabric bandwidth for a physical server should be large enough to allow each tenant sharing the server to have sufficient bandwidth for the communication activity generated by their respective workloads. However, when several workloads all attempt to perform large data transfers, it is very likely that multiple tenants will be able to utilize the full link bandwidth—even if it is over 100 Gb / s. To address this scenario, multiple tenants on the same physical server... It is necessary that provisioning be carried out in a way that ensures each tenant is guaranteed to receive at least a given minimum percentage of the available bandwidth. However, in RDMA-based communications, the ability to enforce limits on how much bandwidth a tenant can generate in the egress direction does not mean that ingress bandwidth can be limited in the same way. That is, multiple remote communication peers may all send data to the same destination in a way that completely overloads the receiver, even though each sender is limited by its maximum transmit bandwidth. Also, an RDMA read operation can originate from a local tenant using only a small amount of egress bandwidth. This has the potential to result in catastrophic ingress bandwidth if a bulk RDMA read operation occurs for multiple remote peers. Therefore, imposing a maximum limit on egress bandwidth to limit the total fabric bandwidth used by a single tenant on a single server is not sufficient.
[0394] According to one embodiment, the system and method can configure an average bandwidth limit for a tenant that ensures the tenant never exceeds its relative portion of the relevant link bandwidth in either the ingress or egress direction, regardless of the use of RDMA read operations, the number of remote peers with active data traffic, and the bandwidth limitations of remote peers. (Methods for achieving this are discussed in the “Long-Term Goals” section below.) In one embodiment, unless the system and method can implement all aspects of communication bandwidth limiting, the highest level of communication SLA for a tenant can only be achieved by restricting that a tenant cannot share a physical server with other tenants, or potentially, that a tenant cannot share a physical HCA with other tenants (i.e., in the case of a server with multiple physical HCAs). If a physical HCA can operate with full link bandwidth utilization for both HCA ports in active-active mode, it is also conceivable to use a restriction that grants a given tenant exclusive access to one of the HCA ports under normal circumstances. Nevertheless, due to HA constraints, failure of the entire HCA (in the case of multiple HCAs per server) or a single HCA port may mean reconfiguration and sharing that no longer guarantees the expected communication SLA for a given tenant.
[0395] In one embodiment, in addition to constraints on overall bandwidth utilization for a single link, each tenant's ability to achieve QoS between different communication flows or flow types depends on it not experiencing severe congestion contention for fabric-level buffer resources or arbitration due to communication activity by other tenants. In particular, this means that if a tenant is using a specific "QoS ID" to achieve low-latency messaging, that tenant should not find itself "competing" with high-volume data traffic from other tenants, depending on how other tenants are using the "QoS ID" and / or how the fabric implementation enforces the use of the "QoS ID" and / or how this maps to buffer allocation and / or packet bandwidth arbitration within the fabric. Therefore, if a tenant communication SLA means that the tenant's internal QoS assumption cannot be met without relying on other tenants sharing the same fabric link "behaving well," this may impose the requirement that tenants must be provisioned without HCA (or HCA port) sharing with other tenants.
[0396] According to one embodiment, shared constraints apply to both the basic bandwidth allocation and QoS issues described above, as well as to the fabric internal links and server-local HCA port links. Therefore, depending on the nature and strictness of the communication SLA for a given tenant, The deployment of VMs for a given instance may be constrained by the sharing of physical servers and / or HCAs, as well as the sharing of ISLs within the fabric. To avoid ISL sharing, both routing restrictions and restrictions on where VMs can provision to each other within a private fabric topology may be applied.
[0397] Topology, routing, and blocking scenario considerations: In one embodiment, as described above, there can be no HCA / HCA ports or any fabric ISLs shared with other tenants to ensure that a tenant can achieve the expected communication performance between sets of VMs communicating without depending on the operation of VMs belonging to other tenants. Therefore, the highest class of SLA offered would typically have this implicitly realized. This is, in principle, the same scheme as the current provisioning models for many traditional systems in the cloud. However, with shared leaf switches, this SLA would require a guarantee that there is no ISL sharing with other tenants. Furthermore, it would be necessary to be able to "optimize non-blocking" in an explicit manner in order for tenants to achieve the best possible balance between flows and the use of available fabric resources (i.e., the communication SW infrastructure must give tenants a way to ensure that communication occurs in a way that different flows do not compete for the same link bandwidth). This would include a way to ensure that communication that may occur through a single leaf switch is actually realized in this way. Also, if communication must involve an ISL, it should be possible to balance traffic across the available ISLs to maximize throughput.
[0398] According to one embodiment, from a single HCA port, it is not meaningful to attempt to balance traffic across multiple ISLs, as long as the maximum available bandwidth is the same for all links in the fabric. From this perspective, it would make sense to use a dedicated "next-hop" ISL for each transmitting HCA port, as long as the available ISLs represent a non-blocking subtopology for the transmitting side. However, a scheme with a dedicated next-hop ISL for each transmitting port is not actually sustainable unless the relevant ISLs represent connectivity only between two leaf switches, because at some point, if communication is between multiple remote peer HCA ports connected to different leaf switches, multiple ISLs would have to be used.
[0399] In one embodiment, in a non-blocking InfiniBand fat tree topology, the prevailing routing algorithm uses "dedicated down paths," which means that in a non-blocking topology, there are an equal number of switch ports at each layer of the fat tree. This means that each end port can have a dedicated port chain from one root switch, through each intermediate switch layer, to the egress leaf switch port connecting the associated HCA port. Thus, all traffic targeting a single HCA port uses this dedicated down path, and there is no traffic on these links to other destination ports (downstream). However, upstream, there can be no dedicated path to each destination, and as a result, some links upstream must be shared by traffic to different destinations. In the next round, this can lead to congestion when different flows to different destinations all try to utilize the full bandwidth on a shared intermediate link. Similarly, if multiple senders are sending to the same destination simultaneously, this can cause congestion on the dedicated down path, which can then quickly spread to other unrelated flows.
[0400] According to one embodiment, as long as a single destination port belongs to a single tenant, there is no risk of congestion between multiple tenants on a dedicated down path. However, in order to reach the root switch (or intermediate switch) representing the dedicated down path, different tenants must be on the up path. The need to use the same link remains a problem. By dedicating as many different route switches as possible to specific tenants, this system and method will reduce the need for different tenants to share routes uplink. However, from a single leaf switch, this scheme may reduce the number of available uplinks to the associated route switch. Therefore, in order to maintain non-blocking bimodal bandwidth between servers (or rather HCA ports) belonging to the same tenant, the number of servers allocated to a single tenant on a particular leaf switch (i.e., within a single rack) will need to be less than or equal to the number of uplinks to the route switch used by that tenant. On the other hand, it makes sense to allocate as many servers as possible to the same tenant within the same rack in order to maximize the ability to communicate through a single crossbar.
[0401] According to one embodiment, this inherently implies a contradiction between being able to utilize guaranteed bandwidth within a single leaf switch and being able to utilize guaranteed bandwidth toward communication peers in different racks. To address this dilemma, the best approach is probably to use a scheme in which tenant VMs are grouped based on which leaf switches (i.e., leaf switch pairs) they are directly connected to, and then there must be an attribute that defines the available bandwidth between such groups. However, again, there is a trade-off between being able to maximize bandwidth between two such groups (e.g., between the same tenant in two racks) and being able to guarantee bandwidth toward multiple remote groups. Furthermore, in the special case of only two layers of switches (i.e., leaf layers interconnected by a single spine layer), a non-blocking topology means that it is always possible for N HCA ports to have N dedicated uplinks between leaf switches and N spine ports belonging to the same tenant. Thus, the configuration is non-blocking toward that tenant, insofar as these N spine ports represent a spine that "owns" all dedicated down paths for all relevant remote peer ports. However, if the relevant remote peers represent dedicated down paths from more than N spine switches, or if the N uplinks are not distributed among all relevant spine switches, this system and method may result in line competition with other tenants.
[0402] According to one embodiment, among VMs in a single tenant, regardless of non-blocking or blocking connectivity, the possibility of line competition still exists between flows from different sources connected to the same leaf switch. That is, if destinations have dedicated down paths from the same spine, and the number of uplinks from the source leaf switch to that spine is less than the number of such simultaneous flows, there is no way to avoid some kind of blocking / congestion on the uplink, as long as all senders are operating at full link speed. In this case, the only option to maintain bandwidth would be to use a secondary path to one of the destinations via a different spine. This would then be seen as representing potential competition with another dedicated down path, since a standard non-blocking fat tree can only have one dedicated downlink per end port.
[0403] According to one embodiment, in some conventional systems, a blocking factor of 3 can exist between leaf switches and spine switches. Therefore, in multi-rack scenarios where the workload is distributed in such a manner that more than one-third of the communication traffic is between racks rather than within racks, the resulting binary bandwidth becomes blocking. For example, the most common scenario with even distribution of traffic between any pair of nodes in an 8-rack system means that 7 / 8 of the communication is between racks, resulting in a substantial blocking effect.
[0404] According to one embodiment, if the cable costs of over-provisioning are acceptable in the system (i.e., given a fixed switch unit cost), it is possible to use additional links to provide both a “backup” downlink to each leaf switch and spare uplink capacity from each leaf to each spine. In other words, in both cases, it provides at least some potential improvement to dynamic workload distributions that represent a non-uniform distribution of traffic and therefore cannot take advantage of a topology that is inherently non-blocking in the first place.
[0405] According to one embodiment, a higher digit full crossbar switch may also increase the size of each single “leaf domain” and reduce the number of spine switches required for a given system size. For example, with a 128-port switch, two full racks with 32 servers each could be included in a single full crossbar leaf domain, and still provide non-blocking uplink connectivity. Similarly, only eight spines would be required to provide non-blocking connectivity between 16 racks (512 servers, 1024 HCA ports). Thus, there are still only eight links from each leaf to each spine (i.e., in the case of a single, fully connected network). In the extreme case where all HCA ports on one leaf send to a single remote leaf via a single spine, this still means a blocking factor of eight. On the other hand, given the even distribution of dedicated down paths for each leaf switch among all spines, the possibility of such an extreme scenario should be negligible.
[0406] According to one embodiment, in the case of a dual independent network / rail, each leaf switch in a redundant leaf switch pair belongs to a single rail with its own dedicated spine. The same eight spines are divided into two groups of four spines each (one for each rail), and therefore, in this case, each leaf in the rail would only need to connect to four spines. Thus, in this case, the worst-case blocking factor would be only 4. On the other hand, in this scenario, the selection of rails for each communication operation becomes even more important in order to provide load balancing across both rails.
[0407] Dynamic vs. Static Packet Route Selection / Forwarding + Multipathing: In one embodiment, standard InfiniBand uses static routes for each destination address, while there are several standard, intellectual property schemes for dynamic route selection in Ethernet switches. For InfiniBand, there are also various intellectual property schemes for "adaptive routing" (some of which may be standardized).
[0408] According to one embodiment, one advantage of dynamic route selection is that it increases the probability of optimally utilizing the relevant binary bandwidth within the fabric, thereby increasing overall throughput. However, potential drawbacks include the possibility of ordering being disrupted and the possibility that congestion in one region of the fabric may spread more easily to other regions (i.e., in ways that could have been avoided if static route selection had been used).
[0409] In one embodiment, “dynamic routing” or “dynamic route selection” is typically used in reference to forwarding decisions made within and between switches, while “multipathing” is a term used when traffic to a single destination can be spread across multiple paths based on explicit addressing from the sender. Such multipathing “stripe” the transmission of a single message across multiple local HCA ports (i.e., the complete message becomes multiple submessages, each representing an individual forwarding operation). This could include being split into multiple sections, which could mean that different forwardings to the same destination are set up to dynamically use different paths through the fabric.
[0410] In one embodiment, if, in the general case, all forwardings from all sources targeting destinations outside the local leaf domain are divided into (smaller) chunks and then distributed across all possible paths / routes toward that destination, the system will achieve optimal utilization of available binary bandwidth and maximize “inter-leaf throughput.” However, this is only true insofar as the communication workload is also evenly distributed across all possible destinations. Otherwise, the impact will be such that any congestion toward a single destination will immediately affect all concurrent flows.
[0411] In one embodiment, the implications of congestion for dynamic route selection and multipathing are that it makes sense to restrict traffic to a single destination and use only a single route, provided that route does not become a victim of congestion on other targets or any intermediate links. In a two-tier fat tree topology with dedicated down routes, this means that the only possible congestion unrelated to end ports will reside on uplinks targeting the same spine switch. This means that it would make sense to treat all uplinks to the same spine as a group of ports sharing the same static route, except that the individual ports used for a particular target will be dynamically selected. Alternatively, individual ports may be selected based on tenant association.
[0412] According to one embodiment, selecting an uplink port within such a group using tenant associations may be based on fixed associations or on a scheme in which different tenants have a “first priority” for using a certain port but also the ability to use other ports. In that case, the ability to use another port would depend on this not conflicting with the “first priority” traffic of the other ports. Thus, tenants would be able to use all relevant binary bandwidth as long as there is no conflict, but if conflict exists, there would be a guaranteed minimum bandwidth. This minimum guaranteed bandwidth may then reflect all bandwidth for one or more links, or a percentage of the bandwidth for one or more links.
[0413] According to one embodiment, in principle, the same dynamic scheme can also be used in a downward path from a spine to a specific leaf. On the one hand, this would increase the risk of congestion resulting from sharing downlinks between flows targeting different endports, but on the other hand, it could provide a way to utilize additional alternative paths between two sets of nodes connected to two different leaf switches, while still providing a way to prevent congestion from spreading between different tenants.
[0414] In one embodiment, in a scenario where different dedicated down paths from a spine to a leaf already represent a particular tenant, it would be relatively simple to have a scheme that allows these links to be used as "backups" for traffic (belonging to the same tenant) to an end port that has a (primary) dedicated down path from another spine on the associated leaf switch.
[0415] According to one embodiment, a possible model would involve the switch handling dynamic route selection between parallel ISLs connecting a single spine or leaf switch, but having a host-level decision about using explicit multipathing via spines that do not represent a (primary) dedicated down path to the target in question.
[0416] Per-tenant bandwidth allocation control: According to one embodiment, if a single HCA is used by only one tenant, the system and method can limit the bandwidth that may originate from the HCA port. This is particularly true when there is a limited binary bandwidth for that tenant with respect to traffic heading to a remote leaf switch.
[0417] According to one embodiment, one aspect of such bandwidth limiting is ensuring that the limit applies only to targets affected by the limited binary bandwidth. In principle, this would involve a scheme in which different target groups are associated with a particular bandwidth allocation (i.e., either the exact maximum rate and / or the average bandwidth over a certain amount of transferred data).
[0418] According to one embodiment, such limitations would, by definition, have to be implemented at the HCA level. Furthermore, such limitations would more or less directly map to virtualized HCA scenarios where VMs belonging to different tenants share an HCA through different virtual functions. In this case, the various “shared bandwidth allocation groups” described above would require an additional dimension, in that they are associated with groups of one or more VFs, and not just full physical HCA ports.
[0419] Tenant-based bandwidth reservation on ISL: In one embodiment, as described above, it may make sense to reserve some guaranteed bandwidth across one or more ISLs for a tenant (or group of tenants). In one scenario, the link can be reserved for a tenant by restricting which tenant is actually allowed to use the entire link. However, for a more flexible and finer-grained scheme, an alternative approach is to use a switch arbitration mechanism to ensure that (some) ingress ports are allowed to use up to X% of the bandwidth of that egress port, regardless of what other ingress ports are competing for bandwidth on one or more egress ports.
[0420] In one embodiment, all ingress ports can thus utilize up to 100% of the bandwidth of the associated egress port, provided that this does not conflict with any traffic from the preferred ingress port.
[0421] In one embodiment, in a scenario where different tenants "own" different ingress ports (e.g., leaf switch ports connecting HCA ports), this scheme would facilitate a flexible and granular scheme for allocating uplink bandwidth to one or more spine switches.
[0422] According to one embodiment, in a downlink path from a spine to a leaf switch, the usefulness of such a scheme would depend on the extent to which a scheme with a strictly dedicated downlink is used or not. If a strictly dedicated downlink is used and the target end port represents a single tenant, by default there is no potential conflict between different tenants attempting to use the downlink. Therefore, in this case, access to the relevant downlink should typically be set up using a round-robin arbitration scheme that has equal access for all relevant ingress ports.
[0423] According to one embodiment, an ingress port represents traffic belonging to different tenants. Therefore, it should not be a problem that packets belonging to one tenant may be sent and that the associated tenant may consume bandwidth on egress ports to which it is not permitted to send. In this case, the assumption is that strict access control (e.g., VLAN-based restrictions on various ports) rather than arbitration policies will be employed to prevent such packets from wasting bandwidth.
[0424] In one embodiment, a leaf switch may be given more bandwidth toward various end ports to a down port from the spine compared to other local end ports, because a down link can, in principle, represent multiple transmitting HCA ports, while a local end port represents only a single HCA port. If this is not the case, there is a scenario in which several remote servers share a single down path to a target leaf switch, but in the next round, if N-1 HCA ports directly connected to the leaf switch also attempt to send to the same local target port, they will share 1 / N of bandwidth toward a single destination on that leaf switch.
[0425] In one embodiment, when virtualized HCAs represent different tenants, the problem of reserving bandwidth within a fabric ISL (i.e., across different ISLs) can become significantly more complex. For ingress / uplink paths, one simplified approach is that it is up to the HCA to provide bandwidth arbitration between different tenants, and whatever is sent on the HCA port will then be handled by the ingress leaf switch according to the port-level arbitration policy. Thus, in this case, there is no change from the leaf switch's perspective.
[0426] In one embodiment, the situation is different in the downlink path (spine to leaf, and leaf ingress to end port), as arbitration decisions may depend not only on the port attempting to forward the packet, but also on which tenant various pending packets belong. One possible solution is to (again) restrict some ISLs to represent only specific tenants (or groups of tenants), and then reflect this in the port-level arbitration scheme. Alternatively (or additionally), different priority or QoS IDs can be used to represent different tenants, as outlined below. Finally, having a “tenant ID” such as a VLAN ID or partition ID, or any related access control header field, used as part of the arbitration logic would facilitate the level of granularity required for arbitration. However, this could significantly increase the complexity of arbitration logic in switches, which already have considerable “temporal and spatial” complexity. Also, such schemes involve an overload of information that may already play a role in end-to-end wire protocols, so it is important that such extra complexity does not conflict with any existing use or assumptions regarding such header field values.
[0427] Different priorities, flow types, and QoS IDs / classes across ISL and endport links: In one embodiment, for different flow types to proceed simultaneously on the same link, it is crucial that they do not compete for the same packet buffer in the switch and HCA. Furthermore, to distinguish the relative priority between different flow types, the arbitration logic determining which packet to send next on various switch egress ports must take into account which packet type queues have which packets to send on which egress ports. The arbitration result determines to what extent all active flows, according to their relative priority, and any flow control conditions (if any) relating to the relevant flow types on the relevant downstream ports, are currently able to send any packet. Depending on the method, it should be in a forward direction.
[0428] In one embodiment, in principle, traffic flows from different tenants can be isolated from each other by using different QoS IDs, even if different tenants are using the same link. However, the scalability of this approach is very limited because the number of packet queues and independent buffer pools that can be supported for each port is typically limited to less than 10. Furthermore, scalability is further reduced if a single tenant wants to use different QoS IDs to isolate different flow types from each other.
[0429] According to one embodiment, as described above, by logically combining multiple ISLs between a single pair of switches, the system and method can restrict some links to some tenants, and then ensure that different tenants can use different QoS IDs independently of each other on different ISLs. However, here again, this imposes a limit on the total bandwidth available to any single tenant, provided that the independence of other tenants is 100% guaranteed.
[0430] According to one embodiment, in the ideal case, HCA ingress (received) packet processing can always occur at a rate higher than the associated link speed, regardless of what transport level operation the incoming packet represents. This means that there is no need for flows controlling different flow types on the egress port on the leaf switch connected to its last link, (i.e.) the HCA port. However, scheduling different packets from different queues on the leaf switch must still reflect the relevant policies regarding priority, fairness, and forward progress. For example, if one small high-priority packet targets a certain end port, and at the same time N ports also attempt to send a "bulk forwarding packet" of the maximum MTU size to the same target port, the high-priority packet should be scheduled before any of the other ports.
[0431] According to one embodiment, in an egress path, the transmitting HCA can schedule and label packets in many different ways. In particular, the use of the overlay protocol as a "bump in the wire" between the VM + virtual HCA and the physical fabric would allow the switch to encode fabric-specific information that may be relevant without disrupting any aspect of the end-to-end protocol between tenant virtual HCA instances.
[0432] In one embodiment, the switch can provide more buffering and internal queuing than current wire protocols assume. In this way, it would be possible to set up buffering, queuing, and arbitration policies that use different QoS classes for different flow types, taking into account that the link is shared by traffic representing multiple tenants with different SLAs.
[0433] According to one embodiment, in this way, different high-priority tenants may also have more private packet buffer capacity within the switch.
[0434] Lossless packet forwarding vs. Lossed packet forwarding: In one embodiment, high-performance RDMA traffic heavily relies on individual packets not being lost due to insufficient buffer capacity in the switch, and also heavily relies on packets arriving in the correct order for each RDMA connection. In principle, the higher the potential bandwidth, the more critical these aspects become to achieving optimal performance. ru.
[0435] In one embodiment, lossless operation requires explicit flow control, and very high bandwidth implies a trade-off between buffer capacity, MTU size, and flow control update frequency.
[0436] According to the embodiment, a drawback of lossless operation is that it leads to congestion when the total bandwidth generated is higher than the downstream / receive capacity. Congestion then (presumably) spreads, ending up slowing down all competing flows somewhere in the fabric for the same buffer.
[0437] According to one embodiment, as described above, the ability to provide flow isolation based on independent buffer pools is a major scalability issue for switch implementation, depending on both the number of ports, the number of different QoS classes, and (as mentioned above) potentially the number of different tenants as well.
[0438] In one embodiment, an alternative approach could define truly lossless operation (i.e., lossless operation based on guaranteed buffer capacity) as a “premium SLA” attribute, thereby restricting this feature to only tenants who have purchased such a premium SLA.
[0439] In one embodiment, the key issue here is that available buffer capacity can be "over-reserved," and the same buffer can be used for both lossy and lossless flows, but buffers allocated to lossy flows can be forcibly excluded whenever packets from lossless flows arrive and require the use of buffers from the same pool. A very minimal set of buffers may be provided to allow lossy flows to proceed forward, but this is at a much lower bandwidth than what can be achieved with optimal buffer allocation.
[0440] In one embodiment, it is also possible to introduce hybrid lossless / lossy flow classes of different classes with respect to the difference in the maximum time that a buffer can be occupied before it has to be forcibly eliminated (when needed) and given to a more premium SLA type flow class. This would work best in the context of fabric realizations with link-level credits, but could potentially be adapted to work with xon / xoff type flow control (i.e., Ethernet pause-based flow control schemes used in RoCE / RDMA).
[0441] Strict packet ordering vs. (more) relaxed packet ordering: In one embodiment, strict ordering and lossless packet forwarding within the fabric allows an HCA implement to achieve reliable connectivity and RDMA with minimal state overhead at the transport level. However, to better tolerate some out-of-order packet delivery due to accidental route changes (by adaptive / dynamic forwarding decisions within the fabric), and to minimize the overhead and delay associated with packets lost due to lossy or "hybrid lossless / lossy" mode forwarding within the fabric, an efficient transport implement would require sufficient state to allow a large number of individual packets (sequence numbers) to arrive out of order and be individually retried while other packets with later sequence numbers are accepted and acknowledged.
[0442] According to one embodiment, the key point here is that when lost or out-of-order packets cause retries in the current default transport implementation, the average bandwidth The goal is to avoid long delays and losses. Furthermore, by preventing subsequent packets from being dropped in the sequence of packets being delivered, this system and method significantly reduces the waste of fabric bandwidth that could otherwise be consumed by other flows.
[0443] Shared services and shared HCA: According to one embodiment, a shared service on the fabric used by multiple tenants (e.g., a backup device) means that if the service cannot provide an end port that can be dedicated to a particular tenant (or a restricted group of tenants), then several end port links will be shared by different tenants. A similar scenario exists when VMs belonging to multiple tenants share the same server and the same HCA port.
[0444] In one embodiment, finely tuned server and HCA resources can be allocated to different tenants, and it is also possible to ensure that the outgoing data traffic bandwidth from the HCA is fairly divided among different tenants according to the relevant SLA levels.
[0445] In one embodiment, it may also be possible to configure packet buffer allocation, queuing priority, and arbitration policies within the fabric that reflect the relative importance and resulting fairness between data traffic belonging to different tenants. However, even highly fine-tuned buffer allocation and arbitration policies within the fabric may not have sufficient granularity to ensure that relative priorities and bandwidth allocations for different tenants are accurately reflected with respect to ingress bandwidth for shared HCA ports.
[0446] According to one embodiment, in order to achieve such fine-grained bandwidth allocation, a dynamic end-to-end flow control scheme is required that can effectively divide and schedule the available ingress bandwidth among several telecommunications peers belonging to one or more tenants.
[0447] In one embodiment, the goal of such a scheme would be that, at any given time, a set of relevant active remote clients can utilize a fair (not necessarily equal) share of the available ingress bandwidth. Furthermore, this bandwidth utilization should be achieved without causing fabric congestion due to attempts to utilize excessive bandwidth at endports. (However, fabric-level congestion can still occur due to overload on shared links within the rest of the fabric.) According to one embodiment, a high-level model for achieving this goal would be that the receiving end can dynamically allocate and update the available bandwidth for the set of remote clients involved. The current bandwidth value for each remote client would need to be calculated based on what is currently provided to each client and what will be required next.
[0448] According to one embodiment, this means that if a single client is currently allowed to use all available bandwidth, and another client also needs to use ingress bandwidth, an update command must be delivered to the currently active client informing it of the new reduced maximum bandwidth, and the new client must be delivered a command informing it that it can use the maximum bandwidth corresponding to the reduction for the current client.
[0449] According to one embodiment, the same scheme would, in principle, apply to "any" number of concurrent clients. However, there is a huge trade-off between being able to guarantee that available bandwidth is never "over-reserved" at any given time and being able to guarantee that available bandwidth is always fully utilized when needed.
[0450] According to one embodiment, a further challenge with this type of scheme is to ensure that it interoperates well with dynamic congestion control and that congestion related to shared paths for multiple targets is handled in a coordinated manner within each transmitter.
[0451] High availability and failover: In one embodiment, in addition to performance, key attributes of a private fabric may include redundancy and the ability to failover communications following any single point of failure without loss of service for any client application. Furthermore, while "loss of service" represents a binary condition (i.e., service exists or is lost), several equally important, but more scalar, attributes include the extent of the power outage during failover, and if so, how long it is. Another important aspect is the extent to which the expected performance is provided (or re-established) during and after the completion of the failover operation.
[0452] In one embodiment, from the perspective of a single node (server), the goal is that any single point of failure in the fabric communication infrastructure outside the server itself (i.e., including a single local HCA) should not mean that the node becomes unable to communicate. However, from the perspective of the entire fabric, there is also the question of what level of overall fabric throughput and performance impact the loss of one or more components has. For example, in the case of a topology size that can operate with only two spine switches, is it acceptable that if one of the spines goes out of service, 50% of the communication capacity between leaves is lost, given the increased risk of binary bandwidth and congestion? According to one embodiment, another issue with per-tenant SLAs is the extent to which the ability to reserve and / or prioritize fabric resources for tenants with premium SLAs should be reflected, in that tenants with premium SLAs receive a proportionally larger share of the remaining available resources following a failure and subsequent failover operation—that is, thus the impact of a failure is less for premium SLA tenants but comes at the expense of more impact for other tenants.
[0453] In one embodiment, with respect to redundancy, the initial resource provisioning for such tenants could also be a “super-premium SLA attribute” in that any single point of failure would not mean that the associated performance / QoS SLA could not be met, either during or after a failure. However, the fundamental problem with such over-provisioning is that extremely fast failover (and failback / rebalancing) must exist to ensure that available resources are always utilized in the most optimal way, and that communication is never interrupted for more than a very short period as a result of any single point of failure.
[0454] According to one embodiment, an example of such a "super premium" setup could be a system with dual HCA-based servers, where both HCAs are active-active. It operates in a tive configuration, and both HCA ports are also used in an active-active configuration with an APM (Automatic Route Migration) scheme that has very little latency before alternative routes are attempted.
[0455] Route selection: According to one embodiment, when multiple possible paths exist between two endpoints, the selection of the best or "correct" path for the associated RDMA connection should ideally be automatic, so that the communication workload experiences the best possible performance within the constraints of the associated SLA, and so that system-level fabric resources are utilized in the most optimal way.
[0456] In one embodiment, ideally this would mean that the application logic within the VM does not need to deal with which local HCAs and which local HCA ports can or should be used for which communications. This also means node-level rather than port-level addressing schemes, meaning that the underlying fabric infrastructure is used transparently to the application.
[0457] According to one embodiment, the relevant workload can thus be more easily deployed on different infrastructures without requiring explicit handling of different system types or system configurations.
[0458] Features: According to one embodiment, the features of this category are assumed to be supported by existing HCA and / or switch hardware using current firmware and software.
[0459] According to one embodiment, the main objectives in this category are as follows: • The ability to limit the total egress bandwidth generated by a single VM or VF belonging to a tenant for each local physical HCA instance. The ability to ensure that a single VM or VF belonging to a tenant of a local physical HCA will be able to utilize at least a minimum percentage of the available local link bandwidth. The ability to restrict which network (Enet) priority VF can use. ○This could mean that multiple VFs must be allocated in order for a single VM to use multiple priorities (i.e., unless priority restrictions are enabled, a single VF can only be allowed to use a single priority). The ability to restrict which ISLs can be used by groups of flows belonging to a single tenant or a group of tenants.
[0460] In one embodiment, an "HCA Resource Limiting Group" (hereinafter referred to as "HRLG") is established for a tenant that shares a physical HCA with other tenants to control HCA usage by that tenant. The HRLG can be configured with a maximum bandwidth that defines the actual data rate that the HRLG can generate, and can also be configured with a minimum bandwidth share that ensures the HRLG achieves at least a specified percentage of the HCA bandwidth when there is circuit competition with other tenants / HRLGs. Unless there is circuit competition with other HRLGs, the VF within the HRLG can be used indefinitely up to the specified rate (or link capacity if no rate limit is defined).
[0461] According to one embodiment, HRLG is up to the number of V that the HCA instance can support. It may include F. Within the HRLG, each VF is expected to receive a fair share of the “allocation” assigned to the HRLG. For each VF, the associated QP will also receive a fair share of access to the local link, as a function of the available HRLG allocation, and any current flow control limits on the QR (i.e., if the QP receives congestion control feedback ordering it to throttle itself, or if it currently has no “credit” to send with the associated priority, the QP will not be considered for local link access). In one embodiment, it is conceivable that restrictions on the priorities that a VF can use could be enforced within the HRLG. As long as this restriction can only be defined for a single priority allowed for a VF, the implication is that a VM assumed to use multiple priorities (though still limited to only a few priorities) would have to use multiple VFs—one for each required priority. (Note: The use of multiple VFs implies that sharing local memory resources between multiple QPs using different priorities would likely present problems, because it would mean that different VFs would have to be allocated and used by ULPs / applications within the VM, depending on which priority restrictions / enforcement policies are defined.) According to one embodiment, within a single HRLG, there is no difference in bandwidth allocation depending on which priority VF / QP is currently using - they all share the relevant allocation in a fair / equal manner. Therefore, in order to associate different bandwidth allocations with different priorities, it is necessary to define one or more dedicated HRLGs that will contain only VFs that are restricted to using the priority that should be associated with the shared allocation represented by the relevant HRLG. In this way, a tenant with multiple VMs sharing the same physical HCA may be given different bandwidth allocations for different priorities.
[0462] In one embodiment, current hardware priority limits prevent data attempted to be sent with an incorrect priority from being sent over the external link, but do not prevent the fetching of the relevant data from local memory. Therefore, if the local memory bandwidth that the HCA can maintain in the egress direction is approximately the same as the available external link bandwidth, there is still some overall HCA link bandwidth that is wasted. However, if the relevant memory bandwidth is (significantly) larger than the external link bandwidth, attempts to use illegal priorities will waste less external link bandwidth, as long as the HCA pipeline operates at optimal efficiency. Still, unless there is much savings in terms of external link bandwidth, a possible alternative scheme to prevent the use of illegal priorities may be to leverage ACL rule enforcement at the switching progress port. If the relevant tenants can be effectively identified without any possibility of spoofing, this can be used to implement tenant / priority associations without the need to allocate individual VFs for each priority to the same VM. However, ensuring that packet / tenant associations are always clearly defined and cannot be spoofed from the sending VM, and dynamically updating the relevant switch ports to perform the relevant enforcement whenever a VF is set up for use by a VM / tenant, both represent a complexity that is not trivial. One possible scheme would be to use a per-VF port MAC to represent a spoof-proof identity that can be associated with a VM / tenant. However, this is not straightforward if VxLAN or other overlay protocols are used, especially unless the external switch is assumed to be involved in (or aware of) the overlay scheme being used.
[0463] According to one embodiment, in order to restrict which flows can use which ISLs, the switch forwarding logic needs to have policies to identify the relevant flows and configure forwarding accordingly. One example is using VLAN IDs to represent flow groups. Different tenants can map different VLAN IDs on the fabric. In this case, one possible scheme would be for the switch to dynamically achieve LAG type balance based on which VLAN IDs are allowed for various ports in any LAG or other port grouping. Another option would involve explicit packet forwarding based on the combination of destination address and VLAN ID.
[0464] According to one embodiment, if a VxLAN-based overlay is used transparently to the physical switch fabric, it would be possible to map different overlays to different VLAN IDs in order to enable the switch to map VLAN IDs to ISLs, as outlined above.
[0465] According to one embodiment, another possible scheme is that the forwarding of individual endpoint addresses is set up according to a routing scheme that takes into account VLAN membership or some other concept of “tenant” association. However, the VLAN ID must be part of the forwarding decision, insofar as the same endpoint address value is permissible in different VLANs.
[0466] According to one embodiment, the delivery of per-tenant flows to either a shared or exclusive ISL may require a holistic routing scheme to deliver traffic in a globally optimized manner within the fabric (fat tree) topology. While the implementation of such a scheme would typically rely on an SDN-type management interface for the switches, the implementation of holistic routing would not be trivial.
[0467] Short-term and medium-term SLA classes: In one embodiment, the following assumes that a non-blocking two-tier fat tree topology is used for a system size (physical node count) exceeding the cardinality of a single leaf switch. It is also assumed that a single VM on a physical server can utilize all fabric bandwidth (via one or more vHCAs / VFs). Therefore, the number of VMs per tenant per physical server is not a parameter that should be considered a tenant-level SLA factor from an HCA / fabric perspective.
[0468] According to one embodiment, the top tier (for example, Premium Plus) is: • A dedicated server can be used. • VMs (or HA policies) can be assigned to the same leaf domain whenever possible, unless the number and size of VMs impose additional distance. If a tenant uses multiple leaf domains (i.e., relative to the number of servers allocated to this tenant within the same leaf domain), on average, it may have non-blocking uplink bandwidth from the local leaf, but will not have dedicated uplinks or uplink bandwidth. • All "flow groups" (i.e., priorities representing different buffer pools and arbitration groups within the fabric) can be used.
[0469] According to one embodiment, the lower tier (e.g., premium) is: • A dedicated server can be used, but there is no guarantee that it will be the same leaf domain. On average, it can have at least 50% non-blocking uplink bandwidth (i.e., relative to the number of servers allocated to this tenant within the same leaf domain), but will not have dedicated uplinks or uplink bandwidth. • All "flow groups" can be used.
[0470] According to one embodiment, the third tier (e.g., Economy Plus) is: • A shared server may be used, but this will require a dedicated "flow group". These resources will be dedicated to the local HCA and switch ports, but will be shared within the fabric. • It may be possible to use all available egress bandwidth from the local server, but it is guaranteed that it will have at least 50% of the total egress bandwidth. • Only one Economy Plus tenant per physical server. On average, this Economy Plus tenant can have at least 25% non-blocking leaf uplink bandwidth relative to the number of servers used.
[0471] According to one embodiment, the fourth layer (e.g., economy) is: • A shared server can be used. • No dedicated priority • You may be able to use up to 50% of the server egress bandwidth, but this may be shared with up to three other economy tenants. On average, up to 25% of the available leaf uplink bandwidth can be shared with other economy tenants within the same leaf domain.
[0472] According to one embodiment, the lowest layer (for example, standby) is: • It may be possible to use reserve capacity without guaranteed bandwidth. Longer-term characteristics: According to one embodiment, the main features discussed in this section are: the ability to enforce priority restrictions on each VF in a manner that allows a single VF to be restricted to using any subset of the entire set of supported priorities. Whenever a data transfer is attempted using a priority that is not permitted for the starting VF, the wasted memory bandwidth or link bandwidth will be reduced to zero. The ability to limit the egress rate to different individual priorities across multiple HRLGs so that each VF within various HRLGs obtains their fair share of the total minimum bandwidth and / or maximum rate of the associated HRLGs, but subject to constraints defined by various associated priority-based assignments. • The ability to perform sender bandwidth control and congestion adjustment based on both per-target and per-shared path / route, with this aggregated at both the VM / VF (i.e., vHCA port) level and the HCA port level. • Ability to limit the gross average transmit and RDMA write ingress bandwidth to VF based on receiver-side throttling on the cooperating remote transmitter. • Ability to limit the average RDMA read ingress bandwidth to the VF without relying on a collaborating remote RDMA read response side. The ability to limit the total average ingress bandwidth to the VF based on receiver-side throttling by a cooperating remote transmitter, including RDMA read in addition to transmit and RDMA write. • The ability of a tenant VM to observe the available binary bandwidth for a group of different peer VMs. • SDN features for routing control and arbitration policies within the fabric.
[0473] In one embodiment, the HCA VF context can be extended to include a list of legitimate priorities (similar to a set of legitimate SLs for an IBTA IB vPort). Whenever a work request attempts to use a priority that is not legitimate for the VF, that work request should fail before any local or remote data transfer is initiated. In addition, priority mapping can also be used to give applications the illusion that any priority can be used. However, this type of mapping, where multiple priorities can be mapped to the same value before a packet is sent, is problematic for applications in that it associates different flow types with different "QoS classes". A drawback is that it may no longer be possible to control the QoS policy. Such restricted mappings represent SLA attributes (i.e., more privileged SLAs mean more actual priority after mapping). However, it is always important that applications can decide which flow types to associate with which QoS class (priority), in a manner that also represents independent flows in the fabric.
[0474] Enabling ingress RDMA read BW allocation via target groups and dynamic BW allocation updates: In one embodiment, the association of a target group with respect to flows from a “producer / source” node implies bandwidth throttling for all outgoing data packets, including UD transmits, RDMA writes, RDMA transmits, and RDMA reads (i.e., RDMA read responses with data), to the extent that this implies complete control over all ingress bandwidth to the vHCA port. This is independent of whether the VM owning the target vHCA port is generating an “excessive” amount of RDMA read requests to multiple peer nodes.
[0475] According to one embodiment, as discussed above, coupling target groups to both flow-specific and "unclaimed" BECN signaling means that the ingress bandwidth per vHCA port can be dynamically throttled for any number of remote peers.
[0476] In one embodiment, the “unsolicited BECN” messages outlined above can also be used to communicate specific rate values in addition to pure CE flagging / unflaggling for different stage numbers. In this way, it is possible to have a scheme in which an initial incoming packet from a new peer (e.g., a CM packet) can trigger the generation of one or more “unsolicited BECN” messages to both the HCA (i.e., the associated firmware / hyper-privileged software) from which the incoming packet originated and the current communicating peer.
[0477] According to one embodiment, in cases where both ports on the HCA are used simultaneously (i.e., an active-active scheme), it may make sense to share target groups between local HCA ports if concurrent flows are sharing several ISLs or even targeting the same destination port.
[0478] According to one embodiment, another reason for sharing target groups between HCA ports is whether the HCA local memory bandwidth can maintain full-speed link speed for both (all) HCA ports. In this case, the target group can be configured such that the aggregated total link bandwidth never exceeds the local memory bandwidth, regardless of whether each port is involved with the source HCA or the destination HCA.
[0479] In one embodiment, for a fixed route destined for a specific destination, any intermediate target group will typically represent only a single ISL at a particular stage in the path. However, when dynamic forwarding is active, both target groups and ECN processing must take this into account. If dynamic forwarding decisions arise only to balance traffic between parallel ISLs between a pair of switches (e.g., an uplink from a single leaf switch to a single spine switch), then all processing is, in principle, very similar to the case where only a single ISL is used. ECN notifications appear to be based on the state of all ports in the group in question, and signaling can be "aggressive" in the sense that it signals based on congestion indications from any of the ports, or it can be more conservative and based on the size of the shared output queue for all ports in the group. The target group configuration is such that any packet As long as it allows forwarding to select the best output port at any given time, it will typically represent the aggregated bandwidth for all links in a group. However, if there is a concept of strict packet ordering per flow, evaluating bandwidth allocation becomes more complex because some flows "must" use the same ISL at some point. If such a flow ordering scheme is based on well-defined header fields, it may be best to represent each port in a group as an independent target group. In this case, the selection of target groups at the source HCA must be able to evaluate the header fields, which will be associated with RC QP connections or address handles, in the same way that the switch does at runtime for all packets.
[0480] According to one embodiment, by default, the initial target group rate for a new remote target may be set conservatively low. In this way, there is inherent throttling until the target has an opportunity to update the associated rate. Thus, all such rate control is independent of the VMs involved, although a VM could request the hypervisor to update different remote peer assignments for both ingress and egress traffic, but this would only be permitted within the total constraints defined for both local and remote vHCA ports.
[0481] Correlation between peer nodes, routes, and target groups: In one embodiment, in order for a VM to be able to identify the bandwidth limits associated with different peer nodes and different groups of peer nodes, a method would be needed to query which target groups are associated with different communication peers (and associated address / routing information). By correlating a set of communication peers with different target groups, and based on the rate limits represented by the different target groups, the VM could track what bandwidth can be achieved for different communication peers. This would then allow the VM to schedule communication operations in a manner that, in principle, achieves the best possible bandwidth utilization over time by having as many concurrent forwardings as possible that do not involve competing target groups.
[0482] Relationship between HCA resource restriction groups and target groups: In one embodiment, the HRLG concept and the target group concept overlap in some ways, in that they both represent bandwidth limits that can be defined and shared between VMs and tenants in a flexible manner. However, the primary focus of HRLG is to define how different portions of local HCA / HCA port capacity can be allocated to different VFs (and thereby VMs and tenants), while the target group concept focuses on bandwidth limits and flow control constraints that exist outside the local HCA with respect to both the final destination and intermediate fabric topology.
[0483] In one embodiment, HRLG is used as a way to control what share of local HCA capacity various VFs may use, but it is meaningful to ensure that the permitted capacity can only be used in a manner that does not conflict with any fabric or remote target limitations or congestion conditions. These external constraints are then dynamically controlled and reflected via the relevant target group.
[0484] According to one embodiment, from an implementation standpoint, the state of all relevant target groups defines which pending work requests for which local QPs are in a flow control state, where they are more egress days at any given time. It is permitted to generate egress traffic. This state, along with the state of whether what the QP does actually has something to send, can then be aggregated at the VF / vHCA port level regarding which VF is the next candidate to send. The decision of which VF to schedule for next transmission on the HCA port will be based on the various HRLG states and policies in the HRLG hierarchy, the set of "ready to send" VFs, and the recent history of which VFs have generated what kind of egress traffic. For the selected VF, the VF-specific arbitration policy will define which QP will be selected for data transfer.
[0485] According to one embodiment, a set of QPs with pending data transfers includes both QPs with local work requests and QPs with pending RDMA read requests from associated remote peers, so that the scheduling and arbitration described above will handle all pending egress data traffic.
[0486] In one embodiment, ingress traffic (including incoming RDMA read responses) is controlled by the current state of all relevant target groups on the remote peer node. This (remote) state includes both a dynamic flow control state based on congestion conditions and explicit updates from this HCA that reflect changes in ingress bandwidth allocation to local VFs on this HCA. Such ingress bandwidth allocation is based on policies reflected by the HRLG hierarchy. In this way, various VMs may have independent "fine-tuned" bandwidth allocations for both ingress and egress, and also based on priority for both ingress and egress.
[0487] SLA Class: In one embodiment, the following proposal assumes that a non-blocking two-tier fat tree topology is used for a system size (physical node count) exceeding the cardinality of a single leaf switch. It also assumes that a single VM on a physical server can utilize all fabric bandwidth (via one or more HCAs / VFs). Therefore, the number of VMs per tenant per physical server is not a parameter that should be considered an SLA factor.
[0488] According to one embodiment, the highest level tier (e.g., Premium Plus) is: Only dedicated servers can be used. • VMs (or HA policies) can be assigned to the same leaf domain whenever possible, unless the number and size of VMs impose additional distance. • If a tenant uses multiple leaf domains, it can always have non-blocking uplink bandwidth from the local leaf. • All "flow groups" can be used.
[0489] According to one embodiment, a lower-level tier (e.g., premium) can be provided. • Only dedicated servers can be used, and there is no guarantee regarding the same leaf domain. • At least 50% of non-blocking uplink bandwidth can be guaranteed (i.e., relative to the number of servers allocated to this tenant within the same leaf domain). • All "flow groups" can be used.
[0490] According to one embodiment, a third level tier can be provided (e.g., Economy Plus). • A shared server may be used, but this would result in four dedicated "flow groups" (i.e., different buffer pools and arbitration groups within the fabric, each with its own priority). ). These resources will be dedicated to the local HCA and switch ports, but will be shared within the fabric. • It may be possible to use all available bandwidth (egress and ingress) from the local server, but it is guaranteed that at least 50% of the total bandwidth will be available. • Limited to one Economy Plus tenant per physical server. • This Economy Plus tenant can be guaranteed at least 25% of the number of servers used by the tenant, which includes non-blocking leaf uplink bandwidth.
[0491] According to one embodiment, a fourth layer can be provided (e.g., economy). • Using a shared server is the only option. • Does not have a dedicated priority • You may be able to use up to 50% of the server bandwidth (egress and ingress), but this may be shared with up to three other economy tenants. • Up to 25% of the available leaf uplink bandwidth can be shared with other economy tenants within the same leaf domain.
[0492] According to one embodiment, it is possible to provide a bottom layer (e.g., standby) that can use spare capacity without guaranteed bandwidth. While various embodiments of this teaching have been described, it should be understood that these embodiments are presented as examples only and not as limitations. The embodiments have been selected and described to illustrate the principles of this teaching and their practical applications. The embodiments illustrate systems and methods to which this teaching is used to improve the performance of systems and methods by providing new and / or improved features, and / or by providing benefits such as reduced resource utilization, increased capacity, improved efficiency, and reduced latency.
[0493] In some embodiments, the features of this teaching are implemented, in whole or in part, in a computer including a processor, a storage medium such as memory, and a network card for communicating with other computers. In some embodiments, the features of this teaching are implemented in a network such as a Local Area Network (LAN), a Switch Fabric Network (e.g., InfiniBand), or a Wide Area Network (WAN). This is realized in a connected distributed computing environment. The distributed computing environment may have all the computers in one location, or it may have clusters of computers in various remote geographic locations connected by a WAN.
[0494] In some embodiments, the features of this teaching are realized in the cloud as part of or as a service of a cloud computing system, based on shared and flexible resources delivered to users in a coordinated, self-service manner using web technologies, in whole or in part. There are five characteristics of the cloud (as defined by the National Institute of Standards and Technology): on-demand self-service, wide-area network access, resource pooling, high-speed scalability, and measured services. See, for example, the NIST Definition of Cloud Computing (Special Publication 800-145 (2011)), which is incorporated herein by reference. Cloud deployment models include public, private, and hybrid. Cloud service models are based on Software as a Service (SAS). Platform as a Service (PaaS) , Database as a Service (DBaaS) and Infrastructure as a Service (IaaS) This includes. As used herein, "cloud" is a combination of hardware, software, network, and web technologies that deliver shared, flexible resources to users in a self-service, coordinated manner. Unless otherwise specified, as used herein, "cloud" encompasses embodiments of public cloud, private cloud, and hybrid cloud, and all cloud deployment models include, but are not limited to, cloud SaaS, cloud DBaaS, cloud PaaS, and cloud IaaS.
[0495] In some embodiments, the features of the teaching invention are realized using or with the help of hardware, software, firmware, or a combination thereof. In some embodiments, the features of the teaching are realized using a processor configured or programmed to perform one or more functions of the teaching. In some embodiments, the processor may be a single-processor or multi-chip processor, a digital signal processor (DSP), a system on a chip (SOC), an application-specific integrated circuit (ASIC), or a field-programmable gate array. (Field programmable gate array: FPGA) or other programmable logic devices This includes vices, state machines, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. In some implementations, the features described herein may be implemented by circuits specialized for specific functions. In other implementations, these features may be implemented, for example, in a processor configured to perform specific functions using instructions stored on a computer-readable storage medium.
[0496] In some embodiments, the features of the teachings are incorporated into software and / or firmware to control the hardware of a processing system and / or networking system, and to enable the processor and / or network to interact with other systems that utilize the features of the teachings. Such software or firmware may include, but is not limited to, application code, device drivers, operating systems, virtual machines, hypervisors, application programming interfaces, programming languages, and execution environments / containers. Appropriate software coding can be readily prepared by a skilled programmer based on the teachings of the disclosure, as will be apparent to a person skilled in the art familiar with software technology.
[0497] In some embodiments, the method includes a computer program product, such as a computer-readable medium, that carries instructions that can be used to carry out the teaching. In some examples, the computer-readable medium is a storage medium or computer-readable medium that carries instructions by storing those instructions, and a system such as a computer can be programmed or otherwise configured to use these instructions to perform any of the processes or functions of the teaching. The storage medium or computer-readable medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), and any type of medium or device suitable for storing instructions and / or data. In certain embodiments, the storage medium or computer-readable medium is a non-temporary storage medium. or non-transient computer-readable media. Computer-readable media may also, or alternatively, include transient media such as carrier waves or transmission signals that propagate such instructions.
[0498] Accordingly, from one perspective, systems and methods for providing bandwidth congestion control in a private fabric in a high-performance computing environment have been described. An exemplary method may provide a first subnet in one or more microprocessors, the first subnet comprising a plurality of switches and a plurality of host channel adapters, each host channel adapter comprising at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, and the first subnet further comprising a plurality of end nodes. The method may provide an end node ingress bandwidth allocation associated with an end node attached to the host channel adapter. The method may allow an end node of the host channel adapter to receive ingress bandwidth, the ingress bandwidth exceeding the end node's ingress bandwidth allocation.
[0499] The foregoing description is not intended to be exhaustive or to limit the scope to the forms disclosed herein. Furthermore, while embodiments of the invention have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope is not limited to the aforementioned set of transactions and steps. Moreover, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this teaching. Furthermore, while specific combinations of features of the invention have been described in various embodiments, it should be understood that different combinations of these features are also within the scope of this teaching, such that features of one embodiment may be incorporated into another. Finally, it will be apparent to those skilled in the art that various additions, reductions, deletions, modifications, and other changes to the form, details, implementation, and use may be made without departing from the spirit and scope of this invention. The present invention is intended to be defined by a proper interpretation of the claims.
Claims
1. A system for providing bandwidth congestion control in a private fabric in a high-performance computing environment, A computer containing one or more microprocessors, The computer provides a first subnet, and the first subnet is Multiple switches, It includes a plurality of host channel adapters, the plurality of host channel adapters being interconnected via the plurality of switches, Multiple target groups are defined within the first subnet, and each target group is defined in the inter-switch link between two of the multiple switches included in the first subnet. Each target group defines a bandwidth limit on the associated inter-switch link of the first subnet. Each record of the aforementioned multiple target groups is stored in the target group repository of one of the aforementioned multiple host channel adapters. A system in which two or more simultaneous flows, each having its own flow rate, are simultaneously subject to the bandwidth limit of at least one target group.
2. The system according to claim 1, wherein at least two of the plurality of target groups define different bandwidth limits in the inter-switch links associated with each of the first subnets.
3. The system according to claim 2, wherein the different bandwidth limits are defined based on the quality of service class defined in each of the associated inter-switch links of the first subnet.
4. The system according to claim 1, wherein at least a portion of the plurality of target groups includes a hierarchy of target groups.
5. For the hierarchy of the target group, a maximum bandwidth is defined, and the maximum bandwidth is, The system according to claim 4, wherein the rate capacity is at least the most limited among at least some of the plurality of target groups.
6. The system according to any one of claims 1 to 5, wherein at least one of the flow rates of the two or more simultaneous flows is modified based on the bandwidth limit of the at least one target group.
7. A method for providing bandwidth congestion control in a private fabric in a high-performance computing environment, A computer comprising one or more microprocessors includes providing a first subnet, wherein the first subnet is Multiple switches, The method includes a plurality of host channel adapters, the plurality of host channel adapters being interconnected via the plurality of switches, and the method further includes The method further includes defining a plurality of target groups within the first subnet, each target group being defined on an inter-switch link between two of the plurality of switches included in the first subnet, and the method further includes: Each target group defines a bandwidth limit on the associated inter-switch link of the first subnet, This includes storing each record of the plurality of target groups in the target group repository of one of the host channel adapters among the plurality of host channel adapters, A method in which two or more simultaneous flows, each having its own flow rate, are simultaneously subject to the bandwidth limit of at least one target group.
8. The method according to claim 7, wherein at least two of the plurality of target groups define different bandwidth limits in the inter-switch links associated with each of the first subnets.
9. The method according to claim 8, wherein the different bandwidth limits are defined based on the quality of service class defined in each of the associated inter-switch links of the first subnet.
10. The method according to claim 7, wherein at least a portion of the plurality of target groups includes a hierarchy of target groups.
11. The method according to claim 10, wherein a maximum bandwidth is defined for the hierarchy of the target groups, and the maximum bandwidth is at least the most restricted rate capacity among at least some of the plurality of target groups.
12. The method according to any one of claims 7 to 11, wherein at least one of the flow rates of the two or more simultaneous flows is modified based on the bandwidth limit of the at least one target group.
13. A computer-readable program for causing a computer to perform the method described in any one of claims 7 to 12.
Citation Information
Patent Citations
Systems and Methods for Efficient Network Isolation and Load Balancing in a Multi-Tenant Cluster Environment
JP2018537004A
RDMA-over-ethernet storage system with congestion avoidance without ethernet flow control
US20170366460A1