System and method for supporting target groups for congestion control in private fabric in high performance computing environment

The system addresses network congestion in high-performance computing by defining target groups and using dynamic LID allocation in a vSwitch architecture, enhancing live migration and resource utilization in cloud computing.

JP2025170317APending Publication Date: 2025-11-18ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025136320
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-11
Filing Date
2025-08-19
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing cloud computing architectures face performance and management bottlenecks due to complex addressing and routing schemes in high-performance interconnects like InfiniBand, hindering live migration of virtual machines and causing network congestion.

Method used

Implementing a system and method for congestion control in a private fabric by defining target groups with bandwidth limits and using a target group repository, along with dynamic LID allocation and pre-populated LIDs in a vSwitch architecture, to manage and optimize network communication.

Benefits of technology

Enhances live migration transparency and reduces network downtime, ensuring efficient resource utilization and flexible VM placement while maintaining high-performance computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170317000001_ABST
    Figure 2025170317000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for supporting target groups for congestion control in a high performance computing environment.SOLUTION: A method provides, at one or more microprocessors, a first subnet comprising a plurality of switches, a plurality of host channel adapters, and a plurality of end nodes including a plurality of virtual machines. The method defines a target group on one of an inter-switch link or at a port of a switch of the plurality of switches. The target group defines a bandwidth limit on at least one of an inter-switch link between two switches of the plurality of switches or at a port of a switch of the plurality of switches. The method also provides a target group repository stored in a memory of a host channel adapter where the defined target group in the target group repository is recorded.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Copyright Notice A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of this patent document or the patent disclosure, provided that such reproduction is in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.

[0002] Priority claims and cross-references to related applications: This application claims the benefit of priority to U.S. Provisional Patent Application, Application No. 62 / 937,594, entitled "SYSTEM AND METHOD FOR PROVIDING QUALITY-OF-SERVICE AND SERVICE-LEVEL AGREEMENTS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT," filed November 19, 2019, which is incorporated herein by reference in its entirety.

[0003] This application also claims the benefit of priority to the following patent applications, each of which is incorporated herein by reference in its entirety: SYSTEM AND METHOD FOR SUPPORTING SYSTEMS, filed May 11, 2020; RDMA BANDWIDTH RESTRICTIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING U.S. patent application entitled "SYSTEM AND METHOD FOR SUPPORTING RDMA BANDWIDTH LIMITATIONS IN PRIVATE FABRIC IN HIGH-PERFORMANCE COMPUTING ENVIRONMENT," application serial number 16 / 872,035; U.S. patent application Ser. No. 16 / 872,038, entitled "SYSTEM AND METHOD FOR SUPPORTING TARGET GROUPS FOR CONGESTION CONTROL IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT," filed May 11, 2020; U.S. patent application Ser. No. 16 / 872,039, entitled "SYSTEM AND METHOD FOR SUPPORTING TARGET GROUPS FOR CONGESTION CONTROL IN A PRIVATE FABRICS IN A NETWORKING ENVIRONMENT"; and U.S. patent application Ser. No. 16 / 872,039, filed May 11, 2020, entitled "SYSTEM AND METHOD FOR SUPPORTING TARGET GROUPS FOR CONG METHOD FOR SUPPORTING USE OF FORWARD AND BACKWARD CONGESTION NOTIFICATIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT No. 16 / 872,043, entitled "System and Method for Supporting the Use of Forward and Backward Congestion Notification in a Private Fabric in a Networking Environment."

[0004] Field The present teachings relate to systems and methods for enforcing quality of service (QOS) and service level agreements (SLAs) in private, high-performance interconnection fabrics such as InfiniBand (IB) and RoCE (Remote Direct Memory Access (RDMA) over Converged Ethernet). [Background technology]

[0005] background As larger cloud computing architectures are deployed, the performance and management bottlenecks associated with traditional networking and storage become significant issues. There has been growing interest in using high performance, lossless interconnects such as InfiniBand (IB) technology as the foundation for cloud computing fabrics, and this is the general area that embodiments of the present teachings are intended to address. Summary of the Invention [Means for solving the problem]

[0006] overview: Particular aspects are set out in the independent claims. Various optional embodiments are set out in the dependent claims.

[0007] Described herein are systems and methods for supporting target groups for congestion control in a private fabric in a high-performance computing environment. An exemplary method can provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, a plurality of host channel adapters, and a plurality of end nodes including a plurality of virtual machines. The method can define a target group on an inter-switch link or one of a port of a switch of the plurality of switches, the target group defining a bandwidth limit on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches. The method can provide a target group repository stored in a memory of the host channel adapter, and the defined target group in the target group repository is recorded. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example of an InfiniBand environment according to one embodiment. [Figure 2] FIG. 1 illustrates an example of a split cluster environment according to one embodiment. [Figure 3] FIG. 1 illustrates an example of a tree topology in a network environment according to one embodiment. [Figure 4] FIG. 1 illustrates an exemplary shared port architecture according to one embodiment. [Figure 5] FIG. 1 illustrates an exemplary vSwitch architecture according to one embodiment. [Figure 6] FIG. 2 illustrates an exemplary vPort architecture according to one embodiment. [Figure 7] FIG. 2 illustrates an exemplary vSwitch architecture with pre-populated LIDs according to one embodiment. [Figure 8] FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID allocation according to one embodiment. [Figure 9] FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs to the vSwitch, according to one embodiment. [Figure 10] FIG. 1 illustrates an exemplary multi-subnet InfiniBand fabric according to one embodiment. [Figure 11] FIG. 2 illustrates an interconnection between two subnets in a high-performance computing environment, according to one embodiment. [Figure 12] FIG. 1 illustrates an interconnection between two subnets via a dual-port virtual router configuration in a high-performance computing environment, according to one embodiment. [Figure 13] FIG. 1 illustrates a flowchart of a method for supporting a dual-port virtual router in a high-performance computing environment, according to one embodiment. [Figure 14] 1 illustrates a system for servicing RDMA read requests as a restricted feature in a high performance computing environment, according to one embodiment. [Figure 15] 1 illustrates a system for servicing RDMA read requests as a restricted feature in a high performance computing environment, according to one embodiment. [Figure 16] 1 illustrates a system for servicing RDMA read requests as a restricted feature in a high performance computing environment, according to one embodiment. [Figure 17] 1 illustrates a system for providing explicit RDMA read bandwidth limiting in a high performance computing environment, according to an embodiment. [Figure 18] 1 illustrates a system for providing explicit RDMA read bandwidth limiting in a high performance computing environment, according to an embodiment. [Figure 19] 1 illustrates a system for providing explicit RDMA read bandwidth limiting in a high performance computing environment, according to an embodiment. [Figure 20]1 is a flowchart of a method for providing RDMA (Remote Direct Memory Access) read requests as restricted features in a high performance computing environment. [Figure 21] 1 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment. [Figure 22] 1 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment. [Figure 23] 1 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment. [Figure 24] 1 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment. [Figure 25] 1 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment. [Figure 26] 1 is a flowchart of a method for combining multiple shared bandwidth segments in a high performance computing environment, according to one embodiment. [Figure 27] 1 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment. [Figure 28] 1 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment. [Figure 29] 1 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment. [Figure 30]1 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment. [Figure 31] 1 is a flowchart of a method for combining target-specific RDMA write and read bandwidth limits in a high-performance computing environment, according to one embodiment. [Figure 32] 1 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment. [Figure 33] 1 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment. [Figure 34] 1 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment. [Figure 35] 1 is a flowchart of a method for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to one embodiment. [Figure 36] 1 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment. [Figure 37] 1 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment. [Figure 38] 1 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment. [Figure 39] 1 is a flowchart of a method for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description: The present teachings are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to like elements. It should be noted that references in this disclosure to "a" or "one" or "several" embodiments are not necessarily to the same embodiment, and such references mean at least one. While specific implementations are described, it is understood that these specific implementations are provided for illustrative purposes only. One skilled in the art will recognize that other components and configurations can be used without departing from the spirit and scope of the present invention.

[0010] Common reference numbers may be used to denote like elements throughout the figures and detailed description, and thus a reference number used in one figure may or may not be referenced in the detailed description specific to that figure if the element is described elsewhere.

[0011] Described herein are systems and methods for providing quality of service (QOS) and service level agreements (SLAs) in a private fabric in a high performance computing environment.

[0012] According to one embodiment, the following description of the present teachings will be directed to InfiniBand as an example of a high performance network. TM (IB) network. Throughout the following discussion, we will refer to InfiniBand TMThe following description may refer to the InfiniBand Trade Association Architecture Specification (variously referred to as the InfiniBand Specification, IB Specification, or Legacy IB Specification). Such references will be understood to refer to the InfiniBand Trade Association Architecture Specification, Volume 1, Version 1.3, published in March 2015 and available at http: / / www.inifinibanda.org, which is incorporated herein by reference in its entirety. It will be apparent to those skilled in the art that other types of high-performance networks may be used without limitation. The following description also uses a fat-tree topology as an example of a fabric topology. It will be apparent to those skilled in the art that other types of fabric topologies may be used without limitation.

[0013] According to an embodiment, the following description uses RoCE (Remote Direct Memory Access (RDMA) over Converged Ethernet). RDMA over Converged Ethernet (RoCE) is a standard protocol that enables efficient RDMA data transfer over Ethernet networks, allowing offload transport and superior performance through hardware RDMA engine implementations. RoCE is a standard protocol defined in the InfiniBand Trade Association (IBTA) standard. RoCE utilizes UDP (User Datagram Protocol) encapsulation, allowing it to cross Layer 3 networks. RDMA is a key capability used natively by InfiniBand interconnect technology. Both InfiniBand and Ethernet RoCE share a common user API but have different physical and link layers.

[0014] According to one embodiment, various portions of this specification include references to InfiniBand fabrics when describing various implementations, however, those skilled in the art will recognize that the various implementations described herein are not intended to be limiting. It will be readily appreciated that various embodiments may also be implemented in a RoCE fabric.

[0015] To meet the demands of modern clouds (e.g., the exascale era), virtual machines must be able to use low-overhead network communication paradigms such as Remote Direct Memory Access (RDMA). Because RDMA bypasses the OS stack and communicates directly with the hardware, pass-through technologies such as single-root I / O virtualization (SR-IOV) network adapters can be used. In accordance with one embodiment, a virtual switch (vSwitch) SR-IOV architecture can be applied to a high-performance lossless interconnect network. Because network reconfiguration time is critical to making live migration a viable option, a scalable, topology-independent dynamic reconfiguration mechanism can be provided in addition to the network architecture.

[0016] According to one embodiment, a routing strategy for a virtualized environment using a vSwitch can be further provided, and an efficient routing algorithm can be provided for a network topology (e.g., a fat-tree topology). The dynamic reconfiguration mechanism can be further tuned to minimize the overhead imposed on the fat-tree.

[0017] According to one embodiment of the present teachings, virtualization can be beneficial for efficient resource utilization and flexible resource allocation in cloud computing. Live migration enables optimizing resource utilization by moving virtual machines (VMs) between physical servers in a manner that is transparent to applications. Thus, virtualization enables consolidation, on-demand provisioning of resources, and flexibility through live migration.

[0018] InfiniBand TM InfiniBand TM (IB) stands for InfiniBand TM Trade Association (InfiniBand TM It is an open-standard, lossless network technology developed by the IEEE 802.11b Trade Association. The technology is based on a serial point-to-point full-duplex interconnect that provides high-throughput and low-latency communications, especially targeted at high-performance computing (HPC) applications and data centers.

[0019] InfiniBand TM ·Architecture (InfiniBand Architecture: IBA) is 2 Supports layer topology division. At the lower layers, an IB network is called a subnet, and a subnet may contain a set of hosts interconnected using switches and point-to-point links. At higher levels, an IB fabric consists of one or more subnets that may be interconnected using routers.

[0020] Within a subnet, hosts may be connected using switches and point-to-point links. In addition, there may be one master management entity, a subnet manager (SM), that resides on a designated device in the subnet. The subnet manager is responsible for configuring, starting, and maintaining the IB subnet. In addition, the subnet manager (SM) may be responsible for performing routing table calculations in the IB fabric. Here, for example, routing in an IB network aims to provide fair load balancing between all source-destination pairs in the local subnet.

[0021] Through the subnet management interface, the subnet manager exchanges control packets called subnet management packets (SMP) with the subnet management agent (SMA). An agent resides on all IB subnet devices. Using SMP, the subnet manager can discover the fabric, configure end nodes and switches, and receive notifications from the SMA.

[0022] According to one embodiment, routing within a subnet in an IB network is performed using a linear forwarding table (LF) stored in the switch. The LFT can be based on the Host Channel Adapter (Host Channel Adapter) on the end nodes. The LFT is calculated by the SM according to the routing mechanism in use. HCA (Hardware Adapter) ports and switches are addressed using local identifiers (LIDs). Each entry in the LFT contains a destination LID (DLI). D) and an output port. Only one entry per LID in the table is supported. When a packet arrives at a switch, its output port is determined by looking up the DLID in the switch's forwarding table. Routing is deterministic because packets follow the same path in the network between a given source-destination pair (LID pair).

[0023] Generally, all other subnet managers except the master subnet manager operate in standby mode for fault tolerance. However, in the situation where the master subnet manager fails, a new master subnet manager is elected by the standby subnet managers. The master subnet manager also performs periodic sweeps of the subnet to detect any topology changes and updates the network accordingly. Reconfigure the network.

[0024] Additionally, hosts and switches within a subnet may be addressed using a local identifier (LID), and a single subnet may be limited to 49151 unicast LIDs. In addition to the LID, which is a local address valid within the subnet, each IB device may have a 64-bit global unique identifier (GUID). The GUID may be used to form a global identifier (GID), which is an IB Layer 3 (L3) address.

[0025] The SM may calculate routing tables (i.e., connections / routes between each pair of nodes in a subnet) at network initialization time. Additionally, whenever the topology changes, the routing tables may be updated to ensure connectivity and optimal performance. During normal operation, the SM may perform periodic light sweeps of the network to check for topology changes. If a change is discovered during a light sweep, Alternatively, if the SM receives a message (trap) signaling a network change, the SM may reconfigure the network according to the discovered change.

[0026] For example, the SM may reconfigure the network when the network topology changes, such as when a link goes down, a device is added, or a link is removed. The reconfiguration step may include a step performed during network initialization. Furthermore, the reconfiguration may have a local scope that is limited to the subnet where the network change occurred. Also, segmentation of a large fabric using routers may limit the reconfiguration scope.

[0027] An example of an InfiniBand fabric is shown in FIG. 1, which illustrates an example of an InfiniBand environment 100 according to one embodiment. In the example shown in FIG. 1, nodes A101 to E105 are The nodes communicate via respective host channel adapters 111-115 using a two-band fabric 120. According to one embodiment, the various nodes (e.g., nodes A101-E105) may be represented by various physical devices. According to one embodiment, the various nodes (e.g., nodes A101-E105) may be represented by various virtual devices, such as virtual machines.

[0028] Partitioning in InfiniBand According to one embodiment, an IB network may support partitioning as a security mechanism for isolating logical groups of systems that share a network fabric. Each HCA port on a node in the fabric may be a member of one or more partitions. Partition membership is managed by a centralized partition manager, which may be part of the SM. The SM may organize the partition membership information for each port as a table of 16-bit partition keys (P_Key). The SM may also manage these Switch and router ports can be configured with a partition enforcement table that contains P_Key information associated with end nodes that send or receive data traffic through the port. Additionally, in the general case, the partition membership of a switch port may represent the collection of all memberships indirectly associated with LIDs routed through the port in the egress direction (towards the link).

[0029] According to one embodiment, a partition is a logical group of ports, and members of a group can only communicate with other members of the same logical group. Isolation can be enforced in host channel adapters (HCAs) and switches by filtering packets using partition membership information. Packets with invalid partitioning information can be dropped as soon as they reach the ingress port. In a partitioned IB system, partitions can be used to create tenant clusters. With partitions in place, nodes cannot communicate with other nodes that belong to different tenant clusters. In this way, the security of the system can be guaranteed even in the presence of faulty or malicious tenant nodes.

[0030] According to one embodiment, for communication between nodes, queue pairs (QP) and end-to-end contexts (EEC), with the exception of the management queue pair (QP0 and QP1), can be assigned to specific partitions. P_Key information can then be added to all transmitted IB transport packets. When a packet arrives at an HCA port or switch, its P_Key value can be checked against a table configured by the SM. If an invalid P_Key value is found, the packet is immediately discarded. In this way, communication is only allowed between ports that share a partition.

[0031] An example of an IB partition is shown in Figure 2, which illustrates an example of a partitioned cluster environment, according to one embodiment. In the example shown in Figure 2, nodes A 101-E 105 communicate via their respective host channel adapters 111-115 using an InfiniBand fabric 120. Nodes A-E are arranged into partitions: partition 1 130, partition 2 140, and partition 3 150. Partition 1 includes node A 101 and node D 104. Partition 2 includes node A 101, node B 102, and node C 103. Partition 3 includes node C 103 and node E 105. With this arrangement of partitions, node D 104 and node E 105 share a single partition. On the other hand, for example, node A 101 and node C 103 can communicate because they are both members of partition 2 140.

[0032] Virtual Machines on InfiniBand Over the past decade, hardware virtualization support has virtually eliminated CPU overhead, memory overhead has been significantly reduced by virtualizing the memory management unit, storage overhead has been reduced by utilizing high-speed SAN storage or distributed network file systems, and device pass-through technologies such as Single Root Input / Output Virtualization (SR-IOV) have been introduced. The prospects for virtualized High Performance Computing (HPC) environments have improved significantly as network I / O overhead has been reduced through the use of high-performance interconnect solutions. Clouds now support virtual HPC (vHPC) clusters with high-performance interconnect solutions, providing the necessary performance. It is possible to provide

[0033] However, when coupled with lossless networks such as InfiniBand (IB), some cloud features such as live migration of virtual machines (VMs) remain problematic due to the complex addressing and routing schemes used in these solutions.IB is an interconnect network technology that offers high bandwidth and low latency, making it well suited for HPC and other communication-intensive workloads.

[0034] The traditional approach to connecting IB devices to VMs is by using direct-assigned SR-IOV. However, achieving live migration of VMs assigned to IB host channel adapters (HCAs) using SR-IOV has proven challenging. Each IB-attached node has three different addresses (i.e., LID, GUID, and GID). When a live migration occurs, one or more of these addresses change. Other nodes communicating with the migrating VM (VM-in-migration) may lose connectivity. This When this occurs, the IB subnet manager (SM) is notified that it should reconnect by sending a Subnet Administration (SA) record route query. An attempt can be made to restore the lost connection by finding out the new address of the virtual machine.

[0035] IB uses three different types of addresses. The first type of address is a 16-bit local identifier (LID). At least one unique LID is assigned by the SM to each HCA port and each switch. The LID is used to route traffic within a subnet. Because the LID is 16 bits long, 65,536 unique address combinations can be configured, of which only 49,151 (0x0001-0xBFFF) can be used as unicast addresses. As a result, the number of available unicast addresses defines the maximum size of an IB subnet. The second type of address is a 64-bit globally unique identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM may assign an additional subnet-specific GUID to an HCA port, which is useful when SR-IOV is used. The third type of address is a 128-bit global identifier (GID). A GID is a valid IPv6 unicast address, at least one of which is assigned to each HCA port. The GID is formed by combining a globally unique 64-bit prefix assigned by the fabric administrator with the GUID address of each HCA port.

[0036] Fat Tree (FTree) Topology and Routing According to one embodiment, some IB-based HPC systems employ a fat tree topology to take advantage of the useful properties that fat trees offer, including full bisection bandwidth and inherent fault tolerance due to the availability of multiple paths between each source-destination pair. The initial concept behind fat trees was to employ thicker links between nodes with more available bandwidth as the tree approached the root of the topology. The thicker links could help avoid congestion in higher-level switches, preserving bisection bandwidth.

[0037] 3 illustrates an example of a tree topology in a network environment, according to one embodiment. As shown in FIG. 3, one or more end nodes 201-204 may be connected in a network fabric 200. The network fabric 200 may be based on a fat-tree topology including multiple leaf switches 211-214 and multiple spine or root switches 231-234. In addition, the network fabric 200 may include one or more intermediate switches, such as switches 221-224.

[0038] 3, each of end nodes 201-204 may be a multi-homed node, i.e., a single node that is connected to two or more portions of network fabric 200 via multiple ports. For example, node 201 may include ports H1 and H2, node 202 may include ports H3 and H4, node 203 may include ports H5 and H6, and node 204 may include ports H7 and H8.

[0039] Additionally, each switch may have multiple switch ports. For example, root switch 231 may have switch ports 1-2, root switch 232 may have switch ports 3-4, root switch 233 may have switch ports 5-6, and root switch 234 may have switch ports 7-8.

[0040] According to an embodiment, the fat-tree routing mechanism is one of the most popular routing algorithms for IB-based fat-tree topologies. The fat-tree routing mechanism is also integrated into OFED (Open Fabric Enterprise Distribution: a standard software for building and deploying IB-based applications). This is implemented in the OpenSM (hardware stack) subnet manager.

[0041] The goal of a fat-tree routing mechanism is to generate an LFT that uniformly spreads shortest-path routes across links in the network fabric. The mechanism traverses the fabric in indexing order and assigns target LIDs for end nodes, and therefore corresponding routes, to each switch port. For end nodes connected to the same leaf switch, the indexing order may depend on the switch ports to which the end nodes are connected (i.e., the port numbering sequence). For each port, the mechanism may maintain a port usage counter, and each time a new route is added, the port usage counter may be used to select the least frequently used port.

[0042] According to one embodiment, in a partitioned subnet, nodes that are not members of a common partition are not allowed to communicate. In practice, this means that some of the routes assigned by the Fat Tree routing algorithm will not be used for user traffic. A problem arises if the Fat Tree routing mechanism generates LFTs for those routes in the same way as other functional routes. This behavior can degrade balancing on the links because nodes are routed in indexing order. Partition Fat-tree routed subnets generally provide poor isolation between partitions because routing is done without awareness of the

[0043] According to one embodiment, a Fat-Tree is a hierarchical network topology that can scale with available network resources. Furthermore, Fat-Tree is easily constructed using commodity switches arranged at various levels of hierarchy. Furthermore, various variants of Fat-Tree are publicly available, including k-ary-n-tree, Extended Generalized Fat-Tree (XGFT), Parallel Ports Generalized Fat-Tree (PGFT), and Real Life Fat-Tree (RLFT).

[0044] Also, a k-ary-n-tree is an n-level fat tree with k n end nodes and n·k n-1 The XGFT fat tree comprises a k-ary-n-tree and a k-ary-n-tree, each with 2k ports. Each switch has the same number of connections up and down the tree. The XGFT fat tree extends the k-ary-n-tree by allowing both a different number of up and down connections for the switches and a different number of connections at each level in the tree. The PGFT definition further extends the XGFT topology to allow multiple connections between switches. A wide variety of topologies can be defined using XGFT and PGFT. However, for practical purposes, a restricted version of PGFT, RLFT, is introduced to define fat trees commonly found in modern HPC clusters. RLFT uses the same port count switches for all levels in the fat tree.

[0045] Input / Output (I / O) Virtualization According to one embodiment, I / O Virtualization (IOV) can make I / O available by allowing virtual machines (VMs) access to the underlying physical resources. The combination of storage traffic and inter-server communication can place an unbearable strain on a single server's I / O resources, resulting in backlogs and idle processors waiting for data. As the number of I / O requests increases, IOV can provide availability and improve the performance, scalability, and elasticity of (virtualized) I / O resources to rival performance levels seen in modern CPU virtualization.

[0046] According to one embodiment, IOV is desired to enable sharing of I / O resources and to allow protected access to resources from VMs. IOV separates the logical device exposed to a VM from its physical implementation. Currently, emulation, paravirtualization, direct assignment (DA), and single-root I / O There can be various types of IOV technologies, such as virtualization (SR-IOV).

[0047] According to one embodiment, one type of IOV technology is software emulation. Software emulation can enable a separated front-end / back-end software architecture. The front-end can be a device driver located in a VM and communicate with a back-end implemented by a hypervisor to provide I / O access. Physical device sharing ratios are high, and live migration of VMs can be achieved with only milliseconds of network downtime. However, software emulation introduces additional, undesirable computational overhead.

[0048] According to one embodiment, another type of IOV technique is direct device assignment. Direct device assignment requires that an I / O device be attached to a VM, but the device Devices are not shared between VMs. Direct attachment, or device passthrough, offers near-inherent performance with minimal overhead. The physical device is attached directly to the VM, bypassing the hypervisor. However, the drawback of such direct device attachment is that there is no sharing between virtual machines, limiting scalability, such as one physical network card being attached to one VM.

[0049] According to one embodiment, Single Root IOV (SR-IOV) is Hardware virtualization can allow a physical device to appear as multiple independent, lightweight instances of the same device. These instances can be assigned to VMs as pass-through devices and accessed as Virtual Functions (VFs). The hypervisor accesses the device through a unique, fully functional Physical Function (PF) (per device). SR-I OV mitigates the scalability issues of purely direct allocation. However, a problem presented by SR-IOV is that it can impair VM migration. Among these IOV technologies, SR-IOV extends the PCI Express (PCIe) standard with a means to allow multiple VMs to directly access a single physical device while maintaining near-inherent performance. This allows SR-IOV to offer superior performance and scalability.

[0050] SR-IOV allows a PCIe device to expose multiple virtual devices that can be shared among multiple guests by assigning one virtual device to each guest. Each SR-IOV device has at least one physical function (PF) and one or more associated virtual functions (VFs). A PF is a communication function controlled by a virtual machine monitor (VMM) or hypervisor. VFs are lightweight PCIe functions, whereas VFs are regular PCIe functions. Each VF has its own base address (BAR) and is assigned a unique requestor ID. The unique requestor ID is managed by the I / O memory management unit (I / O memory management unit). The IOMMU also applies memory and interrupt translation between PFs and VFs.

[0051] Unfortunately, direct device allocation techniques present a barrier to cloud providers in situations where transparent live migration of virtual machines is desired for data center optimization. The essence of live migration is that the memory contents of a VM are copied to a remote hypervisor. Furthermore, the VM is suspended in the source hypervisor, and the VM's operation is resumed in the destination. When using software emulation methods, network interfaces are virtual so that their internal states are stored in memory and then copied. Therefore, downtime can be reduced to a few milliseconds.

[0052] However, migration becomes more difficult when direct device assignment techniques such as SR-IOV are used. In this situation, the entire internal state of the network interface cannot be copied because it is tied to the hardware. Instead, the SR-IOV VF assigned to the VM is detached and a live migration is performed, and a new VF is attached at the destination. For InfiniBand and SR-IOV, this process can result in downtime on the order of several seconds. Furthermore, in the SR-IOV shared port model, the VM's address changes after migration, which adds overhead to the SM and negatively impacts the performance of the underlying network fabric. This will be the case.

[0053] InfiniBand SR-IOV Architecture - Shared Port There can be various types of SR-IOV models (eg, a shared port model, a virtual switch model, and a virtual port model).

[0054] 4 illustrates an exemplary shared port architecture according to one embodiment. As shown, a host 300 (e.g., a host channel adapter) may interact with a hypervisor 310. The hypervisor 310 may assign various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions may be handled by the hypervisor 310.

[0055] 4, a host (e.g., an HCA) appears as a single port to the network with a single shared LID and shared Queue Pair (QP) space between the physical function 320 and the virtual functions 330, 350, 350. However, each function (i.e., the physical function and the virtual function) may have its own GID.

[0056] As shown in FIG. 4, according to one embodiment, various GIDs can be assigned to virtual and physical functions, and a special queue pair, QP0 and QP1 (i.e., InfiniBand TM A dedicated queue pair used for management packets is owned by the physical function. These QPs are exposed to the VFs as well, but the VFs are not allowed to use QP0 (all SMPs coming from the VF towards QP0 are discarded), and QP1 can act as a proxy for the actual QP1 owned by the PF.

[0057] According to one embodiment, the shared port architecture may enable highly scalable data centers that are not limited by the number of VMs (attached to the network by being assigned to virtual functions) because LID space is only consumed by the physical machines and switches in the network.

[0058] However, a drawback of the shared port architecture is that it cannot provide transparent live migration, thereby hindering the potential for flexible VM placement. Because each LID is associated with a specific hypervisor and shared among all VMs residing on that hypervisor, a migrating VM (i.e., a virtual machine migrating to a destination hypervisor) must change its LID to the LID of the destination hypervisor. Furthermore, as a result of the restricted QP0 access, a subnet manager cannot be run inside a VM.

[0059] InfiniBand SR-IOV Architecture Model - Virtual Switch (vSwitch) 5 illustrates an exemplary vSwitch architecture according to one embodiment. As shown, a host 400 (e.g., a host channel adapter) can interact with a hypervisor 410, which can assign various virtual functions 430, 440, and 450 to several virtual machines. Similarly, physical functions can be handled by hypervisor 410. A virtual switch 415 can also be handled by hypervisor 401.

[0060] According to one embodiment, in the vSwitch architecture, each virtual function 430, 440, 450 is a complete virtual Host Channel Adapter. er:vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM, HCA 400 appears as a switch with additional nodes connected via virtual switch 415. Hypervisor 410 can use PF 420, and VMs (attached to virtual functions) use VFs.

[0061] According to one embodiment, the vSwitch architecture provides transparent virtualization. However, because each virtual function is assigned a unique LID, the available number of LIDs is quickly consumed. Similarly, if many LID addresses are used (i.e., one for each physical function and each virtual function), more communication paths must be computed by the SM and more subnet management packets (SMPs) must be sent to the switch to update their LFTs. For example, computing communication paths can take several minutes in a large network. Because the LID space is limited to 49,151 unicast LIDs and each VM (through a VF) occupies one LID per physical node and switch, the number of active VMs is limited by the number of physical nodes and switches in the network, and vice versa.

[0062] InfiniBand SR-IOV Architecture Model - Virtual Port (vPort) 6 illustrates an exemplary vPort concept according to one embodiment. As shown, a host 300 (e.g., a host channel adapter) can interact with a hypervisor 410 that can allocate various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions can be handled by the hypervisor 310.

[0063] According to one embodiment, the vPort concept is loosely defined to allow vendors implementation freedom (e.g., the definition does not stipulate that implementations should be SRIOV-only), and the purpose of vPort is to standardize how VMs are handled in a subnet. The vPort concept allows for the definition of both an SR-IOV shared port-like architecture and a vSwitch-like architecture, or a combination of these architectures, which may be more scalable in both the spatial and performance domains. Also, vPorts support optional LIDs, and unlike shared ports, the SM is aware of all vPorts available in a subnet, even if the vPorts do not use dedicated LIDs.

[0064] InfiniBand SR-IOV Architecture Model - LID Pre-Populated vSwitch According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs.

[0065] 7 illustrates an exemplary vSwitch architecture with pre-populated LIDs, according to one embodiment. As shown, several switches 501-504 are configured to support InfiniBand TM Communication can be established between members of a fabric, such as a fabric. The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. In addition, host channel adapters 510, 520, and 530 can interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, further interacts with and configures several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. , can be assigned to several virtual machines. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515 and virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0066] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 600.

[0067] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0068] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs. Referring to FIG. 7, LIDs are pre-populated for various physical functions 513, 523, and 533, as well as for virtual functions 514-516, 524-526, and 534-536 (even virtual functions not currently associated with active virtual machines). For example, physical function 513 is pre-populated with LID 1, and virtual function 1 534 is pre-populated with LID 10. When a network is booted, LIDs are pre-populated in the SR-IOV vSwitch-enabled subnet. Populated VFs are assigned LIDs as shown in FIG. 7, even if not all of the VFs are occupied by VMs in the network.

[0069] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0070] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, each hypervisor consumes one LID for itself through the PF and can consume one or more LIDs for each additional VF. The sum of all VFs available across all hypervisors in an IB subnet yields the maximum amount of VMs that can run in the subnet. For example, in an IB subnet with 16 virtual functions per hypervisor in the subnet, each hypervisor consumes 17 LIDs in the subnet (one LID for each of the 16 virtual functions and one LID for the physical function). In such an IB subnet, the theoretical hypervisor limit for a single subnet is defined by the number of available unicast LIDs: 2891 (49151 available LIDs divided by 17 LIDs per hypervisor), and the total number of VMs (i.e., limit) is 46256 (2891 hypervisors multiplied by 16 VFs per hypervisor) (effectively, each switch, router, or (In practice, these numbers will be smaller, as other dedicated SM nodes consume LIDs as well.) Note that the vSwitch does not need to occupy additional LIDs, as it can share them with the PF.

[0071] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, once the network is booted, communication paths are calculated for all LIDs. If a new VM needs to be started, the system does not need to add a new LID in the subnet. Otherwise, operations that may completely reconfigure the network, including recalculating paths, are the most time-consuming part. Instead, available ports for VMs are located in one of the hypervisors (i.e., available virtual functions), and virtual machines are assigned to available virtual functions.

[0072] According to one embodiment, the LID pre-populated vSwitch architecture also enables the ability to compute and use different routes to reach different VMs hosted by the same hypervisor. Essentially, this allows such subnets and networks to use LID-Mask-Control-like (LMC-like) features to provide alternative routes towards one physical machine without being bound by the LMC constraint that requires LIDs to be contiguous. The freedom to use non-contiguous LIDs is particularly useful when a VM needs to migrate and its associated LID needs to be delivered to the destination.

[0073] In accordance with one embodiment, several considerations can be taken into account along with the above-described advantages of a LID pre-populated vSwitch architecture. For example, because LIDs are pre-populated in an SR-IOV vSwitch-enabled subnet when the network is booted, the initial route computation (e.g., at startup) may take longer than if the LIDs were not pre-populated.

[0074] InfiniBand SR-IOV Architecture Model - vSwitch with Dynamic LID Allocation According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID allocation.

[0075] 8 illustrates an exemplary vSwitch architecture with dynamic LID assignment, according to one embodiment. As shown, several switches 501-504 are configured to support InfiniBand TMCommunications can be established between members of a fabric, such as a fabric. The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. Host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515 and virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapters through fully functional physical functions 513, 523 and 533 on each of the host channel adapters.

[0076] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 700.

[0077] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0078] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID assignment. Referring to FIG. 8 , various physical functions 513, 523, and 533 are dynamically assigned LIDs, with physical function 513 receiving LID 1, physical function 523 receiving LID 2, and physical function 533 receiving LID 3. Those virtual functions associated with active virtual machines may also receive dynamically assigned LIDs. For example, because virtual machine 1 550 is active and associated with virtual function 1 514, virtual function 514 may be assigned LID 5. Similarly, virtual function 2 515, virtual function 3 516, and virtual function 1 534 are each associated with an active virtual function. As such, these virtual functions are assigned LIDs: LID 7 is assigned to virtual function 2 515, LID 11 is assigned to virtual function 3 516, and LID 9 is assigned to virtual function 1 534. Unlike a vSwitch, which has pre-populated LIDs, virtual functions that are not currently associated with an active virtual machine do not receive an LID assignment.

[0079] According to one embodiment, dynamic LID allocation can substantially reduce initial path computation: When a network is booting for the first time and no VMs are present, a relatively small number of LIDs can be used for initial path computation and LFT distribution.

[0080] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0081] According to one embodiment, when a new VM is created in a system utilizing a vSwitch with dynamic LID allocation, a free VM slot is discovered and a unique, unused unicast LID is discovered as well to determine on which hypervisor the newly added VM should boot. However, there is no known route in the switch's LFT and network to handle the newly added LID. Computing a new set of routes to handle the newly added VM is undesirable in a dynamic environment where several VMs may be booted every minute. In a large IB subnet, computing a new set of routes could take several minutes, and this procedure would have to be repeated each time a new VM is booted.

[0082] Advantageously, according to one embodiment, since all VFs in the hypervisor share the same uplink with the PF, there is no need to compute a new set of routes: iterate through the LFTs of all physical switches in the network and forward from the LID entries belonging to the PF of the hypervisor (where the VM is created) to the newly added LID. Only a single SMP needs to be sent to copy the ports and update the corresponding LFT block of a particular switch, which eliminates the need for the system and method to compute a new set of routes.

[0083] According to one embodiment, the assigned LIDs in a vSwitch with a dynamic LID allocation architecture do not need to be contiguous. Comparing the assigned LIDs on VMs on each hypervisor between a vSwitch with pre-populated LIDs and a vSwitch with dynamic LID allocation, it can be seen that the assigned LIDs in the dynamic LID allocation architecture are discontinuous, whereas the pre-populated LIDs are essentially contiguous. Furthermore, in a vSwitch dynamic LID allocation architecture, when a new VM is created, the next available LID is used for the lifetime of the VM. Conversely, in a vSwitch with pre-populated LIDs, each VM inherits the LID already assigned to its corresponding VF, and in a network without live migration, VMs assigned consecutively to a given VF get the same LID.

[0084] According to one embodiment, a vSwitch with a dynamic LID allocation architecture can address the shortcomings of a vSwitch with a pre-populated LID architecture model at the expense of some additional network and runtime SM overhead. Each time a VM is created, the LFT of the physical switch in the subnet is updated with the newly added LID associated with the created VM. This operation requires one subnet management packet (SMP) to be sent per switch. Because each VM uses the same route as its host hypervisor, features such as LMC are also unavailable. However, there is no limit on the total number of VFs present on all hypervisors, and the number of VFs may exceed the unicast LID limit. In such a case, of course, not all VFs can be simultaneously granted on active VMs. Having more spare hypervisors and VFs adds flexibility for recovering from and optimizing fragmented network failures when operating near the unicast LID limit.

[0085] InfiniBand SR-IOV Architecture Model - Dynamic LID Allocation and Pre-Populated LID vSwitch 9 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs for the vSwitch, according to one embodiment. As shown, several switches 501-504 are configured to support InfiniBand TM Communication can be established between members of a fabric, such as a fabric. The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. The host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515. Hypervisor 521 can assign virtual machine 3 552 to virtual function 3 526. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 2 535. The hypervisor can access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0086] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 800.

[0087] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0088] According to one embodiment, the present disclosure provides a system and method for providing a hybrid vSwitch architecture with dynamic LID assignment and pre-populated LIDs. Referring to FIG. 9 , hypervisor 511 may be deployed with a vSwitch with a pre-populated LID architecture, and hypervisor 521 may be deployed with a vSwitch with pre-populated LIDs and dynamic LID assignment. Hypervisor 531 may be deployed with a vSwitch with dynamic LID assignment. Thus, physical function 513 and virtual functions 514-516 have their LIDs pre-populated (i.e., even virtual functions that are not assigned to active virtual machines are assigned LIDs). Physical function 523 and virtual function 1 524 may have their LIDs pre-populated, while virtual function 2 525 and virtual function 3 526 have their LIDs dynamically assigned (i.e., virtual function 2 525 is available for dynamic LID assignment, and virtual function 3 526 has been dynamically assigned an LID of 11 because it is attached to virtual machine 3 552). Finally, the functions (physical and virtual functions) associated with hypervisor 3 531 can have their LIDs dynamically assigned. This results in virtual function 1 534 and virtual function 3 536 being available for dynamic LID assignment, while virtual function 2 535 has been dynamically assigned an LID of 9 because it is attached to virtual machine 4 553.

[0089] 9, in which both LID pre-populated vSwitches and dynamic LID allocation vSwitches are utilized (independently or combined within any given hypervisor), the number of pre-populated LIDs per host channel adapter can be defined by a fabric administrator and can be in the range 0<=pre-populated VFs<=total VFs (per host channel adapter). The VFs available for dynamic LID allocation can be found by subtracting the number of pre-populated VFs from the total number of VFs (per host channel adapter).

[0090] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0091] InfiniBand - Inter-subnet communication (Fabric Manager) In accordance with one embodiment, in addition to providing an InfiniBand fabric within one subnet, embodiments of the present disclosure may also provide an InfiniBand fabric that spans two or more subnets. Fabric can also be provided.

[0092] FIG. 10 illustrates an exemplary multi-subnet InfiniBand fabric according to one embodiment. As shown in this figure, multiple switches 1001-1004 within subnet A 1000 can provide communication between members of a fabric, such as an InfiniBand fabric, within subnet A 1000 (e.g., an IB subnet). The fabric can include multiple hardware devices, such as channel adapters 1010. Host channel adapters 1010 can interact with hypervisor 1011. The hypervisor can set up multiple virtual functions 1014 with the host channel adapters it interacts with. In addition, the hypervisor can assign virtual machines to each virtual function. For example, virtual machine 1 1015 is assigned to virtual function 1 1014. The hypervisor can access its associated host channel adapter through a fully functional physical function, such as physical function 1013, on each host channel adapter. A number of switches 1021-1024 can provide communication between members of a fabric, such as an InfiniBand fabric, within subnet B 1040 (e.g., an IB subnet). This fabric can include a number of hardware devices, such as host channel adapter 1030. Host channel adapter 1030 can interact with hypervisor 1031. The hypervisor can set up a number of virtual functions 1034 with the host channel adapters it interacts with. In addition, the hypervisor can assign a virtual machine to each virtual function. For example, virtual machine 2 1035 is assigned to virtual function 2 1034. The hypervisor can access its associated host channel adapter through a fully functional physical function, such as physical function 1033, on each host channel adapter. Note that although only one host channel adapter is shown within each subnet (i.e., subnet A and subnet B), it should be understood that each subnet may include multiple host channel adapters and their corresponding components.

[0093] According to one embodiment, each host channel adapter may further be associated with a virtual switch, such as virtual switch 1012 and virtual switch 1032, and as noted above, each HCA may be set up with a different architectural model. Although both subnets in Figure 10 are shown as using the architectural model of a vSwitch with pre-populated LIDs, this is not intended to suggest that all such subnet configurations may follow a similar architectural model.

[0094] According to one embodiment, at least one switch in each subnet may be associated with a router. For example, switch 1002 in subnet A 1000 is associated with router 1005, and switch 1021 in subnet B 1040 is associated with router 1006.

[0095] According to one embodiment, at least one device (e.g., a switch, a node, etc.) can be associated with a fabric manager (not shown). The fabric manager can be used, for example, to discover inter-subnet fabric topologies, create fabric profiles (e.g., virtual machine fabric profiles), and build virtual machine-related database objects that form the basis for building the virtual machine fabric profiles. Additionally, the fabric manager can define legal inter-subnet connectivity regarding which subnets are allowed to communicate over which router ports and using which partition numbers.

[0096] According to one embodiment, a traffic signal at an originating source, such as Virtual Machine 1 in Subnet A, If traffic is destined for a different subnet, such as virtual machine 2 in subnet B, the traffic can be directed to a router in subnet A, i.e., router 1005, which can then send the traffic to subnet B over a link with router 1006.

[0097] Virtual Dual Port Router According to one embodiment, a dual port router abstraction is implemented to route from a global route header (GRH) to an LR. It is possible to provide a simple method that allows defining an inter-subnet router function based on a switch hardware implementation that has the capability to convert from LRH to H (local route header) in addition to performing normal LRH-based switching.

[0098] According to one embodiment, a Virtual Dual Port Router can be logically connected outside of a corresponding switch port and can present an InfiniBand compliant view to a standard management entity, such as a subnet manager.

[0099] According to one embodiment, the dual port router model connects different subnets, with each subnet handling packet forwarding and address mapping on the ingress path to the subnet. This shows that it is possible to connect subnets in a way that gives complete control over the subnetworks and does not affect the routing and logical connectivity within any of the misconnected subnets.

[0100] According to one embodiment, in situations involving misconnected fabrics, the virtual dual port router abstraction can also be used to allow management entities such as subnet managers and IB diagnostic software to function correctly in the presence of unintended physical connections to remote subnets.

[0101] 11 illustrates an interconnection between two subnets in a high performance computing environment, according to one embodiment. Before being configured with a virtual Dual Port Router, Subnet A Switch 1120 in subnet B 1101 may be connected through switch port 1121 of switch 1120 via physical connection 1110 to switch 1130 in subnet B 1102 through switch port 1131 of switch 1130. In such an embodiment, each of switch ports 1121 and 1131 may function as both a switch port and a router port.

[0102] According to one embodiment, the problem with this configuration is that a management entity, such as a subnet manager, in an InfiniBand subnet cannot distinguish between physical ports that are both switch ports and router ports. In this situation, the SM can treat a switch port as having a router port connected to it. However, if the switch port is connected to another subnet with a different subnet manager, for example, via a physical link, the subnet manager can send discovery messages to the physical link. However, such discovery messages are not allowed in the other subnet.

[0103] FIG. 12 illustrates an interconnection between two subnets via a dual-port virtual router configuration in a high-performance computing environment, according to one embodiment.

[0104] According to one embodiment, after configuration, the dual port virtual router configuration is registered by the subnet manager at the appropriate end node that indicates the edge of the subnet for which the subnet manager is responsible. As can be seen, it can be provided.

[0105] According to one embodiment, a switch port in switch 1220 in subnet A 1201 can be connected (i.e., logically connected) to router port 1211 in virtual router 1210 via virtual link 1223. Virtual router 1210 (e.g., a dual-port virtual router) is shown in the embodiment as being external to switch 1220, but can be logically contained within switch 1220 and can also include a second router port, router port II 1212. According to one embodiment, physical link 1203, which can have two ends, connects subnet A 1201 to subnet B 1202 via a first end of the physical link, via a second end of the physical link, via router port II 1212, and to router port II, which is included in virtual router 1230 in subnet B 1202. 1232. Virtual router 1230 may further include a router port 1231 that may be connected (i.e., logically connected) to a switch port 1241 on switch 1240 via a virtual link 1233.

[0106] According to one embodiment, a subnet manager (not shown) on subnet A can discover router port 1211 on virtual router 1210 as the endpoint of the subnet that the subnet manager controls. The dual-port virtual router abstraction allows the subnet manager on subnet A to treat subnet A in the usual way (e.g., as specified in the InfiniBand standard). At the subnet management agent level, the dual-port virtual router abstraction A fabric abstraction can be provided so that a normal switch port appears to the SM, and then at the SMA level, the abstraction can be provided so that there is another port connected to this switch port, and this port becomes a router port on a dual-port virtual router. The local SM can continue to use the traditional fabric topology (where the SM sees the ports as standard switch ports), and therefore the SM sees the router ports as end ports. A physical connection can be made between two switch ports that are also configured as router ports in two different subnets.

[0107] According to one embodiment, a dual port virtual router can also solve the problem that a physical link may be mistakenly connected to any other switch port in the same subnet, or to a switch port that is not intended to provide connectivity to another subnet. Thus, the methods and systems described herein also represent what is outside of a subnet.

[0108] According to one embodiment, a local SM in a subnet, such as subnet A, determines a switch port and then determines the router port connected to this switch port (e.g., router port 1211 connected to switch port 1221 via virtual link 1223). The SM considers router port 1211 to be the edge of the subnet that it manages, so the SM cannot send discovery and / or management messages further than this point (e.g., to router port II 1212).

[0109] According to one embodiment, the dual port virtual router provides the advantage that the dual port virtual router abstraction is managed entirely by a management entity (e.g., SM or SMA) within the subnet to which the dual port virtual router belongs. By keeping management local, the system does not need to provide an external, independent management entity; that is, each side of the inter-subnet connection is responsible for configuring its own dual port virtual router.

[0110] According to one embodiment, when a packet such as an SMP destined for a remote destination (i.e., outside the local subnet) arrives at a local target port that is not configured through the dual-port virtual router, the local port can return a message indicating that it is not a router port.

[0111] Many features of the present disclosure can be implemented in, using, or with the aid of hardware, software, firmware, or a combination thereof. Thus, features of the present disclosure may be implemented using a processing system (e.g., including one or more processors).

[0112] 13 illustrates a method for supporting a dual-port virtual router in a high-performance computing environment according to one embodiment. In step 1310, the method may provide a first subnet in one or more computers including one or more microprocessors. The first subnet includes a plurality of switches, the plurality of switches including at least leaf switches, each of the plurality of switches including a plurality of switch ports. The first subnet further includes a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, a plurality of end nodes, each of the end nodes associated with at least one host channel adapter of the plurality of host channel adapters, and a subnet manager, the subnet manager executing in one of the plurality of switches and the plurality of host channel adapters.

[0113] In step 1320, the method may configure a switch port of the plurality of switch ports on a switch of the plurality of switches as a router port.

[0114] In the switch 1330, the method can logically connect switch ports configured as router ports to a virtual router, the virtual router including at least two virtual router ports.

[0115] Quality of Service and Service Level Agreements in Private Fabrics High-performance computing environments within the cloud and larger clouds at customer and premises installations, such as switched networks running over InfiniBand or RoCE, according to one embodiment, are capable of deploying virtual machine (VM)-based workloads, and one inherent requirement is the ability to define and control quality of service (QOS) for different types of communication flows. Additionally, workloads belonging to different tenants must run within the bounds of associated service level agreements (SLAs) while minimizing interference between such workloads and maintaining QOS assumptions for different communication types.

[0116] RDMA Read as a Limited Feature (ORA20Q246-US-NP-1) According to one embodiment, when defining bandwidth limits in a system that uses traditional network interfaces (NICs), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.

[0117] However, according to an embodiment, in RDMA-based networking where different nodes can generate RDMA read requests (i.e., egress bandwidth), this may represent a small amount of egress bandwidth. However, such RDMA read requests may potentially represent a very large amount of ingress RDMA traffic in response to such RDMA read requests. In such a situation, limiting the egress bandwidth of every node / VM to control overall traffic generation in the system is no longer sufficient. isn't it.

[0118] According to one embodiment, by making RDMA read operations a restricted feature and only allowing such read requests for nodes / VMs that can be trusted not to generate excessive RDMA read-based ingress bandwidth, it is possible to limit total bandwidth usage while only limiting the outgoing (egress) bandwidth for untrusted nodes / VMs.

[0119] FIG. 14 illustrates a system for servicing RDMA read requests as a limited feature in a high performance computing environment, according to one embodiment.

[0120] More specifically, according to one embodiment, Figure 14 illustrates a host channel adapter 1401 that includes a hypervisor 1411. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 1414-1416, and a physical function (PF) 1413. The host channel adapter can further support or include several ports, such as ports 1402 and 1403, that are used to connect the host channel adapter to a network, such as network 1400. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 1401 to several other nodes, such as a switch, additional separate HCAs, etc.

[0121] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 1450, VM2 1451, and VM3 1452.

[0122] According to one embodiment, host channel adapter 1401 can further support a virtual switch 1412 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0123] According to one embodiment, the host channel adapter may implement a trusted RDMA read limit 1460, whereby the read limit 1460 may be configured to block any of the virtual machines (e.g., VM1, VM2, and / or VM3) from sending any RDMA read requests out onto the network (e.g., via ports 1402 or 1403).

[0124] According to one embodiment, trusted RDMA read restriction 1460 can implement host channel adapter-level blocking of certain types of packets from certain endpoints, such as virtual machines, or other physical nodes utilizing HCA 1401 to connect to the network, from generating (i.e., egressing) RDMA read request packets. This configurable restriction component 1460 can, for example, allow only trusted nodes (e.g., VMs or physical end nodes) to generate such types of packets.

[0125] According to an embodiment, the trusted RDMA read restriction component may be configured based on instructions received, for example, by a host channel adapter, or may be configured directly, for example, by a subnet manager (not shown).

[0126] FIG. 15 illustrates a system for servicing RDMA read requests as a limited feature in a high performance computing environment, according to one embodiment.

[0127] More specifically, according to one embodiment, Figure 15 illustrates a host channel adapter 1501 that includes a hypervisor 1511. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 1514-1516, and a physical function (PF) 1513. The host channel adapter can additionally support or include several ports, such as ports 1502 and 1503, that are used to connect the host channel adapter to a network, such as network 1500. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 1501 to several other nodes, such as a switch, additional separate HCAs, etc.

[0128] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 1550, VM2 1551, and VM3 1552.

[0129] According to one embodiment, host channel adapter 1501 can further support a virtual switch 1512 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0130] According to one embodiment, the host channel adapter may implement a trusted RDMA read limit 1560, whereby the read limit 1560 may be configured to block any of the virtual machines (e.g., VM1, VM2, and / or VM3) from sending any RDMA read requests out onto the network (e.g., via ports 1502 or 1503).

[0131] According to one embodiment, trusted RDMA read restriction 1560 can provide host channel adapter-level blocking of certain types of packets from certain endpoints, such as virtual machines, or other physical nodes utilizing HCA 1501 to connect to the network, from generating (i.e., egressing) RDMA read request packets. This configurable restriction component 1560 can, for example, allow only trusted nodes (e.g., VMs or physical end nodes) to generate such types of packets.

[0132] According to an embodiment, the trusted RDMA read restriction component may be configured based on instructions received, for example, by a host channel adapter, or may be configured directly, for example, by a subnet manager (not shown).

[0133] According to an embodiment, as an example, trusted RDMA read restriction 1560 may be configured to trust VM1 1550 and not trust VM2 1551. Thus, RDMA read request 1554 originating from VM1 may be allowed, while RDMA read request 1555 originating from VM2 may be blocked before it leaves and reaches host channel adapter 1501 (although shown outside the HCA in FIG. 15 for convenience of illustration only).

[0134] FIG. 16 illustrates a system for servicing RDMA read requests as a limited feature in a high performance computing environment, according to one embodiment.

[0135] According to one embodiment, a high-speed network such as a switched network or subnetwork 1600 is Within the performance computing environment, several end nodes 1601 and 1602 can support several virtual machines VM1-VM4 1650-1653 interconnected via several switches, such as leaf switches 1611 and 1612, switches 1621 and 1622, and root switches 1631 and 1632.

[0136] Not shown are the various host channel adapters that provide functionality for the connection of nodes 1601 and 1602, as well as the virtual machines to be connected to the subnetwork, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0137] According to one embodiment, in a typical system, RDMA egress bandwidth is limited from any one virtual machine from an end node to prevent any one virtual machine from monopolizing the bandwidth of any link connecting the end node to a subnet. However, such egress bandwidth limiting, while effective in the general case, does not prevent virtual machines from issuing RDMA read requests, such as RDMA read requests 1654 and 1655, because such RDMA read requests are typically small packets that utilize little egress bandwidth.

[0138] However, according to an embodiment, such an RDMA read request may result in the generation of a large amount of return traffic to the issuing entities, such as VM1 and VM3. In such a situation, the RDMA read request may then lead to link congestion and degraded network performance, for example, when read request 1654 results in a large amount of data traffic flowing back to VM1 as a result of execution of the read request at the destination.

[0139] According to one embodiment, this may lead to a loss of performance for the subnet, especially in situations where multiple tenants share the subnet 1600.

[0140] According to one embodiment, each node (or host channel adapter) can be configured with RDMA read limits 1660 and 1661 that place a block on any VM from issuing RDMA read requests if that VM is not trusted. Such RDMA read limits can vary from a constant block on issuing RDMA read requests to a limit that places a time frame on when a virtual machine configured with an RDMA read request limit can issue an RDMA read request (e.g., during periods of slow network traffic). In addition, RDMA read limits 1660 and 1661 can further enable trusted VMs to issue RDMA read requests.

[0141] According to one embodiment, since it is conceivable to have a scenario in which multiple VMs / tenants share a "modern" HCA, i.e., an HCA that has support for the relevant new feature, but are executing RDMA requests to a remote "legacy" HCA that does not have such support, it would make sense to have a way to limit the ingress bandwidth that such VMs can generate for RDMA read responses without resorting to static rate configuration on the "legacy" RDMA read responding HCA. There is no simple way to do this, as VMs are allowed to generate "arbitrary" RDMA read sizes. Also, since multiple RDMA read requests generated over a period of time may in principle all receive response data simultaneously, it is not possible to guarantee that the ingress bandwidth cannot exceed the maximum bandwidth for more than a very limited period of time, unless there is both a limit on the RDMA read size that can be generated in a single request and a limit on the total number of outstanding RDMA read requests from the same vHCA port.

[0142] Thus, according to one embodiment, if a maximum read size is defined for the vHCA, bandwidth control may be based on an allocation to the sum of all outstanding read sizes, or a simpler scheme may be to simply limit the maximum number of outstanding RDMA reads based on a "worst-case" read size. Thus, in either case, there is no limit on peak bandwidth within a short interval (except for the HCA port maximum link bandwidth), but the duration of such a peak bandwidth "window" will be limited. Additionally, however, the transmission rate of RDMA read requests must also be throttled so that the transmission rate of requests does not exceed the maximum allowed ingress rate, assuming that responses with data are received at the same rate. In other words, the maximum outstanding request limit defines the worst-case short interval bandwidth, and the request transmission rate limit will ensure that new requests cannot be generated immediately once a response is received, but only after an associated delay that represents the acceptable average ingress bandwidth for RDMA read responses. Thus, in the worst case, the allowed number of requests are sent without any responses, and then these responses are all received "at the same time." At this point, the next request can be sent immediately when the first response arrives, but the next request will have to be delayed for the specified delay period. Thus, over time, the average ingress bandwidth cannot exceed what the request rate defines. However, a smaller maximum number of outstanding requests reduces possible "variability."

[0143] Using Explicit RDMA Read Bandwidth Limiting (ORA200246-US-NP-1) According to one embodiment, when defining bandwidth limits in a system that uses traditional network interfaces (NICs), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.

[0144] However, in RDMA-based networking, where different nodes can generate RDMA read requests that represent small request messages but potentially very large response messages, according to certain embodiments, it is no longer sufficient to limit the egress bandwidth of every node / VM to control the total traffic generation in the system.

[0145] According to one embodiment, by defining explicit quotas on how much RDMA read ingress bandwidth a node / VM is allowed to generate, independent of any transmit / egress bandwidth limitations, it is possible to control total traffic generation in the system without resorting to restricting RDMA read usage for untrusted nodes / VMs.

[0146] According to one embodiment, the system and method can support the worst-case maximum link bandwidth burst duration / length (i.e., as a result of RDMA read responses "piling up") in addition to supporting the average ingress bandwidth utilization resulting from locally generated RDMA read requests.

[0147] FIG. 17 illustrates a system for providing explicit RDMA read bandwidth limits in a high performance computing environment, according to an embodiment.

[0148] More specifically, according to one embodiment, Figure 17 illustrates a host channel adapter 1701 that includes a hypervisor 1711. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 1714-1716, and physical functions (PFs) 1713. The host channel adapter includes ports 1714-1716 that are used to connect the host channel adapter to a network, such as network 1700. 1701 may further support or comprise several ports, such as 1702 and 1703. The network may comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect HCA 1701 to several other nodes, such as a switch, additional separate HCAs, etc.

[0149] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 1750, VM2 1751, and VM3 1752.

[0150] According to one embodiment, host channel adapter 1701 can further support a virtual switch 1712 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0151] According to one embodiment, the host channel adapter can implement an RDMA read limit 1760, which can be configured to impose a quota on the amount of ingress bandwidth that any VM (of the HCA 1701) can generate in response to an RDMA read request issued by the particular VM. Such ingress bandwidth limiting is performed locally at the host channel adapter.

[0152] According to an embodiment, the RDMA read limiting component may be configured based on instructions received, for example, by a host channel adapter, or may be configured directly, for example, by a subnet manager (not shown).

[0153] FIG. 18 illustrates a system for providing explicit RDMA read bandwidth limits in a high performance computing environment, according to an embodiment.

[0154] More specifically, according to one embodiment, Figure 18 illustrates a host channel adapter 1801 that includes a hypervisor 1811. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 1814-1816, and a physical function (PF) 1813. The host channel adapter can further support or include several ports, such as ports 1802 and 1803, that are used to connect the host channel adapter to a network, such as network 1800. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 1801 to several other nodes, such as a switch, additional separate HCAs, etc.

[0155] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 1850, VM2 1851, and VM3 1852.

[0156] According to one embodiment, host channel adapter 1801 can further support a virtual switch 1812 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0157] According to one embodiment, the host channel adapter implements an RDMA read limit 1860. The read limit 1860 can be configured to impose a quota on the amount of ingress bandwidth that any VM (of the HCA 1701) can generate in response to an RDMA read request issued by the particular VM. Such ingress bandwidth limiting is performed locally at the host channel adapter.

[0158] According to an embodiment, the RDMA read limiting component may be configured based on instructions received, for example, by a host channel adapter, or may be configured directly, for example, by a subnet manager (not shown).

[0159] According to an embodiment, for example, VM1 may have previously sent at least two RDMA read requests requesting that the read operations be performed on connected nodes. In response, VM1 may be in the process of receiving multiple responses to the RDMA read requests, shown in the figure as RDMA read responses 1855 and 1854. Because these RDMA read responses may be quite large, especially when compared to the RDMA read request originally sent by VM1, these read responses 1854 and 1855 may be subject to RDMA read limit 1860, and the ingress bandwidth may be limited or throttled. This throttling may be based on an explicit ingress bandwidth limit or may be based on VM1's QoS and / or SLA set in RDMA limit 1860.

[0160] FIG. 19 illustrates a system for providing explicit RDMA read bandwidth limiting in a high performance computing environment, according to an embodiment.

[0161] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 1900, several end nodes 1901 and 1902 may support several virtual machines VM1-VM4 1950-1953 that are interconnected via several switches, such as leaf switches 1911 and 1912, switches 1921 and 1922, and root switches 1931 and 1932.

[0162] Not shown are the various host channel adapters that provide functionality for the connection of nodes 1901 and 1902, as well as the virtual machines to be connected to the subnetwork, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0163] According to one embodiment, in a typical system, RDMA egress bandwidth is limited from any one virtual machine to an end node to prevent any one virtual machine from monopolizing the bandwidth of any link connecting the end node to a subnet. However, while such egress bandwidth limiting is effective in the general case, it cannot prevent an influx of RDMA read responses from monopolizing the link between the requesting VM and the network.

[0164] In other words, according to one embodiment, if VM1 sends out several RDMA read requests, VM1 cannot control when responses to such read requests are returned to VM1. This can result in a backup / piling up of responses to the RDMA read requests, each attempting to use the same link to return the requested information to VM1 (via RDMA read response 1954). This results in traffic congestion and backlogs in the network.

[0165] According to one embodiment, RDMA limits 1960 and 1961 can impose a quota on the amount of ingress bandwidth a VM can generate in response to RDMA read requests issued by the particular VM. Such ingress bandwidth limiting is performed locally.

[0166] According to one embodiment, if a maximum read size is defined for the vHCA, bandwidth control may be based on a quota on the sum of all outstanding read sizes, or a simpler scheme may be to simply limit the maximum number of outstanding RDMA reads based on a "worst-case" read size. Thus, in either case, there is no limit on peak bandwidth within a short interval (except for the HCA port maximum link bandwidth), but the duration of such a peak bandwidth "window" would be limited. Additionally, however, the transmission rate of RDMA read requests must also be throttled so that the transmission rate of requests does not exceed the maximum allowed ingress rate, assuming that responses with data are received at the same rate. In other words, the maximum outstanding request limit defines the worst-case short interval bandwidth, and the request transmission rate limit would ensure that new requests cannot be generated immediately once a response is received, but only after an associated delay representing the acceptable average ingress bandwidth for RDMA read responses. Thus, in the worst case, the allowed number of requests are sent without any responses, and then these responses are all received "at the same time." At this point, the next request can be sent immediately when the first response arrives, but the next request will have to be delayed for the specified delay period. Thus, over time, the average ingress bandwidth cannot exceed what the request rate defines. However, a smaller maximum number of outstanding requests reduces possible "variability."

[0167] FIG. 20 is a flowchart of a method for providing RDMA (Remote Direct Memory Access) read requests as a restricted feature in a high performance computing environment, according to one embodiment.

[0168] According to one embodiment, in step 2010, the method can provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches and a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, and the plurality of host channel adapters interconnected via a plurality of switches.

[0169] According to one embodiment, in step 2020, the method may provide a plurality of end nodes including a plurality of virtual machines.

[0170] According to one embodiment, in step 2030, the method can associate a host channel adapter with selective RDMA restriction.

[0171] According to one embodiment, in step 2040, the method can host a virtual machine of a plurality of virtual machines on a host channel adapter that includes selective RDMA restriction.

[0172] Combining Multiple Shared Bandwidth Segments (ORA20Q246-US-NP-3) According to one embodiment, traditional bandwidth / rate limiting schemes for network interfaces typically limit the overall aggregated sending rate and possibly a maximum rate for each individual destination. However, in many cases, there is a shared bottleneck in the intermediate network / fabric topology, limiting the target This means that the total bandwidth available to the set is limited by this shared bottleneck, and therefore, if such a shared bottleneck is not taken into account when determining at what rates the various data flows can be sent, the shared bottleneck is likely to become overloaded, even though the rate limits for each target are respected.

[0173] According to certain embodiments, the systems and methods herein may introduce a "target group" object to which multiple individual flows can be associated, which may represent the rate limits of individual (potentially shared) links or other bottlenecks in the network / fabric path the flows are using. Furthermore, the systems and methods may allow each flow to be associated with a hierarchy of such target groups, representing all link segments and any other (shared) bottlenecks in the path between the source and target for an individual flow.

[0174] According to one embodiment, to limit egress bandwidth, the present system and method can establish groups of destinations that share bandwidth allocations to reduce the likelihood of congestion on a shared ISL (Inter-Switch Link). This requires a destination / route association lookup mechanism that can manage which destinations / routes map to which groups at a logical level. This means that the hyper-privileged communications infrastructure must be aware of the actual location of peer nodes in the fabric topology, as well as associated routing and capacity information that can be mapped to "target groups" (i.e., HCA-level object types) within the local HCA with associated bandwidth allocations. However, it is impractical for the HW to perform a direct lookup of WQE (Work Queue Entry) / packet address information to map to the associated target group. Instead, an HCA implementation can provide an association between an RC (Reliable Connected) QP (Queue Pair) and an address handle, which represents the transmission context for outgoing traffic and the associated target group. In this way, this association is transparent at the verbs level. Alternatively, it can be set up by a hyper-privileged software level and then implemented at the HCA HW (and firmware) level. A significant additional complication associated with this scheme is that live VM migration, in which associated VM or vHCA port address information is maintained across migrations, may still mean that there is a change in target group for different communicating peers. However, target group associations do not need to be synchronously updated, as long as the present system and method tolerates some transient period during which the associated bandwidth allocations are not 100% accurate. Thus, while logical connectivity and the ability to communicate may not change with VM migration, the target groups associated with RC connections and address handles in both the migrated VM and its communicating peer VMs may be “completely wrong” after migration. This can mean both that less bandwidth than is available is utilized (e.g., when a VM is moved from a remote location to the same “leaf group” as its peer) and that excess bandwidth is generated (e.g., when a VM is moved from the same “leaf group” as its peer to a remote location, implying a shared ISL with limited bandwidth).

[0175] According to one embodiment, target group specific bandwidth allocations can in principle also be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target group represents.

[0176] According to one embodiment, a target group may redirect objects from a particular destination address. By decoupling from the target, the present system and method gains the ability to represent intermediate shared links or groups of links that may represent bandwidth limitations that may be more restrictive than the target limit.

[0177] According to one embodiment, the system and method can consider a hierarchy of target groups (bandwidth allocations) that reflect bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most restricted rate in the hierarchy. That is, for example, if the target limit is 30 Gb / s and the intermediate uplink limit is 50 Gb / s, the maximum rate toward the target can never exceed 30 Gb / s. On the other hand, if multiple 30 Gb / s targets share the same 50 Gb / s intermediate limit, using the associated target rate limit for flows toward these targets may mean overrunning the intermediate rate limit. Therefore, to ensure the best possible utilization and throughput within the associated limits, all target groups in the associated hierarchy can be considered in strict order of relevance. This means that packets can be sent toward the associated destination only if each target group in the hierarchy represents available bandwidth. Thus, if a single flow is active toward one of the targets in the above example, this flow would be allowed to operate at 30 Gb / s. However, if another flow becomes active towards another target (via a shared intermediate target group), each flow will be limited to 25 Gb / s. If, in the next round, an additional flow towards one of the two targets becomes active, the two flows to the same target will each be operating at 12.5 Gb / s (i.e., on average, and unless they have some additional bandwidth allocation / limitation).

[0178] According to one embodiment, when multiple tenants share a server / HCA, both the initial egress bandwidth and the actual target bandwidth may be shared in addition to sharing any intermediate ISL bandwidth. On the other hand, in a scenario with a dedicated server / HCA per tenant, the intermediate ISL bandwidth represents the only possible "inter-tenant" bandwidth sharing.

[0179] According to one embodiment, target groups should typically be global for an HCA port, and VF / tenant assignment at the HCA level would represent the maximum local traffic a tenant can generate for any combination of targets, globally or for a particular priority. Additionally, it would be possible to use several tenant-specific target groups alongside a "global" target group within the same hierarchy.

[0180] According to an embodiment, there are several possible ways to realize target groups and represent target group associations (hierarchy) for a particular QP or address handle. However, a 16-bit target group ID space and support for up to four or eight target group associations for each QP and address handle can be provided. Each target group ID value would then represent some HW state that reflects the associated IPD (inter-packet delay) value for the associated rate, as well as timer information that defines when the next packet associated with this target group can be sent.

[0181] According to one embodiment, different flows / paths may use different "QOS IDs" (i.e., service levels, priorities, etc.) on the same shared link segment, and therefore different target groups may be associated with the same link segment such that different target groups represent bandwidth allocations for such different QOS IDs. However, it is also possible to represent both QOS ID specific target groups and a single target group representing physical links of the same link segment.

[0182] Similarly, according to an embodiment, the present system and method can further distinguish between different flow types defined by explicit flow type packet header parameters and / or by taking into account operation type (e.g., RDMA read / write / send) and implement different "sub-allocations" to arbitrate between different such flow types. In particular, this can be useful to distinguish flows representing responder mode bandwidth (i.e., typically RDMA read response traffic) from requester mode traffic originally initiated by the local node itself.

[0183] According to one embodiment, with strict use of target groups and rate limits for all associated sending HCAs that add up to a total maximum rate that does not exceed the capacity of any target or shared ISL segment, it is possible, in principle, to avoid "any" congestion. However, this may imply a hard limit on both the sustained bandwidth for different flows and a low average utilization of the available link bandwidth. Therefore, various rate limits may be set to allow different HCAs to use more optimistic maximum rates. In this case, the aggregated sum may be larger than the sustainable maximum and thus lead to congestion.

[0184] FIG. 21 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment.

[0185] More specifically, according to one embodiment, Figure 21 illustrates a host channel adapter 2101 that includes a hypervisor 2111. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2114-2116, and a physical function (PF) 2113. The host channel adapter can further support or include several ports, such as ports 2102 and 2103, that are used to connect the host channel adapter to a network, such as network 2100. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2101 to several other nodes, such as a switch, additional separate HCAs, etc.

[0186] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 2150, VM2 2151, and VM3 2152.

[0187] According to one embodiment, the host channel adapter 2101 can further support a virtual switch 2112 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0188] According to one embodiment, network 2100 may include several switches, such as switches 2140, 2141, 2142, and 2143, which are interconnected and may be connected to host channel adapter 2101, for example, via leaf switches 2140 and 2141, as shown.

[0189] According to one embodiment, switches 2140-2143 may be interconnected, and may further be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure. ) can be connected to

[0190] According to an embodiment, target groups, such as target groups 2170 and 2171, can be defined along inter-switch links (ISLs), such as the ISL between leaf switch 2140 and switch 2142 and between leaf switch 2141 and switch 2143. These target groups 2170 and 2171 can represent bandwidth allocations as HCA objects stored in target group repository 2161 associated with the HCA, for example, accessible by rate limiting component 2160.

[0191] According to an embodiment, target groups 2170 and 2171 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0192] According to an embodiment, target groups 2170 and 2171 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2151 is set to one threshold, but the destination of packets sent from VM2 would pass through target group 2170, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0193] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2170 represents a higher bandwidth limitation than target group 2171 and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limitation of target group 2171 is the controlling bandwidth limitation factor.

[0194] According to an embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided according to the QoS and SLA associated with each flow. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2170 representing, for example, a 10 Gb / s bandwidth allocation, and each flow has equal QoS and SLA associated with it, target group 2170 would represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.

[0195] FIG. 22 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment.

[0196] More specifically, according to one embodiment, Figure 22 illustrates a host channel adapter 2201 that includes a hypervisor 2211. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2214-2216, and a physical function (PF) 2213. The host channel adapter can further support or include several ports, such as ports 2202 and 2203, that are used to connect the host channel adapter to a network, such as network 2200. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2201 to several other nodes, such as a switch, additional separate HCAs, etc.

[0197] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 2250, VM2 2251, and VM3 2252.

[0198] According to one embodiment, the host channel adapter 2201 can further support a virtual switch 2212 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0199] According to one embodiment, network 2200 may include several switches, such as switches 2240, 2241, 2242, and 2243, which may be interconnected and connected to host channel adapter 2201, for example, via leaf switches 2240 and 2241, as shown.

[0200] According to one embodiment, switches 2240-2243 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0201] According to an embodiment, target groups such as target groups 2270 and 2271 may be defined, for example, at switch ports. As shown, target groups 2270 and 2271 are defined at switch ports of switches 2242 and 2243, respectively. These target groups 2270 and 2271 may represent bandwidth allocations as HCA objects, for example, stored in target group repository 2261 associated with the HCA, accessible by rate limiting component 2260.

[0202] According to an embodiment, target groups 2270 and 2271 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0203] According to an embodiment, target groups 2270 and 2271 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2251 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2270, which sets a lower bandwidth limit, the egress bandwidth from VM2 would be less than the default / original egress limit imposed on VM2. The HCA may be responsible for such throttling / egress bandwidth limit adjustment depending on, for example, the target group involved in routing packets from VM2.

[0204] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages towards different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2270 represents a higher bandwidth limitation than target group 2271, and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limitation of target group 2271 is the controlling bandwidth limitation factor.

[0205] According to an embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided according to the QoS and SLA associated with each flow. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2270, representing a bandwidth allocation of, for example, 10 Gb / s, and each flow has equal QoS and SLA associated with it, target group 2270 would represent a limit of 5 Gb / s for each flow. This sharing or division of the target group bandwidth allocation can be modified based on the QoS and SLA associated with each flow.

[0206] According to one embodiment, Figures 21 and 22 show target groups defined at inter-switch links and switch ports, respectively. Those skilled in the art will readily understand that target groups may be defined at various locations within a subnet, and that any given subnet is not limited to target groups defined only at ISLs and switch ports, although in general such target groups may be defined at both ISLs and switch ports within any given subnet.

[0207] FIG. 23 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment.

[0208] More specifically, according to an embodiment, Figure 23 shows a host channel adapter 2301 that includes a hypervisor 2311. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2314-2316, and physical functions (PFs) 2313. The host channel adapter can further support or include several ports, such as ports 2302 and 2303, that are used to connect the host channel adapter to a network, such as network 2300. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2301 to several other nodes, such as a switch, additional separate HCAs, etc.

[0209] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 2350, VM2 2351, and VM3 2352.

[0210] According to one embodiment, the host channel adapter 2301 can also support a virtual switch 2312 via the hypervisor. This is known as a vSwitch adapter. Although not shown, embodiments of the present disclosure may further support a virtual port (vPort) architecture, as described above.

[0211] According to one embodiment, network 2300 may include several switches, such as switches 2340, 2341, 2342, and 2343, which may be interconnected and connected to host channel adapter 2301, for example, via leaf switches 2340 and 2341, as shown.

[0212] According to one embodiment, switches 2340-2343 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0213] According to an embodiment, target groups, such as target groups 2370 and 2371, can be defined along inter-switch links (ISLs), such as the ISL between leaf switch 2340 and switch 2342 and between leaf switch 2341 and switch 2343. These target groups 2370 and 2371 can represent bandwidth allocations as HCA objects, stored in target group repository 2361 associated with the HCA, for example, accessible by rate limiting component 2360.

[0214] According to an embodiment, target groups 2370 and 2371 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0215] According to an embodiment, target groups 2370 and 2371 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2351 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2370, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0216] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2370 represents a higher bandwidth limit than target group 2371, and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limit of target group 2371 is the controlling bandwidth limit factor.

[0217] According to an embodiment, a target group may also be shared by multiple flows. For example, a target group may be shared by multiple flows depending on the QoS and SLA associated with each flow. The bandwidth allocation represented by the group can be divided. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2370, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has equal QoS and SLA associated with it, target group 2370 would represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth allocation can be changed based on the QoS and SLA associated with each flow.

[0218] According to an embodiment, the target group repository may query 2375 the target group 2370, for example, to determine the bandwidth allocation of the target group. Upon determining the bandwidth allocation of the target group, the target group repository may store the allocation value associated with the target group. This allocation may then be used by the rate limiting component to: a) determine whether the bandwidth allocation of the target group is lower than that of the VM based on QoS or SLA, and b) upon such determination, update 2376 the VM's bandwidth allocation based on the path traversing the target group 2370.

[0219] FIG. 24 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment.

[0220] More specifically, according to one embodiment, Figure 24 illustrates a host channel adapter 2401 that includes a hypervisor 2411. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2414-2416, and a physical function (PF) 2413. The host channel adapter can further support or include several ports, such as ports 2402 and 2403, that are used to connect the host channel adapter to a network, such as network 2400. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2401 to several other nodes, such as a switch, additional separate HCAs, etc.

[0221] According to one embodiment, as described above, each of the virtual functions is VM1 2450, VM2 It can host virtual machines (VMs) such as 2451, VM3 2453, etc.

[0222] According to one embodiment, the host channel adapter 2401 can further support a virtual switch 2412 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0223] According to one embodiment, network 2400 may include several switches, such as switches 2440, 2441, 2442, and 2443, which may be interconnected and connected to host channel adapter 2401, for example, via leaf switches 2440 and 2441, as shown.

[0224] According to one embodiment, switches 2440-2443 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0225] According to one embodiment, target groups such as target groups 2470 and 2471 can be defined, for example, on switch ports. , target groups 2470 and 2471 are defined on the switch ports of switches 2442 and 2443, respectively. These target groups 2470 and 2471 may represent bandwidth allocations as HCA objects stored, for example, in a target group repository 2461 associated with the HCA, which is accessible by rate limiting component 2460.

[0226] According to an embodiment, target groups 2470 and 2471 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0227] According to an embodiment, target groups 2470 and 2471 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2451 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2470, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA can be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0228] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2470 represents a higher bandwidth limitation than target group 2471, and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limitation of target group 2471 is the controlling bandwidth limitation factor.

[0229] According to an embodiment, a target group can also be shared by multiple flows. For example, the bandwidth allocation represented by the target group can be divided according to the QoS and SLA associated with each flow. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2470, representing a bandwidth allocation of, for example, 10 Gb / s, and each flow has equal QoS and SLA associated with it, target group 2470 would represent a limit of 5 Gb / s for each flow. This sharing or division of the target group bandwidth allocation can be varied based on the QoS and SLA associated with each flow.

[0230] According to an embodiment, the target group repository may query 2475 the target groups 2470 to, for example, determine the bandwidth allocation of the target group. Upon determining the bandwidth allocation of the target group, the target group repository may store the allocation value associated with the target group. This allocation may then be used by the rate limiting component to: a) determine whether the bandwidth allocation of the target group is lower than that of the VM based on QoS or SLA, and b) upon such determination, limit the bandwidth allocation of the target group 2470 for paths traversing the target group 2470. Based on this, the VM's bandwidth allocation is updated 2476 .

[0231] FIG. 25 illustrates a system for combining multiple shared bandwidth segments in a high performance computing environment, according to an embodiment.

[0232] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 2500, several end nodes 2501 and 2502 may support several virtual machines VM1-VM4 2550-2553 interconnected via several switches, such as leaf switches 2511 and 2512, switches 2521 and 2522, and root switches 2531 and 2532.

[0233] Not shown are the various host channel adapters that provide functionality for the connection of nodes 2501 and 2502, as well as the virtual machines to be connected to the subnetwork, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0234] According to one embodiment, as discussed above, inherent in such a switched fabric is the concept that while each end node or VM may have its own egress / ingress bandwidth limits that traffic flowing into and out of it must adhere to, there may also be links or ports within a subnet that represent bottlenecks for traffic flowing therein. Thus, when determining what rates traffic should flow into or out of such an end node, such as VM1, VM2, VM3, or VM4, rate limiting components 2560 and 2561 can query various target groups, such as 2550 and 2551, to determine whether such target groups represent bottlenecks for traffic flows. In response to such determinations, rate limiting components 2560 and 2561 can then set different or new bandwidth limits for the endpoints they control.

[0235] Further, according to an embodiment, target groups can be queried in a nested / hierarchical manner such that if traffic from VM1 to VM3 utilizes both target groups 2550 and 2551, rate limit 2560 can take into account limits from both such target groups when determining the bandwidth limit from VM1 to VM3.

[0236] FIG. 26 is a flowchart of a method for supporting target groups for congestion control in a private fabric in a high performance computing environment, according to an embodiment.

[0237] According to one embodiment, in step 2610, the method may provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further including a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via a plurality of switches, and the first subnet further including a plurality of end nodes including a plurality of virtual machines.

[0238] According to an embodiment, in step 2620, the method may define a target group on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the target group defining a bandwidth limit on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches.

[0239] According to an embodiment, in step 2630, the method may provide, at the host channel adapter, a target group repository stored in the host channel adapter's memory.

[0240] According to an embodiment, in step 2640, the method may record the defined target groups in a target group repository.

[0241] Target-specific send / RDMA write and RDMA read bandwidth limit combinations (ORA200246-US-NP-2) According to one embodiment, a node / VM may be a target for incoming data traffic that is both the result of send and RDMA write operations initiated by peer nodes / VMs and the result of RDMA read operations initiated by the local node / VM itself. In such a situation, ensuring that the maximum or average ingress bandwidth of the local node / VM is within required bounds becomes problematic unless all of these flows are throttled with respect to rate limits.

[0242] According to an embodiment, the systems and methods described herein can provide target-specific egress rate control in a manner that allows all flows representing the fetching of data from local memory and the transmission of data to an associated remote target to all be subject to the same shared rate limit and associated flow scheduling and arbitration, and different flow types may be given different priorities and / or different shares of the available bandwidth.

[0243] According to one embodiment, there is full control of all ingress bandwidth to a vHCA port insofar as the target group association for a flow from a "producer / source" node implies bandwidth policing of all outgoing data packets, including UD (Unreliable Datagram) sends, RDMA writes, RDMA sends, and RDMA reads (i.e., RDMA read responses with data). This is regardless of whether the VM that owns the target vHCA port is generating an "excessive" amount of RDMA read requests to multiple peer nodes.

[0244] According to one embodiment, coupling target groups to both flow-specific and "unsolicited" BECN signaling means that ingress bandwidth per vHCA port can be dynamically throttled to any number of remote peers.

[0245] According to one embodiment, "unsolicited BECN" messages can also be used to communicate specific rate values ​​in addition to pure CE flagging / unflaggling for different stage numbers. In this way, it is possible to have a scheme where an initial incoming packet (e.g., a communications management (CM) packet) from a new peer can trigger the generation of one or more "unsolicited BECN" messages to both the HCA (i.e., the associated firmware / hyper-privileged software) from which the incoming packet came and to the current communicating peer.

[0246] According to one embodiment, in the case where both ports on an HCA are used simultaneously (i.e. In an active-active scheme), it may make sense to share target groups between local HCA ports when there is a possibility that concurrent flows may share some ISLs or even target the same destination port.

[0247] According to one embodiment, another reason for sharing target groups between HCA ports is if the HCA local memory bandwidth cannot sustain full link speed for both (all) HCA ports. In this case, the target groups can be configured so that the aggregate total link bandwidth never exceeds the local memory bandwidth, regardless of which ports are involved in either the source or destination HCA.

[0248] According to one embodiment, for a fixed route toward a particular destination, any intermediate target group will typically represent only a single ISL at a particular stage in the path. However, if dynamic forwarding is active, both the target group and ECN processing must take this into account. If dynamic forwarding decisions occur solely to balance traffic between parallel ISLs between a pair of switches (e.g., uplinks from a single leaf switch to a single spine switch), all processing is, in principle, very similar to when only a single ISL is used. FECN notifications are likely based on the state of all ports in the relevant group, and signaling can be "aggressive," in the sense that it is signaled based on congestion indications from one of the ports, or more conservative, based on the size of the shared output queue for all ports in the group. The target group configuration will typically represent the aggregate bandwidth for all links in the group, as long as it allows forwarding of any packet to select the best output port at that time. However, if there is a notion of strict packet ordering per flow, evaluation of bandwidth allocation is more complex, as several flows may "must" use the same ISL at some point. If such a flow ordering scheme is based on well-defined header fields, it may be best to represent each port in the group as a separate target group. In this case, the selection of the target group at the source HCA must be able to evaluate the header fields that will be associated with the RC QP connection or address handle in the same way that the switch performs at runtime for all packets.

[0249] According to one embodiment, by default, the initial target group rate for a new remote target may be conservatively set low. In this way, there is inherent throttling until the target has an opportunity to update its associated rate. Thus, all such rate control is independent of the participating VM itself, although a VM could request the hypervisor to update the allocation of different remote peers for both ingress and egress traffic, but this would only be allowed within the aggregate constraints defined for both the local and remote vHCA ports.

[0250] FIG. 27 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment.

[0251] More specifically, according to one embodiment, Figure 27 illustrates a host channel adapter 2701 that includes a hypervisor 2711. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2714-2716, and physical functions (PFs) 2713. The host channel adapter includes ports 2714-2716 that are used to connect the host channel adapter to a network, such as network 2700. The network may further support or comprise several ports, such as 2702 and 2703. The network may comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 2701 to several other nodes, such as a switch, additional separate HCAs, etc.

[0252] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 2750, VM2 2751, and VM3 2752.

[0253] According to one embodiment, the host channel adapter 2701 can further support a virtual switch 2712 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0254] According to one embodiment, network 2700 may include several switches, such as switches 2740, 2741, 2742, and 2743, which are interconnected and may be connected to host channel adapter 2701, for example, via leaf switches 2740 and 2741, as shown.

[0255] According to one embodiment, switches 2740-2743 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0256] According to an embodiment, target groups such as target groups 2770 and 2771 may be defined on inter-switch links (ISLs), such as the ISL between leaf switch 2740 and switch 2742 and between leaf switch 2741 and switch 2743. These target groups 2770 and 2771 may represent bandwidth allocations as HCA objects stored in target group repository 2761 associated with the HCA, for example, accessible by rate limiting component 2760.

[0257] According to an embodiment, target groups 2770 and 2771 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0258] According to an embodiment, target groups 2770 and 2771 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2751 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2770, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA may be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0259] According to one embodiment, the target groups may also be hierarchical in nature, This allows the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages towards different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2770 represents a higher bandwidth limit than target group 2771, and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limit of target group 2771 is the controlling bandwidth limit factor.

[0260] According to an embodiment, a target group may also be shared by multiple flows. For example, the bandwidth allocation represented by the target group may be divided according to the QoS and SLA associated with each flow. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2770, representing a bandwidth allocation of, for example, 10 Gb / s, and each flow has equal QoS and SLA associated with it, target group 2770 would represent a limit of 5 Gb / s for each flow. This sharing or division of the target group bandwidth allocation may be varied based on the QoS and SLA associated with each flow.

[0261] According to an embodiment, bandwidth allocation and performance issues can arise when a VM, e.g., VM1 2750, receives excessive ingress bandwidth 2790 from multiple sources. This can arise, for example, in a situation where VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, e.g., target group 2770 on an inter-switch link, can be updated, e.g., via query 2775, to reflect a lower bandwidth allocation than would typically be allowed.

[0262] According to an embodiment, the rate limiting component 2760 of the HCA may further include VM-specific rate limits 2762 that may be negotiated with other peer HCAs, for example, to align the ingress bandwidth limit for VM1 with the egress bandwidth limit for the node responsible for generating the ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.

[0263] FIG. 28 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment.

[0264] More specifically, according to one embodiment, Figure 28 illustrates a host channel adapter 2801 that includes a hypervisor 2811. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2814-2816, and a physical function (PF) 2813. The host channel adapter can further support or include several ports, such as ports 2802 and 2803, that are used to connect the host channel adapter to a network, such as network 2800. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2801 to several other nodes, such as a switch, additional separate HCAs, etc.

[0265] According to one embodiment, as described above, each of the virtual functions is VM1 2850, VM It can host virtual machines (VMs) such as VM2 2851 and VM3 2852.

[0266] According to one embodiment, host channel adapter 2801 can further support a virtual switch 2812 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0267] According to one embodiment, network 2800 may include several switches, such as switches 2840, 2841, 2842, and 2843, which are interconnected and may be connected to host channel adapter 2801, for example, via leaf switches 2840 and 2841, as shown.

[0268] According to one embodiment, switches 2840-2843 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0269] According to an embodiment, target groups, such as target groups 2870 and 2871, may be defined, for example, at switch ports. As shown, target groups 2870 and 2871 are defined at switch ports of switches 2842 and 2843, respectively. These target groups 2870 and 2871 may represent bandwidth allocations as HCA objects, stored, for example, in target group repository 2861 associated with the HCA, accessible by rate limiting component 2860.

[0270] According to an embodiment, target groups 2870 and 2871 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0271] According to an embodiment, target groups 2870 and 2871 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2851 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2870, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA may be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0272] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages towards different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, target group 2870 represents a higher bandwidth limit than target group 2871, and a packet is sent between both switches represented by the two target groups. When addressed over a link, the bandwidth limitation of the target group 2871 is the controlling bandwidth limitation factor.

[0273] According to an embodiment, a target group may also be shared by multiple flows. For example, the bandwidth allocation represented by the target group may be divided according to the QoS and SLA associated with each flow. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2870, representing a bandwidth allocation of, for example, 10 Gb / s, and each flow has equal QoS and SLA associated with it, target group 2870 would represent a limit of 5 Gb / s for each flow. This sharing or division of the target group bandwidth allocation may be varied based on the QoS and SLA associated with each flow.

[0274] According to an embodiment, bandwidth allocation and performance issues can arise when a VM, e.g., VM1 2850, receives excessive ingress bandwidth 2890 from multiple sources. This can arise, for example, in a situation where VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, e.g., target group 2870 on an inter-switch link, can be updated, e.g., via query 2875, to reflect a lower bandwidth allocation than would typically be allowed.

[0275] According to an embodiment, the rate limiting component 2860 of the HCA may further include VM-specific rate limits 2862 that may be negotiated with other peer HCAs, for example, to align ingress bandwidth limits for VM1 with egress bandwidth limits for nodes responsible for generating ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.

[0276] FIG. 29 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment.

[0277] More specifically, according to an embodiment, Figure 29 illustrates a host channel adapter 2901 that includes a hypervisor 2911. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 2914-2916, and a physical function (PF) 2913. The host channel adapter can further support or include several ports, such as ports 2902 and 2903, that are used to connect the host channel adapter to a network, such as network 2900. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 2901 to several other nodes, such as a switch, additional separate HCAs, etc.

[0278] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 2950, ​​VM2 2951, and VM3 2952.

[0279] According to one embodiment, the host channel adapter 2901 can also support a virtual switch 2912 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure may ,As ,mentioned ,above, ,a ,virtual ,port ,(vPort) ,architecture ,can ,be ,additionally ,supported.

[0280] According to one embodiment, network 2900 may include several switches, such as switches 2940, 2941, 2942, and 2943, which may be interconnected and connected to host channel adapter 2901, for example, via leaf switches 2940 and 2941, as shown.

[0281] According to one embodiment, switches 2940-2943 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0282] According to one embodiment, target groups such as target group 2971 can be defined along an inter-switch link (ISL), such as the ISL between leaf switch 2941 and switch 2943. Other target groups can be defined, for example, at a switch port. As shown, target group 2970 is defined at a switch port of switch 2952. These target groups 2970 and 2971 can represent bandwidth allocations as HCA objects, stored in target group repository 2961 associated with the HCA, for example, accessible by rate limiting component 2960.

[0283] According to an embodiment, target groups 2970 and 2971 can represent specific (and different) bandwidth allocations, which can be divided into allocations for specific priorities ("QOS classes") to reflect expected bandwidth usage for various priorities within the associated paths in the fabric that the target groups represent.

[0284] According to an embodiment, target groups 2970 and 2971 decouple objects from specific destination addresses, and the present system and method gain the ability to represent, in addition to targets, intermediate shared links or groups of links that may represent bandwidth limits that may be more restrictive than the target limits. That is, for example, if the default / original egress limit for VM2 2951 is set to one threshold, but the destination of a packet sent from VM2 would pass through target group 2970, which sets a lower bandwidth limit, the egress bandwidth from VM2 may be limited to a level lower than the default / original egress limit imposed on VM2. The HCA may be responsible for such throttling / egress bandwidth limit adjustment, for example, depending on the target group involved in routing packets from VM2.

[0285] According to an embodiment, target groups can also be hierarchical in nature, allowing the present system and method to consider a hierarchy of target groups (bandwidth allocation) that reflects bandwidth / link sharing at different stages toward different targets. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2970 represents a higher bandwidth limitation than target group 2971, and a packet is addressed via both inter-switch links represented by the two target groups, the bandwidth limitation of target group 2971 is the controlling bandwidth limitation factor.

[0286] According to an embodiment, a target group may also be shared by multiple flows. For example, a target group may be shared by multiple flows depending on the QoS and SLA associated with each flow. The bandwidth allocation represented by the group can be divided. As an example, if VM1 and VM2 both simultaneously transmit flows that would involve target group 2970, which represents, for example, a 10 Gb / s bandwidth allocation, and each flow has equal QoS and SLA associated with it, target group 2970 would represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth allocation can be changed based on the QoS and SLA associated with each flow.

[0287] According to an embodiment, bandwidth allocation and performance issues can arise when a VM, e.g., VM1 2950, ​​receives excessive ingress bandwidth 2990 from multiple sources. This can arise, for example, in a situation where VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a situation, a target group, e.g., target group 2970 on an inter-switch link, can be updated, e.g., via query 2975, to reflect a lower bandwidth allocation than would typically be granted.

[0288] According to an embodiment, the rate limiting component 2960 of the HCA may further include VM-specific rate limits 2962 that may be negotiated with other peer HCAs, for example, to align the ingress bandwidth limit for VM1 with the egress bandwidth limit for the node responsible for generating the ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.

[0289] FIG. 30 illustrates a system for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment.

[0290] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3000, several end nodes 3001 and 3002 may support several virtual machines VM1-VM4 3050-3053 interconnected via several switches such as leaf switches 3011 and 3012, switches 3021 and 3022, and root switches 3031 and 3032.

[0291] Not shown are the various host channel adapters that provide functionality for the connection of nodes 3001 and 3002, as well as the virtual machines to be connected to the subnetwork, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0292] According to one embodiment, a node such as VM3 3052 may enter a bandwidth limit (e.g., from rate limit 3061) when simultaneously processing an RDMA read response 3050 and an RDMA write request 3051 (incoming bandwidth).

[0293] According to one embodiment, rate limits 3060 and 3061 can be configured to ensure that ingress bandwidth allocations are not violated, for example, by throttling RDMA requests (i.e., messages sent by VM3 to VM4 requesting an RDMA read and resulting in an RDMA read response 3050) and RDMA write operations (e.g., RDMA writes from VM2 to VM3).

[0294] For each individual node, the system and method can have a chain of such target groups so that a flow is always coordinated with all other flows that share link bandwidth in different parts of the fabric represented in the target group.

[0295] FIG. 31 is a flowchart of a method for combining target-specific RDMA write and read bandwidth limits in a high performance computing environment, according to one embodiment.

[0296] According to one embodiment, in step 3110, the method can provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further including a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via a plurality of switches, and the first subnet further including a plurality of end nodes including a plurality of virtual machines.

[0297] According to an embodiment, in step 3120, the method can define a target group on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the target group defining a bandwidth limit on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches.

[0298] According to an embodiment, in step 3130, the method can provide, at the host channel adapter, a target group repository stored in the host channel adapter's memory.

[0299] According to an embodiment, in step 3140, the method may record the defined target groups in a target group repository.

[0300] In one embodiment, in step 3150, the method can receive ingress bandwidth at an end node of a host channel adapter from at least two remote sources, where the ingress bandwidth exceeds the ingress bandwidth limit of the end node.

[0301] According to an embodiment, at 3160, in response to receiving ingress bandwidth from at least two sources, the method can update the bandwidth allocation of the target group.

[0302] Combining Ingress Bandwidth Arbitration with Congestion Feedback (ORA200246-US-NP-2) According to one embodiment, when multiple sender nodes / VMs are each and / or all transmitting to a single receiver node / VM, achieving an optimal balance of fairness among the senders to avoid congestion while simultaneously limiting the ingress bandwidth usage consumed by the receiver node / VM below a maximum limit (sufficiently) below the maximum physical link bandwidth that the associated network interface can provide to the ingress traffic is not straightforward. Moreover, the equation becomes even more complicated when different senders are assumed to be allocated different bandwidth allocations due to different SLA levels.

[0303] According to one embodiment, the systems and methods herein can extend older schemes for end-to-end congestion feedback to include both the initial negotiation of bandwidth allocations, the dynamic adjustment of such bandwidth allocations (e.g., to adapt to changes in the number of sending nodes sharing the available bandwidth or to changes in SLAs), and dynamic congestion feedback to indicate that a sender needs to temporarily slow down its associated egress data rate while the overall bandwidth allocation remains the same. Both explicit unsolicited messages and "piggybacked" information within data packets are used to convey relevant information from the target node to the sending node.

[0304] FIG. 32 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment.

[0305] More specifically, according to an embodiment, Figure 32 shows a host channel adapter 3201 that includes a hypervisor 3211. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 3214-3216, and physical function (PF) 3213. The host channel adapter can further support or include several ports, such as ports 3202 and 3203, that are used to connect the host channel adapter to a network, such as network 3200. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 3201 to several other nodes, such as a switch, additional separate HCAs, etc.

[0306] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 3250, VM2 3251, and VM3 3252.

[0307] According to one embodiment, the host channel adapter 3201 can further support a virtual switch 3212 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0308] According to one embodiment, network 3200 may include several switches, such as switches 3240, 3241, 3242, and 3243, which may be interconnected and connected to host channel adapter 3201, for example, via leaf switches 3240 and 3241, as shown.

[0309] According to one embodiment, switches 3240-3243 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0310] According to an embodiment, bandwidth allocation and performance issues can arise when a VM, for example, VM1 3250, receives excessive ingress bandwidth from multiple sources 3290. This can occur, for example, in a situation where VM1 receives one or more RDMA read responses simultaneously with one or more RDMA write operations, and the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM).

[0311] According to an embodiment, the rate limiting component 3260 of the HCA further generates an ingress bandwidth limit for, for example, VM1 on VM1. It may further include VM-specific rate limits 3261 that may be negotiated with other peer HCAs to coordinate with the egress bandwidth limits for the responsible node. Such initial negotiations may be performed, for example, to accommodate changes in the number of sending nodes sharing the available bandwidth or changes in SLAs. These other HCAs / nodes are not shown in the figure.

[0312] According to an embodiment, the negotiation can be updated based on an explicit, unsolicited feedback message 3291 generated as a result of, for example, ingress bandwidth. Such feedback message 3291 can be sent to multiple remote nodes responsible for generating ingress bandwidth 3290 on VM1, for example. Upon receiving such a feedback message, the sending node (the bandwidth sender responsible for the ingress bandwidth on VM1) can update their associated egress bandwidth limits to avoid overloading, for example, the link connecting VM1, while attempting to maintain QoS and SLAs.

[0313] FIG. 33 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment.

[0314] More specifically, according to an embodiment, Figure 33 shows a host channel adapter 3301 that includes a hypervisor 3311. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 3314-3316, and physical functions (PFs) 3313. The host channel adapter can further support or include several ports, such as ports 3302 and 3303, that are used to connect the host channel adapter to a network, such as network 3300. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 3301 to several other nodes, such as a switch, additional separate HCAs, etc.

[0315] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 3350, VM2 3351, and VM3 3352.

[0316] According to one embodiment, the host channel adapter 3301 can further support a virtual switch 3312 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0317] According to one embodiment, network 3300 may include several switches, such as switches 3340, 3341, 3342, and 3343, which may be interconnected and connected to host channel adapter 3301, for example, via leaf switches 3340 and 3341, as shown.

[0318] According to one embodiment, switches 3340-3343 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0319] According to an embodiment, bandwidth allocation and performance issues can arise when a VM, e.g., VM1 3350, receives excessive ingress bandwidth 3390 from multiple sources. This can occur, for example, when VM1 receives one or more RDMA write operations and one or more RDMA read responses simultaneously, causing the ingress bandwidth on VM1 to be overwhelmed by data from two or more sources (e.g., connections This can occur in situations where multiple RDMA requests come from different VMs (one RDMA read response from a connected VM, and one RDMA write request from another connected VM).

[0320] According to an embodiment, the rate limiting component 3360 of the HCA may further include VM-specific rate limits 3361 that may be negotiated with other peer HCAs, for example, to align ingress bandwidth limits for VM1 with egress bandwidth limits for nodes responsible for generating ingress bandwidth on VM1. Such initial negotiations may be performed, for example, to accommodate changes in the number of sending nodes sharing the available bandwidth or changes in SLAs. These other HCAs / nodes are not shown in the figure.

[0321] According to one embodiment, the negotiation can be updated based on, for example, piggyback messages 3391 (messages that reside on top of normal data or other communication packets sent between end nodes) generated as a result of ingress bandwidth. Such piggyback messages 3391 can be sent to multiple remote nodes responsible for generating ingress bandwidth 3390 on VM1, for example. Upon receiving such feedback messages, the sending nodes (bandwidth senders responsible for the ingress bandwidth on VM1) can update their associated egress bandwidth limits to avoid overloading, for example, links connecting VM1, while attempting to maintain QoS and SLAs.

[0322] FIG. 34 illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to an embodiment.

[0323] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3400, several end nodes 3401 and 3402 may support several virtual machines VM1-VM4 3450-3453 interconnected via several switches, such as leaf switches 3411 and 3412, switches 3421 and 3422, and root switches 3431 and 3432.

[0324] Not shown are the various host channel adapters that provide functionality for the connection of nodes 3401 and 3402, as well as the virtual machines to be connected to the subnets, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0325] According to an embodiment, a node such as VM3 3452 may enter an ingress bandwidth limit (e.g., from rate limit 3461) upon receiving multiple RDMA ingress bandwidth packets (e.g., multiple RDMA writes), such as 3451 and 3452. This may occur, for example, when there is no communication between the various sending nodes to adjust the bandwidth limits.

[0326] According to one embodiment, the systems and methods herein can extend the scheme for end-to-end congestion feedback to include both an initial negotiation of bandwidth allocation (i.e., negotiating with VM3, or all sender nodes whose bandwidth limits are associated with VM3, targeting VM3 in their ingress traffic), dynamic adjustment of such bandwidth allocation (e.g., to adapt to changes in the number of sender nodes sharing the available bandwidth or changes in SLAs), and dynamic congestion feedback to indicate that a sender needs to temporarily slow down its associated egress data rate while the overall bandwidth allocation remains the same. Such dynamic congestion feedback can be, for example, This may occur in return messages (e.g., feedback messages 3470) to the various sending nodes instructing each sending node on updated bandwidth limits to utilize when sending traffic to VM3. Such feedback messages 3460 may take the form of explicit unsolicited messages and "piggybacking" information within data packets to convey relevant information from the target node (i.e., VM3 in the illustrated embodiment) to the sending nodes.

[0327] FIG. 35 is a flowchart of a method for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment, according to one embodiment.

[0328] According to one embodiment, in step 3510, the method may provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further including a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via a plurality of switches, and the first subnet further including a plurality of end nodes including a plurality of virtual machines.

[0329] According to an embodiment, in step 3520, the method may provide, at the host channel adapter, an end node ingress bandwidth allocation associated with an end node attached to the host channel adapter.

[0330] According to one embodiment, in step 3530, the method may negotiate bandwidth allocation between an end node attached to the host channel adapter and a remote end node.

[0331] According to one embodiment, in step 3540, the method can receive ingress bandwidth from a remote source at an end node attached to a host channel adapter, the ingress bandwidth exceeding the ingress bandwidth limit of the end node.

[0332] According to one embodiment, at 3550, in response to receiving ingress bandwidth from at least two sources, the method can send a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress bandwidth allocation of the end node attached to the host channel adapter has been exceeded.

[0333] Use of Multiple Congestion Experience (CE) Flags in Both Forward Explicit Congestion Notification (FECN) and Backward Explicit Congestion Notification (BECN) Signaling (ORA200246-US-NP-4) According to one embodiment, congestion notification is traditionally based on data packets encountering congestion at some point (e.g., some link segment between some node / switch pair along the path from sender to target through the network / fabric topology) being marked with a "congestion" status flag (also called a CE flag), and this status is then reflected in a response packet sent from the target back to the sender.

[0334] According to one embodiment, a problem with this scheme is that it does not allow a sending node to distinguish between flows that are subject to congestion on the same link segment, even though they represent different targets. Also, when multiple paths are available between a sending node and a target node pair, it is difficult to distinguish between flows that are subject to congestion on different alternative paths. Any information in requires that some flow is active for the relevant target over the relevant path.

[0335] According to an embodiment, the systems and methods described herein extend the congestion marking scheme to facilitate multiple CE flags in the same packet, configuring switch ports to represent stage numbers that define which CE flag indexes should be updated. A particular path between a particular sender and a particular target through an ordered sequence of switch ports represents a particular ordered list of unique stage numbers, and thereby also represents a CE flag index number.

[0336] In this way, according to an embodiment, a sender node receiving congestion feedback with multiple CE flags set can map the various CE flags to different "target group" contexts that will represent the relevant congestion condition states and associated dynamic rate reductions. Furthermore, different flows to different targets will share congestion information and dynamic rate reduction states associated with the shared link segment that is represented by the shared "target group" at the sender node.

[0337] According to one embodiment, when congestion occurs, a key issue is that congestion feedback should ideally be associated with all relevant target groups in the tier associated with the flow receiving the congestion feedback. The affected target groups should then dynamically adjust their maximum rates accordingly. Therefore, the HW state of each target group must also include any current congestion status and associated "throttle information."

[0338] According to one embodiment, a key aspect here is that FECN signaling should have the ability to include multiple "Congestion Experienced" (CE) flags so that a switch that detects congestion can mark a flag corresponding to that stage in the topology. In a typical Fat-Tree, each switch has a unique (highest) stage number upwards and another unique (highest) stage number downwards. Thus, a flow using a particular path will then be associated with a particular sequence of stage numbers that will include all or only a subset of the entire set of stage numbers in the complete fabric. However, for that particular flow, the various stage numbers associated with that path can then be mapped to one or more target groups associated with that flow. In this way, a received BECN for a flow can mean that the target groups associated with each CE-flagged stage in the BECN will be updated to indicate congestion, and the dynamic maximum rates for these target groups can then be adjusted accordingly.

[0339] According to one embodiment, while inherently suited to a fat-tree topology, the concept of a switch "stage number" can be generalized to represent almost any topology in which such numbers can be assigned to a switch. However, in this general case, the stage number is not simply a function of the output port, but rather a function of each input / output port number tuple. The required quantity of stage numbers and the path-specific mapping to target groups are also more complex in the general case. Therefore, in this context, the inference assumes only a fat-tree topology.

[0340] According to an embodiment, multiple CE flags in a single packet are not currently a supported feature for standard protocol headers. Therefore, this may be supported based on an extension of the standard header and / or by inserting additional independent FECN packets into the flow. Conceptually, the creation of an additional packet in a flow is similar to the use of an encapsulation scheme within a switch, and the effect is that packets being received at wire speed cannot be forwarded at the same wire speed because more "overhead bytes" must be sent downstream. Inserting an additional packet typically results in more overhead than encapsulation, but as long as this overhead is amortized over multiple data packets (there is no need to send such additional notifications for every data packet), this overhead is likely to be tolerable.

[0341] According to one embodiment, it is also possible to have a scheme whereby the switch firmware can monitor congestion conditions within the switch and, as a result, send an "unsolicited BECN" to the associated sending node. However, this means that the switch firmware must have more state information about the associated senders, as well as the mapping between ports, priorities, and associated senders, which may also include dynamic information about what addresses are involved in packets experiencing congestion.

[0342] According to an embodiment, for RC QP, the "CE flag to target group" mapping will typically be part of the QP context, and any BECN information received in an ACK / response packet will thereby be handled in a simple manner for the relevant QP context and associated target group. However, in the case of "unsolicited BECN" (e.g., as a result of datagram traffic with only application-level responses / ACKs, or as a result of a "congestion warning" being broadcast to multiple potential senders), the reverse mapping is not simple—at least not in terms of being automatically handled by the HW. Therefore, a better approach is to have a scheme in which, although both FECNs can lead to automatic HW-generated BECNs in the case of connected (RC) flows, both FECN events with HW automatic BECN generation and FECN events without HW-generated BECNs can be handled by firmware and / or hyper-privileged software associated with the HCA receiving the FECN. In this way, there can be a FW / SW-generated "unsolicited BECN" sent to one or more potential senders affected by the observed congestion. The FW / SW receiving these "unsolicited BECNs" can then perform mapping to the relevant local target groups based on the payload data in the received "BECN message" and then trigger the local HW to update the target group state, similar to what occurs in full HW controlled processing of RC-associated BECNs.

[0343] According to an embodiment, an RC ACK / Response packet without a BECN notification, or a subset of stage numbers with the CE flag set that differs (fewer) from the previously recorded state, may lead to a corresponding update of the associated target group in the local HCA. Similarly, an "Unsolicited BECN" may be sent by the responding HCA (i.e., the associated sw / fw) to indicate that the previously signaled congestion no longer exists.

[0344] According to one embodiment, as described above, the target group concept combined with dynamic congestion feedback at either the HW level or the FW / SW level provides flexible control of egress bandwidth generated by HCAs as well as by tenants sharing individual vHCAs and physical HCAs.

[0345] According to one embodiment, target groups are identified completely independently of the associated remote address and route information at the VM level, so that the use of target groups and the extent to which communications from VMs are based on overlays or other virtual networking schemes can be easily managed. There is no dependency between them. The only requirement is that the hyper-privileged software controlling the HCA resources be able to define the relevant mappings. It would also be possible to use a scheme where at the VM / vHCA level, there is a "logical target group ID" that is mapped by the HCA to an actual target group. However, it is not obvious that this would be useful, except to hide the actual target group ID from the tenant. - If you need to change what target group is associated with a particular destination because the underlying route has changed, this may not involve other destinations. Therefore, in the general case, updating a target group should involve updating all involved QPs and address handles, rather than simply updating the logical-to-physical target group ID mapping.

[0346] According to one embodiment, for a virtualized target HCA, it is possible to represent individual vHCA ports as the final destination target group, rather than the physical HCA port. In this way, the target group hierarchy for a remote peer node can include both a target group representing the destination physical HCA port and an additional target group representing the final destination for the vHCA port. In this way, the present system and method have the ability to limit the ingress bandwidth of individual vHCA ports (VFs), while the bandwidth per physical HCA port and associated sender target group means that the sum of the ingress bandwidth allocations per vHCA port need not be kept less than the physical HCA port bandwidth (or associated bandwidth allocation).

[0347] According to one embodiment, within a sending HCA, target groups can be used to represent the sharing of a physical HCA port in the egress direction by assigning different target groups to different tenants. Also, to facilitate multiple VMs from the same tenant sharing a tenant-level target group for a physical HCA port, different target groups can be assigned to different such VMs. Such a target group will then be set up as the initial target group for all egress communications from that VM.

[0348] FIG. 36 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment.

[0349] More specifically, according to an embodiment, Figure 36 shows a host channel adapter 3601 that includes a hypervisor 3611. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 3614-3616, and physical function (PF) 3613. The host channel adapter can further support or include several ports, such as ports 3602 and 3603, that are used to connect the host channel adapter to a network, such as network 3600. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 3601 to several other nodes, such as a switch, additional separate HCAs, etc.

[0350] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 3650, VM2 3651, and VM3 3652.

[0351] According to one embodiment, the host channel adapter 3601 can further support a virtual switch 3612 via a hypervisor, for situations where a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can also support a virtual port (vPort) architecture, as described above. can.

[0352] According to one embodiment, network 3600 may include several switches, such as switches 3640, 3641, 3642, and 3643, which are interconnected and may be connected to host channel adapter 3601, for example, via leaf switches 3640 and 3641, as shown.

[0353] According to one embodiment, switches 3640-3643 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0354] According to an embodiment, an ingress packet 3690 may experience congestion at any stage along its path while traversing the network, and the packet can be marked by a switch upon detecting such congestion at any of the stages. In addition to marking the packet as having experienced congestion, the marking switch can further indicate the stage at which the packet experienced congestion. Upon arriving at a destination node, e.g., VM1 3650, VM1 can (e.g., automatically) send a response packet via an explicit feedback message 3691 that can indicate to the sending node that the packet experienced congestion and at which stage the packet experienced congestion.

[0355] According to one embodiment, ingress packets may include a bit field that is updated to indicate where the packet experienced congestion, and explicit feedback messages may mirror / represent this bit field when notifying the sending node of such congestion.

[0356] According to one embodiment, each switch port represents a stage in the entire subnet. Therefore, each packet transmitted in the subnet can traverse a maximum number of stages. To identify where congestion was detected (which could be in multiple places), the congestion marking (e.g., CE flag) is extended from a simple binary flag (experienced congestion) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages would be three. If the system has a path from A to B and the routing is known, each end node can determine through which switch port the packet traversed at any given stage of the path. By doing so, each end node can determine at which distinct switch port the packet experienced congestion by correlating the routing with the received congestion message.

[0357] According to an embodiment, the system may provide congestion feedback back indicating at what stage congestion is detected, and then if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer granularity of information regarding congestion.

[0358] According to one embodiment, by providing such finer granularity, an end node can then use alternate paths when routing future packets. Or, for example, if an end node has multiple flows that all go to different destinations, but congestion is detected at a common stage in the path, rerouting can be triggered. Rather than having ten different congestion notifications, the present system and method provides an immediate reaction in terms of throttling associated with it. This allows for more efficient handling of congestion notifications. This is logical.

[0359] FIG. 37 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment.

[0360] More specifically, according to an embodiment, Figure 37 shows a host channel adapter 3701 that includes a hypervisor 3711. The hypervisor can host or be associated with several virtual functions (VFs), such as VFs 3714-3716, and physical functions (PFs) 3713. The host channel adapter can further support or include several ports, such as ports 3702 and 3703, that are used to connect the host channel adapter to a network, such as network 3700. The network can comprise, for example, a switched network, such as an InfiniBand network or a RoCE network, that can connect HCA 3701 to several other nodes, such as a switch, additional separate HCAs, etc.

[0361] According to one embodiment, as described above, each of the virtual functions can host a virtual machine (VM), such as VM1 3750, VM2 3751, and VM3 3752.

[0362] According to one embodiment, the host channel adapter 3701 can further support a virtual switch 3712 via a hypervisor, for situations in which a vSwitch architecture is implemented. Although not shown, embodiments of the present disclosure can further support a virtual port (vPort) architecture, as described above.

[0363] According to one embodiment, network 3700 may include several switches, such as switches 3740, 3741, 3742, and 3743, which are interconnected and may be connected to host channel adapter 3701, for example, via leaf switches 3740 and 3741, as shown.

[0364] According to one embodiment, switches 3740-3743 may be interconnected and may further be connected to other switches and other end nodes (eg, other HCAs) not shown in the figure.

[0365] According to an embodiment, an ingress packet 3790 may experience congestion at any stage of its path while traversing the network, and the packet can be marked by a switch upon detecting such congestion at any of the stages. In addition to marking the packet as having experienced congestion, the marking switch can further indicate the stage at which the packet experienced congestion. Upon arriving at a destination node, e.g., VM1 3750, VM1 can (e.g., automatically) send a response packet via a piggyback message (a message on top of another message / packet sent from the receiving node to the sending node) 3791 that can indicate to the sending node that the packet experienced congestion and at which stage the packet experienced congestion.

[0366] According to one embodiment, ingress packets may include a bit field that is updated to indicate where the packet experienced congestion, and explicit feedback messages may mirror / represent this bit field when notifying the sending node of such congestion.

[0367] According to one embodiment, each switch port represents a stage in the entire subnet. Therefore, each packet transmitted in the subnet can traverse a maximum number of stages. To identify where congestion was detected (which could be in multiple places), the congestion marking (e.g., CE flag) is extended from a simple binary flag (experienced congestion) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages would be three. If the system has a path from A to B and the routing is known, each end node can determine through which switch port the packet traversed at any given stage of the path. By doing so, each end node can determine at which distinct switch port the packet experienced congestion by correlating the routing with the received congestion message.

[0368] According to an embodiment, the system may provide congestion feedback back indicating at what stage congestion is detected, and then if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer granularity of information regarding congestion.

[0369] According to one embodiment, by providing such finer granularity, an end node can then use alternative paths when routing future packets. Or, for example, if an end node has multiple flows that all go to different destinations, but congestion is detected at a common stage in the path, rerouting can be triggered. Rather than having 10 different congestion notifications, the present system and method provides an immediate reaction in terms of the associated throttling. This is a more efficient handling of congestion notifications.

[0370] FIG. 38 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment.

[0371] According to one embodiment, within a high-performance computing environment such as a switched network or subnet 3800, several end nodes 3801 and 3802 may support several virtual machines VM1-VM4 3850-3853 that are interconnected via several switches, such as leaf switches 3811 and 3812, switches 3821 and 3822, and root switches 3831 and 3832.

[0372] Not shown are the various host channel adapters that provide functionality for the connection of nodes 3801 and 3802, as well as the virtual machines to be connected to the subnetwork, according to one embodiment. The discussion of such an embodiment is described above with respect to SR-IOV, and each virtual machine may be associated with a hypervisor virtual function on a host channel adapter.

[0373] According to one embodiment, packet 3851 sent from node VM3 3852 to VM1 3850 may traverse subnet 3800 via several links, or stages, such as stage 1 through stage 6, as shown. While traversing the subnet, packet 3851 may experience congestion at any of these stages and may be marked by a switch upon detecting such congestion at any of the stages. In addition to marking the packet as having experienced congestion, the marking switch may further indicate the stage at which the packet experienced congestion. Upon reaching destination node VM1, VM1 may indicate that the packet has experienced congestion and at which stage A response packet may be sent (eg, automatically) via a feedback message 3870 that may indicate to VM3 3852 if the packet experienced congestion at that stage.

[0374] According to one embodiment, each switch port represents a stage in the entire subnet. Therefore, each packet transmitted in the subnet can traverse a maximum number of stages. To identify where congestion was detected (which could be in multiple places), the congestion marking (e.g., CE flag) is extended from a simple binary flag (experienced congestion) to a bit field containing multiple bits. Each bit in the bit field can then be associated with a stage number that can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages would be three. If the system has a path from A to B and the routing is known, each end node can determine through which switch port the packet traversed at any given stage of the path. By doing so, each end node can determine at which distinct switch port the packet experienced congestion by correlating the routing with the received congestion message.

[0375] According to an embodiment, the system may provide congestion feedback back indicating at what stage congestion is detected, and then if an end node has congestion caused by a shared link segment, congestion control is applied to that segment rather than a different end port. This provides finer granularity of information regarding congestion.

[0376] According to one embodiment, by providing such finer granularity, an end node can then use alternative paths when routing future packets. Or, for example, if an end node has multiple flows that all go to different destinations, but congestion is detected at a common stage in the path, rerouting can be triggered. Rather than having 10 different congestion notifications, the present system and method provides an immediate reaction in terms of the associated throttling. This is a more efficient handling of congestion notifications.

[0377] FIG. 39 is a flowchart of a method for using multiple CE flags in both FECN and BECN in a high performance computing environment, according to an embodiment.

[0378] According to one embodiment, in step 3910, the method may provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further including a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via a plurality of switches, and the first subnet further including a plurality of end nodes including a plurality of virtual machines.

[0379] According to one embodiment, in step 3920, the method may receive, at an end node attached to the host channel adapter, an ingress packet from a remote end node, the ingress packet traversing at least a portion of a first subnet before being received at the end node, and the ingress packet includes marking indicating that the ingress packet experienced congestion while traversing at least a portion of the first subnet.

[0380] According to one embodiment, upon receiving an ingress packet, in step 3930 The method may include sending, by the end node, a response message from the end node attached to the host channel adapter to a remote end node, the response message indicating that the ingress packet experienced congestion during traversal of at least a portion of the first subnet, the response message including a bit field.

[0381] QOS and SLA in switched fabrics, including private fabrics According to one embodiment, private network fabrics within clouds and larger clouds at customer and premise installations (e.g., private fabrics such as those used to build dedicated distributed appliances or general-purpose high-performance computing resources) desire the ability to deploy VM-based workloads, and one inherent requirement is the ability to define and control quality of service (QOS) for different types of communication flows. Additionally, workloads belonging to different tenants must run within the bounds of associated service level agreements (SLAs) while minimizing interference between such workloads and maintaining QOS assumptions for different communication types.

[0382] In accordance with one embodiment, the following sections discuss relevant problem scenarios, goals, and potential solutions.

[0383] According to one embodiment, the initial scheme for provisioning fabric resources to cloud customers (also known as "tenants") is that tenants can be allocated a dedicated portion of a rack (e.g., an allotted rack) or one or more full racks. This granularity means that each tenant is guaranteed to always have its communication SLAs met, as long as the allocated resources are fully operational. This is also true when a single rack is divided into multiple portions, because the granularity is always a full physical server with an HCA. Connectivity between different such servers within a single rack can, in principle, always be via a single full crossbar switch. In this case, communication traffic between sets of servers belonging to the same tenant does not result in resources being shared in a way that could lead to contention or congestion between flows belonging to different tenants.

[0384] However, according to embodiments, since the redundant switches are shared, it is crucial that traffic generated by workloads on one server cannot target a server belonging to another tenant. While such traffic does not facilitate any inter-tenant communication or data leakage / observation, the result could be severe interference or even DOS-like effects on communication flows belonging to other tenants.

[0385] Despite the fact that, according to one embodiment, a full crossbar leaf switch inherently means that all communication between local servers can occur only through the local switch, there are some cases where this may not be possible or cannot be achieved due to other practical issues. According to one embodiment, it is important that only one HCA port is being used for data traffic at any given time, for example, if the host bus (PCIe) generation can only sustain the bandwidth of one fabric link at a time. Therefore, if all servers do not agree on which local switch to use for data traffic, some traffic will have to traverse the inter-switch links (ISLs) between the local leaf switches. According to one embodiment, if one or more servers lose connectivity to one of the switches, all communication must go through the other switch. Thus, again, if all pairs of servers cannot agree on using the same single switch, In this case, some data traffic will have to go through the ISL. According to one embodiment, if a server can use both HCA ports (and therefore both leaf switches), but it is not possible to enforce that connections are only established through HCA ports that connect to the same switch, some data traffic may traverse an ISL. One reason for ending up with this scheme is that the lack of socket / port numbers in the Fabric host stack means that a process can only establish one socket to accept incoming connections. This socket can then only be associated with a single HCA port at a time. As long as that same single socket is also used when establishing outgoing connections, the system will end up with a number of connections requiring ISLs, even though multiple processes have their single sockets evenly distributed among the local HCA ports.

[0386] According to one embodiment, in addition to the special case single-rack scenario imposing ISL usage / sharing outlined above, when the provisioning granularity is extended to a multi-rack configuration where leaf switches in each rack are interconnected by spine switches, then the communication SLAs for different tenants become highly dependent on which servers are allocated to which tenants and how different communication flows are mapped by the fabric-level routing scheme onto different switch-to-switch links. A key issue in this scenario is that the two optimization aspects are somewhat contradictory. According to one embodiment, on the one hand, in order to provide the best possible performance, all concurrent flows targeting different destination ports should use as many different paths through the fabric (i.e., different switches - switch links - ISLs) as possible. According to an embodiment, on the other hand, in order to provide predictable QOS and SLAs for different tenants, it is important that flows belonging to different tenants do not compete for bandwidth on the same ISL at the same time. In the general case, this means that there must be restrictions on which paths can be used by different tenants.

[0387] However, according to an embodiment, in some situations, depending on the size of the system, the number of tenants, and how the servers are provisioned for different tenants, it may not be possible to avoid flows belonging to different tenants from competing for bandwidth on the same ISL. In this situation, there are two main approaches that can be used from a fabric perspective to address the problem and reduce the possibility of line contention. According to an embodiment, limiting which switch buffer resources can be occupied by different tenants (or groups of tenants) to ensure that flows from different tenants have forward progress independent of other tenants despite competing for the same ISL bandwidth. According to one embodiment, an "admission control" mechanism is implemented that limits the maximum bandwidth that can be consumed by one tenant at the expense of other tenants.

[0388] According to one embodiment, for a physical fabric configuration, one consideration is that the bisection bandwidth be as high as possible, ideally non-blocking or even over-provisioned. However, even with non-blocking bisection bandwidth, there may be scenarios where it is difficult to achieve the desired SLAs for one or more tenants from the current allocation of servers to those different tenants. In such situations, the best approach would generally be to re-provision at least some of the servers for the different tenants to reduce the need for separate ISLs and bisection bandwidth.

[0389] According to one embodiment, some multi-rack systems have a blocking fat-tree topology, the assumption behind which is that workloads will be provisioned such that the associated communication servers are located to a large extent within the same rack, meaning that a significant portion of bandwidth usage is only between ports on the local leaf switches. Also, in traditional workloads, the majority of data traffic is from one fixed set of nodes to another fixed set of nodes. However, in next-generation servers with non-volatile memory and newer communication and storage middleware, according to one embodiment, communication workloads will become even more demanding and less predictable, as different servers may provide multiple functions simultaneously.

[0390] According to one embodiment, a goal is to provide per-tenant provisioning granularity at the VM level as opposed to the physical server level, and the goal is to be able to deploy up to tens of VMs on the same physical server, where different sets of VMs on the same physical server may belong to different tenants, and where various tenants may each represent multiple workloads with different characteristics.

[0391] Additionally, according to embodiments, while current fabric deployments have used different Type of Service (TOS) associations to provide basic QOS (traffic isolation) for different flow types (e.g., to prevent lock messages from being "stalled" after large bulk data transfers), it is desirable to also provide communication SLAs for different tenants. These SLAs are assumed to ensure that tenants experience workload throughput and response times that conform to expectations, even if the workloads are provisioned on physical infrastructure shared by other tenants. The relevant SLAs for a tenant shall be met regardless of concurrent activity by workloads belonging to other tenants.

[0392] According to one embodiment, a workload may be provisioned with a fixed (minimum) set of CPU cores / threads and physical memory on a fixed (minimum and / or maximum) set of physical servers, but provisioning fixed / guaranteed networking resources is generally not so straightforward, insofar as deployment implies sharing of HCAs / NICs in the servers. Sharing of HCAs also inherently implies that at least the ingress and egress links to the fabric are shared by different tenants. Thus, while different CPU cores / threads can operate truly in parallel, there is no way to divide the capacity of a single fabric link except through some kind of bandwidth multiplexing or “time sharing.” This basic bandwidth sharing may or may not be combined with the use of different “QOS IDs” (e.g., service level, priority, DSCP, traffic class, etc.) that are taken into account when implementing buffer selection / allocation and bandwidth arbitration within the fabric.

[0393] According to one embodiment, the overall server memory bandwidth should be very high compared to the typical memory bandwidth needs of any individual CPU core to prevent memory-intensive workloads on some cores from imposing latencies on other cores. Similarly, in the ideal case, the available fabric bandwidth for a physical server should be large enough to allow each tenant sharing the server to have sufficient bandwidth for the communication activity generated by their associated workloads. However, when several workloads all attempt to perform large data transfers, it is very likely that multiple tenants may utilize the entire link bandwidth—even 100 Gb / s or more. To address this scenario, provisioning multiple tenants on the same physical server requires each to allocate at least a given minimum percentage of the available bandwidth. It is required that this be done in a way that ensures that tenants are guaranteed to get what they need. However, in RDMA-based communications, the ability to enforce limits on how much bandwidth a tenant can generate in the egress direction does not mean that ingress bandwidth can be limited in the same way. That is, multiple remote communication peers may all send data to the same destination in a way that completely overloads the receiver, even though each sender is limited by a maximum send bandwidth. Also, RDMA read operations may originate from a local tenant with very little egress bandwidth. This can potentially result in catastrophic ingress bandwidth when bulk RDMA read operations originate for multiple remote peers. Therefore, it is not sufficient to impose a maximum limit on egress bandwidth to limit the total fabric bandwidth used by a single tenant on a single server.

[0394] According to an embodiment, the present system and method can configure an average bandwidth limit for a tenant that will ensure that the tenant will never exceed its relative portion of the associated link bandwidth in either the ingress or egress direction, regardless of its use of RDMA read operations, regardless of the number of remote peers with active data traffic, and regardless of the bandwidth limits of the remote peers. (Methods for achieving this are discussed in the "Long-Term Goals" section below.) According to one embodiment, unless the present system and method can enforce all aspects of communication bandwidth restriction, the highest level of communication SLA for a tenant can only be achieved by restricting the tenant from sharing a physical server with other tenants, or potentially restricting the tenant from sharing a physical HCA with other tenants (i.e., in the case of a server with multiple physical HCAs). If the physical HCAs can operate in active-active mode with full link bandwidth utilization for both HCA ports, it is conceivable to use restrictions that give a given tenant exclusive access to one of the HCA ports in the normal case. Nevertheless, due to HA constraints, failure of a complete HCA (in the case of multiple HCAs per server) or a single HCA port may mean reconfiguration and sharing that no longer guarantees the expected communication SLA for a given tenant.

[0395] According to one embodiment, in addition to constraints on overall bandwidth utilization for a single link, the ability of each tenant to achieve QOS between different communication flows or flow types depends on it not experiencing severe congestion contention for fabric-level buffer resources or arbitration due to communication activity by other tenants. In particular, this means that if a tenant uses a particular “QOS ID” to achieve low-latency messaging, that tenant should not find itself “competing” with high-volume data traffic from another tenant, depending on how the other tenant is using the “QOS ID” and / or how the fabric implementation enforces the use of the “QOS ID” and / or how this maps to buffer allocation and / or packet bandwidth arbitration within the fabric. Thus, if tenant communication SLAs mean that intra-tenant QOS assumptions cannot be met without relying on other tenants sharing the same fabric link “behaving well,” this may mandate that tenants must be provisioned without HCA (or HCA port) sharing with other tenants.

[0396] According to one embodiment, for both the basic bandwidth allocation and QOS issues discussed above, sharing constraints are applied to fabric internal links and server-local HCA port links. Thus, depending on the nature and stringency of the communication SLA for a given tenant, VM deployment for a tenant may involve sharing of physical servers and / or HCAs as well as fabric To prevent ISL sharing, both routing restrictions and restrictions on where VMs can be provisioned relative to each other within the private fabric topology may be applied.

[0397] Topology, routing and blocking scenario considerations: According to one embodiment, as described above, to guarantee that a tenant can achieve the expected communication performance between a set of communicating VMs without any dependency on the operation of VMs belonging to other tenants, no HCA / HCA ports or any fabric ISLs can be shared with other tenants. Therefore, the highest SLA class offered would typically have this as an implicit implementation. This is, in principle, the same scheme as the current provisioning model for many traditional systems in the cloud. However, with shared leaf switches, this SLA would require a guarantee of no ISL sharing with other tenants. Also, for tenants to achieve the best possible balance between their flows and the utilization of available fabric resources, they would need to be able to “optimize non-blocking” in an explicit manner (i.e., the communications SW infrastructure must provide tenants with a way to ensure that communications occur in a way that different flows do not compete for the same link bandwidth). This would include a way to ensure that communications that can occur through a single leaf switch are actually realized in this manner. Furthermore, if communications must involve ISLs, it should be possible to balance traffic across the available ISLs to maximize throughput.

[0398] According to one embodiment, from a single HCA port, as long as the maximum available bandwidth is the same for all links in the fabric, there is no point in attempting to balance traffic across multiple ISLs. From this perspective, it would make sense to use a dedicated "next-hop" ISL per sending HCA port, as long as the available ISLs represent a non-blocking subtopology for the sender. However, unless the associated ISL represents only connectivity between two leaf switches, a scheme with a dedicated next-hop ISL per sending port is not sustainable in practice because, at some point, multiple ISLs would have to be used if communication was with multiple remote peer HCA ports connected to different leaf switches.

[0399] According to one embodiment, in a non-blocking InfiniBand Fat Tree topology, popular routing algorithms use a "dedicated down path," which means that in a non-blocking topology, there are the same number of switch ports at each layer of the Fat Tree. This means that each end port can have a dedicated port chain from one root switch through each intermediate switch layer to the egress leaf switch port connecting the associated HCA port. Thus, all traffic targeting a single HCA port uses this dedicated down path, and there is no traffic to other destination ports (in the downward direction) on these links. However, in the upward direction, there may not be a dedicated path to each destination, and as a result, some links in the upward direction must be shared by traffic to different destinations. In the next round, this can lead to congestion when different flows to different destinations all try to utilize the full bandwidth on the shared intermediate links. Similarly, if multiple senders are simultaneously sending to the same destination, this can cause congestion on the dedicated down path, which can then quickly spread to other unrelated flows.

[0400] According to one embodiment, there is no risk of congestion between multiple tenants in a dedicated down path as long as a single destination port belongs to a single tenant. However, it is still a problem that different tenants may need to use the same link in the up direction to reach the root switch (or intermediate switch) that represents the dedicated down path. By dedicating many different root switches to specific tenants, the present system and method will reduce the need for different tenants to share paths in the upward direction. However, from a single leaf switch, this scheme may reduce the number of available uplinks toward the associated root switch. Therefore, to maintain non-blocking bisection bandwidth between servers (or rather HCA ports) belonging to the same tenant, the number of servers allocated to a single tenant on a particular leaf switch (i.e., within a single rack) will need to be less than or equal to the number of uplinks toward the root switch used by that tenant. On the other hand, to maximize the ability to communicate through a single crossbar, it makes sense to allocate as many servers as possible to the same tenant in the same rack.

[0401] According to one embodiment, this inherently represents a conflict between the availability of guaranteed bandwidth within a single leaf switch and the availability of guaranteed bandwidth toward communicating peers in different racks. To address this dilemma, perhaps the best approach is to use a scheme in which tenant VMs are grouped based on which leaf switch (i.e., leaf switch pair) they are directly connected to, and then there must be an attribute that defines the available bandwidth between such groups. However, again, there is a trade-off between being able to maximize bandwidth between two such groups (e.g., between the same tenant in two racks) and being able to guarantee bandwidth to multiple remote groups. Furthermore, in the special case of only two tiers of switches (i.e., leaf tiers interconnected by a single spine tier), a non-blocking topology means that it is always possible for N HCA ports to have N dedicated uplinks between leaf switches and N spine ports belonging to the same tenant. Thus, as long as these N spine ports represent spines that "own" all dedicated downpaths for all associated remote peer ports, the configuration is non-blocking for that tenant. However, if the associated remote peers represent dedicated downlink paths from more than N spine switches, or if the N uplinks are not distributed among all associated spine switches, the present system and method may result in contention for other tenants.

[0402] According to one embodiment, between VMs of a single tenant, regardless of non-blocking or blocking connectivity, there is still a possibility of contention between flows from different sources connected to the same leaf switch. That is, if a destination has a dedicated downlink path from the same spine and the number of uplinks from the source leaf switch to that spine is less than the number of such concurrent flows, there is no way to avoid any kind of blocking / congestion on the uplinks as long as all senders operate at full link speed. In this case, the only option to preserve bandwidth would be to use a secondary path to one of the destinations through a different spine. This would then represent a potential conflict with another dedicated downlink path, since a standard non-blocking fat-tree can only have one dedicated downlink per end port.

[0403] According to one embodiment, in some conventional systems, there can be a blocking factor of three between leaf switches and spine switches. Thus, in multi-rack scenarios where workloads are distributed in a manner that means more than one-third of the communication traffic is between racks rather than within a rack, the resulting bisection bandwidth becomes blocking. For example, the most common scenario with an even distribution of traffic between any pair of nodes in an eight-rack system means that seven-eighths of the communication is between racks, making the blocking effect substantial.

[0404] According to one embodiment, the cable costs of overprovisioning are reduced in the system. If acceptable (i.e., given a fixed switch unit cost), additional links can be used both to provide "backup" downlinks to each leaf switch and to provide spare uplink capacity from each leaf to each spine—both cases representing uneven distribution of traffic and thus offering at least some potential improvement over dynamic workload distribution that cannot take advantage of an inherently non-blocking topology in the first place.

[0405] According to certain embodiments, higher-radix full crossbar switches also have the potential to increase the size of each single "leaf domain" and reduce the number of spine switches required for a given system size. For example, with a 128-port switch, two full racks with 32 servers could be included in a single full crossbar leaf domain and still provide non-blocking uplink connectivity. Similarly, to provide non-blocking connectivity between 16 racks (512 servers, 1024 HCA ports), only eight spines would be required. Thus, there are still only eight links from each leaf to each spine (i.e., in the case of a single, fully connected network). In the extreme case where all HCA ports on one leaf send to a single remote leaf through a single spine, this still implies a blocking factor of 8. On the other hand, given the even distribution of dedicated downlink paths for each leaf switch among all spines, the likelihood of such an extreme scenario should be negligible.

[0406] According to one embodiment, in the case of dual independent networks / rails, each leaf switch in a redundant leaf switch pair belongs to a single rail with a dedicated spine. The same eight spines are divided into two groups of four spines (one for each rail), so in this case, each leaf in a rail would only need to connect to four spines. Therefore, in this case, the worst-case blocking factor would be only four. On the other hand, in this scenario, the selection of the rail for each communication operation becomes even more important to provide load balancing across both rails.

[0407] Dynamic vs. Static Packet Route Selection / Forwarding + Multipathing: According to one embodiment, standard InfiniBand uses static routes per destination address, but there are several standard, proprietary schemes for dynamic route selection in Ethernet switches. For InfiniBand, there are also various proprietary schemes for "adaptive routing" (some of which may be standardized).

[0408] According to one embodiment, one advantage of dynamic route selection is that it increases the probability of optimally utilizing the associated bisection bandwidth within the fabric, thereby increasing overall throughput. However, a potential disadvantage is that ordering may be violated and congestion in one region of the fabric may more easily spread to other regions (i.e., in a manner that would have been avoided if static route selection had been used).

[0409] According to one embodiment, "dynamic routing" or "dynamic route selection" are typically used in reference to forwarding decisions made within and between switches, while "multipathing" is a term used when traffic to a single destination can be spread across multiple paths based on explicit addressing from the sender. Such multipathing may involve "striping" the transmission of a single message across multiple local HCA ports (i.e., the complete message is split into multiple sub-messages, each representing an individual forwarding operation), so that different forwarding operations to the same destination can be performed across the fabric. This can mean that the network is set up to dynamically use different routes along which it can travel.

[0410] According to one embodiment, in the general case, if all transfers from all sources targeting a destination outside the local leaf domain are split into smaller(er) chunks and then distributed across all possible paths / routes towards that destination, the system will achieve optimal utilization of the available bisection bandwidth and also maximize "leaf-to-leaf throughput." Yet, this is only true as long as the communication workload is also evenly distributed across all possible destinations. Otherwise, the effect will be that any congestion towards a single destination will immediately affect all simultaneous flows.

[0411] According to one embodiment, the implications of congestion on dynamic route selection and multipathing are that it makes sense to restrict traffic to a single destination to use only a single path / route, unless that path / route falls victim to congestion on other targets or on any intermediate links. In a two-tier fat-tree topology with a dedicated down path, this means that the only possible congestion not related to end ports will be on uplinks targeting the same spine switch. This means that it would make sense to treat all uplinks to the same spine as a group of ports sharing the same static route, except that the individual port used for a particular target would be dynamically selected. Alternatively, the individual port could be selected based on tenant association.

[0412] According to an embodiment, selecting uplink ports within such groups using tenant associations may be based on fixed associations or on a scheme where different tenants have a "first priority" to use certain ports but also the ability to use other ports. In that case, the ability to use another port would depend on it not competing with the "first priority" traffic of other ports. In this way, tenants would be able to use all of the associated bisection bandwidth unless there was contention, but if there was contention, there would be a guaranteed minimum bandwidth. This minimum guaranteed bandwidth could then reflect the total bandwidth for a single or several links, or a percentage of the bandwidth of one or more links.

[0413] According to an embodiment, in principle, the same dynamic scheme can also be used on the downward path from the spine to a specific leaf. On the one hand, this would increase the risk of congestion due to sharing a downlink between flows targeting different end ports, but on the other hand, it may provide a way to utilize additional alternative paths between two sets of nodes connected to two different leaf switches, but still prevent congestion from spreading between different tenants.

[0414] According to one embodiment, in a scenario where different dedicated down paths from the spine to the leaf already represent a particular tenant, it would be relatively simple to have a scheme that allows these links, on the associated leaf switch, to be used as "spare" for traffic (belonging to the same tenant) to end ports that have a (primary) dedicated down path from another spine.

[0415] According to one embodiment, one possible model would be for the switch to handle dynamic route selection between parallel ISLs connecting a single spine or leaf switch, but have a host-level decision to use explicit multipathing via spines that do not represent a (primary) dedicated down path to the associated target.

[0416] Bandwidth admission control per pel tenant: According to one embodiment, when a single HCA is used only by a single tenant, the present system and method can limit the bandwidth that can be generated from the HCA port, particularly if there is a limited bisection bandwidth for that tenant for traffic destined for a remote leaf switch.

[0417] According to an embodiment, one aspect of such bandwidth limitation is to ensure that the limitation is applied only to targets that are affected by the limited bisection bandwidth. In principle, this would involve a scheme where different target groups are associated with specific bandwidth allocations (i.e., either strict maximum rates and / or average bandwidths over a certain amount of transferred data).

[0418] According to one embodiment, such restrictions would, by definition, have to be implemented at the HCA level. Also, such restrictions would map more or less directly to a virtualized HCA scenario where VMs belonging to different tenants share an HCA through different virtual functions. In this case, the various "shared bandwidth allocation groups" introduced above would require an additional dimension in that they are associated with groups of one or more VFs, not just complete physical HCA ports.

[0419] Per-tenant bandwidth reservation on ISLs: According to one embodiment, as described above, it may make sense to reserve some guaranteed bandwidth across one or more ISLs for a tenant (or group of tenants). In some scenarios, links can be reserved for tenants by restricting which tenants are truly allowed to use the complete link. However, to have a more flexible and finer-grained scheme, an alternative approach is to use a switch arbitration mechanism to ensure that (some) ingress ports are allowed to use up to X% of one or more egress ports' bandwidth, regardless of what other ingress ports are competing for bandwidth on those egress ports.

[0420] In this way, according to one embodiment, every ingress port can use up to 100% of the bandwidth of its associated egress port, but only as long as this does not conflict with any traffic from a preferred ingress port.

[0421] According to one embodiment, in a scenario where different tenants "own" different ingress ports (e.g., leaf switch ports connecting HCA ports), this scheme would facilitate a flexible and fine-grained scheme for allocation of uplink bandwidth to one or more spine switches.

[0422] According to an embodiment, for downlink paths from spine to leaf switches, the usefulness of such a scheme will depend on the extent to which a scheme with strictly dedicated down paths is or is not used. If a strictly dedicated down path is used and the target end ports represent a single tenant, by default there is no potential conflict between different tenants trying to use the downlink. Therefore, in this case, access to the associated downlink should typically be set up using a round-robin arbitration scheme with equal access for all associated ingress ports.

[0423] According to an embodiment, an ingress port can represent traffic belonging to different tenants, so that packets belonging to one tenant can be sent to the associated tenant. It should never be a problem that packets can consume bandwidth on egress ports that they are not allowed to transmit on. In this case, the assumption is that strict access control (e.g., VLAN-based restrictions on various ports) is employed rather than an arbitration policy to prevent such packets from wasting bandwidth.

[0424] According to one embodiment, in a leaf switch, downlinks from the spine may be given more bandwidth towards various end ports compared to other local end ports because a downlink can in principle represent multiple sending HCA ports, while a local end port only represents a single HCA port. If this were not the case, it would be possible to have a scenario where several remote servers share a single downpath to a target leaf switch, but in the next round, if N-1 HCA ports directly connected to a leaf switch also try to send to the same local target port, they would end up sharing 1 / N of the bandwidth towards a single destination on that leaf switch.

[0425] According to one embodiment, when virtualized HCAs represent different tenants, the problem of reserving bandwidth within fabric ISLs (i.e., across various ISLs) can become significantly more complex. For the ingress / uplink path, one simplified approach is that it is up to the HCA to provide bandwidth arbitration between different tenants, and then whatever is sent out on the HCA port will be handled by the ingress leaf switch according to the port-level arbitration policy. Thus, in this case, nothing changes from the leaf switch's perspective.

[0426] According to one embodiment, in the downlink path (spine to leaf and leaf ingress to end port), the situation is different, as arbitration decisions may depend not only on the port attempting to forward the packet but also on which tenant the various pending packets belong to. One possible solution is (again) to restrict some ISLs to only represent specific tenants (or groups of tenants) and then reflect this in the port-level arbitration scheme. Alternatively (or additionally), different priority or QOS IDs could be used to represent different tenants, as outlined below. Finally, having a "tenant ID" such as a VLAN ID or partition ID or any related access control header field used as part of the arbitration logic would facilitate the necessary level of granularity for arbitration. However, this could significantly increase the complexity of the arbitration logic, which already has significant "time and space" complexity in the switch. Also, because such a scheme involves an overload of information that may already play a role in the end-to-end wire protocol, it is important that such extra complexity does not conflict with any existing uses or assumptions regarding such header field values.

[0427] Different priorities, flow types and QOS IDs / classes across ISLs and end port links: According to one embodiment, for different flow types to travel simultaneously on the same link, it is important that they do not compete for the same packet buffers in the switch and HCA. Also, to distinguish relative priority between different flow types, the arbitration logic that decides which packet to send next on the various switch egress ports must take into account what packet type queues have to send on which egress ports. The outcome of the arbitration should be that all active flows are making forward progress according to their relative priority and according to how much the flow control conditions (if any) for the associated flow types on the associated downstream ports currently allow for transmitting any packets.

[0428] According to one embodiment, in principle, different QOS IDs can be used to isolate traffic flows from different tenants from each other, even if the different tenants use the same link. However, the scalability of this approach is very limited because the number of packet queues and independent buffer pools that can be supported for each port is typically limited to less than 10. Furthermore, if a single tenant wants to use different QOS IDs to isolate different flow types from each other, scalability is further reduced.

[0429] According to one embodiment, by logically combining multiple ISLs between a single pair of switches as described above, the present system and method can restrict some links to some tenants and then ensure that different tenants can use different QOS IDs on different ISLs independently from each other. However, again, this imposes a limit on the total bandwidth available to any single tenant if the independence of other tenants is guaranteed 100%.

[0430] According to one embodiment, in the ideal case, HCA ingress (receive) packet processing can always occur at a rate higher than the associated link speed, regardless of which transport-level operation the incoming packet represents. This means there is no need for flow control of different flow types on its last link, i.e., the egress port on the leaf switch connecting to the HCA port. However, the scheduling of different packets from different queues in the leaf switch must still reflect the relevant policies regarding priority, fairness, and forward progress. For example, if one small high-priority packet targets an end port, and N ports also attempt to send maximum MTU-sized "bulk forwarding packets" to the same target port, the high-priority packet should be scheduled before any of the other ports.

[0431] According to one embodiment, on the egress path, the sending HCA can schedule and label packets in many different ways. In particular, the use of an overlay protocol as a "bump in the wire" between the VM+virtual HCA and the physical fabric would allow the encoding of fabric-specific information that the switch may be concerned with without perturbing any aspect of the end-to-end protocol between tenant virtual HCA instances.

[0432] According to one embodiment, the switch may provide more buffering and internal queuing than current wire protocols allow. In this way, it may be possible to set up buffering, queuing, and arbitration policies that take into account that a link may be shared by traffic representing multiple tenants with different SLAs and that use different QOS classes for different flow types.

[0433] In this way, different high priority tenants may also have more private packet buffer capacity within the switch, according to one embodiment.

[0434] Lossless vs. lossy packet transfer: According to one embodiment, high performance RDMA traffic depends heavily on individual packets not being lost due to lack of buffer capacity in the switch, and also on packets arriving in the correct order for each RDMA connection. As a general rule, the higher the potential bandwidth, the more important these aspects are to achieving optimal performance.

[0435] According to one embodiment, lossless operation requires explicit flow control, and very high bandwidth implies a trade-off between buffer capacity, MTU size, and flow control update frequency.

[0436] According to an embodiment, a drawback of lossless operation is that if the total bandwidth generated is higher than the downstream / receive capacity, it will lead to congestion, which will then (possibly) spread and end up slowing down all flows competing for the same buffer somewhere in the fabric.

[0437] According to one embodiment, as mentioned above, the ability to provide flow separation based on independent buffer pools is a key scalability issue for a switch implementation, which depends on both the number of ports, the number of different QOS classes, and potentially (as introduced above) also the number of different tenants.

[0438] According to one embodiment, an alternative approach may be to make truly lossless operation (i.e., lossless based on guaranteed buffer capacity) a "premium SLA" attribute, thereby restricting this feature only to tenants that have purchased such a premium SLA.

[0439] According to one embodiment, the key issue here is being able to "oversubscribe" available buffer capacity - the same buffers can be used for both lossy and lossless flows, but buffers allocated to lossy flows can be preempted whenever a packet from a lossless flow arrives and needs to use a buffer from the same pool. A very minimal set of buffers can be provided to allow lossy flows to proceed in the forward direction, but at a (much) lower bandwidth than what can be achieved with optimal buffer allocation.

[0440] According to one embodiment, it is also possible to introduce hybrid lossless / lossy flow classes that differ in terms of the maximum time they can occupy a buffer before it has to be evicted (when needed) and given to a more premium SLA type flow class - this would work best in the context of a fabric implementation with link level credits, but could potentially be adapted to work with xon / xoff type flow control (i.e., the Ethernet pause-based flow control scheme used for RoCE / RDMA).

[0441] Strict packet ordering vs. relaxed packet ordering: According to one embodiment, strict ordering and lossless packet forwarding within the fabric allows an HCA implementation to provide reliable connectivity and RDMA with minimal state overhead at the transport level. However, to better tolerate some amount of out-of-order packet delivery due to occasional route changes (due to adaptive / dynamic forwarding decisions within the fabric), as well as to minimize the overhead and delays associated with lost packets due to lossy or "hybrid lossless / lossy" mode forwarding within the fabric, an efficient transport implementation would need to keep sufficient state to allow many individual packets (sequence numbers) to arrive out of order and be individually retried while other packets with later sequence numbers are accepted and acknowledged.

[0442] According to one embodiment, the key point here is to avoid long delays and loss of average bandwidth when a lost or out-of-order packet causes a retry in the current default transport implementation. Also, subsequent packets in the sequence of packets to be delivered are not delivered. By avoiding dropped packets, the present system and method also significantly reduces wasted fabric bandwidth that could otherwise be consumed by other flows.

[0443] Shared Services and Shared HCAs: According to one embodiment, a shared service on the fabric (e.g., a backup device) that is used by multiple tenants means that some end port links will be shared by different tenants unless the service can provide end ports that can be dedicated to a specific tenant (or a restricted group of tenants). A similar scenario exists when VMs belonging to multiple tenants share the same server and the same HCA port.

[0444] According to one embodiment, different tenants can be allocated fine-tuned server and HCA resources, and it can also be ensured that outgoing data traffic bandwidth from the HCA is divided fairly among the different tenants according to their associated SLA levels.

[0445] According to an embodiment, it may also be possible to set packet buffer allocation and queuing priority and arbitration policies within the fabric that reflect the relative importance, and therefore fairness, between data traffic belonging to different tenants. However, even with very fine-tuned buffer allocation and arbitration policies within the fabric, the granularity may not be fine enough to ensure that the relative priorities and bandwidth allocations for different tenants are accurately reflected in terms of ingress bandwidth to a shared HCA port.

[0446] According to one embodiment, to achieve such fine-grained bandwidth allocation, a dynamic end-to-end flow control scheme is needed that can effectively divide and schedule available ingress bandwidth among several telecommunication peers belonging to one or more tenants.

[0447] According to one embodiment, the goal of such a scheme would be to allow a set of associated active remote clients at any given time to utilize their fair (not necessarily equal) share of the available ingress bandwidth. This bandwidth utilization should also occur without causing congestion in the fabric due to attempts to utilize excess bandwidth at end ports. (Even so, fabric-level congestion may still occur due to overload on shared links in the rest of the fabric.) According to one embodiment, a high-level model for achieving this goal would be that the receiver could dynamically allocate and update available bandwidth for a set of associated remote clients. The current bandwidth value for each remote client would need to be calculated based on what is currently offered to each client and what is next required.

[0448] According to one embodiment, this means that if a single client is currently allowed to use all available bandwidth and another client also needs to use ingress bandwidth, an update instruction must be delivered to the currently active client informing it about the new reduced maximum bandwidth, and the new client must be delivered an instruction that it can use the maximum bandwidth corresponding to the reduction for the current client.

[0449] According to one embodiment, in principle the same scheme would apply to an "arbitrary" number of simultaneous clients. However, there is of course a huge trade-off between being able to guarantee that the available bandwidth is never "oversubscribed" at any one time and ensuring that the available bandwidth is always fully utilized when needed.

[0450] According to an embodiment, a further challenge with this type of scheme is to ensure that it interoperates well with dynamic congestion control, and also to ensure that congestion related to a shared path for multiple targets is handled in a coordinated manner within each sender.

[0451] High availability and failover: According to one embodiment, in addition to performance, a key attribute of a private fabric may be redundancy and the ability to fail over communications following any single point of failure without loss of service for any client applications. Furthermore, while "loss of service" represents a binary condition (i.e., service is present or lost), some equally important, but more scalar, attributes are the extent to which there is a power outage during failover, and if so, how long it is. Another important aspect is the extent to which expected performance is provided (or re-established) during and after completion of the failover operation.

[0452] According to one embodiment, from the perspective of a single node (server), the goal is that any single point of failure in the fabric communication infrastructure outside the server itself (i.e., including a single local HCA) should not mean that the node will be unable to communicate. However, from the perspective of the complete fabric, there is also the question of what level of fabric-wide throughput and performance impact the loss of one or more components means. For example, for a topology size that can operate with only two spine switches - is it then acceptable to lose 50% of leaf-to-leaf communication capacity if one of the spines is out of service, in terms of increased bisection bandwidth and risk of congestion? According to one embodiment, another issue for per-tenant SLAs is the extent to which the ability to reserve and / or prioritize fabric resources for tenants with premium SLAs should be reflected in that such tenants will get a proportionally larger share of the remaining available resources following a failure and subsequent failover operation - i.e., the impact of a failure will thus be less for premium SLA tenants, but at the expense of a greater impact for other tenants.

[0453] According to one embodiment, it may also be a "super premium SLA attribute" that initial resource provisioning for such tenants with regard to redundancy would ensure that no single point of failure would mean that the associated performance / QOS SLAs could not be met either during or after a failure. However, the fundamental problem with such over-provisioning is that there must be extremely fast failover (and failback / rebalancing) to ensure that available resources are always utilized in the most optimal way and that communications are never stopped for more than a very short period of time as a result of any single point of failure.

[0454] According to one embodiment, an example of such a "super premium" setup may be a system with dual HCA-based servers, where both HCAs are operating in an active-active fashion, and both HCA ports are also configured to provide a delay before an alternate path is tried. It is also used in active-active systems with a very short APM (Automatic Path Migration) scheme.

[0455] Route Selection: According to one embodiment, when multiple possible paths exist between two endpoints, the selection of the best or "correct" path for an associated RDMA connection should ideally be automatic so that the communications workload experiences the best possible performance within the constraints of the associated SLA, and so that system-level fabric resources are utilized in the most optimal manner.

[0456] In accordance with an embodiment, in the ideal case, this would mean that application logic within a VM does not need to deal with which local HCA and which local HCA port can or should be used for which communications. This also means a node-level rather than port-level addressing scheme, and that the underlying fabric infrastructure is used transparently to the application.

[0457] In this way, according to an embodiment, related workloads may be more easily deployed on different infrastructures without requiring explicit handling of different system types or system configurations.

[0458] Features: According to one embodiment, features in this category are expected to be supported by existing HCA and / or switch hardware using current firmware and software.

[0459] According to one embodiment, the main goals in this category are: Ability to limit the total egress bandwidth generated by a single VM or VF belonging to a tenant per local physical HCA instance. The ability to guarantee that a VF belonging to a single VM or tenant on a local physical HCA will be able to utilize at least a minimum percentage of the available local link bandwidth. Ability to restrict which network (Enet) priorities a VF can use. ○ This may mean that multiple VFs must be allocated in order for a single VM to use multiple priorities (i.e., unless priority restriction is enabled and a single VF is only allowed to use a single priority). The ability to restrict which ISLs can be used by a group of flows belonging to a single tenant or group of tenants.

[0460] According to one embodiment, to control HCA usage by a tenant that shares a physical HCA with other tenants, an "HCA Resource Limit Group" (herein referred to as "HRLG") is established for that tenant. HRLGs can be configured with a maximum bandwidth that defines the actual data rate that can be generated by the HRLG, and can also be configured with a minimum bandwidth share that ensures that the HRLG achieves at least a specified percentage of the HCA bandwidth when there is contention with other tenants / HRLGs. As long as there is no contention with other HRLGs, VFs within the HRLG can persistently use up to the specified rate (or link capacity if no rate limit is defined).

[0461] According to one embodiment, an HRLG may contain up to the number of VFs that an HCA instance can support. Within an HRLG, each VF receives a fair share of the "quota" assigned to the HRLG. For each VF, the associated QP will also get a fair share of access to the local link as a function of the available HRLG allocation and any current flow control limits on the QR (i.e., if a QP receives congestion control feedback commanding it to throttle itself, or if it currently has no "credit" to transmit at the associated priority, the QP will not be considered for local link access). According to one embodiment, within the HRLG, it is conceivable to enforce a restriction on the priority that a VF can use. To the extent that this restriction can only be defined in terms of a single priority allowed for a VF, the implication is that a VM that is supposed to use multiple priorities (while still being limited to only a few priorities) will have to use multiple VFs—one for each required priority. (Note: The use of multiple VFs means that sharing local memory resources between multiple QPs using different priorities is likely to present a problem, since it means that different VFs will have to be allocated and used by ULPs / applications within the VM depending on which priority restriction / enforcement policy is defined.) According to one embodiment, within a single HRLG, there is no difference in bandwidth allocation depending on which priority a VF / QP is currently using - they all share the associated allocation in a fair / equal manner. Therefore, to associate different bandwidth allocations with different priorities, it is necessary to define one or more dedicated HRLGs that will contain only VFs that are restricted to using the priority to be associated with the shared allocation represented by the associated HRLG. In this way, a VM or tenants with multiple VMs sharing the same physical HCA can be given different bandwidth allocations for different priorities.

[0462] According to one embodiment, current hardware priority limits prevent data attempted to be sent at an illegal priority from being sent to the external link, but do not prevent the fetching of associated data from local memory. Thus, if the local memory bandwidth that the HCA can maintain in the egress direction is approximately the same as the available external link bandwidth, there is still overall wasted HCA link bandwidth. However, if the associated memory bandwidth is (significantly) greater than the external link bandwidth, attempts to use illegal priorities will waste less external link bandwidth as long as the HCA pipeline operates at optimal efficiency. Nevertheless, unless there is much savings in terms of external link bandwidth, a possible alternative scheme for preventing the use of illegal priorities may be to leverage ACL rule enforcement at the switch ingress port. If the associated tenants can be effectively identified without any possibility of spoofing, this may be used to enforce tenant / priority association without the need to allocate individual VFs for each priority for the same VM. However, both ensuring that packet / tenant associations are always well-defined and cannot be spoofed from the sending VM, and dynamically updating the associated switch ports to perform the associated enforcement whenever a VF is set up for use by a VM / tenant, represent nontrivial complexities. One possible scheme would be to use a per-VF port MAC to represent an unspoofable ID that can be associated with a VM / tenant. However, if VxLAN or other overlay protocols are used, this is not straightforward—especially unless the external switch is assumed to be involved in (or aware of) the overlay scheme being used.

[0463] According to one embodiment, to restrict which flows can use which ISLs, the switch forwarding logic must have policies to identify related flows and configure forwarding accordingly. One example is to use VLAN IDs to represent flow groups. If different tenants map to different VLAN IDs on the fabric, one possible scheme is to define which VLAN IDs are used by any LAG or other port. A switch could dynamically achieve LAG type balancing based on what is allowed for various ports in the port grouping. Another option would involve explicit forwarding of packets based on a combination of destination address and VLAN ID.

[0464] According to one embodiment, if a VxLAN-based overlay is used transparently to the physical switch fabric, it may be possible to map different overlays to different VLAN IDs, allowing the switch to map VLAN IDs to ISLs as outlined above.

[0465] According to one embodiment, another possible scheme is for forwarding of individual endpoint addresses to be set up according to a routing scheme that takes into account VLAN membership or some other notion of "tenant" association. However, to the extent that the same endpoint address value is allowed in different VLANs, the VLAN ID needs to be part of the forwarding decision.

[0466] According to an embodiment, distribution of per-tenant flows to either shared or exclusive ISLs may require a holistic routing scheme to distribute traffic in a globally optimized manner within the fabric (fat-tree) topology. Implementation of such a scheme would typically rely on SDN-type management interfaces for the switches, but implementing holistic routing would not be trivial.

[0467] Short and medium term SLA classes: According to one embodiment, the following assumes that a non-blocking two-tier fat-tree topology is used for system sizes (physical node counts) that exceed the cardinality of a single leaf switch. It also assumes that a single VM on a physical server can use all of the fabric bandwidth (through one or more vHCAs / VFs). Therefore, the number of VMs per tenant per physical server is not a parameter that needs to be considered as a tenant-level SLA factor from an HCA / fabric perspective.

[0468] According to one embodiment, the top tier (e.g., Premium Plus) includes: · Dedicated servers can be used. · VMs can be allocated to the same leaf domain whenever possible, except when the number and size of the VMs (or HA policies) imply additional distance. If a tenant uses multiple leaf domains (i.e., relative to the number of servers allocated to this tenant within the same leaf domain), then on average, they will have non-blocking uplink bandwidth from the local leaf, but no dedicated uplink or uplink bandwidth. All "flow groups" (i.e. priorities representing different buffer pools and arbitration groups within the fabric) can be used.

[0469] According to one embodiment, a lower tier (e.g., premium) may: You can use a dedicated server, but there is no guarantee of the same leaf domain. On average, they will have at least 50% non-blocking uplink bandwidth (i.e., relative to the number of servers allocated to this tenant within the same leaf domain), but will have no dedicated uplink or uplink bandwidth. All "flow groups" can be used.

[0470] According to one embodiment, a third tier (e.g., Economy Plus) includes: · You can use a shared server, but you will have a dedicated "flow group." These resources will be dedicated to the local HCA and switch ports, but will be shared within the fabric. · It may have the ability to use all available egress bandwidth from the local server, but is guaranteed to have at least 50% of the total egress bandwidth. Only one Economy Plus tenant per physical server. On average, you will have at least 25% non-blocking leaf uplink bandwidth for the number of servers used by this Economy Plus tenant.

[0471] According to one embodiment, the fourth tier (e.g., economy) includes: ·Shared servers can be used No dedicated priority You can be granted up to 50% of the server egress bandwidth, which may be shared with up to three other economy tenants. On average, up to 25% of the available leaf uplink bandwidth can be shared with other economy tenants within the same leaf domain.

[0472] According to one embodiment, the lowest layer (e.g., standby) May use spare capacity without guaranteed bandwidth Longer term features: According to one embodiment, the main features discussed in this section are: · The ability to enforce priority restrictions per VF in a way that allows a single VF to be restricted to use any subset of the entire set of supported priorities. · Zero wasted memory or link bandwidth whenever a data transfer is attempted with a priority that is not permitted for the initiating VF. Ability to limit egress rates for different individual priorities across multiple HRLGs so that each VF in the various HRLGs gets its fair share of the associated HRLG total minimum bandwidth and / or maximum rate, but subject to the constraints defined by the allocations per the various associated priorities. Ability to perform egress bandwidth control and congestion adjustment on both a per-target and per-shared path / route basis, and this is aggregated at both the VM / VF (i.e. vHCA port) level and the HCA port level. Ability to limit aggregate average send and RDMA write ingress bandwidth for a VF based on receiver throttling of the cooperating remote sender. Ability to limit the average RDMA read ingress bandwidth for a VF without relying on a cooperating remote RDMA read responder. Ability to include RDMA reads in addition to sends and RDMA writes when limiting the aggregate average ingress bandwidth for a VF based on receiver throttling of the cooperating remote sender. Ability for tenant VMs to observe available bisection bandwidth to different groups of peer VMs. SDN features for routing control and arbitration policies within the fabric.

[0473] According to one embodiment, the HCA VF context can be extended to include a list of legal priorities (similar to the set of legal SLs for an IBTA IB vPort). Whenever a work request attempts to use a priority that is not legal for the VF, the work request should fail before any local or remote data transfer is initiated. In addition, priority mapping can also be used to give applications the illusion that any priority can be used. However, this type of mapping, in which multiple priorities can be mapped to the same value before a packet is sent, has the drawback that an application may no longer have control over its own QOS policy, in that it associates different flow types with different "QOS classes." Such a restricted mapping represents SLA attributes (i.e., a more privileged SLA means more actual priority after mapping). However, it is always important for an application to be able to decide which flow types to associate with which QOS classes (priorities), in a way that also represents independent flows in the fabric.

[0474] Achieving ingress RDMA read BW allocation via target groups and dynamic BW allocation updates: According to one embodiment, there is full control of all ingress bandwidth to a vHCA port insofar as the target group association for a flow from a "producer / source" node implies bandwidth policing of all outgoing data packets, including UD sends, RDMA writes, RDMA sends, and RDMA reads (i.e., RDMA read responses with data). This is regardless of whether the VM that owns the target vHCA port is generating an "excessive" amount of RDMA read requests to multiple peer nodes.

[0475] According to one embodiment, as discussed above, coupling target groups to both flow-specific and "unsolicited" BECN signaling means that ingress bandwidth per vHCA port can be dynamically throttled to any number of remote peers.

[0476] According to one embodiment, the "unsolicited BECN" messages outlined above can also be used to communicate specific rate values ​​in addition to pure CE flagging / unflaggling for different stage numbers. In this way, it is possible to have a scheme where an initial incoming packet (e.g., a CM packet) from a new peer can trigger the generation of one or more "unsolicited BECN" messages to both the HCA (i.e., the associated firmware / hyper-privileged software) from which the incoming packet came and to the current communicating peer.

[0477] According to one embodiment, in cases where both ports on an HCA are used simultaneously (i.e., an active-active scheme), it may make sense to share target groups between local HCA ports when concurrent flows may share some ISLs or even target the same destination port.

[0478] According to one embodiment, another reason for sharing target groups between HCA ports is if the HCA local memory bandwidth cannot sustain full link speed for both (all) HCA ports. In this case, the target groups can be configured so that the aggregate total link bandwidth never exceeds the local memory bandwidth, regardless of which ports are involved in either the source or destination HCA.

[0479] According to one embodiment, for a fixed route towards a particular destination, any intermediate target group will typically represent only a single ISL at a particular stage in the path. However, if dynamic forwarding is active, both the target group and ECN processing must take this into account. If the dynamic forwarding decision occurs solely to balance traffic between parallel ISLs between a pair of switches (e.g., uplinks from a single leaf switch to a single spine switch), all processing is, in principle, very similar to when only a single ISL is used. FECN notification is likely based on the state of all ports in the relevant group, and signaling can be "aggressive," in the sense that it is signaled based on congestion indications from any of the ports, or it can be more conservative and based on the size of the shared output queue for all ports in the group. Target group configuration typically involves only a single ISL within the group, as long as it allows forwarding for any packet to select the best output port at that time. would represent the aggregate bandwidth for all links. However, if there is a notion of strict packet ordering per flow, then the evaluation of bandwidth allocation is more complex, since several flows may "must" use the same ISL at some point. If such a flow ordering scheme is based on well-defined header fields, then it may be best to represent each port in the group as a separate target group. In this case, the selection of target groups at the source HCA must be able to evaluate the header fields that will be associated with the RC QP connection or address handle in the same way that the switch performs for all packets at runtime.

[0480] According to one embodiment, by default, the initial target group rate for a new remote target may be conservatively set low. In this way, there is inherent throttling until the target has an opportunity to update its associated rate. Thus, all such rate control is independent of the participating VM itself, although a VM could request the hypervisor to update the allocation of different remote peers for both ingress and egress traffic, but this would only be allowed within the aggregate constraints defined for both the local and remote vHCA ports.

[0481] Correlation of peer nodes, paths and target groups: According to an embodiment, a way would be needed for a VM to query which target groups are associated with various communicating peers (and associated address / route information) to be able to identify the bandwidth limits associated with different peer nodes and different groups of peer nodes. Based on correlating the set of communicating peers with the various target groups and the rate limits that the various target groups represent, the VM would be able to track what bandwidth it can achieve for the various communicating peers. This would then, in principle, allow the VM to schedule communication operations in a manner that achieves the best possible bandwidth utilization over time by having as many simultaneous transfers as possible without conflicting target groups.

[0482] The relationship between HCA resource limit groups and target groups: According to an embodiment, the HRLG concept and the target group concept overlap in some ways in that they both represent bandwidth limits that can be defined and shared among VMs and tenants in a flexible manner. However, the primary focus of the HRLG is to define how different portions of the local HCA / HCA port capacity can be allocated to different VFs (and thereby VMs and tenants), while the target group concept focuses on bandwidth limits and flow control constraints that exist outside the local HCA with respect to both the final destination and the intermediate fabric topology.

[0483] In this way, according to one embodiment, it makes sense to use HRLGs as a way to control what share of local HCA capacity various VFs can use, but ensure that the allowed capacity can only be used in a manner that does not conflict with any fabric or remote target limits or congestion conditions - these external constraints are then dynamically controlled and reflected via the associated target groups.

[0484] According to one embodiment, from an implementation point of view, the state of all relevant target groups will define which pending work requests for which local QPs will be in a flow control state, where they are allowed to generate more egress data traffic at any given time. This state determines what the QPs are actually sending. This can then be aggregated at the VF / vHCA port level as to which VF is a candidate for transmitting next, along with the state regarding which VF has something to transmit. The decision as to which VF to schedule for transmission next on the HCA port will be based on the state and policies of the various HRLGs in the HRLG hierarchy, the set of "ready to transmit" VFs, and the recent history of which VFs have generated what egress traffic. For a selected VF, a VF-specific arbitration policy will define which QP is selected for data forwarding.

[0485] According to one embodiment, the set of QPs with pending data transfers includes both QPs with local work requests and QPs with pending RDMA read requests from associated remote peers, so that the above scheduling and arbitration will address all pending egress data traffic.

[0486] According to one embodiment, ingress traffic (including incoming RDMA read responses) will be controlled by the current state of all associated target groups at the remote peer node. This (remote) state includes both dynamic flow control state based on congestion conditions and explicit updates from this HCA that reflect changes in ingress bandwidth allocations for local VFs on this HCA. Such ingress bandwidth allocations will be based on policies reflected by the HRLG hierarchy. In this way, various VMs may have "fine-tuned" independent bandwidth allocations for both ingress and egress, and even on a per-priority basis for both ingress and egress.

[0487] SLA Class: According to one embodiment, the following proposal assumes that a non-blocking two-tier fat-tree topology is used for system sizes (physical node counts) that exceed the cardinality of a single leaf switch. It is also assumed that a single VM on a physical server can use all of the fabric bandwidth (through one or more HCAs / VFs). Therefore, the number of VMs per tenant per physical server is not a parameter that needs to be considered as an SLA factor.

[0488] According to one embodiment, the highest level tier (e.g., Premium Plus) includes: Only dedicated servers are available. · VMs can be allocated to the same leaf domain whenever possible, except when the number and size of the VMs (or HA policies) imply additional distance. If a tenant uses multiple leaf domains, it will always have non-blocking uplink bandwidth from the local leaf. All "flow groups" can be used.

[0489] According to some embodiments, lower level tiers (e.g., premium) may be offered. Only dedicated servers can be used and there is no guarantee of the same leaf domain. Can be guaranteed at least 50% of non-blocking uplink bandwidth (i.e., relative to the number of servers allocated to this tenant within the same leaf domain). All "flow groups" can be used.

[0490] According to some embodiments, a third level tier may be provided (e.g., Economy Plus). · Shared servers may be used, but will have four dedicated "flow groups" (i.e. priorities representing different buffer pools and arbitration groups within the fabric, etc.). These resources will be dedicated to the local HCA and switch ports, but will be shared within the fabric. · You can have the ability to use all available bandwidth (egress and ingress) from the local server, but you are guaranteed to have at least 50% of the total bandwidth. Limited to one Economy Plus tenant per physical server. At least 25% non-blocking leaf uplink bandwidth can be guaranteed for the number of servers used by this Economy Plus tenant.

[0491] According to some embodiments, a fourth tier can be provided (e.g., economy) -Only available on shared servers No dedicated priority You may be allowed to use up to 50% of the server bandwidth (egress and ingress), which may be shared with up to three other economy tenants. Up to 25% of the available leaf uplink bandwidth can be shared with other economy tenants within the same leaf domain.

[0492] According to some embodiments, a bottom layer (e.g., standby) can be provided that can use spare capacity without guaranteed bandwidth. While various embodiments of the present teachings have been described, it should be understood that the embodiments are presented by way of example and not limitation. The embodiments were selected and described in order to explain the principles of the present teachings and their practical application. The embodiments illustrate systems and methods in which the present teachings are utilized to enhance the performance of the systems and methods by providing new and / or improved features and / or by providing benefits such as reduced resource utilization, increased capacity, improved efficiency, and reduced latency.

[0493] In some embodiments, features of the present teachings are implemented, in whole or in part, in a computer that includes a processor, a storage medium such as memory, and a network card for communicating with other computers. In some embodiments, features of the present teachings are implemented in a computer system where one or more clusters of computers are connected by a network, such as a local area network (LAN), a switched fabric network (e.g., InfiniBand), or a wide area network (WAN). The distributed computing environment may have all the computers in one location, or may have a cluster of computers in various remote geographic locations connected by a WAN.

[0494] In some embodiments, features of the present teachings are implemented in the cloud as part of a cloud computing system or as a service, based in whole or in part on shared, elastic resources delivered to users in a self-service, coordinated manner using web technologies. There are five characteristics of the cloud (as defined by the National Institute of Standards and Technology): on-demand self-service, wide-area network access, resource pooling, rapid elasticity, and measured service. See, e.g., "The NIST Definition of Cloud Computing," Special Publication 800-145 (2011), incorporated herein by reference. Cloud deployment models include public, private, and hybrid. Cloud service models include Software as a Service (Sa). Platform as a Service (PaaS) , Database as a Service (DBaaS) and Infrastructure as a Service (IaaS) As used herein, cloud is a combination of hardware, software, network, and web technologies that delivers shared, elastic resources to users in a self-service, coordinated manner. Unless otherwise specified, cloud, as used herein, encompasses public cloud, private cloud, and hybrid cloud embodiments, and all cloud deployment models, including, but not limited to, cloud SaaS, cloud DBaaS, cloud PaaS, and cloud IaaS.

[0495] In some embodiments, features of the present teachings are implemented using or with the aid of hardware, software, firmware, or a combination thereof. In some embodiments, features of the present teachings are implemented using a processor configured or programmed to perform one or more functions of the present teachings. The processor, in some embodiments, is a single or multi-chip processor, a digital signal processor (DSP), a system on a chip (SOC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a state machine, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. In some implementations, features of the present teachings may be implemented by circuitry specialized for a particular function. In other implementations, these features may be implemented in a processor configured to perform a particular function using instructions stored on a computer-readable storage medium, for example.

[0496] In some embodiments, features of the present teachings are incorporated into software and / or firmware for controlling the hardware of a processing system and / or networking system and for enabling the processor and / or network to interact with other systems that utilize features of the present teachings. Such software or firmware may include, but is not limited to, application code, device drivers, operating systems, virtual machines, hypervisors, application programming interfaces, programming languages, and execution environments / containers. Suitable software coding can be readily prepared by skilled programmers based on the teachings of the present disclosure, as will be apparent to those skilled in the software arts.

[0497] In some embodiments, the present approach includes a computer program product, such as a computer-readable medium bearing instructions that can be used to implement the present teachings. In some examples, the computer-readable medium is a storage medium or computer-readable medium having instructions stored thereon, thereby carrying instructions, that can be used to program or otherwise configure a system, such as a computer, to perform any of the processes or functions of the present teachings. The storage medium or computer-readable medium can include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), and any type of medium or device suitable for storing instructions and / or data. In certain embodiments, the storage medium or computer-readable medium is a non-transitory storage medium or non-transitory computer-readable medium. The computer-readable medium can also or alternatively be a Alternatively, it may include a transient medium such as a carrier wave or transmission signal that propagates such instructions.

[0498]

[0013] Thus, from one aspect, a system and method for supporting target groups for congestion control in a private fabric in a high-performance computing environment have been described. An exemplary method can provide, in one or more microprocessors, a first subnet, the first subnet including a plurality of switches, a plurality of host channel adapters, and a plurality of end nodes including a plurality of virtual machines. The method can define a target group on an inter-switch link or one of a port of a switch of the plurality of switches, the target group defining a bandwidth limit on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches. The method can provide a target group repository stored in a memory of the host channel adapter, the defined target group in the target group repository being recorded.

[0499] The foregoing description is not intended to be exhaustive or to limit the scope of the present teachings to the precise form disclosed. Moreover, while embodiments of the present teachings have been described using a particular series of transactions and steps, it will be apparent to those skilled in the art that the scope is not limited to the series of transactions and steps described above. Furthermore, while embodiments have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are within the scope of the present teachings. Furthermore, while specific combinations of features of the present invention have been described in various embodiments, it should be understood that different combinations of these features are within the scope of the present teachings, such as features from one embodiment being incorporated into another embodiment. Furthermore, it will be apparent to those skilled in the art that various additions, deductions, deletions, modifications, and other changes in form, details, implementation, and application can be made without departing from the spirit and scope of the present teachings. It is intended that the present invention be defined by appropriate interpretation of the following claims.

Claims

1. 1. A system for supporting target groups for congestion control in a private fabric in a high performance computing environment, comprising: one or more microprocessors; a first subnet, the first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, the first subnet further comprising: a plurality of end nodes each including a plurality of virtual machines; a target group is defined on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the target group defining a bandwidth limit on the at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches; the host channel adapter includes a target group repository stored in a memory of the host channel adapter; The defined target groups are recorded in the target group repository.

2. the target group is defined on the inter-switch link between two switches of the plurality of switches; a first end node of the plurality of end nodes disposed in the host channel adapter that includes the target group repository; 2. The system of claim 1, wherein the first end node is associated with a first egress bandwidth limit defined in the host channel adapter that includes the target group repository, the first egress bandwidth limit being associated with a first quality of service (QoS) agreement of the first end node.

3. 3. The system of claim 2, wherein packets egressing from the first end node are routed via the inter-switch link between two switches of the plurality of switches where the target group is defined.

4. The system of claim 3 , wherein the bandwidth limit defined by the target group is less than the first egress bandwidth limit.

5. The system of claim 4 , wherein the first egress bandwidth limit associated with the first end node is updated to be less than or equal to the bandwidth limit defined by the target group.

6. 10. The system of claim 1, wherein the bandwidth limits defined by the target group include a plurality of bandwidth limits, each of the plurality of bandwidth limits being associated with a different QoS agreement.

7. 10. A system according to any one of the preceding claims, wherein the target group is decoupled from any particular destination address.

8. Congestion Control in Private Fabrics for High-Performance Computing Environments 1. A method for supporting a target group for providing a first subnet in one or more microprocessors, said first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, the first subnet further comprising: a plurality of end nodes including a plurality of virtual machines, the method further comprising: defining a target group on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the target group defining a bandwidth limit on the at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the method further comprising: providing, in a host channel adapter, a target group repository stored in a memory of the host channel adapter; and recording the defined target group in the target group repository.

9. the target group is defined on the inter-switch link between two switches of the plurality of switches; a first end node of the plurality of end nodes disposed in the host channel adapter that includes the target group repository; 9. The method of claim 8, wherein the first end node is associated with a first egress bandwidth limit defined in the host channel adapter that includes the target group repository, the first egress bandwidth limit being associated with a first quality of service (QoS) agreement of the first end node.

10. 10. The method of claim 9, wherein packets egressing from the first end node are routed via the inter-switch link between two switches of the plurality of switches where the target group is defined.

11. The method of claim 10 , wherein the bandwidth limit defined by the target group is less than the first egress bandwidth limit.

12. The method of claim 11 , wherein the first egress bandwidth limit associated with the first end node is updated to be less than or equal to the bandwidth limit defined by the target group.

13. The method of any one of claims 8 to 12, wherein the bandwidth limit defined by the target group comprises a plurality of bandwidth limits, each of the plurality of bandwidth limits being associated with a different QoS agreement.

14. The method of any one of claims 8 to 13, wherein the target group is decoupled from any particular destination address.

15. 1. A computer-readable medium having instructions for supporting target groups for congestion control in a private fabric in a high performance computing environment, the instructions, when read and executed, causing a computer to perform the steps of: providing a first subnet in one or more microprocessors, said first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each of the host channel adapters including at least one host channel adapter port, the plurality of host channel adapters interconnected via the plurality of switches, the first subnet further comprising: a plurality of end nodes including a plurality of virtual machines, said step further comprising: defining a target group on at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the target group defining a bandwidth limit on the at least one of an inter-switch link between two switches of the plurality of switches or a port of a switch of the plurality of switches, the step further comprising: providing, in a host channel adapter, a target group repository stored in a memory of the host channel adapter; and recording the defined target groups in the target group repository.

16. the target group is defined on the inter-switch link between two switches of the plurality of switches; a first end node of the plurality of end nodes disposed in the host channel adapter that includes the target group repository; 16. The computer-readable medium of claim 15, wherein the first end node is associated with a first egress bandwidth limit defined in the host channel adapter that includes the target group repository, the first egress bandwidth limit being associated with a first quality of service (QoS) agreement of the first end node.

17. 17. The computer-readable medium of claim 16, wherein packets egressing from the first end node are routed via the inter-switch link between two switches of the plurality of switches where the target group is defined.

18. 20. The computer-readable medium of claim 17, wherein the bandwidth limit defined by the target group is less than the first egress bandwidth limit.

19. 20. The computer-readable medium of claim 18, wherein the first egress bandwidth limit associated with the first end node is updated to be less than or equal to the bandwidth limit defined by the target group.

20. 20. The computer-readable medium of claim 15, wherein the bandwidth limits defined by the target group include a plurality of bandwidth limits, each of the plurality of bandwidth limits being associated with a different QoS agreement.