Systems and methods for supporting the use of forward and reverse congestion notifications in a private architecture in a high-performance computing environment
Through the InfiniBand network architecture and FECN/BECN mechanism, multiple CE flags are used to inform congestion, solving the complexity of network resource management and virtual machine migration problems in high-performance computing environments, and achieving efficient resource allocation and performance optimization.
Patent Information
- Application Number
- CN202080080798.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-11
- Filing Date
- 2020-08-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-08-06
AI Technical Summary
In high-performance computing environments, prior art is difficult to effectively manage and optimize network resource allocation, especially when virtual machine migration and network topology changes, resulting in performance bottlenecks and management complexity.
The InfiniBand network architecture is adopted, combined with FECN and BECN mechanisms, and multiple CE flags are used for congestion notification, real-time routing and bandwidth management are implemented, and real-time migration and resource optimization in a virtualized environment.
It realizes efficient network resource allocation and management, reduces downtime during virtual machine migration, improves network performance and resource utilization, and supports elasticity and scalability in cloud computing environments.
Smart Images

Figure CN116057913B_ABST
Abstract
Description
[0001] Copyright Notice
[0002] Part of the disclosure of this patent document contains material which is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.
[0003] Priority Claim and Cross - Reference to Related Applications:
[0004] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 937,594, filed on November 19, 2019, entitled "SYSTEM AND METHOD FOR PROVIDING QUALITY-OF-SERVICE AND SERVICE-LEVEL AGREEMENTS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT", which is incorporated herein by reference in its entirety.
[0005] This application also claims the benefit of priority of the following applications, each of which is incorporated herein by reference in its entirety: U.S. Patent Application No. 16 / 872,035, filed May 11, 2020, entitled "SYSTEM AND METHOD FOR SUPPORTING RDMA BANDWIDTH RESTRICTIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT"; U.S. Patent Application No. 16 / 872,038, filed May 11, 2020, entitled "SYSTEM AND METHOD FOR PROVIDING BANDWIDTH CONGESTION CONTROL IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT"; U.S. Patent Application No. 16 / 872,039, filed May 11, 2020, entitled "SYSTEM AND METHOD FOR SUPPORTING TARGET GROUPS FOR CONSTION CONTROL IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT"; and U.S. Patent Application No. 16 / 872,043, filed May 11, 2020, entitled "SYSTEM AND METHOD FOR SUPPORTING USE OF FORWARD AND BACKWARD CONCESTION NOTIFICATIONS IN A PRIVATE FABRIC IN A HIGH PERFORMANCE COMPUTING ENVIRONMENT". TECHNICAL FIELD
[0006] The present teachings relate to systems and methods for implementing Quality of Service (QoS) and Service Level Agreements (SLAs) in private high-performance interconnect architectures such as InfiniBand (IB) and RoCE (RDMA (Remote Direct Memory Access) over Converged Ethernet). BACKGROUND ART
[0007] With the introduction of larger cloud computing architectures, performance and management bottlenecks associated with traditional networks and storage devices have become important issues. There is increasing interest in using high-performance lossless interconnects such as InfiniBand (IB) technology as the basis for cloud computing architectures. This is the general area that embodiments of the present invention are directed to solving. Summary of the Invention
[0008] Specific aspects are set forth in the appended independent claims. Various alternative embodiments are set forth in the dependent claims.
[0009] Described herein are systems and methods for using multiple CE (Congestion Experienced) flags in both FECN (Forward Explicit Congestion Notification) and BECN (Backward Explicit Congestion Notification) in a high-performance computing environment. An example method may provide a first subnet that includes a plurality of switches, a plurality of host channel adapters, and a plurality of end nodes. The method may receive an ingress packet from a remote end node at an end node attached to a host channel adapter, wherein the ingress packet traverses at least a portion of the first subnet before being received at the end node. The method may, upon receiving the ingress packet, send a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet. Brief Description of the Drawings
[0010] Figure 1 A diagram illustrating an InfiniBand environment according to an embodiment is shown.
[0011] Figure 2 A diagram illustrating a partitioned cluster environment according to an embodiment is shown.
[0012] Figure 3 A diagram illustrating a tree topology in a network environment according to an embodiment is shown.
[0013] Figure 4 A diagram illustrating an example shared port architecture according to an embodiment is shown.
[0014] Figure 5 A diagram illustrating an example vSwitch architecture according to an embodiment is shown.
[0015] Figure 6 A diagram illustrating an example vPort architecture according to an embodiment is shown.
[0016] Figure 7 A diagram illustrating an example vSwitch architecture with pre-populated LIDs according to an embodiment is shown.
[0017] Figure 8Shows an example vSwitch architecture with dynamic LID allocation according to an embodiment.
[0018] Figure 9 Shows an example vSwitch architecture with a vSwitch having dynamic LID allocation and pre-populated LIDs according to an embodiment.
[0019] Figure 10 Shows an example multi-subnet InfiniBand architecture according to an embodiment.
[0020] Figure 11 Shows an interconnection between two subnets in a high-performance computing environment according to an embodiment.
[0021] Figure 12 Shows an interconnection between two subnets configured via a dual-port virtual router in a high-performance computing environment according to an embodiment.
[0022] Figure 13 Shows a flowchart of a method for supporting a dual-port virtual router in a high-performance computing environment according to an embodiment.
[0023] Figure 14 Shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to an embodiment.
[0024] Figure 15 Shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to an embodiment.
[0025] Figure 16 Shows a system for providing an RDMA read request as a restricted feature in a high-performance computing environment according to an embodiment.
[0026] Figure 17 Shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment.
[0027] Figure 18 Shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment.
[0028] Figure 19 Shows a system for providing an explicit RDMA read bandwidth limit in a high-performance computing environment according to an embodiment.
[0029] Figure 20 Is a flowchart of a method for providing an RDMA (Remote Direct Memory Access) read request as a restricted feature in a high-performance computing environment.
[0030] Figure 21 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0031] Figure 22 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0032] Figure 23 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0033] Figure 24 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0034] Figure 25 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0035] Figure 26 A flowchart of a method for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0036] Figure 27 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment according to an embodiment is shown.
[0037] Figure 28 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment according to an embodiment is shown.
[0038] Figure 29 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment according to an embodiment is shown.
[0039] Figure 30 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment according to an embodiment is shown.
[0040] Figure 31 A flowchart of a method for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment according to an embodiment is shown.
[0041] Figure 32 A system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment according to an embodiment is shown.
[0042] Figure 33Illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment according to an embodiment.
[0043] Figure 34 Illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment according to an embodiment.
[0044] Figure 35 Is a flowchart of a method for combining ingress bandwidth arbitration and congestion feedback in a high performance computing environment according to an embodiment.
[0045] Figure 36 Illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment according to an embodiment.
[0046] Figure 37 Illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment according to an embodiment.
[0047] Figure 38 Illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment according to an embodiment.
[0048] Figure 39 Is a flowchart of a method for using multiple CE flags in both FECN and BECN in a high performance computing environment according to an embodiment. Detailed Description
[0049] The present teachings are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate similar elements. It should be noted that references to "one," "a," or "some" embodiments in this disclosure are not necessarily to the same embodiment, and such references mean at least one. Although specific implementations are discussed, it should be understood that the specific implementations are provided for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope.
[0050] Common reference numerals may be used throughout the drawings and the detailed description to indicate the same elements; thus, the reference numerals used in the figures may or may not be referenced in the specific description particular to that figure if the element has been described elsewhere.
[0051] Described herein are systems and methods for providing quality of service (QOS) and service level agreement (SLA) in a private architecture in a high performance computing environment.
[0052] According to an embodiment, the following description of the present teachings uses InfiniBand TM(IB) network as an example of a high - performance network. Throughout the following description, reference may be made to InfiniBand TM specification (also variously referred to as the InfiniBand specification, the IB specification, or the legacy IB specification). Such reference is understood to be to the Trade Association Architecture Specification published by the InfiniBand Trade Association in March 2015 and available at http: / / www.inifinibandta.org, Volume 1, Version 1.3, the entire content of which is incorporated herein by reference. It will be apparent to those skilled in the art that other types of high - performance networks may be used without limitation. The following description also uses a fat - tree topology as an example of an architectural topology. It will be apparent to those skilled in the art that other types of architectural topologies may be used without limitation.
[0053] According to an embodiment, the following description uses RoCE (RDMA (Remote Direct Memory Access) over Converged Ethernet). RDMA over Converged Ethernet (RoCE) is a standard protocol that enables efficient data transfer of RDMA over Ethernet, thus allowing transmission offloading to be achieved using a hardware RDMA engine and having excellent performance. RoCE is a standard protocol defined in the InfiniBand Trade Association (IBTA) standards. RoCE uses UDP (User Datagram Protocol) encapsulation, allowing it to operate beyond layer 3 networks. RDMA is a key capability used by the InfiniBand interconnect technology itself. Both InfiniBand and Ethernet RoCE share a common user API but have different physical and link layers.
[0054] According to an embodiment, although parts of the specification include references to the InfiniBand architecture in describing various embodiments, those of ordinary skill in the art will readily understand that the various embodiments described herein may also be implemented in a RoCE architecture.
[0055] To meet the requirements of clouds in the current era (e.g., the Exascale era), it is desirable for virtual machines to utilize low-overhead network communication paradigms such as Remote Direct Memory Access (RDMA). RDMA bypasses the OS stack and communicates directly with the hardware, so that direct-pass technologies such as Single Root I / O Virtualization (SR-IOV) network adapters can be used. According to an embodiment, a virtual switch (vSwitch) SR-IOV architecture can be provided for applicability in high-performance lossless interconnection networks. Since network reconfiguration time is crucial for making live migration a practical option, in addition to the network architecture, a scalable and topology-independent dynamic reconfiguration mechanism can also be provided.
[0056] According to an embodiment, and in addition, a routing policy for a virtualized environment using a vSwitch can be provided, and an efficient routing algorithm for a network topology (e.g., a fat-tree topology) can be provided. The dynamic reconfiguration mechanism can be further tuned to minimize the overhead introduced in the fat tree.
[0057] According to embodiments of the present teachings, virtualization can be beneficial for efficient resource utilization and elastic resource allocation in cloud computing. Live migration makes it possible to optimize resource usage by moving virtual machines (VMs) between physical servers in an application-transparent manner. Thus, virtualization can achieve consolidation, on-demand provisioning of resources, and elasticity through live migration.
[0058] InfiniBand TM
[0059] InfiniBand TM (IB) is an open standard lossless network technology developed by the InfiniBand TM Trade Association. The technology is based on a serial point-to-point full-duplex interconnection that provides high throughput and low latency communication, specifically targeting high-performance computing (HPC) applications and data centers.
[0060] InfiniBand TM architecture (IBA) supports a two-tier topology division. At the lower tier, the IB network is called a subnet, where a subnet can include a collection of hosts interconnected using switches and point-to-point links. At the upper tier, the IB architecture consists of one or more subnets that can be interconnected using routers.
[0061] Within a subnet, switches and point-to-point links can be used to connect hosts. Additionally, there can be a primary management entity, namely the Subnet Manager (SM), which resides on a designated device within the subnet. The Subnet Manager is responsible for configuring, activating, and maintaining the IB subnet. Additionally, the Subnet Manager (SM) can be responsible for performing routing table calculations within the IB fabric. Here, for example, the routing of the IB network is intended to perform proper load balancing between all source and destination pairs within the local subnet.
[0062] Through the subnet management interface, the Subnet Manager exchanges control packets with the Subnet Management Agent (SMA), and these control packets are referred to as Subnet Management Packets (SMP). The Subnet Management Agent resides on each IB subnet device. By using SMP, the Subnet Manager is able to discover the fabric, configure end nodes and switches, and receive notifications from the SMA.
[0063] According to an embodiment, the intra-subnet routing in the IB network can be based on a Linear Forwarding Table (LFT) stored in the switch. The LFT is calculated by the SM according to the routing mechanism in use. Within the subnet, the Host Channel Adapter (HCA) ports on end nodes and switches are addressed using a Local Identifier (LID). Each entry in the Linear Forwarding Table (LFT) includes a Destination LID (DLID) and an output port. Only one entry per LID is supported in the table. When a packet arrives at the switch, its output port is determined by looking up the DLID in the switch's forwarding table. The routing is deterministic because the packet takes the same path in the network between a given source-destination pair (LID pair).
[0064] Generally, all other Subnet Managers other than the primary Subnet Manager operate in standby mode for fault tolerance. However, in the event of a failure of the primary Subnet Manager, a new primary Subnet Manager is negotiated by the standby Subnet Managers. The primary Subnet Manager also performs periodic scans of the subnet to detect any topology changes and reconfigure the network accordingly.
[0065] Additionally, Local Identifiers (LIDs) can be used to address hosts and switches within a subnet, and a single subnet can be limited to 49151 unicast LIDs. In addition to the LID which is a local address valid within the subnet, each IB device can also have a 64-bit Globally Unique Identifier (GUID). The GUID can be used to form a Global Identifier (GID) which serves as an IB Layer 3 (L3) address.
[0066] At network initialization, the SM can compute the routing table (i.e., the connections / routes between every pair of nodes within the subnet). Additionally, whenever the topology changes, the routing table can be updated to ensure connectivity and optimal performance. During normal operation, the SM can perform periodic light scans of the network to check for topology changes. If a change is detected during a light scan, or if the SM receives information (a trap) signaling a network change, then the SM can reconfigure the network based on the detected change.
[0067] For example, when the network topology changes, such as when a link is disconnected, when a device is added, or when a link is removed, the SM can reconfigure the network. The reconfiguration steps can include the steps performed during network initialization. Additionally, the reconfiguration can have a local scope limited to the subnet in which the network change occurs. Additionally, segmenting a large architecture with routers can limit the scope of reconfiguration.
[0068] Figure 1 An example InfiniBand architecture is shown, which depicts a diagram of an InfiniBand environment 100 according to an embodiment. In Figure 1 the example shown, nodes A - E 101 - 105 communicate using an InfiniBand architecture 120 via respective host channel adapters 111 - 115. According to an embodiment, the various nodes (e.g., nodes A - E 101 - 105) can be represented by various physical devices. According to an embodiment, the various nodes (e.g., nodes A - E 101 - 105) can be represented by various virtual devices such as virtual machines.
[0069] Partitioning in InfiniBand
[0070] According to an embodiment, the IB network can support partitioning as a security mechanism to provide isolation for logical groups of systems sharing a network architecture. Each HCA port on a node in the architecture can be a member of one or more partitions. Partition membership is managed by a centralized partition manager, which can be part of the SM. The SM can configure the partition membership information on each port as a table of 16 - bit partition keys (P_Keys). The SM can also configure switch and router ports with a partition enforcement table containing P_Key information associated with the end nodes sending or receiving data traffic through those ports. Additionally, in general, the partition membership of a switch port can represent the union of all memberships indirectly associated with the LID routed via the port in the egress (towards the link) direction.
[0071] According to an embodiment, a partition is a logical group of ports such that members of the group can only communicate with other members of the same logical group. At a host channel adapter (HCA) and a switch, partition membership information can be used to filter packets to enforce isolation. Once a packet arrives at an incoming port, a packet with invalid partition information can be discarded. In a partitioned IB system, partitions can be used to create tenant clusters. With a partition implementation in place, a node cannot communicate with other nodes belonging to different tenant clusters. In this way, the security of the system can be ensured even when there are compromised or malicious tenant nodes.
[0072] According to an embodiment, for communication between nodes, queue pairs (QPs) and end-to-end contexts (EECs), in addition to management queue pairs (QP0 and QP1), can be assigned to a specific partition. Then, P_Key information can be added to each IB transmission packet sent. When a packet arrives at an HCA port or a switch, its P_Key value can be verified against a table configured by the SM. If an invalid P_Key value is found, the packet is immediately discarded. In this way, communication is only allowed between ports sharing a partition.
[0073] Figure 2 An example of an IB partition is shown, which shows a diagram of a partitioned cluster environment according to an embodiment. In Figure 2 In the example shown, nodes A - E 101 - 105 communicate via respective host channel adapters 111 - 115 using the InfiniBand architecture 120. Nodes A - E are arranged into partitions, namely partition 1 130, partition 2 140, and partition 3 150. Partition 1 includes node A 101 and node D 104. Partition 2 includes node A 101, node B 102, and node C 103. Partition 3 includes node C 103 and node E 105. Due to the partition arrangement, node D 104 and node E 105 are not allowed to communicate because these nodes do not share a partition. Meanwhile, for example, node A 101 and node C 103 are allowed to communicate because these nodes are both members of partition 2 140.
[0074] Virtual Machines in InfiniBand
[0075] In the past decade, the prospects for virtualizing high-performance computing (HPC) environments have improved considerably, as CPU overhead has been virtually eliminated through hardware virtualization support; memory overhead has been significantly reduced by virtualizing the memory management unit; storage device overhead has been reduced by using fast SAN storage devices or distributed network file systems; and network I / O overhead has been reduced by using device passthrough techniques such as single root I / O virtualization (SR-IOV). Now, the cloud has the potential to use high-performance interconnect solutions to host virtual HPC (vHPC) clusters and provide the necessary performance.
[0076] However, when coupled with lossless networks such as InfiniBand (IB), some cloud functions, such as live migration of virtual machines (VMs), remain a problem due to the complex addressing and routing schemes used in these solutions. IB is an interconnect network technology that provides high bandwidth and low latency and is thus well-suited for HPC and other communication-intensive workloads.
[0077] The traditional method for connecting an IB device to a VM is by leveraging SR-IOV with direct assignment. However, implementing live migration of VMs allocated with IB host channel adapters (HCAs) using SR-IOV has proven challenging. Each IB-connected node has three different addresses: LID, GUID, and GID. When a live migration occurs, one or more of these addresses change. Other nodes communicating with the migrating VM may lose connectivity. When this happens, attempts can be made to update the lost connection by locating the new address of the virtual machine to reconnect to by sending a subnet management (SA) path record query to the IB subnet manager (SM).
[0078] IB uses three different types of addresses. The first type of address is the 16-bit local identifier (LID). The SM assigns at least one unique LID to each HCA port and each switch. The LID is used to route traffic within a subnet. Since the LID is 16 bits long, 65,536 unique address combinations can be made, of which only 49,151 (0x0001 - 0xBFFF) can be used as unicast addresses. Therefore, the number of available unicast addresses limits the maximum size of the IB subnet. The second type of address is the 64-bit globally unique identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM can assign additional subnet-unique GUIDs to HCA ports, which are useful when using SR-IOV. The third type of address is the 128-bit global identifier (GID). The GID is a valid IPv6 unicast address, and at least one GID is assigned to each HCA port. The GID is formed by combining a globally unique 64-bit prefix assigned by the fabric administrator and the GUID address of each HCA port.
[0079] Fat - Tree (FTree) Topology and Routing
[0080] According to an embodiment, some of the IB-based HPC systems employ a fat-tree topology to take advantage of the useful features provided by the fat tree. These features include full bisection bandwidth and inherent fault tolerance since multiple paths are available between each source-destination pair. The original idea behind the fat tree was to use fatter links with more available bandwidth between nodes as the tree moves towards the root of the topology. The fatter links can help avoid congestion in the upper-level switches and maintain bisection bandwidth.
[0081] Figure 3 A diagram of a tree topology in a network environment according to an embodiment is shown. As Figure 3 shown, one or more end nodes 201 - 204 can be connected in the network architecture 200. The network architecture 200 can be based on a fat-tree topology including multiple leaf switches 211 - 214 and multiple backbone switches or root switches 231 - 234. Additionally, the network architecture 200 can include one or more intermediate switches, such as switches 221 - 224.
[0082] Also as Figure 3 shown, each of the end nodes 201 - 204 can be a multi-homed node, i.e., a single node connected to two or more parts of the network architecture 200 through multiple ports. For example, node 201 can include ports H1 and H2, node 202 can include ports H3 and H4, node 203 can include ports H5 and H6, and node 204 can include ports H7 and H8.
[0083] In addition, each switch can have multiple switch ports. For example, root switch 231 can have switch ports 1 - 2, root switch 232 can have switch ports 3 - 4, root switch 233 can have switch ports 5 - 6, and root switch 234 can have switch ports 7 - 8.
[0084] According to an embodiment, the fat - tree routing mechanism is one of the most popular routing algorithms for IB - based fat - tree topologies. The fat - tree routing mechanism is also implemented in the OFED (Open Fabric Enterprise Distribution - a standard software stack for building and deploying IB - based applications) subnet manager OpenSM.
[0085] The fat - tree routing mechanism aims to generate LFTs that evenly disperse the shortest - path routes across the links in the network architecture. The mechanism traverses the architecture in index order and assigns the destination LIDs (and thus the corresponding routes) of the end nodes to each switch port. For end nodes connected to the same leaf switch, the index order can depend on the switch port (i.e., the port - number order) to which the end node is connected. For each port, the mechanism can maintain a port - usage counter and can use this port - usage counter to select the least - used port each time a new route is added.
[0086] According to an embodiment, in a partitioned subnet, nodes that are not members of a common partition are not allowed to communicate. In practice, this means that some of the routes assigned by the fat - tree routing algorithm are not used for user traffic. This problem occurs when the fat - tree routing mechanism generates LFTs for those routes in the same way it does for other functional paths. Since the nodes are routed in index order, this behavior can lead to degraded balance on the links. Since routing can be performed independently of partitioning, in general, fat - tree - routed subnets provide poor isolation between partitions.
[0087] According to an embodiment, a fat - tree is a hierarchical network topology that can scale using available network resources. Moreover, it is easy to construct a fat - tree using commercial switches placed at different levels of the hierarchy. Different variants of fat - trees can generally be obtained, including k - ary - n - trees, Extended Generalized Fat - Trees (XGFTs), Parallel Port Generalized Fat - Trees (PGFTs), and Real - Life Fat - Trees (RLFTs).
[0088] A k - ary - n - tree has k n end nodes and n·k n-1An n - level fat - tree of switches, where each switch has 2k ports. Each switch has the same number of up - and down - links in the tree. The XGFT fat - tree extends the k - ary - n - tree by allowing switches to have different numbers of up - and down - links and different numbers of links at each level in the tree. The PGFT definition further broadens the XGFT topology and allows multiple connections between switches. A wide variety of topologies can be defined using XGFT and PGFT. However, for practical purposes, the RLFT, which is a restricted version of the PGFT, is introduced to define the fat - trees common in today's HPC clusters. The RLFT uses switches with the same port count at all levels of the fat - tree.
[0089] Input / Output (I / O) virtualization
[0090] According to an embodiment, I / O virtualization (IOV) can provide the availability of I / O by allowing virtual machines (VMs) to access underlying physical resources. The combination of storage traffic and inter - server communication brings an increased load on I / O resources that can overwhelm a single server, resulting in backlogs and idle processors due to the processor waiting for data. As the number of I / O requests increases, IOV can provide availability; and can improve the performance, scalability, and flexibility of (virtualized) I / O resources to match the performance levels seen in modern CPU virtualization.
[0091] According to an embodiment, IOV is desirable because it can allow sharing of I / O resources and provide protected access of VMs to resources. IOV decouples the logical device exposed to the VM from the physical implementation of that logical device. Currently, there can be different types of IOV technologies, such as emulation, paravirtualization, direct assignment (DA), and single - root I / O virtualization (SR - IOV).
[0092] According to an embodiment, one type of IOV technology is software emulation. Software emulation can allow a decoupled front - end / back - end software architecture. The front - end can be a device driver placed in the VM that communicates with the back - end implemented by the hypervisor to provide I / O access. The physical device sharing ratio is high, and live migration of VMs may require only a few milliseconds of network downtime. However, software emulation introduces additional, unwanted computational overhead.
[0093] According to an embodiment, another type of IOV technology is direct device assignment. Direct device assignment involves coupling an I / O device to a VM, where there is no device sharing between VMs. Direct assignment or device - passthrough provides near - native performance with minimal overhead. The physical device bypasses the hypervisor and attaches directly to the VM. However, the drawback of this direct device assignment is limited scalability because there is no sharing between virtual machines - one physical network card is coupled to one VM.
[0094] According to an embodiment, a single root IOV (SR-IOV) can allow a physical device to appear as multiple independent lightweight instances of the same device through hardware virtualization. These instances can be assigned to a VM as a passthrough device and accessed as virtual functions (VFs). The hypervisor accesses the device through a unique, full-featured physical function (PF) (per device). SR-IOV eases the scalability issues of pure direct assignment. However, an issue presented by SR-IOV is that it may affect VM migration. Among these IOV technologies, SR-IOV can extend the PCI Express (PCIe) specification by means that allow direct access to a single physical device from multiple VMs while maintaining near-native performance. Therefore, SR-IOV can provide good performance and scalability.
[0095] SR-IOV allows a PCIe device to expose multiple virtual devices that can be shared among multiple guests by assigning a virtual device to each guest. Each SR-IOV device has at least one physical function (PF) and one or more associated virtual functions (VFs). The PF is a normal PCIe function controlled by a virtual machine monitor (VMM) or hypervisor, while the VF is a lightweight PCIe function. Each VF has its own base address register (BAR) and is assigned a unique requester ID, which enables the I / O memory management unit (IOMMU) to distinguish traffic flows to / from different VFs. The IOMMU also applies memory and interrupt translation between the PF and VFs.
[0096] However, unfortunately, direct device assignment techniques pose an obstacle to cloud providers in cases where transparent live migration of virtual machines is expected for data center optimization. The essence of live migration is to copy the memory content of the VM to a remote hypervisor. Then the VM is paused at the source hypervisor and its operation is resumed at the destination. When using the software emulation method, the network interface is virtual, so its internal state is stored in memory and also copied. Thus, the downtime can be reduced to a few milliseconds.
[0097] However, when using direct device assignment techniques such as SR-IOV, migration becomes more difficult. In this case, the complete internal state of the network interface cannot be copied because it is hardware-bound. Instead, the SR-IOV VF assigned to the VM is detached, the live migration will run, and a new VF will be attached at the destination. In the case of InfiniBand and SR-IOV, this process can introduce downtime on the order of seconds. Additionally, in the SR-IOV shared port model, after migration, the address of the VM will change, resulting in additional overhead in the SM and negatively affecting the performance of the underlying network architecture.
[0098] InfiniBand SR - IOV Architecture - Shared Port
[0099] There can be different types of SR-IOV models, for example, a shared port model, a virtual switch model, and a virtual port model.
[0100] Figure 4 An example shared port architecture according to an embodiment is shown. As shown, a host 300 (e.g., a host channel adapter) can interact with a hypervisor 310, and the hypervisor 310 can allocate various virtual functions 330, 340, 350 to multiple virtual machines. Similarly, the physical function can be processed by the hypervisor 310.
[0101] According to an embodiment, when using a shared port architecture such as Figure 4 the shared port architecture shown, the host (e.g., an HCA) appears as a single port in the network with a single shared LID and a shared queue pair (QP) space between the physical function 320 and the virtual functions 330, 350, 350. However, each function (i.e., the physical function and the virtual functions) can have its own GID.
[0102] As Figure 4 shown, according to an embodiment, different GIDs can be assigned to the virtual functions and the physical function, and special queue pairs QP0 and QP1 (i.e., dedicated queue pairs for InfiniBand management packets) are owned by the physical function. These QPs are also exposed to the VF, but the VF is not allowed to use QP0 (all SMPs from the VF towards QP0 are discarded), and QP1 can act as a proxy for the actual QP1 owned by the PF.
[0103] According to an embodiment, the shared port architecture can allow a highly scalable data center that is not limited by the number of VMs (attached to the network by being assigned to virtual functions), because the LID space is only consumed by the physical machines and switches in the network.
[0104] However, a disadvantage of the shared port architecture is that it cannot provide transparent live migration, thus hindering the possibility of flexible VM placement. Since each LID is associated with a specific hypervisor and is shared among all VMs residing on the hypervisor, the migrating VM (i.e., the virtual machine migrating to the destination hypervisor) must change its LID to the LID of the destination hypervisor. In addition, due to restricted QP0 access, the subnet manager cannot run inside the VM.
[0105] InfiniBand SR - IOV Architecture Model - Virtual Switch (vSwitch)
[0106] Figure 5Shows an example vSwitch architecture according to an embodiment. As shown, a host 400 (e.g., a host channel adapter) can interact with a hypervisor 410, and the hypervisor 410 can allocate various virtual functions 430, 440, 450 to multiple virtual machines. Similarly, physical functions can be processed by the hypervisor 410. The virtual switch 415 can also be processed by the hypervisor 401.
[0107] According to an embodiment, in the vSwitch architecture, each virtual function 430, 440, 450 is a complete virtual host channel adapter (vHCA), which means that the VMs assigned to the VF are assigned a complete set of IB addresses (e.g., GID, GUID, LID) and dedicated QP space in the hardware. To the rest of the network and the SM, the HCA 400 appears as a switch with additional nodes connected to it via the virtual switch 415. The hypervisor 410 can use the PF 420, and the VMs (attached to the virtual functions) use the VF.
[0108] According to an embodiment, the vSwitch architecture provides transparent virtualization. However, since each virtual function is assigned a unique LID, the number of available LIDs is quickly consumed. Similarly, in the case of using many LID addresses (i.e., one LID address for each physical function and each virtual function), more communication paths must be calculated by the SM, and more subnet management packets (SMPs) must be sent to the switch to update its LFT. For example, the calculation of communication paths may take several minutes in a large network. Since the LID space is limited to 49151 unicast LIDs, and since each VM (via the VF), physical node, and switch each occupy one LID, the number of physical nodes and switches in the network limits the number of active VMs, and vice versa.
[0109] InfiniBand SR - IOV Architecture Model - Virtual Port (vPort)
[0110] Figure 6 Shows an example vPort concept according to an embodiment. As shown, a host 300 (e.g., a host channel adapter) can interact with a hypervisor 410, and the hypervisor 410 can allocate various virtual functions 330, 340, 350 to multiple virtual machines. Similarly, physical functions can be processed by the hypervisor 310.
[0111] According to an embodiment, the vPort concept is loosely defined to give vendors freedom in implementation (e.g., the definition does not specify that the implementation must be SRIOV - specific), and the goal of vPort is to standardize the way VMs are handled in a subnet. Using the vPort concept, an SR - IOV - like shared port and a vSwitch - like architecture or a combination of both that can be more scalable in both the spatial and performance domains can be defined. vPort supports an optional LID, and different from a shared port, even if a vPort does not use a dedicated LID, the SM knows all vPorts available in the subnet.
[0112] InfiniBand SR - IOV Architecture Model – vSwitch with Pre - populated LID
[0113] According to an embodiment, the present disclosure provides systems and methods for providing a vSwitch architecture with pre - populated LIDs.
[0114] Figure 7 An example vSwitch architecture with pre - populated LIDs according to an embodiment is shown. As shown, multiple switches 501 - 504 can provide communication between members of an architecture (such as an InfiniBand architecture) within a network switching environment 600 (e.g., an IB subnet). The architecture can include multiple hardware devices, such as host channel adapters 510, 520, 530. Each of the host channel adapters 510, 520, 530 can in turn interact with hypervisors 511, 521, and 531 respectively. Each hypervisor can in turn establish and assign a plurality of virtual functions 514, 515, 516, 524, 525, 526, 534, 535, 536 to a plurality of virtual machines in combination with the host channel adapter with which it interacts. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515, and assign virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can in turn assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapter through full - featured physical functions 513, 523, 533 on each of the host channel adapters.
[0115] According to an embodiment, each of the switches 501 - 504 can include a plurality of ports (not shown) for setting up a linear forwarding table to direct traffic within the network switching environment 600.
[0116] According to an embodiment, virtual switches 512, 522, and 532 can be processed by their respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), which means that the VMs assigned to the VF are assigned a complete set of IB addresses (e.g., GID, GUID, LID) and dedicated QP space in the hardware. To the rest of the network and the SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected to them via the virtual switches.
[0117] According to an embodiment, the present disclosure provides systems and methods for providing a vSwitch architecture with pre-populated LIDs. Referring to Figure 7 , LIDs are pre-populated to various physical functions 513, 523, 533 and virtual functions 514 - 516, 524 - 526, 534 - 536 (even those virtual functions that are not currently associated with active virtual machines). For example, physical function 513 is pre-populated with LID1, while virtual function 1 534 is pre-populated with LID 10. When the network is started, LIDs are pre-populated in the subnet enabled with SR-IOV vSwitch. Even if not all VFs are occupied by VMs in the network, the populated VFs are assigned LIDs, as Figure 7 shown.
[0118] According to an embodiment, very similar to a physical host channel adapter can have more than one port (two ports are common for redundancy), a virtual HCA can also be represented by two ports and connected to an external IB subnet via one, two, or more virtual switches.
[0119] According to an embodiment, in a vSwitch architecture with pre-populated LIDs, each hypervisor can consume one LID for itself through the PF and one additional LID for each attached VF. The sum of all VFs available across all hypervisors in an IB subnet gives the maximum number of VMs allowed to run in the subnet. For example, in an IB subnet with 16 virtual functions per hypervisor in the subnet, each hypervisor then consumes 17 LIDs in the subnet (one LID for each of the 16 virtual functions plus one LID for the physical function). In such an IB subnet, the theoretical hypervisor limit for a single subnet is determined by the number of available unicast LIDs and is: 2,891 (49,151 available LIDs divided by 17 LIDs per hypervisor), and the total number of VMs (i.e., the limit) is 46,256 (2,891 hypervisors multiplied by 16 VFs per hypervisor). (In practice, these numbers are actually smaller because each switch, router, or dedicated SM node in the IB subnet also consumes LIDs). Note that the vSwitch does not need to consume additional LIDs because it can share LIDs with the PF.
[0120] According to an embodiment, in a vSwitch architecture with pre-populated LIDs, when the network is first started up, communication paths are calculated for all LIDs. When a new VM needs to be started, the system does not have to add a new LID in the subnet, an action that would otherwise result in a complete reconfiguration of the network, including path recalculation, which is the most time-consuming part. Instead, an available port (i.e., an available virtual function) for the VM is located in one of the hypervisors, and the virtual machine is attached to the available virtual function.
[0121] According to an embodiment, a vSwitch architecture with pre-populated LIDs also allows the ability to calculate and use different paths to reach different VMs hosted by the same hypervisor. In essence, this allows such subnets and networks to provide alternative paths towards a physical machine using features similar to LID Masking Control (LMC), without being restricted by the LMC requirement that LIDs must be contiguous. The ability to freely use non-contiguous LIDs is particularly useful when a VM needs to be migrated and its associated LID carried to the destination.
[0122] According to an embodiment, and the advantages of the vSwitch architecture with pre-populated LIDs shown above, certain considerations can be taken into account. For example, since the LIDs are pre-populated in a subnet with an SR-IOV vSwitch enabled when the network is started up, the initial path calculation (e.g., at startup) may take longer than if the LIDs were not pre-populated.
[0123] InfiniBand SR - IOV Architecture Model - vSwitch with Dynamic LID Allocation
[0124] According to an embodiment, the present disclosure provides systems and methods for providing a vSwitch architecture with dynamic LID allocation.
[0125] Figure 8 An example vSwitch architecture with dynamic LID allocation according to an embodiment is shown. As shown, multiple switches 501 - 504 can provide communication between members of an architecture (such as an InfiniBand architecture) within a network switching environment 700 (e.g., an IB subnet). The architecture can include multiple hardware devices, such as host channel adapters 510, 520, 530. Each of the host channel adapters 510, 520, 530 can interact with hypervisors 511, 521, 531 respectively. Each hypervisor can establish multiple virtual functions 514, 515, 516, 524, 525, 526, 534, 535, 536 in conjunction with the host channel adapter with which it interacts and assign the multiple virtual functions to multiple virtual machines. For example, virtual machine 1550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2515, and assign virtual machine 3 552 to virtual function 3 516. Hypervisor 531 can, in turn, assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapter through full-featured physical functions 513, 523, 533 on each of the host channel adapters.
[0126] According to an embodiment, each of the switches 501 - 504 can include multiple ports (not shown) that are used to set up a linear forwarding table to direct traffic within the network switching environment 700.
[0127] According to an embodiment, virtual switches 512, 522, and 532 can be processed by their respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), which means that the VMs assigned to the VF are assigned a complete set of IB addresses (e.g., GID, GUID, LID) and dedicated QP space in the hardware. To the rest of the network and the SM (not shown), the HCAs 510, 520, and 530 appear as switches with additional nodes connected to them via the virtual switches.
[0128] According to an embodiment, the present disclosure provides systems and methods for providing a vSwitch architecture with dynamic LID allocation. Refer to Figure 8, LIDs are dynamically assigned to various physical functions 513, 523, 533, where physical function 513 receives LID 1, physical function 523 receives LID 2, and physical function 533 receives LID 3. Those virtual functions associated with active virtual machines can also receive dynamically assigned LIDs. For example, since virtual machine 1550 is active and associated with virtual function 1514, virtual function 514 can be assigned LID 5. Similarly, virtual function 2515, virtual function 3516, and virtual function 1534 are each associated with an active virtual function. Thus, these virtual functions are assigned LIDs, where LID 7 is assigned to virtual function 2515, LID 11 is assigned to virtual function 3516, and LID 9 is assigned to virtual function 1534. Different from vSwitches with pre-populated LIDs, those virtual functions not currently associated with active virtual machines do not receive LID assignments.
[0129] According to an embodiment, with dynamic LID assignment, the initial path calculation can be significantly reduced. When the network is first started and there are no VMs, then a relatively small number of LIDs can be used for the initial path calculation and LFT distribution.
[0130] According to an embodiment, very similar to a physical host channel adapter that can have more than one port (two ports are common for redundancy), a virtual HCA can also be represented with two ports and connected to an external IB subnet via one, two, or more virtual switches.
[0131] According to an embodiment, when creating a new VM in a system using a vSwitch with dynamic LID assignment, an idle VM time slot is found to decide on which hypervisor to start the newly added VM, and also a unique unused unicast LID is found. However, there are no known paths and LFTs of switches in the network for handling the newly added LID. Calculating a new set of paths to handle the newly added VM is not desirable in a dynamic environment where several VMs can be started per minute. In a large IB subnet, calculating a new set of routes can take several minutes, and this process will have to be repeated every time a new VM is started.
[0132] Advantageously, according to an embodiment, since all VFs in the hypervisor share the same uplink with the PF, there is no need to calculate a new set of routes. It is only necessary to traverse the LFTs of all physical switches in the network, copy the LID entries of the forwarding ports subordinate to the PF of the hypervisor (where the VM is created) to the newly added LID, and send a single SMP to update the corresponding LFT block of a specific switch. Thus, the system and method avoid the need to calculate a new set of routes.
[0133] According to an embodiment, the LIDs allocated in a vSwitch with a dynamic LID allocation architecture are not necessarily consecutive. When comparing the LIDs allocated to VMs on each hypervisor in a vSwitch with pre-populated LIDs with those in a vSwitch with dynamic LID allocation, it should be noted that the LIDs allocated in the dynamic LID allocation architecture are non-consecutive, while those pre-populated are essentially consecutive. In the vSwitch dynamic LID allocation architecture, when a new VM is created, the next available LID is used throughout the VM's life cycle. In contrast, in a vSwitch with pre-populated LIDs, each VM inherits the LID that has been allocated to the corresponding VF, and in a network without live migration, VMs continuously attached to a given VF obtain the same LID.
[0134] According to an embodiment, a vSwitch with a dynamic LID allocation architecture can utilize the pre-populated LID architecture model to address the drawbacks of the vSwitch at the cost of some additional network and runtime SM overhead. Each time a VM is created, the LFT of the physical switch in the subnet is updated with the newly added LID associated with the created VM. For this operation, a per-switch one subnet management packet (SMP) needs to be sent. Since each VM is using the same path as its host hypervisor, functions similar to LMC are also not available. However, there is no limit to the total number of VFs present in all hypervisors, and the number of VFs can exceed the number limited by the unicast LID. Of course, if this is the case, then not all VFs are allowed to be attached to active VMs simultaneously, but when the operation approaches the unicast LID limit, having more idle hypervisors and VFs increases the flexibility for disaster recovery and segmented network optimization.
[0135] InfiniBand SR - IOV Architecture Model - vSwitch with Dynamic LID Allocation and Pre - populated LID
[0136] Figure 9Shows an example vSwitch architecture with a vSwitch having dynamic LID allocation and pre-populated LIDs according to an embodiment. As shown, multiple switches 501 - 504 can provide communication between members of an architecture (such as an InfiniBand architecture) within a network switching environment 800 (e.g., an IB subnet). The architecture can include multiple hardware devices, such as host channel adapters 510, 520, 530. Each of the host channel adapters 510, 520, 530 can in turn interact with hypervisors 511, 521, and 531 respectively. Each hypervisor can in turn establish multiple virtual functions 514, 515, 516, 524, 525, 526, 534, 535, 536 in conjunction with the host channel adapter with which it interacts and assign the multiple virtual functions to multiple virtual machines. For example, virtual machine 1 550 can be assigned by hypervisor 511 to virtual function 1 514. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515. Hypervisor 521 can assign virtual machine 3 552 to virtual function 3526. Hypervisor 531 can in turn assign virtual machine 4 553 to virtual function 2 535. The hypervisors can access the host channel adapters through full-featured physical functions 513, 523, 533 on each of the host channel adapters.
[0137] According to an embodiment, each of the switches 501 - 504 can include multiple ports (not shown) that are used to set up a linear forwarding table to direct traffic within the network switching environment 800.
[0138] According to an embodiment, the virtual switches 512, 522, and 532 can be processed by their respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter (vHCA), which means that the VMs assigned to the VF are assigned a complete set of IB addresses (e.g., GID, GUID, LID) and dedicated QP space in the hardware. To the rest of the network and the SM (not shown), the HCAs 510, 520, and 530 appear as switches with additional nodes connected to them via the virtual switches.
[0139] According to an embodiment, the present disclosure provides systems and methods for providing a hybrid vSwitch architecture with dynamic LID allocation and pre-populated LIDs. Refer to Figure 9, the hypervisor 511 can be arranged with a vSwitch having a pre-populated LID architecture, and the hypervisor 521 can be arranged with a vSwitch having a pre-populated LID and dynamic LID allocation. The hypervisor 531 can be arranged with a vSwitch having dynamic LID allocation. Thus, the physical function 513 and virtual functions 514 - 516 have their LIDs pre-populated (i.e., even those virtual functions not attached to an active virtual machine are assigned LIDs). The physical function 523 and virtual function 1 524 can have their LIDs pre-populated, while virtual functions 2 and 3, 525 and 526 have their LIDs dynamically allocated (i.e., virtual function 2 525 is available for dynamic LID allocation, and since virtual machine 3 552 is attached, virtual function 3 526 has a dynamically allocated LID 11). Finally, the functions (physical and virtual) associated with hypervisor 3 531 can have their LIDs dynamically allocated. This makes virtual functions 1 and 3, 534 and 536 available for dynamic LID allocation, while virtual function 2 535 has a dynamically allocated LID 9 since virtual machine 4 553 is attached there.
[0140] According to embodiments such as Figure 9 shown, in which both a vSwitch with pre-populated LIDs and a vSwitch with dynamic LID allocation are utilized (independently or in combination within any given hypervisor), the number of pre-populated LIDs per host channel adapter can be defined by the architecture administrator and can be in the range of 0 <= pre-populated VFs <= total VFs (per host channel adapter), and the VFs available for dynamic LID allocation can be found by subtracting the number of pre-populated VFs from the total number of VFs (per host channel adapter).
[0141] According to an embodiment, very similar to how a physical host channel adapter can have more than one port (two ports are common for redundancy), a virtual HCA can also be represented with two ports and is connected to an external IB subnet via one, two, or more virtual switches.
[0142] InfiniBand - Inter - Subnet Communication (Fabric Manager)
[0143] According to an embodiment, in addition to providing an InfiniBand architecture within a single subnet, embodiments of the present disclosure can also provide an InfiniBand architecture spanning two or more subnets.
[0144] Figure 10Shows an example multi - subnet InfiniBand architecture according to an embodiment. As shown, within subnet A 1000, multiple switches 1001 - 1004 can provide communication between members of an architecture (such as an InfiniBand architecture) within subnet A 1000 (e.g., an IB subnet). The architecture can include multiple hardware devices, such as, for example, channel adapter 1010. The host channel adapter 1010 can in turn interact with a hypervisor 1011. The hypervisor can establish multiple virtual functions 1014 in conjunction with the host channel adapter with which it interacts. The hypervisor can additionally assign virtual machines to each virtual function, such as virtual machine 1 1015 being assigned to virtual function 1 1014. The hypervisor can access its associated host channel adapter through a full - featured physical function (such as physical function 1013) on each of the host channel adapters. Within subnet B 1040, multiple switches 1021 - 1024 can provide communication between members of an architecture (such as an InfiniBand architecture) within subnet B 1040 (e.g., an IB subnet). The architecture can include multiple hardware devices, such as, for example, channel adapter 1030. The host channel adapter 1030 can in turn interact with a hypervisor 1031. The hypervisor can establish multiple virtual functions 1034 in conjunction with the host channel adapter with which it interacts. The hypervisor can additionally assign virtual machines to each of the virtual functions, such as virtual machine 2 1035 being assigned to virtual function 2 1034. The hypervisor can access its associated host channel adapter through a full - featured physical function (such as physical function 1033) on each of the host channel adapters. It should be noted that while only one host channel adapter is shown within each subnet (i.e., subnet A and subnet B), it should be understood that multiple host channel adapters and their corresponding components can be included within each subnet.
[0145] According to an embodiment, each of the host channel adapters can additionally be associated with a virtual switch (such as virtual switches 1012 and 1032), and each HCA can be established with a different architecture model, as described above. While Figure 10 both subnets within are shown as using a vSwitch with a pre - populated LID architecture model, this is not meant to imply that all such subnet configurations can follow a similar architecture model.
[0146] According to an embodiment, at least one switch within each subnet can be associated with a router, such as switch 1002 within subnet A 1000 being associated with router 1005, and switch 1021 within subnet B 1040 being associated with router 1006.
[0147] According to an embodiment, at least one device (e.g., a switch, a node, etc.) can be associated with an architecture manager (not shown). The architecture manager can be used to, for example, discover the inter-subnet architecture topology, create architecture profiles (e.g., virtual machine architecture profiles), and build database objects related to virtual machines, which form the basis for building virtual machine architecture profiles. In addition, the architecture manager can define legal inter-subnet connectivity based on which subnets are allowed to communicate using which partition numbers via which router ports.
[0148] According to an embodiment, when traffic at a source origin (such as virtual machine 1 within subnet A) is addressed to a destination in a different subnet (such as virtual machine 2 within subnet B), the traffic can be addressed to a router within subnet A, i.e., router 1005, which can then forward the traffic to subnet B via its link with router 1006.
[0149] Virtual Dual - Port Router
[0150] According to an embodiment, the dual-port router abstraction can provide a simple way to enable the definition of subnet-to-subnet router functionality based on switch hardware implementations that have the ability to perform GRH (Global Routing Header) to LRH (Local Routing Header) conversion in addition to normal LRH-based switching.
[0151] According to an embodiment, a virtual dual-port router can be logically connected external to the corresponding switch port. Such a virtual dual-port router can provide a view compliant with the InfiniBand specification to standard management entities (such as subnet managers).
[0152] According to an embodiment, the dual-port router model means that different subnets can be connected in such a way that each subnet has complete control over the forwarding of packets and the address mapping in the ingress path to the subnet and does not affect the routing and logical connectivity within any incorrectly connected subnets.
[0153] According to an embodiment, in the case of an incorrectly connected architecture, using the virtual dual-port router abstraction can also allow management entities such as subnet managers and IB diagnostic software to behave correctly in the presence of an unexpected physical connectivity to a remote subnet.
[0154] Figure 11Illustrates the interconnection between two subnets in a high-performance computing environment according to an embodiment. Before being configured with a virtual dual-port router, the switch 1120 in subnet A 1101 can be connected to the switch 1130 in subnet B 1102 via the switch port 1121 of the switch 1120 through the physical connection 1110 and via the switch port 1131 of the switch 1130. In such an embodiment, each of the switch ports 1121 and 1131 can act as both a switch port and a router port.
[0155] According to an embodiment, the problem with this configuration is that management entities such as subnet managers in an InfiniBand subnet cannot distinguish physical ports that are both switch ports and router ports. In this case, the SM may regard the switch port as having a router port connected to the switch port. However, if the switch port is connected to another subnet via, for example, a physical link to another subnet manager, the subnet manager may be able to send discovery messages on the physical link. However, such discovery messages are not allowed at the other subnet.
[0156] Figure 12 Illustrates the interconnection between two subnets configured via a dual-port virtual router in a high-performance computing environment according to an embodiment.
[0157] According to an embodiment, after configuration, a dual-port virtual router configuration can be provided such that the subnet manager sees the correct end nodes, which indicate the ends of the subnets for which the subnet manager is responsible.
[0158] According to an embodiment, at the switch 1220 in subnet A 1201, the switch port can be connected (i.e., logically connected) to the router port 1211 in the virtual router 1210 via the virtual link 1223. The virtual router 1210 (e.g., a dual-port virtual router), which can be logically included within the switch 1220 when shown outside the switch 1220 in the embodiment, can also include a second router port: router port II 1212. According to an embodiment, the physical link 1203 that can have two ends can connect subnet A 1201 to subnet B 1202 via the second end of the physical link through the router port II 1212 and the router port II 1232 in the virtual router 1230 included in subnet B 1202. The virtual router 1230 can additionally include a router port 1231, which can be connected (i.e., logically connected) to the switch port 1241 on the switch 1240 via the virtual link 1233.
[0159] According to an embodiment, a subnet manager (not shown) on subnet A may detect router port 1211 on virtual router 1210 as an endpoint of the subnet controlled by the subnet manager. The dual-port virtual router abstraction may allow the subnet manager on subnet A to handle subnet A in a normal manner (e.g., as defined by the InifiniBand specification). At the subnet management agent level, the dual-port virtual router abstraction may be provided such that the SM sees a normal switch port and then, at the SMA level, sees that there is another port connected to the switch port and this port is an abstraction of the router port on the dual-port virtual router. In the local SM, the conventional architecture topology may continue to be used (the SM treats the port as a standard switch port in the topology), and thus the SM treats the router port as an end port. The physical connection may be made between two switch ports, which are also configured as router ports in two different subnets.
[0160] According to an embodiment, the dual-port virtual router may also solve the problem that a physical link may be incorrectly connected to some other switch port in the same subnet or to a switch port not intended to provide a connection to another subnet. Thus, the methods and systems described herein also provide a representation of what is external to the subnet.
[0161] According to an embodiment, within a subnet (such as subnet A), the local SM determines the switch port and then determines the router port connected to that switch port (e.g., router port 1211 connected to switch port 1221 via virtual link 1223). Since the SM treats router port 1211 as the end of the subnet managed by the SM, the SM cannot send discovery messages and / or management messages beyond this point (e.g., to router port II 1212).
[0162] According to an embodiment, the benefit provided by the above dual-port virtual router is that the dual-port virtual router abstraction is completely managed by a management entity (e.g., SM or SMA) within the subnet to which the dual-port virtual router belongs. By allowing management only on the local side, the system does not have to provide an external independent management entity. That is, each side of the subnet-to-subnet connection can be responsible for configuring its own dual-port virtual router.
[0163] According to an embodiment, in the case where a packet such as an SMP addressed to a remote destination (i.e., outside the local subnet) arrives at a local destination port not configured by the dual-port virtual router as described above, then the local port may return a message indicating that it is not a router port.
[0164] Many features of the present teachings may be implemented in hardware, software, firmware, or combinations thereof, executed with hardware, software, firmware, or combinations thereof, or facilitated by hardware, software, firmware, or combinations thereof. Accordingly, features of the present teachings may be implemented using a processing system (e.g., including one or more processors).
[0165] Figure 13 A method for supporting a dual-port virtual router in a high-performance computing environment according to an embodiment is shown. At step 1310, the method may provide, at one or more computers including one or more microprocessors, a first subnet including: a plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports; a plurality of host channel adapters, each host channel adapter including at least one host channel adapter port; a plurality of end nodes, wherein each end node of the end nodes is associated with at least one host channel adapter of the plurality of host channel adapters; and a subnet manager that runs on one of the plurality of switches and the plurality of host channel adapters.
[0166] At step 1320, the method may configure a switch port of the plurality of switch ports on a switch of the plurality of switches as a router port.
[0167] At step 1330, the method may logically connect the switch port configured as a router port to a virtual router that includes at least two virtual router ports.
[0168] Quality of Service and Service Level Agreement in a Private Fabric
[0169] According to an embodiment, high-performance computing environments such as switched networks running on InfiniBand or RoCE within a cloud and within a larger cloud installed at a customer's and on-premise have the ability to deploy virtual machine (VM)-based workloads, where an inherent requirement is that quality of service (QOS) can be defined and controlled for different types of traffic flows. Additionally, workloads belonging to different tenants must execute within the boundaries of relevant service level agreements (SLAs), while minimizing interference between such workloads and maintaining QOS assumptions for different communication types.
[0170] RDMA Read as a Restricted Feature (ORA200246 - US - NP - 1)
[0171] According to an embodiment, when defining bandwidth limits in a system using a conventional network interface card (NIC), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.
[0172] However, according to an embodiment, for RDMA-based networking where different nodes can generate RDMA read requests (i.e., egress bandwidth), this can represent a small amount of egress bandwidth. However, such RDMA read requests may potentially represent a very large amount of ingress RDMA traffic in response to such RDMA read requests. In such a case, to control the total traffic generation in the system, restricting the egress bandwidth of all nodes / VMs is no longer sufficient.
[0173] According to an embodiment, by making the RDMA read operation a restricted feature and only allowing such read requests from trusted nodes / VMs that do not generate excessive RDMA read-based ingress bandwidth, the total bandwidth utilization can be restricted while only restricting the transmission (egress) bandwidth of untrusted nodes / VMs.
[0174] Figure 14 A system for providing an RDMA read request as a restricted feature in a high-performance computing environment is shown according to an embodiment.
[0175] More specifically, according to an embodiment, Figure 14 A host channel adapter 1401 including a hypervisor 1411 is shown. The hypervisor can host / associate multiple virtual functions (VFs) such as VF 1414 - 1416 and a physical function (PF) 1413. The host channel adapter can additionally support / include multiple ports such as ports 1402 and 1403, which are used to connect the host channel adapter to a network such as network 1400. The network can include, for example, a switched network such as an InfiniBand network or a RoCE network, which can connect the HCA 1401 to multiple other nodes such as switches, additional and separate HCAs, etc.
[0176] According to an embodiment, as described above, each virtual function can host virtual machines (VMs) such as VM1 1450, VM2 1451, and VM3 1452.
[0177] According to an embodiment, the host channel adapter 1401 can additionally support a virtual switch 1412 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture, as described above.
[0178] According to an embodiment, the host channel adapter can implement a trusted RDMA read restriction 1460, whereby the read restriction 1460 can be configured to prevent any virtual machine (e.g., VM1, VM2, and / or VM3) from issuing any RDMA read requests to the network (e.g., via port 1402 or 1403).
[0179] According to an embodiment, the trusted RDMA read restriction 1460 can implement a host channel adapter-level block on generating (i.e., issuing) RDMA read request packets of certain types from certain endpoints such as virtual machines or other physical nodes connected to the network using the HCA 1401. The configurable restriction component 1460 can, for example, only allow trusted nodes (e.g., VMs or physical end nodes) to generate such packets.
[0180] According to an embodiment, the trusted RDMA read restriction component can be configured, for example, based on instructions received by the host channel adapter, or it can be directly configured, for example, by a subnet manager (not shown).
[0181] Figure 15 A system for providing RDMA read requests as a restricted feature in a high-performance computing environment is shown according to an embodiment.
[0182] More specifically, according to an embodiment, Figure 15 A host channel adapter 1501 including a hypervisor 1511 is shown. The hypervisor can host / associate multiple virtual functions (VFs) such as VF 1514 - 1516 and a physical function (PF) 1513. The host channel adapter can additionally support / include multiple ports such as ports 1502 and 1503 that are used to connect the host channel adapter to a network such as network 1500. The network can include, for example, a switched network such as an InfiniBand network or a RoCE network that can connect the HCA 1501 to multiple other nodes such as switches, additional and separate HCAs, etc.
[0183] According to an embodiment, as described above, each virtual function can host a virtual machine (VM) such as VM1 1550, VM2 1551, and VM3 1552.
[0184] According to an embodiment, the host channel adapter 1501 can additionally support a virtual switch 1512 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture as described above.
[0185] According to an embodiment, the host channel adapter can implement a trusted RDMA read restriction 1560, whereby the read restriction 1560 can be configured to block any virtual machine (e.g., VM1, VM2, and / or VM3) from issuing any RDMA read requests to the network (e.g., via ports 1502 or 1503).
[0186] According to an embodiment, the trusted RDMA read restriction 1560 can implement a host channel adapter-level block on generating (i.e., issuing) RDMA read request packets of certain types from certain endpoints such as virtual machines or other physical nodes connected to the network using the HCA 1501. The configurable restriction component 1560 can, for example, only allow trusted nodes (e.g., VMs or physical end nodes) to generate such packets.
[0187] According to an embodiment, the trusted RDMA read restriction component can be configured, for example, based on instructions received by the host channel adapter, or it can be directly configured by a subnet manager (not shown), for example.
[0188] According to an embodiment, as an example, the trusted RDMA read restriction 1560 can be configured to trust VM1 1550 and not trust VM2 1551. Thus, an RDMA read request 1554 initiated from VM1 can be allowed, while an RDMA read request 1555 initiated from VM2 can be blocked before the read request leaves the host channel adapter 1501 (but shown outside the HCA in the figure for drawing convenience only). Figure 15 within the HCA, which is shown merely for drawing convenience).
[0189] Figure 16 A system for providing RDMA read requests as a restricted feature in a high-performance computing environment according to an embodiment is shown.
[0190] According to an embodiment, within a high-performance computing environment such as a switched network or subnet 1600, multiple end nodes 1601 and 1602 can support multiple virtual machines VM1-VM4 1650-1653, which are interconnected via multiple switches such as leaf switches 1611 and 1612, switches 1621 and 1622, and root switches 1631 and 1632.
[0191] According to an embodiment, not shown in the figure are various host channel adapters that provide the functionality for connecting the nodes 1601 and 1602 and the virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine can be associated with a virtual function of the hypervisor on the host channel adapter.
[0192] According to an embodiment, in a general system, the RDMA egress bandwidth of any one virtual machine from an end node is restricted to prevent any one virtual machine from monopolizing the bandwidth of any link connecting the end node to a subnet. However, although this egress bandwidth restriction is effective in general, it does not prevent a virtual machine from issuing RDMA read requests, such as RDMA read requests 1654 and 1655. This is because such RDMA read requests are typically small packets and are of little use in utilizing the egress bandwidth.
[0193] However, according to an embodiment, such RDMA read requests may result in a large amount of return traffic being generated to publishing entities such as VM1 and VM3. In this case, for example, when read request 1654 causes a large amount of data traffic to flow back to VM1 due to the execution of the read request at the destination, then the RDMA read request may cause link congestion and a degradation in network performance.
[0194] According to an embodiment, especially in the case where more than one tenant shares subnet 1600, this may result in a performance loss of the subnet.
[0195] According to an embodiment, each node (or host channel adapter) may be configured with RDMA read restrictions 1660 and 1661, which will prevent any VM from issuing an RDMA read request when the VM is not trusted. Such RDMA read restrictions may vary from a permanent block of issuing an RDMA read request to setting a time range limit on when a virtual machine configured with an RDMA read request restriction can issue an RDMA read request (e.g., during slow network traffic). Additionally, the RDMA read restrictions 1660 and 1661 may additionally allow trusted VMs to issue RDMA read requests.
[0196] According to an embodiment, since there may be a scenario where multiple VMs / tenants are sharing a "new" HCA - that is, an HCA that supports relevant new features, but is executing an RDMA request to a remote "old" HCA that does not have such support, it makes sense to find a way to limit the ingress bandwidth that such VMs can generate in terms of RDMA read responses, without relying on a static rate configuration on the "old" RDMA read responder HCA. As long as VMs are allowed to generate "any" RDMA read size, there is no straightforward way to do this. Additionally, since multiple RDMA read requests generated over a period of time may in principle all receive response data simultaneously, it is not possible to ensure that the ingress bandwidth cannot exceed the maximum bandwidth within a very limited amount of time, unless there are restrictions on the RDMA read size that can be generated in a single request and on the total number of outstanding RDMA read requests from the same vHCA port.
[0197] Thus, according to an embodiment, assuming a maximum read size is defined for the vHCA, bandwidth control can be based on a quota for the sum of all outstanding read sizes, or a simpler scheme could be to limit the maximum number of outstanding RDMA reads based only on the "worst-case" read size. Thus, in either case, there is no limit to the peak bandwidth within a short interval (except for the maximum link bandwidth of the HCA port), but the duration of such a peak bandwidth "window" will be limited. However, in addition, the send rate of RDMA read requests must also be throttled so that the send rate of requests does not exceed the allowed maximum ingress rate (assuming responses with data are received at the same rate). In other words, the maximum outstanding request limit defines the worst-case short-interval bandwidth, and the request send rate limit will ensure that new requests cannot be generated immediately upon receipt of a response, but only after a relevant delay representing the acceptable average ingress bandwidth of the RDMA read response. Thus, in the worst case, the allowed number of requests have been sent without any responses, and then all of these responses are received "simultaneously". At this point, the next request can be sent immediately upon arrival of the first response, but subsequent requests will have to be delayed by the specified delay period. Thus, over time, the average ingress bandwidth cannot exceed the bandwidth defined by the request rate. However, a smaller maximum number of outstanding requests will reduce the possible "variance".
[0198] Using Explicit RDMA Read Bandwidth Limitation (ORA200246 - US - NP - 1)
[0199] According to an embodiment, when defining bandwidth limits in a system using a conventional network interface card (NIC), it is generally sufficient to control the egress bandwidth that each node / VM is allowed to generate on the network.
[0200] However, according to an embodiment, for RDMA-based networking where different nodes can generate RDMA read requests that represent few request messages but potentially very many response messages, limiting the egress bandwidth of all nodes / VMs to control the total traffic generation in the system is no longer sufficient.
[0201] According to an embodiment, the total traffic generation in the system can be controlled by defining an explicit quota for how much RDMA read ingress bandwidth a node / VM is allowed to generate independent of any send / egress bandwidth limits, without relying on restricting the use of RDMA reads by untrusted nodes / VMs.
[0202] According to an embodiment, in addition to supporting the average ingress bandwidth utilization due to locally generated RDMA read requests, the system and method can also support (i.e., as a result of "stacking" of RDMA read responses) the worst-case duration / length of a maximum link bandwidth burst.
[0203] Figure 17Illustrates a system for providing explicit RDMA read bandwidth limits in a high performance computing environment, according to an embodiment.
[0204] More specifically, according to an embodiment, Figure 17 Illustrates a host channel adapter 1701 that includes a hypervisor 1711. The hypervisor may host / associate multiple virtual functions (VFs) (such as VFs 1714 - 1716) and a physical function (PF) 1713. The host channel adapter may additionally support / include multiple ports, such as ports 1702 and 1703, which are used to connect the host channel adapter to a network, such as network 1700. The network may include, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 1701 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0205] According to an embodiment, as described above, each virtual function may host virtual machines (VMs), such as VM1 1750, VM2 1751, and VM3 1752.
[0206] According to an embodiment, the host channel adapter 1701 may additionally support a virtual switch 1712 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0207] According to an embodiment, the host channel adapter may implement an RDMA read limit 1760, whereby the read limit 1760 may be configured to set a quota on the amount of ingress bandwidth that any VM (of the HCA 1701) may generate in response to an RDMA read request issued by a particular VM. Limiting such ingress bandwidth is performed locally at the host channel adapter.
[0208] According to an embodiment, the RDMA read limit component may be configured, for example, based on instructions received by the host channel adapter, or it may be directly configured, for example, by a subnet manager (not shown).
[0209] Figure 18 Illustrates a system for providing explicit RDMA read bandwidth limits in a high performance computing environment, according to an embodiment.
[0210] More specifically, according to an embodiment, Figure 18Shows a host channel adapter 1801 including a hypervisor 1811. The hypervisor can host / associate multiple virtual functions (VFs) such as VFs 1814 - 1816 and a physical function (PF) 1813. The host channel adapter can additionally support / include multiple ports such as ports 1802 and 1803 which are used to connect the host channel adapter to a network such as network 1800. The network can include for example a switched network such as an InfiniBand network or a RoCE network which can connect the HCA 1801 to multiple other nodes such as switches, additional and separate HCAs, etc.
[0211] According to an embodiment, as described above, each virtual function can host a virtual machine (VM) such as VM1 1850, VM2 1851, and VM3 1852.
[0212] According to an embodiment, the host channel adapter 1801 can additionally support a virtual switch 1812 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture as described above.
[0213] According to an embodiment, the host channel adapter can implement an RDMA read limit 1860, whereby the read limit 1860 can be configured to set a quota on the amount of ingress bandwidth that any VM (of HCA 1701) can generate in response to an RDMA read request issued by a particular VM. Limiting such ingress bandwidth is performed locally at the host channel adapter.
[0214] According to an embodiment, the RDMA read limit component can be configured, for example, based on instructions received by the host channel adapter, or it can be directly configured by a subnet manager (not shown) for example.
[0215] According to an embodiment, for example, VM1 may have previously issued at least two RDMA read requests to request read operations to be performed on a connected node. In response, VM1 may be in the process of receiving multiple responses to the RDMA read requests (shown as RDMA read responses 1855 and 1854 in the figure). Since these RDMA read responses may be very large, especially compared to the RDMA read requests initially sent by VM1, these read responses 1854 and 1855 may be subject to the RDMA read limit 1860 and the ingress bandwidth may be limited or throttled. This throttling can be based on an explicit ingress bandwidth limit, or it can be based on the QoS and / or SLA of VM1 set within the RDMA limit 1860.
[0216] Figure 19A system for providing explicit RDMA read bandwidth limits in a high performance computing environment, according to an embodiment, is shown.
[0217] According to an embodiment, within a high performance computing environment such as a switched network or subnet 1900, multiple end nodes 1901 and 1902 may support multiple virtual machines VM1-VM4 1950-1953, which are interconnected via multiple switches such as leaf switches 1911 and 1912, switches 1921 and 1922, and root switches 1931 and 1932.
[0218] Not shown in the figure, according to an embodiment, are various host channel adapters that provide the functionality for connecting nodes 1901 and 1902 and virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of a hypervisor on a host channel adapter.
[0219] According to an embodiment, in a general system, the RDMA egress bandwidth of any one virtual machine from an end node is restricted to prevent any one virtual machine from monopolizing the bandwidth of any link connecting the end node to the subnet. However, while this egress bandwidth restriction is effective in general, it does not prevent an influx of RDMA read responses from monopolizing the link between the requesting VM and the network.
[0220] According to an embodiment, in other words, if VM1 issues multiple RDMA read requests, then VM1 has no control over when the responses to these read requests are returned to VM1. This can lead to a backup / accumulation of responses to RDMA read requests, each of which attempts to return the requested information (via RDMA read response 1954) to VM1 using the same link. This results in traffic congestion and backlog in the network.
[0221] According to an embodiment, RDMA limits 1960 and 1961 may set quotas on the amount of ingress bandwidth that a VM may generate with respect to responses to RDMA read requests issued by a particular VM. Restricting such ingress bandwidth is performed locally.
[0222] According to an embodiment, assuming that a maximum read size is defined for the vHCA, bandwidth control can be based on a quota for the sum of all outstanding read sizes, or a simpler scheme can be to limit the maximum number of outstanding RDMA reads only based on the "worst-case" read size. Thus, in either case, there is no limit to the peak bandwidth within a short interval (except for the maximum link bandwidth of the HCA port), but the duration of such a peak bandwidth "window" will be limited. However, in addition, the sending rate of RDMA read requests must also be throttled so that the sending rate of requests does not exceed the allowed maximum ingress rate (assuming that responses with data are received at the same rate). In other words, the maximum outstanding request limit defines the worst-case short-interval bandwidth, and the request sending rate limit will ensure that new requests cannot be generated immediately upon receiving a response, but only after a relevant delay representing the acceptable average ingress bandwidth of the RDMA read response. Thus, in the worst case, the allowed number of requests has been sent without any responses, and then all these responses are received "simultaneously". At this point, the next request can be sent immediately upon the arrival of the first response, but subsequent requests will have to be delayed by the specified delay time. Therefore, over time, the average ingress bandwidth cannot exceed the bandwidth defined by the request rate. However, a smaller maximum number of outstanding requests will reduce the possible "variance".
[0223] Figure 20 is a flowchart of a method for providing RDMA (Remote Direct Memory Access) read requests as a restricted feature in a high-performance computing environment according to an embodiment.
[0224] According to an embodiment, at step 2010, the method can provide a first subnet at one or more microprocessors, the first subnet including a plurality of switches, a plurality of host channel adapters, where each of the host channel adapters includes at least one host channel adapter port, and where the plurality of host channel adapters are interconnected via the plurality of switches.
[0225] According to an embodiment, at step 2020, the method can provide a plurality of end nodes, the plurality of end nodes including a plurality of virtual machines.
[0226] According to an embodiment, at step 2030, the method can associate the host channel adapter with a selective RDMA restriction.
[0227] According to an embodiment, at step 2040, the method can host a virtual machine among the plurality of virtual machines at the host channel adapter including the selective RDMA restriction.
[0228] Combining Multiple Shared Bandwidth Segments (ORA200246 - US - NP - 3)
[0229] According to an embodiment, conventional bandwidth / rate limiting schemes for network interfaces are typically limited to the combined total transmission rate and possibly the maximum rate for each individual destination. However, in many cases, there are shared bottlenecks in the intermediate network / structural topology, which means that the total bandwidth available to a group of destinations is limited by this shared bottleneck. Therefore, unless such shared bottlenecks are taken into account when deciding at what rate to send various data streams, although individual destination rate limits are observed, the shared bottleneck is also likely to become overloaded.
[0230] According to an embodiment, the systems and methods herein can introduce an object "destination group", to which multiple individual streams can be associated, and wherein the destination group can represent the rate limits of separate (possibly shared) links or other bottlenecks within the network / architecture path being used by the individual streams. Additionally, the systems and methods can allow each stream to be associated with a hierarchy of such destination groups, such that all link segments and any other (shared) bottlenecks in the path between the sender and the destination of the individual stream can be represented.
[0231] According to an embodiment, to limit the egress bandwidth, the system and method can establish destination groups that share bandwidth quotas in order to reduce the likelihood of congestion on shared ISLs (Inter-Switch Links). This requires a lookup mechanism related to destinations / paths, which can be managed based on which destinations / paths will map to which groups at the logical level. This means that the hyper-privileged communication infrastructure must know the actual locations of peer nodes in the architecture topology and the relevant routing and capacity information that can be mapped to "destination groups" (i.e., HCA-level object types) within a local HCA with an associated bandwidth quota. However, it is impractical to have the hardware platform directly look up WQE (Work Queue Entry) / packet address information to map to the relevant destination group(s). Instead, the HCA implementation can provide an association between RC (Reliable Connection) QP (Queue Pair) and address handles, which represents the sending context of the outgoing traffic and the relevant destination group. In this way, the association can be transparent at the verb level and can alternatively be set at the hyper-privileged software level and subsequently enforced at the HCA hardware (and firmware) level. An important additional complexity associated with this scenario is that live VM migration, which maintains the real-time VM or vHCA port address information throughout the migration process, may still mean that the destination groups of different communication peers change. However, as long as the system and method tolerate some transient periods where the relevant bandwidth quotas are not 100% correct, the destination group associations do not have to be updated synchronously. Therefore, although the logical connections and communication capabilities may not change due to VM migration, the destination group(s) associated with the RC connections and address handles in the migrating VM and its communicating peer VMs may be "completely wrong" after migration. This can mean using less bandwidth than is available (e.g., when a VM moves from a remote location to the same "leaf group" as its peer) and can also mean generating excessive bandwidth (e.g., when a VM moves from the same "leaf group" as its peer to a remote location, which means the shared ISL has limited bandwidth).
[0232] According to an embodiment, the bandwidth quota specific to a destination group can also be divided into quotas for specific priorities ("QoS classes") in principle, in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the destination group.
[0233] According to an embodiment, the destination group decouples the object from the specific destination address, and the system and method obtain the ability to represent intermediate, shared links or link groups, which can represent bandwidth limitations other than the destination and may be more restrictive than the destination limitations.
[0234] According to an embodiment, the system and method may consider a hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if the target is limited to 30 Gb / s and the intermediate uplink is limited to 50 Gb / s, then the maximum rate for the target will never exceed 30 Gb / s. On the other hand, if multiple 30 Gb / s targets share the same 50 Gb / s intermediate limit, then using the relevant target rate limits for the flows towards these targets may mean exceeding the intermediate rate limit. Therefore, to ensure optimal utilization and throughput within the relevant limits, all target groups in the relevant hierarchy can be considered in a relevant strict order. This means that packets can be sent towards the relevant destination if and only if each target group in the hierarchy represents available bandwidth. Thus, if a single flow is active for one of the targets in the above example, then that flow will be allowed to operate at 30 Gb / s. However, once another flow becomes active for another target (via a shared intermediate target group), then each flow will be limited to 25 Gb / s. If in the next round, an additional flow towards one of the two targets becomes active, then the two flows towards the same target will each operate at 12.5 Gb / s (i.e., on average, unless they have any additional bandwidth quotas / limits).
[0235] According to an embodiment, when multiple tenants share a server / HCA, in addition to sharing any intermediate ISL bandwidth, the initial egress bandwidth and the actual target bandwidth can also be shared. On the other hand, in a scenario where each tenant has a dedicated server / HCA, the intermediate ISL bandwidth represents the only possible "inter-tenant" bandwidth sharing.
[0236] According to an embodiment, the target group should generally be global for an HCA port, and the VF / tenant quota at the HCA level will represent the maximum local traffic that a tenant can generate globally or for any combination of specific priorities for the target. Nevertheless, it is still possible to use target groups specific to some tenants as well as "global" target groups in the same hierarchy.
[0237] According to an embodiment, there are several possible ways to implement the target group and the target group association (hierarchy) representing a specific QP or address handle. However, a 16-bit target group ID space and support for up to 4 or 8 target group associations can be provided for each QP and address handle. Then, each target group ID value will represent some hardware state reflecting the relevant IPD (inter-packet delay) value for the relevant rate and timer information defining when the next packet associated with that target group can be sent.
[0238] According to an embodiment, since different flows / paths can use different "QOS IDs" (i.e., service levels, priorities, etc.) on the same shared link segment, different target groups can also be associated with the same link segment, such that different target groups represent bandwidth quotas for such different QOS IDs. However, it is also possible to represent a target group specific to a QOS ID and a single target group representing the physical link of the same link segment.
[0239] According to an embodiment, similarly, the system and method can additionally distinguish different flow types defined by explicit flow type packet header parameters and / or by considering the type of operation (e.g., RDMA read / write / send) in order to implement different "sub-quotas" for arbitration between different such flow types. In particular, this may be useful for distinguishing flows representing responder mode bandwidth (i.e., typically RDMA read response traffic) from requester mode traffic initially initiated by the local node itself.
[0240] According to an embodiment, in the case of strictly using target groups and the rate limits of all relevant sender HCAs adding up to a total maximum rate that does not exceed the capacity of any target or shared ISL segment, "any" congestion can be avoided in principle. However, this may imply severe restrictions on the low average utilization of the continuous bandwidth of different flows and the available link bandwidth. Therefore, various rate limits can be set to allow different HCAs to use more optimistic maximum rates. In this case, the aggregated sum is higher than the sustainable maximum and may therefore result in congestion.
[0241] Figure 21 A system for combining multiple shared bandwidth segments in a high-performance computing environment according to an embodiment is shown.
[0242] More specifically, according to an embodiment, Figure 21 A host channel adapter 2101 including a hypervisor 2111 is shown. The hypervisor can host / associate multiple virtual functions (VFs) (such as VFs 2114 - 2116) and a physical function (PF) 2113. The host channel adapter can additionally support / include multiple ports, such as ports 2102 and 2103, which are used to connect the host channel adapter to a network, such as network 2100. The network can include, for example, a switched network, such as an InfiniBand network or a RoCE network, which can connect the HCA 2101 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0243] According to an embodiment, as described above, each virtual function can host virtual machines (VMs), such as VM1 2150, VM2 2151, and VM3 2152.
[0244] According to an embodiment, the host channel adapter 2101 may additionally support a virtual switch 2112 via a hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0245] According to an embodiment, as shown, the network 2100 may include a plurality of switches, such as switches 2140, 2141, 2142, and 2143, which are interconnected and may be connected to the host channel adapter 2101, for example, via leaf switches 2140 and 2141.
[0246] According to an embodiment, switches 2140-2143 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0247] According to an embodiment, target groups, such as target groups 2170 and 2171, may be defined along inter-switch links (ISLs), such as the ISLs between leaf switch 2140 and switch 2142 and between leaf switch 2141 and switch 2143. For example, these target groups 2170 and 2171 may represent bandwidth quotas as HCA objects, which are stored at a target group repository 2161 associated with the HCA and may be accessed by a rate limiting component 2160.
[0248] According to an embodiment, target groups 2170 and 2171 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be divided into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0249] According to an embodiment, target groups 2170 and 2171 decouple objects from specific destination addresses, and the system and method obtain the ability to represent intermediate, shared links or link groups, which may represent bandwidth limitations other than targets and may be more restrictive than target limitations. That is, for example, if the default / original egress limit on VM2 2151 is set to a threshold, but the destination of packets sent from VM2 will pass through target group 2170, which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be limited to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA may be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in the routing of packets from VM2.
[0250] According to an embodiment, the target groups can also be hierarchical in nature, whereby the system and method can consider a hierarchy of target groups (bandwidth quotas) that reflects bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2170 represents a higher bandwidth limit than the bandwidth limit of target group 2171, and a packet is addressed across the two switch - to - switch links represented by the two target groups, then the bandwidth limit of target group 2171 is the controlling bandwidth - limiting factor.
[0251] According to an embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flows, the bandwidth quota represented by the target group can be divided. As an example, if both VM1 and VM2 are simultaneously sending flows that will involve target group 2170, and target group 2170 represents, for example, a 10 Gb / s bandwidth quota, and the corresponding flows have equal QoS and SLA associated with them, then target group 2170 will represent a 5 Gb / s limit for each flow. This sharing or division of the target - group bandwidth quota can vary based on the QoS and SLA associated with the respective flows.
[0252] Figure 22 A system for combining multiple shared - bandwidth segments in a high - performance computing environment according to an embodiment is shown.
[0253] More specifically, according to an embodiment, Figure 22 A host channel adapter 2201 including a hypervisor 2211 is shown. The hypervisor can host / associate multiple virtual functions (VFs) (such as VF 2214 - 2216) as well as a physical function (PF) 2213. The host channel adapter can additionally support / include multiple ports, such as ports 2202 and 2203, which are used to connect the host channel adapter to a network, such as network 2200. The network can include, for example, a switched network, such as an InfiniBand network or a RoCE network, which can connect the HCA 2201 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0254] According to an embodiment, as described above, each virtual function can host a virtual machine (VM), such as VM1 2250, VM2 2251, and VM3 2252.
[0255] According to an embodiment, the host channel adapter 2201 can additionally support a virtual switch 2212 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture, as described above.
[0256] According to an embodiment, as shown, network 2200 may include multiple switches, such as switches 2240, 2241, 2242, and 2243, which are interconnected and may be connected to host channel adapter 2201 via, for example, leaf switches 2240 and 2241.
[0257] According to an embodiment, switches 2240 - 2243 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0258] According to an embodiment, target groups, such as target groups 2270 and 2271, may be defined at, for example, switch ports. As shown, target groups 2270 and 2271 are defined at the switch ports of switches 2242 and 2243, respectively. For example, these target groups 2270 and 2271 may represent bandwidth quotas as HCA objects, which are stored at target group repository 2261 associated with the HCA and accessible by rate limiting component 2260.
[0259] According to an embodiment, target groups 2270 and 2271 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be divided into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0260] According to an embodiment, target groups 2270 and 2271 decouple objects from specific destination addresses, and the system and method obtain the ability to represent intermediate, shared links or link groups, which may represent bandwidth limitations other than the target and may be more restrictive than the target limitations. That is, for example, if the default / original egress limit on VM2 2251 is set to a threshold, but the destination of packets sent from VM2 will pass through target group 2270 with a lower bandwidth limit set, then the egress bandwidth from VM2 may be restricted to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA may be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in the routing of packets from VM2.
[0261] According to an embodiment, target groups may also be hierarchical in nature, whereby the system and method may consider the hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2270 represents a higher bandwidth limit than the bandwidth limit of target group 2271, and a packet is addressed across the two switch - to - switch links represented by the two target groups, then the bandwidth limit of target group 2271 is the controlling bandwidth - limiting factor.
[0262] According to an embodiment, a target group may also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flow, the bandwidth quota represented by the target group may be divided. As an example, both VM1 and VM2 simultaneously send flows that will involve target group 2270, which represents a bandwidth quota of, for example, 10 Gb / s, and the respective flows have equal QoS and SLA associated with them. Then target group 2270 will represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth quota may vary based on the QoS and SLA associated with the respective flow.
[0263] According to an embodiment, Figure 21 and Figure 22 respectively illustrate target groups defined at the inter-switch link and the switch port. Those of ordinary skill in the art will readily understand that target groups may be defined at different locations within a subnet, and any given subnet is not limited to defining target groups only at the ISL and the switch port. Generally speaking, such target groups may be defined at the ISL and the switch port within any given subnet.
[0264] Figure 23 Illustrates a system for combining multiple shared bandwidth segments in a high-performance computing environment according to an embodiment.
[0265] More specifically, according to an embodiment, Figure 23 Illustrates a host channel adapter 2301 including a hypervisor 2311. The hypervisor may host / associate multiple virtual functions (VFs) (such as VFs 2314 - 2316) and a physical function (PF) 2313. The host channel adapter may additionally support / include multiple ports, such as ports 2302 and 2303, which are used to connect the host channel adapter to a network, such as network 2300. The network may include, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 2301 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0266] According to an embodiment, as described above, each virtual function may host a virtual machine (VM), such as VM1 2350, VM2 2351, and VM3 2352.
[0267] According to an embodiment, the host channel adapter 2301 may additionally support a virtual switch 2312 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0268] According to an embodiment, as shown in the figure, network 2300 may include a plurality of switches, such as switches 2340, 2341, 2342, and 2343, which are interconnected and may be connected to host channel adapter 2301 via, for example, leaf switches 2340 and 2341.
[0269] According to an embodiment, switches 2340 - 2343 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0270] According to an embodiment, target groups, such as target groups 2370 and 2371, may be defined along inter - switch links (ISLs), such as the ISLs between leaf switch 2340 and switch 2342 and between leaf switch 2341 and switch 2343. For example, these target groups 2370 and 2371 may represent bandwidth quotas as HCA objects, which are stored at target group repository 2361 associated with the HCA and may be accessed by rate limiting component 2360.
[0271] According to an embodiment, target groups 2370 and 2371 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be divided into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0272] According to an embodiment, target groups 2370 and 2371 decouple objects from specific destination addresses, and the system and method obtain the ability to represent intermediate, shared links or link groups, which may represent bandwidth limitations other than the target and may be more restrictive than the target limitations. That is, for example, if the default / original egress limit on VM2 2351 is set to a threshold, but the destination of the packets sent from VM2 will pass through target group 2370, which has a lower bandwidth limit set, then the egress bandwidth from VM2 may be limited to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA may be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in the routing of the packets from VM2.
[0273] According to an embodiment, the target groups can also be hierarchical in nature, whereby the system and method can consider a hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2370 represents a higher bandwidth limit than the bandwidth limit of target group 2371, and a packet is addressed across the two switch - to - switch links represented by the two target groups, then the bandwidth limit of target group 2371 is the controlling bandwidth - limiting factor.
[0274] According to an embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flows, the bandwidth quota represented by the target group can be divided. As an example, if both VM1 and VM2 are simultaneously sending flows that will involve target group 2370, and target group 2370 represents a bandwidth quota of, for example, 10 Gb / s, and the corresponding flows have equal QoS and SLA associated with them, then target group 2370 will represent a 5 Gb / s limit for each flow. This sharing or division of the target - group bandwidth quota can vary based on the QoS and SLA associated with the respective flows.
[0275] According to an embodiment, the target - group repository can query target group 2370, for example, to determine the bandwidth quota of the target group. After determining the bandwidth quota of the target group, the target - group repository can store the quota value associated with the target group. Then, the rate - limiting component can use this quota to a) determine whether the bandwidth quota of the target group is lower than the bandwidth quota for the VM based on QoS or SLA, and b) when making such a determination, update the bandwidth quota for the VM based on the path through target group 2370.
[0276] Figure 24 A system for combining multiple shared - bandwidth segments in a high - performance computing environment according to an embodiment is shown.
[0277] More specifically, according to an embodiment, Figure 24 A host - channel adapter 2401 including a hypervisor 2411 is shown. The hypervisor can host / associate multiple virtual functions (VFs) (such as VFs 2414 - 2416) and a physical function (PF) 2413. The host - channel adapter can additionally support / include multiple ports, such as ports 2402 and 2403, which are used to connect the host - channel adapter to a network, such as network 2400. The network can include, for example, a switched network, such as an InfiniBand network or a RoCE network, which can connect the HCA 2401 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0278] According to an embodiment, as described above, each virtual function may host a virtual machine (VM), such as VM1 2450, VM2 2451, and VM3 2453.
[0279] According to an embodiment, the host channel adapter 2401 may additionally support a virtual switch 2412 via a hypervisor. This is for the case of implementing the vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0280] According to an embodiment, as shown, the network 2400 may include a plurality of switches, such as switches 2440, 2441, 2442, and 2443, which are interconnected and may be connected to the host channel adapter 2401 via, for example, leaf switches 2440 and 2441.
[0281] According to an embodiment, switches 2440 - 2443 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0282] According to an embodiment, target groups, such as target groups 2470 and 2471, may be defined at, for example, switch ports. As shown, target groups 2470 and 2471 are defined at the switch ports of switches 2442 and 2443, respectively. For example, these target groups 2470 and 2471 may represent bandwidth quotas as HCA objects, which are stored at a target group repository 2461 associated with the HCA and may be accessed by a rate limiting component 2460.
[0283] According to an embodiment, target groups 2470 and 2471 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be divided into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0284] According to an embodiment, target groups 2470 and 2471 decouple the object from a specific destination address, and the system and method obtain the ability to represent intermediate, shared links or link groups, which may represent bandwidth limitations other than the target and may be more restrictive than the target limitations. That is, for example, if the default / original egress limit on VM2 2451 is set to a threshold, but the destination of the packets sent from VM2 will pass through target group 2470, which sets a lower bandwidth limit, then the egress bandwidth from VM2 may be restricted to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA may be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in the routing of the packets from VM2.
[0285] According to an embodiment, the target groups may also be hierarchical in nature, whereby the system and method may consider a hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2470 represents a higher bandwidth limit than the bandwidth limit of target group 2471, and a packet is addressed across an inter-switch link represented by both target groups, then the bandwidth limit of target group 2471 is the controlling bandwidth limiting factor.
[0286] According to an embodiment, a target group may also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flows, the bandwidth quota represented by the target group may be divided. As an example, both VM1 and VM2 are simultaneously sending flows that will involve target group 2470, which represents a bandwidth quota of, for example, 10 Gb / s, and the respective flows have equal QoS and SLA associated with them, then target group 2470 will represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth quota may be based on changes in the QoS and SLA associated with the respective flows.
[0287] According to an embodiment, the target group repository may query target group 2470 to determine, for example, the bandwidth quota of the target group. After determining the bandwidth quota of the target group, the target group repository may store the quota value associated with the target group. Then, the rate limiting component may use this quota to a) determine whether the bandwidth quota of the target group is lower than the bandwidth quota for the VM based on QoS or SLA, and b) when making such a determination, update the bandwidth quota for the VM based on the path through target group 2470.
[0288] Figure 25 A system for combining multiple shared bandwidth segments in a high performance computing environment according to an embodiment is shown.
[0289] According to an embodiment, within a high performance computing environment such as a switched network or subnet 2500, multiple end nodes 2501 and 2502 may support multiple virtual machines VM1-VM4 2550-2553, which are interconnected via multiple switches such as leaf switches 2511 and 2512, switches 2521 and 2522, and root switches 2531 and 2532.
[0290] According to an embodiment, not shown in the figure are the various host channel adapters that provide the functionality for connecting nodes 2501 and 2502 as well as the virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0291] According to an embodiment, as discussed above, the concept inherent in such a switching fabric is that while each end node or VM may have its own egress / ingress bandwidth limits that the incoming and outgoing traffic must abide by, there may also be links or ports within the subnet that represent bottlenecks for the traffic flowing therein. Thus, when determining at what rate traffic should flow in and out of such end nodes (such as VM1, VM2, VM3, or VM4), the rate limiting components 2560 and 2561 can query various destination groups, such as 2550 and 2551, to determine whether such destination groups represent bottlenecks for the traffic flow. After such a determination, the rate limiting components 2560 and 2561 can then set different or new bandwidth limits at the endpoints that the rate limiting components can control.
[0292] In addition, according to an embodiment, the destination groups can be queried in a nested / hierarchical manner such that if traffic from VM1 to VM3 is to utilize the destination groups 2550 and 2551, then the rate limiting 2560 can consider the limits from both such destination groups when determining the bandwidth limit from VM1 to VM3.
[0293] Figure 26 is a flowchart of a method according to an embodiment for supporting destination groups in a high performance computing environment for congestion control in a private fabric.
[0294] According to an embodiment, at step 2610, the method can provide a first subnet at one or more microprocessors, the first subnet including: a plurality of switches, the plurality of switches including at least one leaf switch, wherein each switch of the plurality of switches includes a plurality of switch ports; a plurality of host channel adapters, wherein each host channel adapter of the plurality of host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches; and a plurality of end nodes, including a plurality of virtual machines.
[0295] According to an embodiment, at step 2620, the method can define a destination group on an inter-switch link between two switches of the plurality of switches or on at least one of the ports of a switch of the plurality of switches, wherein the destination group defines a bandwidth limit on an inter-switch link between two switches of the plurality of switches or on at least one of the ports of a switch of the plurality of switches.
[0296] According to an embodiment, at step 2630, the method can provide a destination group repository stored in a memory of the host channel adapter at the host channel adapter.
[0297] According to an embodiment, at step 2640, the method can record the defined destination group in the destination group repository.
[0298] Combining Destination - Specific Send / RDMA Write and RDMA Read Bandwidth Limitation (ORA200246 - US - NP - 2)
[0299] According to an embodiment, a node / VM can be the target of incoming data traffic, which is the result of both send and RDMA-write operations initiated by peer nodes / VMs and RDMA-read operations initiated by the local node / VM itself. In such a case, ensuring that the maximum or average ingress bandwidth of the local node / VM is within the desired bounds can be problematic unless all these flows are coordinated in terms of rate limiting.
[0300] According to an embodiment, the systems and methods described herein can be implemented to achieve target-specific egress rate control in a manner that allows all flows representing data fetched from local memory and sent to associated remote targets to be subject to the same shared rate limit and associated flow scheduling and arbitration. Additionally, different flow types can be given different priorities and / or different shares of available bandwidth.
[0301] According to an embodiment, as long as the target group association of flows from a "producer / sender" node implies bandwidth regulation for all outgoing data packets - including UD (unreliable datagram) sends, RDMA writes, RDMA sends, and RDMA reads (i.e., RDMA read responses with data) - there is full control over all ingress bandwidth of the vHCA port. This is independent of whether the VM owning the target vHCA port is generating an "excessive" number of RDMA read requests to multiple peer nodes.
[0302] According to an embodiment, the coupling of a target group with flow- and "unsolicited" BECN signaling specific means that the ingress bandwidth of each vHCA port can be dynamically throttled for any number of remote peers.
[0303] According to an embodiment, in addition to conveying pure CE marking / unmarking for different phase numbers, "unsolicited BECN" messages can also be used to convey specific rate values. In this way, it is possible to have a scheme where initial incoming packets from a new peer (e.g., communication management (CM) packets) can trigger the generation of one or more "unsolicited BECN" messages to the HCA from which the incoming packet originated (i.e., the associated firmware / hyper-privileged software) and to the current communication peer.
[0304] According to an embodiment, in the case of using two ports on an HCA simultaneously (i.e., an active-active scenario), it may make sense to share target groups between local HCA ports if concurrent flows are likely to share some ISLs or even target the same destination port.
[0305] According to an embodiment, another reason to share a target group among HCA ports is if the HCA local memory bandwidth cannot sustain the full link speed of two (all) HCA ports. In this case, the target group can be set so that the total aggregated link bandwidth never exceeds the local memory bandwidth, regardless of which port on the source or destination HCA is involved.
[0306] According to an embodiment, in the case of having a fixed route towards a specific destination, any (one or more) intermediate target groups will generally represent only a single ISL at a specific stage in the path. However, when dynamic forwarding is active, both the target group and the ECN handling must take this into account. In the case where the dynamic forwarding decision will occur only to balance the traffic between parallel ISLs between a pair of switches (e.g., uplinks from a single leaf switch to a single spine switch), then all handling is in principle very similar to the handling when only a single ISL is used. FECN notification will occur based on the status of all ports in the relevant group, and the signaling can be "aggressive" in the sense that it is signaled based on a congestion indication from any port, or it can be more conservative and based on the size of the shared output queue of all ports in the group. The target group configuration generally represents the aggregated bandwidth of all links in the group, as long as the forwarding allows any packet to select the best output port at that point in time. However, if there is a strict concept of packet order preservation for each flow, then the evaluation of the bandwidth quota is more complex because some flows may "have to" use the same ISL at a certain point in time. If such a flow order scheme is based on well-defined header fields, then it is better to represent each port in the group as an independent target group. In this case, the selection of the target group on the sender-side HCA must be able to perform the same evaluation of the header fields that will be associated with the RC QP connection or address handle, just as the switch will perform for each packet at runtime.
[0307] According to an embodiment, by default, the initial target group rate for a new remote target can be conservatively set low. In this way, there is an inherent throttling before the target has a chance to update the relevant rate. Thus, all such rate control is independent of the VM itself involved, but the VM will be able to request the hypervisor to update the quotas for the ingress and egress traffic for different remote peers, but this can only be granted within the total constraints defined for the local and remote vHCA ports.
[0308] Figure 27 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high-performance computing environment according to an embodiment is shown.
[0309] More specifically, according to an embodiment, Figure 27A host channel adapter 2701 including a hypervisor 2711 is shown. The hypervisor can host / associate multiple virtual functions (VFs) such as VF 2714 - 2716 and a physical function (PF) 2713. The host channel adapter can additionally support / include multiple ports such as ports 2702 and 2703 which are used to connect the host channel adapter to a network such as network 2700. The network can include for example a switched network such as an InfiniBand network or a RoCE network which can connect the HCA 2701 to multiple other nodes such as switches, additional and separate HCAs, etc.
[0310] According to an embodiment, as described above, each virtual function can host a virtual machine (VM) such as VM1 2750, VM2 2751, and VM3 2752.
[0311] According to an embodiment, the host channel adapter 2701 can additionally support a virtual switch 2712 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture as described above.
[0312] According to an embodiment, as shown, the network 2700 can include multiple switches such as switches 2740, 2741, 2742, and 2743 which are interconnected and can be connected to the host channel adapter 2701 via for example leaf switches 2740 and 2741.
[0313] According to an embodiment, switches 2740 - 2743 can be interconnected and can additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0314] According to an embodiment, target groups such as target groups 2770 and 2771 can be defined at an inter - switch link (ISL) such as the ISL between leaf switch 2740 and switch 2742 and the ISL between leaf switch 2741 and switch 2743. For example, these target groups 2770 and 2771 can represent bandwidth quotas as HCA objects stored at a target group repository 2761 associated with the HCA and accessible by a rate limiting component 2760.
[0315] According to an embodiment, target groups 2770 and 2771 can represent specific (and different) bandwidth quotas. These bandwidth quotas can be divided into quotas of specific priorities (“QOS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0316] According to an embodiment, the destination groups 2770 and 2771 decouple the object from a specific destination address, and the system and method obtain the ability to represent an intermediate, shared link or link group, which links or link groups can represent bandwidth limitations other than the destination and may be more restrictive than the destination limitations. That is, for example, if the default / original egress limit on VM2 2751 is set to a threshold, but the destination of the packets sent from VM2 will pass through the destination group 2770 that sets a lower bandwidth limit, then the egress bandwidth from VM2 can be restricted to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA can be responsible for such throttling / egress bandwidth limit adjustment, depending on the destination group involved in the packet routing from VM2.
[0317] According to an embodiment, the destination groups can also be hierarchical in nature, whereby the system and method can consider a hierarchy of destination groups (bandwidth quotas) that reflect the bandwidth / link sharing for different destinations at different stages. In principle, this means that a particular flow should be associated with the destination group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if the destination group 2770 represents a higher bandwidth limit than the bandwidth limit of the destination group 2771, and the packets are addressed through two inter-switch links represented by the two destination groups, then the bandwidth limit of the destination group 2771 is the controlling bandwidth limiting factor.
[0318] According to an embodiment, the destination groups can also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flows, the bandwidth quota represented by the destination group can be divided. As an example, both VM1 and VM2 simultaneously send flows that will involve the destination group 2770, which represents a bandwidth quota of, for example, 10 Gb / s, and the corresponding flows have equal QoS and SLA associated with them. Then the destination group 2770 will represent a 5 Gb / s limit for each flow. This sharing or division of the destination group bandwidth quota can be based on changes in the QoS and SLA associated with the respective flows.
[0319] According to an embodiment, bandwidth quota and performance issues can arise when a VM (e.g., VM1 2750) is subjected to excessive ingress bandwidth 2790 from multiple sources. For example, this can occur when VM1 is subjected to one or more RDMA read responses while being subjected to one or more RDMA write operations, where the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a case, for example, the destination group, such as the destination group 2770 on the inter-switch link, can be updated to reflect a lower bandwidth quota via, for example, query 2775 than is normally allowed.
[0320] Additionally, according to an embodiment, the rate limiting component 2760 of the HCA may additionally include VM-specific rate limiting 2762, which may negotiate with other peer HCAs to coordinate, for example, the ingress bandwidth limit for VM1, and the egress bandwidth limit for the node responsible for generating the ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.
[0321] Figure 28 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment is shown, according to an embodiment.
[0322] More specifically, according to an embodiment, Figure 28 A host channel adapter 2801 including a hypervisor 2811 is shown. The hypervisor may host / associate multiple virtual functions (VFs) (such as VFs 2814 - 2816) and a physical function (PF) 2813. The host channel adapter may additionally support / include multiple ports, such as ports 2802 and 2803, which are used to connect the host channel adapter to a network, such as network 2800. The network may include, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 2801 to multiple other nodes, such as switches, attached and separate HCAs, etc.
[0323] According to an embodiment, as described above, each virtual function may host a virtual machine (VM), such as VM1 2850, VM2 2851, and VM3 2852.
[0324] According to an embodiment, the host channel adapter 2801 may additionally support a virtual switch 2812 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0325] According to an embodiment, as shown, the network 2800 may include multiple switches, such as switches 2840, 2841, 2842, and 2843, which are interconnected and may be connected to the host channel adapter 2801, for example, via leaf switches 2840 and 2841.
[0326] According to an embodiment, switches 2840 - 2843 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0327] According to an embodiment, target groups, such as target groups 2870 and 2871, may be defined at, for example, switch ports. As shown, target groups 2870 and 2871 are defined at the switch ports of switches 2842 and 2843, respectively. For example, these target groups 2870 and 2871 may represent bandwidth quotas as HCA objects that are stored at a target group repository 2861 associated with the HCA and that are accessible by a rate limiting component 2860.
[0328] According to an embodiment, target groups 2870 and 2871 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be partitioned into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of the various priorities within the relevant paths in the architecture represented by the target groups.
[0329] According to an embodiment, target groups 2870 and 2871 decouple the object from a specific destination address, and the system and method obtain the ability to represent intermediate, shared links or link groups that may represent bandwidth limitations other than the target and that may be more restrictive than the target limitations. That is, for example, if the default / original egress limit on VM2 2851 is set to a threshold, but the destination of packets sent from VM2 will pass through target group 2870 that has a lower bandwidth limit set, then the egress bandwidth from VM2 may be limited to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA may be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in routing the packets from VM2.
[0330] According to an embodiment, target groups may also be hierarchical in nature, whereby the system and method may consider a hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limiting rate in the hierarchy. That is, for example, if target group 2870 represents a higher bandwidth limit than the bandwidth limit of target group 2871 and a packet is addressed across two inter-switch links represented by the two target groups, then the bandwidth limit of target group 2871 is the controlling bandwidth limiting factor.
[0331] According to an embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flow, the bandwidth quota represented by the target group can be divided. As an example, both VM1 and VM2 simultaneously send flows that will involve target group 2870, which represents a bandwidth quota of, for example, 10 Gb / s, and the respective flows have equal QoS and SLA associated with them. Then target group 2870 will represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth quota can vary based on the QoS and SLA associated with the respective flow.
[0332] According to an embodiment, bandwidth quota and performance issues can arise when a VM (e.g., VM1 2850) experiences excessive ingress bandwidth 2890 from multiple sources. For example, this can occur when VM1 experiences one or more RDMA read responses while undergoing one or more RDMA write operations, where the ingress bandwidth on VM1 is from two or more sources (e.g., one RDMA read response is from a connected VM and one RDMA write request is from another connected VM). In such a case, for example, a target group, such as target group 2870 on an inter-switch link, can be updated to reflect a lower bandwidth quota via, for example, query 2875 than is normally allowed.
[0333] Additionally, according to an embodiment, the rate limiting component 2860 of the HCA can additionally include VM-specific rate limiting 2862, which can negotiate with other peer HCAs to coordinate, for example, the ingress bandwidth limit for VM1, as well as the egress bandwidth limit for the node responsible for generating the ingress bandwidth on VM1. These other HCAs / nodes are not shown in this figure.
[0334] Figure 29 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment is shown, according to an embodiment.
[0335] More specifically, according to an embodiment, Figure 29 A host channel adapter 2901 including a hypervisor 2911 is shown. The hypervisor can host / associate multiple virtual functions (VFs) (such as VFs 2914 - 2916) as well as a physical function (PF) 2913. The host channel adapter can additionally support / include multiple ports, such as ports 2902 and 2903, which are used to connect the host channel adapter to a network, such as network 2900. The network can include, for example, a switched network, such as an InfiniBand network or a RoCE network, which can connect the HCA 2901 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0336] According to an embodiment, as described above, each virtual function may host a virtual machine (VM), such as VM1 2950, VM2 2951, and VM3 2952.
[0337] According to an embodiment, the host channel adapter 2901 may additionally support a virtual switch 2912 via a hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0338] According to an embodiment, as shown, the network 2900 may include a plurality of switches, such as switches 2940, 2941, 2942, and 2943, which are interconnected and may be connected to the host channel adapter 2901 via, for example, leaf switches 2940 and 2941.
[0339] According to an embodiment, switches 2940 - 2943 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0340] According to an embodiment, a target group, such as target group 2971, may be defined along an inter-switch link (ISL), such as the ISL between leaf switch 2941 and switch 2943. Other target groups may be defined, for example, at a switch port. As shown, target group 2970 is defined at a switch port of switch 2952. For example, these target groups 2970 and 2971 may represent bandwidth quotas as HCA objects, which are stored at a target group repository 2961 associated with the HCA and accessible by a rate limiting component 2960.
[0341] According to an embodiment, target groups 2970 and 2971 may represent specific (and different) bandwidth quotas. These bandwidth quotas may be divided into quotas of specific priorities (“QoS classes”) in order to reflect the expected bandwidth usage of various priorities within the relevant paths in the architecture represented by the target groups.
[0342] According to an embodiment, target groups 2970 and 2971 decouple an object from a specific destination address, and the system and method obtain the ability to represent an intermediate, shared link or link group, which can represent bandwidth limitations other than the target and may be more restrictive than the target limitation. That is, for example, if the default / original egress limit on VM2 2951 is set to a threshold, but the destination of the packets sent from VM2 will pass through target group 2970 which sets a lower bandwidth limit, then the egress bandwidth from VM2 can be restricted to a level lower than the level of the default / original egress limit set on VM2. For example, the HCA can be responsible for such throttling / egress bandwidth limit adjustment, depending on the target group involved in the packet routing from VM2.
[0343] According to an embodiment, target groups can also be hierarchical in nature, whereby the system and method can consider a hierarchy of target groups (bandwidth quotas) that reflect bandwidth / link sharing for different targets at different stages. In principle, this means that a particular flow should be associated with the target group (maximum rate) that represents the most limited rate in the hierarchy. That is, for example, if target group 2970 represents a higher bandwidth limit than the bandwidth limit of target group 2971, and the packet is addressed through two inter-switch links represented by the two target groups, then the bandwidth limit of target group 2971 is the controlling bandwidth limiting factor.
[0344] According to an embodiment, a target group can also be shared by multiple flows. For example, depending on the QoS and SLA associated with the respective flows, the bandwidth quota represented by the target group can be divided. As an example, both VM1 and VM2 simultaneously send flows that will involve target group 2970, which represents a bandwidth quota of, for example, 10 Gb / s, and the respective flows have equal QoS and SLA associated with them, then target group 2970 will represent a 5 Gb / s limit for each flow. This sharing or division of the target group bandwidth quota can be based on changes in the QoS and SLA associated with the respective flows.
[0345] According to an embodiment, bandwidth quota and performance issues can occur when a VM (e.g., VM1 2950) is subjected to excessive ingress bandwidth 2990 from multiple sources. For example, this can occur when VM1 is subjected to one or more RDMA read responses while being subjected to one or more RDMA write operations, where the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM). In such a case, for example, a target group, such as target group 2970 on an inter-switch link, can be updated to reflect a lower bandwidth quota via, for example, query 2975 than is normally allowed.
[0346] Additionally, according to an embodiment, the rate limiting component 2960 of the HCA may additionally include VM-specific rate limiting 2962, which may negotiate with other peer HCAs to coordinate, for example, the ingress bandwidth limit for VM1, as well as the egress bandwidth limit for the node responsible for generating the ingress bandwidth on VM1. These other HCAs / nodes are not shown in the figure.
[0347] Figure 30 A system for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment is shown, according to an embodiment.
[0348] According to an embodiment, within a high performance computing environment such as a switched network or subnet 3000, multiple end nodes 3001 and 3002 may support multiple virtual machines VM1-VM4 3050-3053, which are interconnected via multiple switches such as leaf switches 3011 and 3012, switches 3021 and 3022, and root switches 3031 and 3032.
[0349] What is not shown in the figure, according to an embodiment, are the various host channel adapters that provide the functionality for connecting nodes 3001 and 3002 and the virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of a hypervisor on a host channel adapter.
[0350] According to an embodiment, when a node (such as VM3 3052) simultaneously processes an RDMA read response 3050 and an RDMA write request 3051 (bandwidth on the ingress), it may encounter a bandwidth limit (e.g., from rate limiting 3061).
[0351] According to an embodiment, rate limiters 3060 and 3061 may be configured to, for example, ensure that the ingress bandwidth quota is not violated by coordinating RDMA requests (i.e., requests for RDMA reads sent by VM3 to VM4, messages that result in the RDMA read response 3050), and RDMA write operations (e.g., an RDMA write from VM2 to VM3).
[0352] For each individual node, system, and method, there may be such a chain of target groups such that the flow will always coordinate with all other flows sharing link bandwidth in different parts of the architecture represented in the target groups.
[0353] Figure 31 Is a flowchart of a method for combining target-specific RDMA-write and RDMA-read bandwidth limits in a high performance computing environment, according to an embodiment.
[0354] According to an embodiment, at step 3110, the method may provide a first subnet at one or more microprocessors, the first subnet including: a plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports; a plurality of host channel adapters, wherein each host channel adapter of the host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches; and a plurality of end nodes, including a plurality of virtual machines.
[0355] According to an embodiment, at step 3120, the method may define a target group on an inter-switch link between two switches of the plurality of switches or on at least one of the ports of a switch of the plurality of switches, wherein the target group defines a bandwidth limit on an inter-switch link between two switches of the plurality of switches or on at least one of the ports of a switch of the plurality of switches.
[0356] According to an embodiment, at step 3130, the method may provide a target group repository stored in a memory of the host channel adapter at the host channel adapter.
[0357] According to an embodiment, at step 3140, the method may record the defined target group in the target group repository.
[0358] According to an embodiment, at step 3150, the method may receive ingress bandwidth from at least two remote sources at an end node of the host channel adapter, the ingress bandwidth exceeding an ingress bandwidth limit of the end node.
[0359] According to an embodiment, at 3160, in response to receiving ingress bandwidth from at least two sources, the method may update a bandwidth quota of the target group.
[0360] Combining Ingress Bandwidth Arbitration and Congestion Feedback (ORA200246 - US - NP - 2)
[0361] According to an embodiment, when multiple sender nodes / VMs each and / or all send to a single receiver node / VM, it is not straightforward to achieve fairness among the senders, avoid congestion, and at the same time limit the ingress bandwidth usage consumed by the receiver node / VM to (far) below the maximum limit that the associated network interface can provide for ingress traffic. Additionally, when different bandwidth quotas should be allocated to different senders due to different SLA levels, then the equation becomes even more complex.
[0362] According to embodiments, the systems and methods herein can extend traditional schemes for end-to-end congestion feedback to include an initial negotiation of bandwidth quotas, a dynamic adjustment of such bandwidth quotas (e.g., to accommodate changes in the number of sender nodes sharing the available bandwidth, or changes in the SLA), and dynamic congestion feedback to indicate that a sender needs to temporarily reduce the associated egress data rate, although the overall bandwidth quota remains unchanged. Explicit, proactive messages and "piggybacked" information are used in data packets to convey relevant information from the destination node to the sender node.
[0363] Figure 32 A system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment is shown, according to an embodiment.
[0364] More specifically, according to an embodiment, Figure 32 A host channel adapter 3201 including a hypervisor 3211 is shown. The hypervisor can host / associate multiple virtual functions (VFs) such as VFs 3214 - 3216 and a physical function (PF) 3213. The host channel adapter can additionally support / include multiple ports such as ports 3202 and 3203 that are used to connect the host channel adapter to a network such as network 3200. The network can include, for example, a switched network such as an InfiniBand network or a RoCE network that can connect the HCA 3201 to multiple other nodes such as switches, additional and separate HCAs, etc.
[0365] According to an embodiment, as described above, each virtual function can host a virtual machine (VM) such as VM1 3250, VM2 3251, and VM3 3252.
[0366] According to an embodiment, the host channel adapter 3201 can additionally support a virtual switch 3212 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture, as described above.
[0367] According to an embodiment, as shown, the network 3200 can include multiple switches such as switches 3240, 3241, 3242, and 3243 that are interconnected and can be connected to the host channel adapter 3201, for example, via leaf switches 3240 and 3241.
[0368] According to an embodiment, switches 3240 - 3243 can be interconnected and can additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0369] According to an embodiment, bandwidth quota and performance issues may occur when a VM (e.g., VM1 3250) is subjected to excessive ingress bandwidth 3290 from multiple sources. For example, this may occur when VM1 is subjected to one or more RDMA read responses while being subjected to one or more RDMA write operations, where the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM).
[0370] According to an embodiment, the rate limiting component 3260 of the HCA may additionally include VM-specific rate limiting 3261, which may negotiate with other peer HCAs to coordinate, for example, the ingress bandwidth limit of VM1, as well as the egress bandwidth limit of the nodes responsible for generating the ingress bandwidth on VM1. For example, such an initial negotiation may be performed to accommodate changes in the number of sender nodes sharing the available bandwidth or changes in the SLA. These other HCAs / nodes are not shown in the figure.
[0371] According to an embodiment, the above negotiation may be updated based on, for example, explicit and proactive feedback messages 3291 generated as a result of the ingress bandwidth. For example, such feedback messages 3291 may be sent to multiple remote nodes responsible for generating the ingress bandwidth 3290 on VM1. After receiving such feedback messages, the sender nodes (bandwidth senders responsible for the ingress bandwidth on VM1) may update their relevant egress bandwidth limits on the sender nodes so as not to overload, for example, the link connected to VM1 while still attempting to maintain QoS and SLA.
[0372] Figure 33 A system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to an embodiment is shown.
[0373] More specifically, according to an embodiment, Figure 33 A host channel adapter 3301 including a hypervisor 3311 is shown. The hypervisor may host / associate multiple virtual functions (VFs) (such as VF 3314 - 3316) and a physical function (PF) 3313. The host channel adapter may additionally support / include multiple ports, such as ports 3302 and 3303, which are used to connect the host channel adapter to a network, such as network 3300. The network may include, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 3301 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0374] According to an embodiment, as described above, each virtual function may host virtual machines (VMs), such as VM1 3350, VM2 3351, and VM3 3352.
[0375] According to an embodiment, the host channel adapter 3301 may additionally support a virtual switch 3312 via a hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0376] According to an embodiment, as shown, the network 3300 may include multiple switches, such as switches 3340, 3341, 3342, and 3343, which are interconnected and may be connected to the host channel adapter 3301 via, for example, leaf switches 3340 and 3341.
[0377] According to an embodiment, switches 3340 - 3343 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0378] According to an embodiment, when a VM (e.g., VM1 3350) experiences excessive ingress bandwidth 3390 from multiple sources, bandwidth quota and performance issues may occur. For example, this may occur when VM1 experiences one or more RDMA read responses while undergoing one or more RDMA write operations, where the ingress bandwidth on VM1 comes from two or more sources (e.g., one RDMA read response from a connected VM and one RDMA write request from another connected VM).
[0379] According to an embodiment, the rate limiting component 3360 of the HCA may additionally include VM - specific rate limiting 3361, which may negotiate with other peer HCAs to coordinate, for example, the ingress bandwidth limit of VM1, as well as the egress bandwidth limit of the nodes responsible for generating the ingress bandwidth on VM1. For example, such an initial negotiation may be performed to accommodate changes in the number of sender nodes sharing the available bandwidth or changes in the SLA. These other HCAs / nodes are not shown in the figure.
[0380] According to an embodiment, the above negotiation may be updated based on, for example, a piggyback message 3391 (a message residing on regular data or other communication packets sent between end nodes) generated as a result of the ingress bandwidth. For example, such a piggyback message 3391 may be sent to multiple remote nodes responsible for generating the ingress bandwidth 3390 on VM1. After receiving such feedback messages, the sender nodes (bandwidth senders responsible for the ingress bandwidth on VM1) may update their relevant egress bandwidth limits on the sender nodes so as not to overload, for example, the links connected to VM1 while still attempting to maintain QoS and SLA.
[0381] Figure 34Illustrates a system for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to an embodiment.
[0382] According to an embodiment, within a high-performance computing environment such as a switched network or a subnet 3400, multiple end nodes 3401 and 3402 may support multiple virtual machines VM1-VM4 3450-3453, which are interconnected via multiple switches such as leaf switches 3411 and 3412, switches 3421 and 3422, and root switches 3431 and 3432.
[0383] Not shown in the figure, according to an embodiment, are various host channel adapters that provide the functionality for connecting nodes 3401 and 3402 and the virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine may be associated with a virtual function of the hypervisor on the host channel adapter.
[0384] According to an embodiment, when a node (such as VM3 3452) receives, for example, multiple RDMA ingress bandwidth packets (e.g., multiple RDMA writes), such as 3451 and 3452, it may encounter an ingress bandwidth limit (e.g., from rate limiter 3461). This may occur, for example, when there is no communication between the respective sender nodes to coordinate the bandwidth limits.
[0385] According to an embodiment, the systems and methods herein may extend the end-to-end congestion feedback scheme to include an initial negotiation of bandwidth quotas (i.e., VM3 negotiates, or the bandwidth limit associated with VM3, with all sender nodes targeting VM3 within the ingress traffic), a dynamic adjustment of such bandwidth quotas (e.g., to accommodate changes in the number of sender nodes sharing the available bandwidth, or changes in the SLA), and dynamic congestion feedback to indicate that the sender needs to temporarily reduce the associated egress data rate, although the overall bandwidth quota remains unchanged. For example, such dynamic congestion feedback may occur in a return message (e.g., feedback message 3470) to the respective sender nodes, indicating to each sender node the updated bandwidth limit to be utilized when sending traffic to VM3. Such a feedback message 3460 may take the form of an explicit, proactive message as well as "piggybacked" information within data packets to convey the relevant information from the target node (i.e., VM3 in the depicted embodiment) to the sender nodes.
[0386] Figure 35 Is a flowchart of a method for combining ingress bandwidth arbitration and congestion feedback in a high-performance computing environment according to an embodiment.
[0387] According to an embodiment, at step 3510, the method may provide a first subnet at one or more microprocessors, the first subnet including: a plurality of switches, the plurality of switches including at least leaf switches, wherein each switch in the plurality of switches includes a plurality of switch ports; a plurality of host channel adapters, wherein each host channel adapter in the host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches; and a plurality of end nodes, including a plurality of virtual machines.
[0388] According to an embodiment, at step 3520, the method may provide an end node ingress bandwidth quota at the host channel adapter associated with an end node attached to the host channel adapter.
[0389] According to an embodiment, at step 3530, the method may negotiate a bandwidth quota between an end node attached to the host channel adapter and a remote end node.
[0390] According to an embodiment, at step 3540, the method may receive ingress bandwidth from a remote source at an end node attached to the host channel adapter, the ingress bandwidth exceeding the ingress bandwidth limit of the end node.
[0391] According to an embodiment, at 3550, in response to receiving ingress bandwidth from at least two sources, the method may send a response message from an end node attached to the host channel adapter to a remote end node, the response message indicating that the ingress bandwidth quota of the end node attached to the host channel adapter has been exceeded.
[0392] Using Multiple CE (Congestion Experienced) Flags in FECN (Forward Explicit Congestion Notification) and BECN (Backward Explicit Congestion Notification) Signaling (ORA200246 - US - NP - 4) Figure 36
[0393] According to an embodiment, conventionally, congestion notification is based on data packets that encounter congestion at a certain point (e.g., a certain link segment between a certain node / switch pair along the path from the sender to the target through the network / architecture topology) being marked with a "congestion" status flag (also known as the CE flag), and then this status is reflected in the response packet sent back from the target to the sender.
[0394] According to an embodiment, the problem with this scheme is that it does not allow the sender node to distinguish between flows that experience congestion at the same link segment, even though they represent different targets. Additionally, when there are multiple paths available between the sender and target node pair, any information about congestion on different alternative paths requires that a certain flow be active via the relevant path for the relevant target.
[0395] According to an embodiment, the systems and methods described herein extend the congestion marking scheme to facilitate multiple CE flags in the same packet and configure switch ports to represent a phase number that defines what CE flag index it should update. Between a particular sender and a particular destination, a particular path through an ordered sequence of switch ports will then represent a particular ordered list of unique phase numbers, and thus also represent CE flag index numbers.
[0396] According to an embodiment, in this way, a sender node receiving congestion feedback with multiple CE flags set can map the various CE flags to different "destination group" contexts, which will then represent the associated congestion condition status and associated dynamic rate reduction. Additionally, different flows for different destinations will share congestion information and dynamic rate reduction status associated with a shared link segment represented by a shared "destination group" in the sender node.
[0397] According to an embodiment, when congestion does occur, the key issue is that the congestion feedback should ideally be associated with all relevant destination groups in the hierarchy associated with the flow receiving the congestion feedback. Then, the affected destination groups should dynamically adjust their maximum rates accordingly. Therefore, the hardware state of each destination group must also include any current congestion status and associated "throttling information".
[0398] According to an embodiment, an important aspect here is that the FECN signaling should have the ability to include multiple "Congestion Experienced" (CE) flags such that a switch detecting congestion can mark the flags corresponding to its phase in the topology. In a regular fat tree, each switch has a unique (maximum) phase number in the upward direction and another unique (maximum) phase number in the downward direction. Thus, a flow using a particular path will then be associated with a sequence of particular phase numbers, which will include all or only a subset of the total set of phase numbers in the entire architecture. However, for a particular flow, the individual phase numbers associated with the path can then be mapped to one or more destination groups associated with that flow. In this way, the BECN of the received flow can imply that the (one or more) destination groups associated with the phase of each CE mark in the BECN will be updated to indicate congestion, and the dynamic maximum rates of these destination groups can then be adjusted accordingly.
[0399] According to an embodiment, while inherently applicable to a fat tree topology, the "phase number" concept of a switch can be generalized to represent almost any topology to which such numbers can be assigned to switches. However, in this general case, the phase number is not just a function of the output port, but a function of each input / output port number tuple. In the general case, the number of required phase numbers and the path-specific mapping to destination groups are also more complex. Therefore, in this context, the reasoning assumes only a fat tree topology.
[0400] According to an embodiment, multiple CE flags in a single packet are not currently a feature supported by standard protocol headers. Thus, this can be supported based on an extension of the standard header, and / or it can be supported by inserting additional independent FECN packets in the flow – conceptually, the generation of additional packets in the flow is similar to using an encapsulation scheme within a switch, and the effect is that packets received at line speed cannot be forwarded at the same line speed because more “overhead bytes” must be transmitted downstream. Inserting additional packets will generally be more than the encapsulation overhead, but as long as this overhead is spread across multiple data packets (no such additional notification needs to be sent for each data packet), the overhead may be acceptable.
[0401] According to an embodiment, there can also be a scenario where switch firmware can monitor the congestion condition within the switch and thus send an “active BECN” to the relevant sender node. However, this means that the switch firmware must have more state information about the relevant sender, as well as the mapping between ports and priorities and the relevant sender, which can also include dynamic information about which addresses the packets experiencing congestion involve.
[0402] According to an embodiment, for an RC QP, the “CE flag to target group” mapping is typically part of the QP context, and any BECN information received in an ACK / response packet will thus be handled in a straightforward manner for the relevant QP context and the associated target group. However, in the case of “active BECN” (e.g., because the datagram traffic only has application-level responses / ACKs, or because “congestion warnings” are broadcast to multiple potential senders), the reverse mapping is not straightforward – at least not straightforward in terms of being automatically handled by hardware. Thus, a better approach is to formulate a scenario where FECN can both result in automatically hardware-generated BECN in the case of a connected (RC) flow, but both FECN events with hardware-automatic BECN generation and FECN events without hardware-generated BECN can be handled by the firmware and / or hyper-privileged software associated with the HCA that receives the FECN. In this way, there can be FW / SW-generated “active BECN” sent to one or more potential senders affected by the observed congestion. The FW / SW that receives these “active BECN” can then perform the mapping to the relevant local target group based on the payload data in the received “BECN message”, and then can trigger the local hardware to update the target group state, similar to what happens in the fully hardware-controlled handling of RC-related BECN.
[0403] According to an embodiment, a subset of stage numbers without any BECN notification or with a CE flag set that is different from (less than) the RC ACK / response packet of an earlier-recorded state may cause a corresponding update within the associated target group in the local HCA. Similarly, the responder HCA (i.e., the associated software / firmware) may send an “active BECN” to indicate that the congestion signaled earlier no longer exists.
[0404] According to an embodiment, as described above, the target group concept combined with dynamic congestion feedback at the hardware or firmware / software level provides flexible control over the egress bandwidth generated by the HCA as well as by individual vHCAs and tenants sharing a physical HCA.
[0405] According to an embodiment, since target groups are identified completely independently of the associated remote address and path information at the VM level, there is no dependency between the use of target groups and the extent to which communication from the VM is based on an overlay or other virtual networking scheme. The only requirement is that the hyper-privileged software controlling the HCA resources be able to define the relevant mapping. Additionally, a scheme with a “logical target group ID” at the VM / vHCA level can be used, which is then mapped by the HCA to an actual target group. However, it is not clear whether this is useful, other than that it hides the actual target group ID from the tenant. In cases where it is necessary to change which target groups are associated with a particular destination due to a change in the underlying path, then this may not involve other destinations. Thus, in general, updating a target group must involve updating all the QPs and address handles involved, rather than just updating the logical-to-physical target group ID mapping.
[0406] According to an embodiment, for a virtualized target HCA, individual vHCA ports rather than physical HCA ports can be represented as the final destination target groups. In this way, the target group hierarchy of the remote peer can include a target group representing the destination physical HCA port as well as additional target groups representing the final destination according to the vHCA ports. In this way, the system and method have the ability to limit the ingress bandwidth of individual vHCA ports (VFs), while the bandwidth of each physical HCA port and the associated sender target group means that the sum of the ingress bandwidth quotas for each vHCA port need not remain less than the physical HCA port bandwidth (or the associated bandwidth quota).
[0407] According to an embodiment, within the sender HCA, by assigning different target groups to different tenants, target groups can be used to represent the sharing of physical HCA ports in the egress direction. Additionally, to facilitate tenant-level target groups for multiple VMs from the same tenant to share a physical HCA port, different target groups can be assigned to different such VMs. Such target groups are then set as the initial target groups for all egress communication from that VM.
[0408] Figure 36 illustrates a system for using multiple CE flags in both FECN and BECN in a high performance computing environment according to an embodiment.
[0409] More specifically, according to an embodiment, Figure 37 illustrates a host channel adapter 3601 including a hypervisor 3611. The hypervisor may host / associate multiple virtual functions (VFs) (such as VFs 3614 - 3616) and a physical function (PF) 3613. The host channel adapter may additionally support / include multiple ports, such as ports 3602 and 3603, which are used to connect the host channel adapter to a network, such as network 3600. The network may include, for example, a switched network, such as an InfiniBand network or a RoCE network, which may connect the HCA 3601 to multiple other nodes, such as switches, additional and separate HCAs, etc.
[0410] According to an embodiment, as described above, each virtual function may host a virtual machine (VM), such as VM1 3650, VM2 3651, and VM3 3652.
[0411] According to an embodiment, the host channel adapter 3601 may additionally support a virtual switch 3612 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure may additionally support a virtual port (vPort) architecture, as described above.
[0412] According to an embodiment, as shown, the network 3600 may include multiple switches, such as switches 3640, 3641, 3642, and 3643, which are interconnected and may be connected to the host channel adapter 3601, for example, via leaf switches 3640 and 3641.
[0413] According to an embodiment, switches 3640 - 3643 may be interconnected and may additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0414] According to an embodiment, when an ingress packet 3690 traverses the network, it may experience congestion at any stage of its path, and when such congestion is detected at any stage, the switch may mark the packet. In addition to marking the packet as having experienced congestion, the switch that performs the marking may additionally indicate the stage at which the packet experienced congestion. Upon reaching the destination node, e.g., VM1 3650, VM1 may send a response packet via an explicit feedback message 3691 (e.g., automatically), which may indicate to the sending node that the packet experienced congestion and at which stage(s) the packet experienced congestion.
[0415] According to an embodiment, the ingress grouping may include a bit field that is updated to indicate where the grouping has experienced congestion, and the explicit feedback message may mirror / represent that bit field when notifying the sender node of such congestion.
[0416] According to an embodiment, each switch port represents a stage in the overall subnet. Thus, each packet sent in the subnet can traverse a maximum number of stages. To identify where congestion has been detected (possibly at multiple locations), the congestion marker (e.g., the CE flag) is extended from a simple binary flag (has experienced congestion) to a bit field that contains multiple bits. Then, each bit of the bit field can be associated with a stage number that can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages is three. When the system has a path from A to B and the routing is known, then each end node can determine which switch port the packet traverses at any given stage of the path. By doing so, each end node can determine at which different switch ports the packet has experienced congestion by correlating the routing with the received congestion message.
[0417] According to an embodiment, the system may provide a returned congestion feedback that indicates which congestion has been detected at which stage, and then if the end node has congestion due to a shared link segment, then congestion control is applied to that segment rather than to different end ports. This provides more fine-grained information about congestion.
[0418] According to an embodiment, by providing such finer granularity, the end node can then use an alternative path when routing future packets. Or, for example, if the end node has multiple flows all going to different destinations but congestion is detected at a common stage of the path, then rerouting can be triggered. The system and method provide an immediate reaction in terms of associated throttling - rather than having 10 different congestion notifications. This can dispose of congestion notifications more efficiently.
[0419] Figure 37 A system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to an embodiment is shown.
[0420] More specifically, according to an embodiment, Figure 38Shows a host channel adapter 3701 including a hypervisor 3711. The hypervisor can host / associate multiple virtual functions (VFs) (such as VFs 3714 - 3716) and a physical function (PF) 3713. The host channel adapter can additionally support / include multiple ports, such as ports 3702 and 3703, which are used to connect the host channel adapter to a network, such as network 3700. The network can include, for example, a switched network, such as an InfiniBand network or a RoCE network, which can connect the HCA 3701 to multiple other nodes, such as switches, attached and separate HCAs, etc.
[0421] According to an embodiment, as described above, each virtual function can host a virtual machine (VM), such as VM1 3750, VM2 3751, and VM3 3752.
[0422] According to an embodiment, the host channel adapter 3701 can additionally support a virtual switch 3712 via the hypervisor. This is for the case of implementing a vSwitch architecture. Although not shown, embodiments of the present disclosure can additionally support a virtual port (vPort) architecture, as described above.
[0423] According to an embodiment, as shown, the network 3700 can include multiple switches, such as switches 3740, 3741, 3742, and 3743, which are interconnected and can be connected to the host channel adapter 3701, for example, via leaf switches 3740 and 3741.
[0424] According to an embodiment, switches 3740 - 3743 can be interconnected and can additionally be connected to other switches and other end nodes (e.g., other HCAs) not shown in the figure.
[0425] According to an embodiment, when the ingress packet 3790 traverses the network, it may experience congestion at any stage of its path, and when such congestion is detected at any stage, the switch can mark the packet. In addition to marking the packet as having experienced congestion, the switch that performs the marking can additionally indicate the stage at which the packet experienced congestion. When arriving at a destination node, such as VM1 3750, VM1 can send a response packet (e.g., automatically) via a piggyback message (a message residing on top of another message / packet sent from the receiving node to the sender node) 3791, which can indicate to the sending node that the packet experienced congestion and at which (which) stage(s) the packet experienced congestion.
[0426] According to an embodiment, the ingress packet can include a bit field that is updated to indicate where the packet experienced congestion, and the explicit feedback message can mirror / represent the bit field when notifying the sender node of such congestion.
[0427] According to an embodiment, each switch port represents a stage in the overall subnet. Thus, each packet sent in the subnet can traverse the maximum number of stages. To identify where congestion is detected (which may be at multiple locations), the congestion marker (e.g., the CE flag) is extended from a simple binary flag (experienced congestion) to a bit field that contains multiple bits. Then, each bit of the bit field can be associated with a stage number that can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages is three. When the system has a path from A to B and the routing is known, then each end node can determine which switch port the packet traverses at any given stage of the path. By doing so, each end node can determine at which different switch ports the packet experienced congestion by correlating the routing with the received congestion message.
[0428] According to an embodiment, the system can provide returned congestion feedback that indicates which congestion is detected at which stage, and then if the end node has congestion due to a shared link segment, then congestion control is applied to that segment rather than to different end ports. This provides more fine-grained information about congestion.
[0429] According to an embodiment, by providing this finer granularity, the end node can then use an alternative path when routing future packets. Or, for example, if the end node has multiple flows all going to different destinations but congestion is detected at a common stage of the path, then rerouting can be triggered. The system and method provide an immediate reaction in terms of associated throttling - rather than having 10 different congestion notifications. This can dispose of congestion notifications more efficiently.
[0430] Figure 39 A system for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to an embodiment is shown.
[0431] According to an embodiment, in a high-performance computing environment such as a switched network or subnet 3800, multiple end nodes 3801 and 3802 can support multiple virtual machines VM1-VM4 3850-3853, which are interconnected via multiple switches such as leaf switches 3811 and 3812, switches 3821 and 3822, and root switches 3831 and 3832.
[0432] Not shown in the figure, according to an embodiment, are various host channel adapters that provide the functionality for connecting nodes 3801 and 3802 as well as the virtual machines to be connected to the subnet. Such embodiments were discussed above with respect to SR-IOV, where each virtual machine can be associated with a virtual function of the hypervisor on the host channel adapter.
[0433] According to an embodiment, for a packet 3851 sent from VM3 3852 to VM1 3850, it can traverse subnet 3800 via multiple links or stages (such as stages 1 to 6) as shown in the figure. When packet 3851 traverses the subnet, it may experience congestion at any of these stages, and when such congestion is detected at any stage, it can be marked by a switch. In addition to marking the packet as having experienced congestion, the switch that performs the marking can additionally indicate the stage at which the packet experienced congestion. When reaching the destination node VM1, VM1 can send a response packet via feedback message 3870 (e.g., automatically), which can indicate to VM3 3852 that the packet experienced congestion and at which stage(s) the packet experienced congestion.
[0434] According to an embodiment, each switch port represents a stage in the entire subnet. Thus, each packet sent in the subnet can traverse the maximum number of stages. To identify the location(s) where congestion is detected (which may be at multiple locations), the congestion marking (e.g., CE flag) is extended from a simple binary flag (experienced congestion) to a bit field containing multiple bits. Then, each bit of the bit field can be associated with a stage number, which can be assigned to each switch port. For example, in a three-stage fat tree, the maximum number of stages is three. When the system has a path from A to B and the routing is known, then each end node can determine which switch port the packet traverses at any given stage of the path. By doing so, each end node can determine at which different switch ports the packet experienced congestion by associating the routing with the received congestion message.
[0435] According to an embodiment, the system can provide returned congestion feedback that indicates at which stage congestion is detected, and then if the end node has congestion due to a shared link segment, then congestion control is applied to that segment rather than to different end ports. This provides more fine-grained information about congestion.
[0436] According to an embodiment, by providing such finer granularity, the end node can then use an alternative path when routing future packets. Or, for example, if the end node has multiple flows all going to different destinations but congestion is detected at a common stage of the path, then rerouting can be triggered. The system and method provide an immediate reaction in terms of associated throttling - rather than having 10 different congestion notifications. This can handle congestion notifications more efficiently.
[0437] QOS and SLA in a Switching Fabric Such as a Private Fabric is a flowchart of a method for using multiple CE flags in both FECN and BECN in a high-performance computing environment according to an embodiment.
[0438] According to an embodiment, at step 3910, the method may provide a first subnet at one or more microprocessors, the first subnet including: a plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports; a plurality of host channel adapters, wherein each host channel adapter of the host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches; and a plurality of end nodes, including a plurality of virtual machines.
[0439] According to an embodiment, at step 3920, the method may receive an ingress packet from a remote end node at an end node attached to a host channel adapter, wherein the ingress packet traverses at least a portion of the first subnet before being received at the end node, and wherein the ingress packet includes a marker indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet.
[0440] According to an embodiment, upon receiving the ingress packet, at step 3930, the method may send a response message from the end node attached to the host channel adapter to the remote end node by the end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet, wherein the response message includes a bit field.
[0441] Considerations Regarding Topology, Routing, and Blocking Scenarios:
[0442] According to an embodiment, in a cloud and in a private network architecture in a larger cloud installed by a customer and on-premises (e.g., a private architecture such as those for building dedicated distributed devices or general high-performance computing resources), there is a desire to deploy VM-based workloads, where an essential requirement is that quality of service (QoS) can be defined and controlled for different types of traffic flows. In addition, workloads belonging to different tenants must execute within the boundaries of relevant service level agreements (SLAs), while minimizing interference between such workloads and maintaining QoS assumptions for different communication types.
[0443] According to an embodiment, the following sections discuss related problem scenarios, objectives, and potential solutions.
[0444] According to an embodiment, an initial scenario for provisioning fabric resources to cloud customers (also referred to as "tenants") is that a dedicated portion of a rack (e.g., a quarter rack), or one or more complete racks, can be allocated to a tenant. This granularity means that as long as the allocated resources are fully operational, each tenant can be guaranteed a communication SLA that is always met. This is also the case when a single rack is partitioned into multiple sections, since the granularity is always a complete physical server with an HCA. In principle, the connections between different such servers in a single rack always occur through a single full crossbar switch. In this case, there are no resources shared in a way that could cause contention or congestion between flows belonging to different tenants due to the communication traffic between the sets of servers belonging to the same tenant.
[0445] However, according to an embodiment, since the redundant switches are shared, it is crucial that traffic generated by the workload on one server cannot be targeted at a server belonging to another tenant. Even if such traffic is not detrimental to any communication or data leakage / observation between tenants, the result could be a severe disruption to the communication flows belonging to other tenants or even a DOS (Denial of Service)-like effect.
[0446] According to an embodiment, although the fact that there is a full crossbar leaf switch essentially means that all communication between local servers can occur only via the local switch, this may not be possible or may not be implemented due to other practical issues in several cases:
[0447] · According to an embodiment, for example, if the host bus (PCIe) generates a bandwidth that can only sustain one fabric link at a time, it is crucial that only one HCA port is used for data traffic at any point in time. Thus, if not all servers agree on which local switch to use for data traffic, some traffic will have to pass through the inter-switch link (ISL) between the local leaf switches.
[0448] · According to an embodiment, if one or more servers have lost their connection to one of the switches in the switch fabric, then all communication must occur via the other switch. Thus, again, if not all server pairs agree to use the same single switch, some data traffic will have to pass through the (one or more) ISLs.
[0449] · According to an embodiment, if a server can use two HCA ports (and thus can use two leaf switches), but cannot be forced to establish connections only via the HCA ports connected to the same switch, then some data traffic may pass through the ISL.
[0450] ο The reason for finally adopting this solution is that the lack of socket / port numbers in the fabric host stack means that a process can only establish one socket to accept incoming connections. Then, this socket can only be associated with a single HCA port at a time. As long as the same single socket is also used when establishing an outgoing connection, the system will ultimately end up with multiple connections that require ISLs, even though a large number of processes evenly distribute their single socket among the local HCA ports.
[0451] According to an embodiment, in addition to the special single-rack scenario for implementing ISL usage / sharing outlined above, once the provisioning granularity is extended to a multi-rack configuration where the leaf switches in each rack are interconnected by a backbone switch, then the communication SLAs of different tenants become highly dependent on which servers are assigned to which tenants, and how different communication flows are mapped to different switch-switch links through the fabric-level routing scheme. The key issue in this scenario is that the two optimization aspects are somewhat contradictory:
[0452] · According to an embodiment, on the one hand, all concurrent flows to different destination ports should use as many different paths through the fabric (i.e., different switch-switch links - ISLs) as possible in order to provide the best possible performance.
[0453] · According to an embodiment, on the other hand, in order to provide predictable QoS and SLAs for different tenants, it is important that flows belonging to different tenants do not compete for bandwidth on the same ISL simultaneously. In general, this means that it is necessary to restrict which paths different tenants can use.
[0454] However, according to an embodiment, in some cases, depending on the size of the system, the number of tenants, and how servers are provisioned to different tenants, it may not be possible to avoid flows belonging to different tenants competing for bandwidth on the same ISL. In such cases, from an architectural perspective, there are mainly two approaches that can be used to solve the problem and reduce possible contention:
[0455] · According to an embodiment, restrict which switch buffer resources different tenants (or tenant groups) can occupy to ensure that flows from different tenants have forward progress independent of (one or more) other tenants, despite competing for the same ISL bandwidth.
[0456] · According to an embodiment, implement a "permission control" mechanism that will limit the maximum bandwidth that one tenant can consume at the expense of other tenants.
[0457] According to an embodiment, for a physical architecture configuration, the problem is that the bisection bandwidth is as high as possible and ideally non-blocking or even over-provisioned. However, even with non-blocking bisection bandwidth, considering the current allocation of servers to different tenants, there may be scenarios where it is difficult to achieve the desired SLA for one or more tenants. In such cases, the best approach is to re-provision at least some of the servers in the server for different tenants to reduce the need for individual ISLs and generally for the bisection bandwidth.
[0458] According to an embodiment, some multi-rack systems have a blocking fat-tree topology, and the underlying assumption is that the workload will be provisioned such that the associated communication servers are largely within the same rack, which means that a large portion of the bandwidth utilization is only between ports on the local leaf switches. Additionally, in traditional workloads, most of the data traffic is from one fixed set of nodes to another fixed set of nodes. However, according to an embodiment, for next-generation servers with non-volatile memory and newer communication and storage middleware, since different servers may provide multiple functions simultaneously, the communication workload will be even higher and more difficult to predict.
[0459] According to an embodiment, the goal is to provide provisioning granularity at the VM level rather than at the physical server level for each tenant. Additionally, the goal is to be able to deploy up to dozens of VMs on the same physical server, where different sets of VMs on the same physical server may belong to different tenants, and each tenant may individually represent multiple workloads with different characteristics.
[0460] Furthermore, according to an embodiment, although current architecture deployments already use different type-of-service (TOS) associations to provide basic QOS (traffic separation) for different flow types (e.g., to prevent lock messages from becoming "stagnant" after a large volume of data transfer), it is also desirable to provide communication SLAs for different tenants. These SLAs should ensure that tenants experience workload throughput and response times that meet expectations, even when the workload is provisioned on a physical infrastructure shared by other tenants. The relevant SLAs for a tenant should be achieved independently of the concurrent activities of workloads belonging to other tenants.
[0461] According to an embodiment, while it is possible to provision a fixed (minimum) set of CPU cores / threads and physical memory for a workload on a fixed (minimum and / or maximum) set of physical servers, provisioning fixed / guaranteed networking resources is generally less straightforward whenever the deployment implies sharing of HCAs / NICs at the server. The shared nature of the HCA also inherently means that at least the ingress and egress links to / from the fabric are shared by different tenants. Thus, while different CPU cores / threads can operate truly in parallel, there is no way to partition the capacity of a single fabric link other than some form of bandwidth multiplexing or “time sharing”. This basic bandwidth sharing may or may not be combined with the use of different “QOS IDs” (e.g., service level, priority, DSCP, traffic class, etc.) that will be taken into account when buffer selection / assignment and bandwidth arbitration are implemented within the fabric.
[0462] According to an embodiment, the overall server memory bandwidth should be very high relative to the typical memory bandwidth requirements of any individual CPU core in order to prevent memory-intensive workloads on some cores from causing latency to other cores. Similarly, in an ideal scenario, the available fabric bandwidth of a physical server should be large enough to allow each tenant sharing the server to have sufficient bandwidth for the communication activities generated relative to the (one or more) relevant workloads. However, when several workloads all attempt to perform large amounts of data transfer, it is very likely that more than one tenant can utilize the full link bandwidth – even 100 Gb / s and above. To address this scenario, provisioning multiple tenants onto the same physical server needs to be done in a way that ensures that each tenant is guaranteed to receive at least a given minimum percentage of the available bandwidth. However, for RDMA-based communication, the ability to enforce a limit on how much bandwidth a tenant can generate in the egress direction does not mean that the ingress bandwidth can be limited in the same way. That is, although each sender is limited by the maximum send bandwidth, multiple remote communication peers can potentially send data to the same destination in a way that completely overloads the receiver. Additionally, an RDMA read operation can be generated from a local tenant using only trivial egress bandwidth. If a large number of RDMA read operations are generated for multiple remote peers, this can potentially result in disruptive ingress bandwidth. Thus, in order to limit the total fabric bandwidth used by a single tenant on a single server, it is not sufficient to enforce a maximum limit on the egress bandwidth.
[0463] According to an embodiment, the system and method can configure an average bandwidth limit for a tenant, which will ensure that the tenant never exceeds a relative portion of its associated link bandwidth in either the ingress or egress direction, independent of the use of RDMA read operations and independent of the number of remote peers with active data traffic, and independent of the bandwidth limits of the remote peers. (How this is achieved will be discussed in the “Long-Term Goals” section below).
[0464] According to an embodiment, as long as the system and method cannot implement all aspects of communication bandwidth limitation, the highest level communication SLA of a tenant can only be achieved by restricting it from sharing physical servers with other tenants, or potentially, it will not share physical HCAs with other tenants (i.e., in the case of a server with multiple physical HCAs). In the case where the physical HCA can operate in an active-active mode with full link bandwidth utilization of two HCA ports, restrictions can also be used in the case where a given tenant is granted exclusive access to one of the HCA ports under normal circumstances. Nevertheless, due to HA constraints, the failure of a complete HCA (in the case of multiple HCAs per server) or a single HCA port may mean that the reconfiguration and sharing of the expected communication SLA of a given tenant are no longer guaranteed.
[0465] According to an embodiment, in addition to the limitation on the overall bandwidth utilization of a single link, the ability of each tenant to implement QoS between different communication flows or flow types depends on it not experiencing severe congestion conflicts due to the communication activities of other tenants with respect to the architecture-level buffering resources or arbitration. In particular, this means that if a tenant is using a specific "QoS ID" to implement low-latency messaging, then it should not find itself "competing" with a large amount of data traffic from another tenant due to how another tenant is using the "QoS ID" and / or due to how the architecture implementation enforces the use of the "QoS ID" and / or how it maps to packet buffer allocation and / or bandwidth arbitration within the architecture. Therefore, if the tenant communication SLA means that the tenant's internal QoS assumptions cannot be achieved without relying on other tenants sharing the same architecture link(s) to "behave well", then this may force the tenant to be provisioned without sharing an HCA (or HCA port) with other tenants.
[0466] According to an embodiment, for the above basic bandwidth allocation and QoS issues, the sharing constraints apply to the internal architecture links as well as the server-local HCA port links. Therefore, depending on the nature and strictness of the communication SLA of a given tenant, deploying a VM for the tenant may have restrictions on the sharing of physical servers and / or HCAs as well as the sharing of internal architecture ISLs. To avoid ISL sharing, routing restrictions and restrictions on the locations where VMs can be provisioned relative to each other within a private architecture topology can be applied.
[0467] Dynamic vs. Static Packet Routing / Forwarding and Multipath:
[0468] According to an embodiment, as described above, in order to ensure that a tenant can achieve the expected communication performance among a group of communication VMs without relying on the operation of VMs belonging to other tenants, there may be no multiple HCAs / one HCA port or any fabric ISL shared with other tenants. Thus, the highest SLA category provided will typically imply this implementation. This is in principle the same scenario as the current provisioning model of many traditional systems in the cloud. However, for a shared leaf switch, this SLA will need to ensure that there is no ISL shared with other tenants. Additionally, in order for a tenant to achieve the best possible balance of flows and utilization of available fabric resources, it needs to be able to "optimize non-blocking" in an explicit manner (i.e., the communication software infrastructure must provide a way for the tenant to ensure that communication occurs in such a way that different flows do not compete for the same link bandwidth). This will include a way to ensure that communication that can occur via a single leaf switch is actually achieved in this manner. Additionally, in cases where communication must involve an ISL, then it should be possible to balance the traffic among the available ISLs to maximize throughput.
[0469] According to an embodiment, from a single HCA port, as long as the maximum available bandwidth is the same for all links in the fabric, there is no point in attempting to balance the traffic among multiple ISLs. From this perspective, as long as the available ISLs represent a non-blocking sub-topology with respect to the sender, it makes sense to use a dedicated "next-hop" ISL for each sending HCA port. However, unless the relevant ISL represents only the connection between two leaf switches, the scenario of each sender port having a dedicated next-hop ISL is not really sustainable because at some point, if the communication is with multiple remote peer HCA ports connected to different leaf switches, then more than one ISL must be used.
[0470] According to an embodiment, in a non-blocking Infiniband fat-tree topology, popular routing algorithms use a "dedicated down-path", which means that in a non-blocking topology, there are the same number of switch ports in each layer of the fat tree. This means that each end port can have a dedicated port chain from a root switch, through each intermediate switch layer, until the egress leaf switch port connecting the associated HCA port. Thus, all traffic for a single HCA port will use this dedicated down-path, and there will be no traffic to any other destination port (in the down direction) on these links. However, in the up direction, there cannot be a dedicated path to each destination, and as a result, some links in the up direction will have to be shared by traffic to different destinations. In the next round, this can lead to congestion when different flows to different destinations all try to utilize the full bandwidth on the shared intermediate links. Similarly, if multiple senders are sending to the same destination simultaneously, then this can cause congestion in the dedicated down-path, which can then quickly spread to other unrelated flows.
[0471] According to an embodiment, as long as a single destination port belongs to a single tenant, then there is no congestion risk among multiple tenants in the dedicated down-path. However, it is still a problem that different tenants may need to use the same links in the up direction to reach the root switch (or intermediate switch) representing the dedicated down-path. By dedicating as many different root switches as possible to specific tenants, the system and method will reduce the need for different tenants to share paths in the up direction. However, from a single leaf switch, this scheme can reduce the number of available up-links towards the associated root switch(es). Therefore, in order to maintain non-blocking bipartite bandwidth between servers (or rather HCA ports) belonging to the same tenant, the number of servers assigned to a single tenant on a specific leaf switch (i.e., in a single rack) will need to be less than or equal to the number of up-links used by that tenant towards the root switch(es). On the other hand, in order to maximize the ability to communicate via a single cross-switch, it makes sense to assign as many servers as possible to the same tenant within the same rack.
[0472] According to an embodiment, this essentially implies being able to exploit the conflict between the guaranteed bandwidth within a single leaf switch and the guaranteed bandwidth to communication peers in different racks. To resolve this dilemma, the best approach might be to use a scheme in which tenant VMs are grouped based on which leaf switches (i.e., leaf switch pairs) they are directly connected to, and then the properties of the available bandwidth between these groups need to be defined. However, again, there is a trade-off between being able to maximize the bandwidth between two such groups (e.g., between the same tenant in two racks) and being able to guarantee bandwidth to multiple remote groups. Nevertheless, in the special case of having only two tiers of switches (i.e., the leaf tier is interconnected by a single spine tier), a non-blocking topology means that there can always be N dedicated uplinks between the leaf switch with N HCA ports belonging to the same tenant and the N spine ports. Thus, as long as these N spine ports represent the spine that "owns" all the dedicated down paths to all relevant remote peer ports, this configuration is non-blocking for that tenant. However, if the relevant remote peers represent dedicated down paths from more than N spine switches, or if the N uplinks are not distributed among all relevant spine switches, then the system and method have possible contention conflicts with other tenants.
[0473] According to an embodiment, among the VMs of a single tenant, independent of non-blocking or blocking connections, there is still potential contention between flows from different sources connected to the same leaf switch. That is, if the destination has a dedicated down path from the same spine, and the number of uplinks from the source leaf switch to that spine is less than the number of such concurrent flows, then as long as all senders operate at full link speed, some kind of blocking / congestion on the uplink cannot be avoided. In this case, to maintain bandwidth, the only option is to use an auxiliary path via a different spine to one of the destinations. This then would represent a potential conflict with another dedicated down path, since a standard non-blocking fat tree can only have one dedicated downlink per end port.
[0474] According to an embodiment, in the case of some traditional systems, there may be a blocking factor of three (3) between the leaf switch and the spine switch. Thus, in a multi-rack scenario where the workload is distributed in a way that means more than one-third of the communication traffic is between racks rather than within racks, the resulting fragmented bandwidth will be blocked. For example, in an 8-rack system, the most general scenario with uniform traffic distribution between any pair of nodes means that 7 / 8 of the communication is between racks, and the blocking effect will be significant.
[0475] According to an embodiment, if the cost of over-provisioned cables can be tolerated in the system (i.e., given a fixed switch unit cost), then additional links can be used to provide a "backup" downlink for each leaf switch and also provide idle uplink capacity from each leaf to each backbone. That is, in both cases, at least some potential remedies are provided for the dynamic workload distribution representing the uneven distribution of traffic, and thus a topology that is essentially non-blocking first cannot be exploited.
[0476] According to an embodiment, a higher-radix full crossbar switch also has the potential to increase the size of each individual "leaf domain" and reduce the number of backbone switches required for a given system size. For example, in the case of a 128-port switch, two full racks of 32 servers can be contained within a single full crossbar switch leaf domain and still provide non-blocking uplink connections. Similarly, only 8 backbones are required to provide non-blocking connections between 16 racks (512 servers, 1024 HCA ports). Thus, there are still only 8 uplinks from each leaf to each backbone (i.e., in the case of a single fully-connected network). In the extreme case where all HCA ports on a single leaf are sent via a single backbone to a single remote leaf, this still implies a blocking factor of 8. On the other hand, assuming that the dedicated downlink paths of each leaf switch are evenly distributed among all backbones, the likelihood of this extreme scenario should be negligible.
[0477] According to an embodiment, in the case where each leaf switch in a redundant pair of leaf switches belongs to a dual-independent network / track with a dedicated backbone for each track. The same 8 backbones will be divided into two groups of 4 each (one for each track), so in this case each leaf in the track only needs to be connected to 4 backbones. Thus, in this case, the worst-case blocking factor will only be 4. On the other hand, in this scenario, it becomes even more important to select the track for each communication operation to provide load balancing across the two tracks.
[0478] Per - Tenant Bandwidth Permission Control:
[0479] According to an embodiment, while standard InfiniBand uses static routing for each destination address, there are several standard and proprietary schemes for dynamic routing in Ethernet switches. For InfiniBand, there are also various proprietary schemes for "adaptive routing" (some of which may be standardized).
[0480] According to an embodiment, one advantage of dynamic routing is the higher likelihood of optimal utilization of the associated bisection bandwidth within the architecture, and thus also a higher total throughput. However, a potential disadvantage is that ordering may be disrupted and congestion in one area of the architecture may be more likely to spread to other areas (i.e., in a way that could have been avoided if static routing had been used).
[0481] According to an embodiment, while "dynamic routing" or "dynamic routing selection" is generally used for forwarding decisions that occur within and between switches, "multi-path" is the term used when traffic to a single destination can be distributed across multiple paths based on explicit addressing from the sender(s). Such multi-pathing can include "striping" a single message across multiple local HCA ports (i.e., the complete message is divided into multiple sub-messages, each sub-message representing a separate transmission operation), and it can mean that different transmissions to the same destination are set up to use different paths through the architecture in a dynamic manner.
[0482] According to an embodiment, in general, if all transmissions from all sources targeting destinations outside the local leaf domain are divided into (smaller) chunks and then these chunks are distributed across all possible paths / routes towards the destination, the system will achieve optimal utilization of the available bisection bandwidth and will also maximize the "inter-leaf throughput". However, this only holds if the communication workload is also evenly distributed across all possible destinations. If not, then the effect is that any congestion towards a single destination will quickly affect all concurrent flows.
[0483] According to an embodiment, the impact of congestion on dynamic routing and multi-pathing is that it makes sense to limit the traffic to a single destination to using only a single path / route, as long as that route / path is not a victim of congestion for other destinations or any intermediate links. In a two-tier fat tree topology with dedicated down paths, this means that the only possible congestion independent of the end ports will exist on the upstream links to the same backbone switch. This means that it makes sense to treat all upstream links to the same backbone as a group of ports sharing the same static route, except for the individual ports that are dynamically selected for specific destinations. Alternatively, the individual ports can be selected based on tenant association.
[0484] According to an embodiment, using tenant associations to select uplink ports within such groups can be based on fixed associations or on a scenario where different tenants have a "first priority" to use a certain (or certain) port but are also able to use other ports. The ability to use another port will then depend on this not conflicting with the "first priority" traffic of another port. In this way, as long as there are no conflicts, it is possible for tenants to use all relevant half-duplex bandwidth, but when there are conflicts, there will be a guaranteed minimum bandwidth. This minimum guaranteed bandwidth can then reflect all the bandwidth of a single or several links, or a percentage of the bandwidth of one or more links.
[0485] According to an embodiment, in principle, the same dynamic scheme can also be used for the downward path from the backbone to a specific leaf. On the one hand, this increases the risk of congestion caused by sharing the downlink between flows for different end ports, but on the other hand, it can provide a way to utilize additional alternative paths between two sets of nodes connected to two different leaf switches, while still providing a way to prevent congestion propagation between different tenants.
[0486] According to an embodiment, in a scenario where different dedicated downward paths from the backbone to the leaf already represent a specific tenant, a scheme that allows these links to be used as "alternatives" for traffic (belonging to the same tenant) to end ports on the relevant leaf switch with its (primary) dedicated downward path from another backbone will be relatively straightforward.
[0487] According to an embodiment, a possible model would be to have the switch handle dynamic routing between parallel ISLs connecting a single backbone or leaf switch, but with a host-level decision regarding the use of explicit multipath via a (one or more) backbone that does not represent the (primary) dedicated downward path to the relevant target.
[0488] Per - Tenant Bandwidth Reservation on ISLs:
[0489] According to an embodiment, in the case where a single HCA is used only by a single tenant, the system and method can limit the bandwidth that can be generated from the HCA port. In particular, this applies to cases where there is limited half-duplex bandwidth for the tenant for traffic flowing to a (one or more) remote leaf switch.
[0490] According to an embodiment, one aspect of such bandwidth limitation is to ensure that the limitation is only applied to the targets affected by the restricted half-duplex bandwidth. In principle, this would involve a scheme where different target groups are associated with specific bandwidth quotas (i.e., strict maximum rates and / or average bandwidths for transmitting a certain amount of data).
[0491] According to an embodiment, by definition, such a restriction must be implemented at the HCA level. Additionally, such a restriction will more or less directly map to a virtualized HCA scenario where VMs belonging to different tenants share the HCA via different virtual functions. In such a case, the various "shared bandwidth quota groups" introduced above will require additional dimensions in terms of being associated with a group of one or more VFs rather than just a complete physical HCA port.
[0492] Different Priorities, Flow Types, and QOS ID / Classes across ISLs and End - Port Links:
[0493] According to an embodiment, as described above, it can make sense to reserve some guaranteed bandwidth across one or more ISLs for a tenant (or group of tenants). In one scenario, a complete link can be reserved for a (one or more) tenant by restricting which tenants are allowed to use the link. However, for a more flexible and finer-grained solution, an alternative approach is to use a switch arbitration mechanism to ensure that up to X% of the bandwidth of one or more egress ports will be allowed to be used by a certain (certain) ingress port, regardless of which other ingress ports are competing for bandwidth on the same egress port.
[0494] According to an embodiment, in this way, all ingress ports can use up to 100% of the bandwidth of the (one or more) associated egress ports, as long as this does not conflict with any traffic from a prioritized ingress port.
[0495] According to an embodiment, in a scenario where different tenants "own" different ingress ports (e.g., the leaf switch ports connected to the HCA ports), then this solution will facilitate a flexible and fine-grained solution for allocating uplink bandwidth to one or more backbone switches.
[0496] According to an embodiment, in the downlink path from the backbone to the leaf switches, the usefulness of this solution will depend on the extent to which a solution with a strictly dedicated downlink path is used or not. If a strictly dedicated downlink path is used and the target end port represents a single tenant, then by default, there is no potential conflict between different tenants trying to use the downlink. Therefore, in such a case, access to the relevant downlink should generally be set to use a round-robin arbitration scheme with equal access rights for all relevant ingress ports.
[0497] According to an embodiment, since the ingress ports can represent traffic belonging to different tenants, it should never be a problem that packets belonging to one tenant can be sent to an egress port that the relevant tenant is not allowed to send to and consume the bandwidth on that egress port. In such a case, assume that strict access control (e.g., VLAN-based restrictions on various ports) rather than an arbitration policy is adopted to prevent such packets from wasting any bandwidth.
[0498] According to an embodiment, on a leaf switch, a downstream port from the backbone may be given more bandwidth towards various end ports relative to other local end ports because, in principle, the downstream link can represent multiple sender HCA ports while the local end port represents only a single HCA port. If this were not the case, then a scenario could occur where several remote servers share a single downstream path to a target leaf switch, but then in the next round, in the case where the N-1 HCA ports directly connected to the leaf switch also attempt to send to the same local target port, the bandwidth towards a single destination would be shared 1 / N on that leaf.
[0499] According to an embodiment, in the case where virtualized HCAs represent different tenants, the problem of reserving bandwidth within the architecture (i.e., across various) ISLs can become much more complex. For the ingress / upstream link path, a simplified approach is for the HCA to provide bandwidth arbitration between different tenants, and then anything sent out on the HCA port will be disposed of by the ingress leaf switch according to the port-level arbitration policy. Thus, from the perspective of the leaf switch, there is no change in this case.
[0500] According to an embodiment, in the downstream link path (from the backbone to the leaf and from the leaf ingress to the end port), the situation is different because the arbitration decision may depend not only on the port attempting to forward the packet but also on which tenant the various outstanding packets belong to. One possible solution is to (again) limit some ISLs to represent only a specific tenant (or group of tenants) and then reflect this in the port-level arbitration scheme. Alternatively (or additionally), different priorities or QoS IDs can be used to represent different tenants, as described below. Finally, using the "tenant ID" or any relevant access control header field (such as VLAN ID or partition ID) as part of the arbitration logic will help achieve the required level of granularity for arbitration. However, this may significantly increase the complexity of the arbitration logic in a switch that already has significant "time and space" complexity. Additionally, since such schemes involve an overload of information that may already be in use in the end-to-end wired protocol, it is important that such additional complexity does not conflict with any existing use or assumptions about the values of such header fields.
[0501] Lossless vs. Lossy Packet Forwarding:
[0502] According to an embodiment, in order for different flow types to make progress simultaneously on the same link, it is crucial that they do not compete for the same packet buffers in the switch and the HCA. Additionally, in order to distinguish the relative priorities between different flow types, the arbitration logic that determines what packets to send next on each switch egress port must take into account which packet type queues have something to emit on which egress port. The result of the arbitration should be that all active flows progress forward according to their relative priorities and the extent to which the flow control conditions of the relevant flow types on the relevant downstream ports (if any) currently allow them to send any packets.
[0503] According to an embodiment, in principle, different QOS IDs can be used to keep traffic flows from different tenants independent of each other, even if they are using the same link. However, since the number of packet queues and independent buffer pools that each port can support is typically limited to less than 10, the scalability of this approach is very limited. Additionally, when a single tenant wants to use different QOS IDs to keep different flow types independent of each other, then the scalability is further reduced.
[0504] According to an embodiment, as described above, by logically combining multiple ISLs between a single pair of switches, the system and method can then restrict some links to some tenants and then ensure that different tenants can use different QOS IDs on different ISLs independently of each other. However, again, if 100% independence from other tenants is to be guaranteed, this places a limit on the total bandwidth available to any single tenant.
[0505] According to an embodiment, ideally, HCA ingress (receive) packet processing can always occur at a rate higher than the relevant link speed, regardless of what transport level operation the incoming packet represents. This means that there is no need for flow control of different flow types on the egress ports of the leaf switch that is the last link (i.e., connected to the HCA port). However, the scheduling of different packets in different queues in the leaf switch must still reflect the relevant policies of priority, fairness, and forward progress. For example, if a small high-priority packet is destined for an end port while N ports are also trying to send "bulk transfer packets" of the maximum MTU size to the same destination port, then the high-priority packet should be scheduled before any others.
[0506] According to an embodiment, in the egress path, the sending HCA can schedule and label packets in many different ways. In particular, using an overlay protocol as a "bump in the wire" between the VM + virtual HCA and the physical architecture will allow encoding of architecture-specific information that may be relevant to the switch without messing up any aspect of the end-to-end protocol between tenant virtual HCA instances.
[0507] According to an embodiment, a switch can provide more buffering and internal queuing than currently assumed by wired protocols. In this way, buffering, queuing, and arbitration policies can be set that take into account that the link is shared by traffic from multiple tenants with different SLAs and use different QOS classes for different flow types.
[0508] According to an embodiment, in this way, different high-priority tenants may also have more private packet buffer capacity within the switch.
[0509] Strict vs. (More) Relaxed Packet Ordering:
[0510] According to an embodiment, high-performance RDMA traffic heavily depends on each packet not being lost due to lack of buffer capacity in the switch and also on the packets arriving in the correct order for each individual RDMA connection. In principle, the higher the potential bandwidth, the more critical these aspects are for achieving optimal performance.
[0511] According to an embodiment, lossless operation requires explicit flow control, and very high bandwidth means a trade-off between buffer capacity, MTU size, and flow control update frequency.
[0512] According to an embodiment, a disadvantage of lossless operation is that when the total bandwidth generated is higher than the downstream / receiving capacity, it will cause congestion. The congestion will then (most likely) spread and will ultimately slow down all flows competing for the same buffers somewhere within the architecture.
[0513] According to an embodiment, as described above, the ability to provide flow separation based on independent buffer pools is a major scalability issue, which for switch implementations depends on the number of ports, the number of different QOS classes, and (as described above) may also depend on the number of different tenants.
[0514] According to an embodiment, an alternative approach could be to make true lossless operation (i.e., lossless based on guaranteed buffer capacity) an "advanced SLA" attribute and thereby limit this feature to tenants that have purchased such an advanced SLA.
[0515] According to an embodiment, the key issue here is the ability to "overbook" the available buffer capacity so that the same buffers can be used for lossy and lossless flows, but the buffers allocated to lossy flows can be preempted at any time when packets from lossless flows arrive and need to use the buffers from the same pool. A very small set of buffers can be provided to allow the lossy flow to progress, but at a bandwidth that is much lower than what could be achieved with optimized buffer allocation.
[0516] According to an embodiment, different classes of hybrid lossless / lossy flow classes can also be introduced based on the difference in the maximum time a buffer can be occupied before it must be pre-empted and given to a higher SLA type of flow class (when this is required). This works best in the context of an architecture implementation with link-level credits, but can potentially also apply to working with xon / xoff type flow control (i.e., Ethernet pause-based flow control schemes for RoCE / RDMA).
[0517]
[0518] According to an embodiment, through strict ordering and lossless packet forwarding within the architecture, the HCA implementation can achieve reliable connections and RDMA at the transport level with minimal state overhead. However, to better tolerate a certain amount of out-of-order packet delivery due to occasional changes in routing (due to adaptive / dynamic forwarding decisions within the architecture), and also to minimize the overhead and latency associated with (one or more) packet losses due to lossy or "hybrid lossless / lossy" mode forwarding within the architecture, an efficient transport implementation will need to maintain sufficient state to allow a large number of individual packets (sequence numbers) to arrive out of order and be individually retried while other packets with later sequence numbers are being accepted and acknowledged.
[0519] According to an embodiment, the key point here is to avoid long delays and average bandwidth losses when lost or out-of-order packets cause retries using the current default transport implementation. Additionally, by preventing subsequent packets in a series of posted packets from being dropped, the system and method also significantly reduce waste of the architecture bandwidth, which otherwise might consume a large amount of bandwidth that could have been consumed by other flows.
[0520] Shared services and shared HCAs:
[0521] According to an embodiment, shared services on an architecture (e.g., a backup device) used by multiple tenants means that some end-port links will be shared by different tenants, unless the service can provide (one or more) end-ports that can be dedicated to a specific tenant (or restricted tenant group). A similar scenario exists when VMs belonging to multiple tenants are sharing the same server and (one or more) the same HCA ports.
[0522] According to an embodiment, finely tuned server and HCA resources can be allocated to different tenants, and it can also be ensured that the outgoing data traffic bandwidth from the HCA is fairly divided among different tenants according to the relevant SLA levels.
[0523] According to an embodiment, within the architecture, it is also possible to set up packet buffer allocation and queuing priorities, as well as arbitration policies, that reflect relative importance and thus reflect fairness between data traffic belonging to different tenants. However, even with very finely tuned buffer allocation and arbitration policies within the architecture, the granularity may not be fine enough to ensure that the relative priorities and bandwidth quotas of different tenants are accurately reflected in terms of the ingress bandwidth to the shared HCA port.
[0524] According to an embodiment, to achieve such fine-grained bandwidth allocation, a dynamic end-to-end traffic control scheme is needed that can effectively partition and schedule the available ingress bandwidth among multiple remote communication peers belonging to one or more tenants.
[0525] According to an embodiment, the goal of such a scheme would be that at any point in time, the relevant set of active remote clients is able to utilize a fair (not necessarily equal) share of their available ingress bandwidth. Additionally, this bandwidth utilization should occur without causing congestion in the architecture due to attempts to use too much bandwidth at the end ports. (Nonetheless, architecture-level congestion may still occur due to overloading of shared links within the rest of the architecture).
[0526] According to an embodiment, a high-level model for achieving this goal would be that the receiver is able to dynamically allocate and update the available bandwidth for the relevant set of remote clients. The current bandwidth value for each remote client needs to be calculated based on what is currently being provided to each client and what is needed next.
[0527] According to an embodiment, this means that if a single client is currently allowed to use all the available bandwidth and then another client also needs to use the ingress bandwidth, then an update instruction must be passed to the currently active client that informs the client of the new reduced maximum bandwidth, and an instruction must be passed to the new client indicating that it can use the maximum bandwidth corresponding to the reduction of the current client.
[0528] According to an embodiment, the same scheme would then in principle apply to "any" number of concurrent clients. However, there is of course a significant trade-off between being able to ensure that the available bandwidth is not "overbooked" at any point in time and ensuring that the available bandwidth is always fully utilized when needed.
[0529] According to an embodiment, an additional challenge of such a scheme is to ensure that it interoperates well with dynamic congestion control and also to ensure that congestion related to the shared path for multiple destinations is handled in a coordinated manner within each sender.
[0530] High availability and failover:
[0531] According to an embodiment, in addition to performance, key attributes of a private architecture can be redundancy and the ability to failover communication without loss of service for any client applications after any single point of failure. Nevertheless, while "service loss" represents a binary condition (i.e., the service either exists or is lost), some equally important but more scalar attributes are the extent of any power-down time during failover and, if so, for how long. Another key aspect is the extent to which expected performance is provided (or re-established) during and after the completion of the failover operation.
[0532] According to an embodiment, while from the perspective of a single node (server), the goal that there should be no single point of failure in the architecture communication infrastructure outside the server itself (i.e., including a single local HCA) should mean that the node becomes unable to communicate. However, from the perspective of the complete architecture, there is also a question about the extent to which the loss of one or more components means that the throughput and performance across the architecture are affected. For example, in a case where the topology size can only operate with two backbone switches. Then, in terms of bisection bandwidth and the increased risk of congestion, if one of the backbones stops service and 50% of the leaf-leaf communication capacity is lost, is this acceptable?
[0533] According to an embodiment, another issue related to the SLA per tenant is the extent to which the ability to reserve and / or prioritize architecture resources for tenants with a premium SLA after a failure and subsequent failover operations should be reflected in such tenants obtaining a proportionally larger share of the remaining available resources? That is, in this way, the impact of the failure on premium SLA tenants will be smaller but at the cost of a greater impact on other tenants.
[0534] According to an embodiment, in terms of redundancy, it can also be a "super-premium SLA attribute", that is, the initial resource provisioning for such tenants will ensure that no single point of failure will mean that the relevant performance / QOS SLA will not be met either during or after a failure. However, the fundamental problem with this over-provisioning is that there must be extremely fast failover (and failure recovery / rebalancing) in order to ensure that available resources are always utilized in the best possible way and that no communication will stop for more than a very trivial period due to any single point of failure.
[0535] According to an embodiment, an example of such a "super-premium" setup can be a system with a server based on dual HCAs, where both HCAs are operating in an active-active manner and where both HCA ports are also utilized in an active-active manner using an APM (Automatic Path Migration) scheme with a very short latency before attempting alternative paths.
[0536] Path selection:
[0537] According to an embodiment, when there are multiple possible paths between two endpoints, ideally, the best or "correct" path should be automatically selected for the associated RDMA connection such that the communication workload experiences the best possible performance within the constraints of the associated SLA and also such that the system-level architecture resources are utilized in an optimized manner.
[0538] According to an embodiment, ideally, this would mean that the application logic within the VM would not have to deal with which local HCAs and which local HCA ports can or should be used for what communication. This also means an addressing scheme at the node level rather than the port level and that the underlying architecture infrastructure is used transparently to the application.
[0539] According to an embodiment, in this way, the associated workload can be more easily deployed on different infrastructures without having to explicitly handle different system types or system configurations.
[0540] Features:
[0541] According to an embodiment, it is assumed that the features in this category are supported by existing HCA and / or switch hardware using current firmware and software.
[0542] According to an embodiment, the main objectives in this category are as follows:
[0543] · To be able to limit the total egress bandwidth generated by the VFs belonging to a single VM or tenant of each local physical HCA instance.
[0544] · To be able to ensure that the VFs belonging to a single VM or tenant of a local physical HCA will be able to utilize at least a minimum percentage of the available local link bandwidth.
[0545] · To be able to limit the network (Enet) priorities that the VFs can use.
[0546] ο This may mean that multiple VFs must be allocated in order for a single VM to use multiple priorities (i.e., as long as only a single VF is allowed to use a single priority when priority limiting is enabled).
[0547] · To be able to limit which ISLs can be used by the flow groups belonging to a single tenant or tenant group.
[0548] According to an embodiment, in order to control the HCA usage of a tenant that is sharing a physical HCA with other tenants, a "HCA Resource Limit Group" (referred to herein as "HRLG") will be established for the tenant. The HRLG can be set with a maximum bandwidth, which defines the actual data rate that the HRLG can generate, and the HRLG can also be set with a minimum bandwidth share, which will ensure that the HRLG will achieve at least a specified percentage of the HCA bandwidth when there is contention with other tenants / HRLGs. As long as there is no contention with other HRLGs, the VFs in the HRLG can permanently use up to the specified rate (or link capacity if no rate limit is defined).
[0549] According to an embodiment, the HRLG can contain up to the number of VFs that the HCA instance can support. Within the HRLG, it is desirable for each VF to receive a fair share of the "quota" allocated to the HRLG. For each VF, the associated QP will also obtain its fair access share of the local link based on the available HRLG quota and any current flow control limits of the QP. (That is, if the QP has received congestion control feedback indicating that it should throttle itself, or if there is no "credit" to send on the relevant priority currently, then the QP will not be considered for local link access).
[0550] According to an embodiment, within the HRLG, restrictions can be imposed on which priorities the VF can use. As long as this restriction can only be defined based on a single priority allowed for the VF, it means that a VM that should use multiple priorities (but still limited to only some priorities) will have to use multiple VFs - one VF for each required priority. (Note: Using multiple VFs means that sharing local memory resources among multiple QPs using different priorities may pose problems, as it means that depending on which priority restrictions / execution policies are defined, the ULP / application in the VM must allocate and use different VFs).
[0551] According to an embodiment, within a single HRLG, there is no difference in bandwidth allocation depending on the priority that the VF / QP is currently using. They all share the relevant quota in a fair / equal manner. Therefore, in order to associate different bandwidth quotas with different priorities, one or more dedicated HRLGs need to be defined, and the one or more dedicated HRLGs will only include VFs restricted to using the priorities associated with the shared quota represented by the relevant HRLG. In this way, different bandwidth quotas can be given to a VM or a tenant with multiple VMs sharing the same physical HCA for different priorities.
[0552] According to an embodiment, current hardware priority restrictions prevent data sent with an illegal priority from being sent onto an external link, but do not prevent the relevant data from being fetched from local memory. Thus, there is still wasted overall HCA link bandwidth in cases where the local memory bandwidth that the HCA can sustain in the egress direction is roughly the same as the available external link bandwidth. However, if the relevant memory bandwidth is (significantly) greater than the external link bandwidth, then as long as the HCA pipeline is operating at optimal efficiency, attempting to use an illegal priority will waste less external link bandwidth. Nevertheless, as long as there is not much to be saved in terms of external link bandwidth, a possible alternative to preventing the use of illegal priorities could be to enforce it using ACL rules in the switch ingress port. If the relevant tenant can be effectively identified without any possibility of spoofing, then this can be used to enforce tenant / priority association without having to allocate a separate VF for each priority of the same VM. However, ensuring that the packet / tenant association is always well-defined and impossible to spoof from the sending VM, and dynamically updating the relevant switch ports whenever a VF is set for use by a VM / tenant to perform the relevant enforcement presents non-trivial complexity. One possible solution is to use each VF port MAC to represent a non-spoofable ID that can be associated with a VM / tenant. However, if VxLAN or other overlay protocols are being used, then this is not straightforward - especially as long as the external switch is not supposed to participate (or know) in the overlay scheme being used.
[0553] According to an embodiment, to limit which flows can use which ISLs, the switch forwarding logic needs to have a policy for identifying the relevant flows and setting forwarding accordingly. An example is using VLAN IDs to represent flow groups. In cases where different tenants are mapped to different VLAN IDs on the fabric, one possible solution is that the switch can dynamically implement LAG type balancing based on which VLAN IDs are allowed on each port in any LAG or other port group. Another option would involve explicitly forwarding packets based on a combination of the destination address and VLAN ID.
[0554] According to an embodiment, in the case of transparently using VxLAN-based overlay on a physical switch fabric, different overlays can then be mapped to different VLAN IDs to allow the switch to map VLAN IDs to ISLs as described above.
[0555] According to an embodiment, another possible solution is to set the forwarding of each endpoint address according to a routing scheme that takes into account some other concept such as VLAN membership or "tenant" association. However, as long as the same endpoint address value is allowed in different VLANs, the VLAN ID needs to be part of the forwarding decision.
[0556] According to an embodiment, distributing each tenant flow to a shared or exclusive ISL may require an overall routing scheme to distribute traffic in a globally optimized manner within the architecture (fat tree) topology. The implementation of such a scheme typically depends on the SDN type management interface of the switch, but the implementation of overall routing is not easy.
[0557] Short-term and medium-term SLA categories:
[0558] According to an embodiment, the following assumes that a non-blocking two-layer fat tree topology is being used for system sizes (physical node counts) beyond the base of a single leaf switch. Additionally, it is assumed that a single VM on a physical server will be able to use all of the architecture bandwidth (via one or more vHCA / VFs). Therefore, from the perspective of the HCA / architecture, the number of VMs per tenant per physical server is not a parameter that needs to be considered as a tenant-level SLA factor.
[0559] According to an embodiment, the top layer (e.g., Premium Plus)
[0560] · Can use dedicated servers.
[0561] · Can be allocated on the same leaf domain as much as possible, unless the number and size of the VMs (or HA policies) imply additional distance.
[0562] · If a tenant (i.e., relative to the number of servers allocated to that tenant within the same leaf domain) is using more than one leaf domain, then on average, non-blocking uplink bandwidth can be obtained from the local leaf, but there will be no dedicated uplink or uplink bandwidth.
[0563] · Can be able to use all "flow groups" (i.e., representing different buffer pools and arbitration group priorities, etc. within the architecture).
[0564] According to an embodiment, the lower layer (e.g., Premium)
[0565] · Can use dedicated servers, but the same leaf domain is not guaranteed.
[0566] · Can have, on average (i.e., relative to the number of servers allocated to this tenant within the same leaf domain), at least 50% non-blocking uplink bandwidth, but there will be no dedicated uplink or uplink bandwidth.
[0567] · Can be able to use all "flow groups".
[0568] According to an embodiment, the third layer (e.g., Economy Plus)
[0569] · May use shared servers, but will have dedicated "flow groups".
[0570] ο These resources will be dedicated to the local HCA and switch ports, but will be shared within the architecture.
[0571] · It can have the ability to use all available egress bandwidth from the local server, but is guaranteed to have at least 50% of the total egress bandwidth.
[0572] · Each physical server has only one Economy Plus tenant.
[0573] · On average, it can have at least 25% non-blocking leaf uplink bandwidth relative to the number of servers being used by the Economy Plus tenant.
[0574] According to an embodiment, the fourth layer (e.g., Economy)
[0575] · Can use shared servers
[0576] · Has no dedicated priority
[0577] · Can be allowed to use up to 50% of the server egress bandwidth, but can share this bandwidth with at most 3 other Economy tenants.
[0578] · On average, it can share up to 25% of the available leaf uplink bandwidth with other Economy tenants within the same leaf domain.
[0579] According to an embodiment, the lowest layer (e.g., Spare)
[0580] · Can use idle capacity with no guaranteed bandwidth
[0581] Long-term features:
[0582] According to an embodiment, the main features discussed in this section are as follows:
[0583] · It can enforce per-VF priority limits in a way that allows restricting any subset of the total set of supported priorities for a single VF.
[0584] · Whenever an attempt is made to transfer data using a priority not allowed for the initial VF, the wasted memory bandwidth or link bandwidth is zero.
[0585] · It can limit the egress rate of different individual priorities across multiple HRLGs such that each VF in each HRLG will receive its fair share of the total minimum bandwidth and / or maximum rate of the associated HRLG, but subject to the constraints defined by the various associated per-priority quotas.
[0586] · It can perform sender bandwidth control and congestion adjustment based on per-destination and per-shared path / routing, and perform this aggregation at the VM / VF (i.e., vHCA port) level and the HCA port level.
[0587] · Capable of restricting the total average send and RDMA write ingress bandwidth to the VF based on receiver throttling of cooperative remote senders.
[0588] · Capable of restricting the average RDMA read ingress bandwidth to the VF without relying on cooperative remote RDMA read responders.
[0589] · When restricting the total average ingress bandwidth to the VF based on receiver throttling of cooperative remote senders, capable of including RDMA reads in addition to sends and RDMA writes.
[0590] · Tenant VMs are able to observe the available segmented bandwidth for different groups relative to peer VMs.
[0591] · SDN features for routing control and arbitration policies within the architecture.
[0592] According to an embodiment, the HCA VF context can be extended to include a legal priority list (similar to the legal SL set of an IBTA IBvPort). Whenever a work request is attempting to use a priority that is not legal for the VF, the work request should fail before any local or remote data transfer is initiated. Additionally, a priority mapping can be used so as to give an application the illusion that it can use any priority. However, the drawback of this mapping that can map multiple priorities to the same value before sending a packet is that in terms of associating different flow types with different "QOS classes", the application may no longer be able to control its own QOS policy. This restricted mapping represents SLA attributes (i.e., more privileged SLA means more actual priorities after mapping). However, it is always important for the application to be able to decide which flow types are associated with which QOS classes (priorities) in a way that will also represent independent flows within the architecture.
[0593] Inbound RDMA read BW quota via target group and dynamic BW quota updates:
[0594] According to an embodiment, as long as the destination group association of the flow from the "producer / sender" node implies bandwidth regulation for all outgoing data packets - including UD sends, RDMA writes, RDMA sends, and RDMA reads (i.e., RDMA read responses with data) - full control of all ingress bandwidth of the vHCA port can be achieved. This is independent of whether the VM owning the destination vHCA port is generating an "excessive" number of RDMA read requests to multiple peer nodes.
[0595] According to an embodiment, as described above, the coupling of the destination group with flow - specific and "active" BECN signaling means that the ingress bandwidth of each vHCA port can be dynamically throttled for any number of remote peers.
[0596] According to an embodiment, in addition to the pure CE marking / unmarking for different stage numbers, the above "active BECN" message can also be used to convey a specific rate value. In this way, there may be a scenario where an initial incoming packet (e.g., a CM packet) from a new peer can trigger the generation of one or more "active BECN" messages to the HCA (i.e., the relevant firmware / hyper-privileged software) from which the incoming packet originated and to the current communication peer.
[0597] According to an embodiment, in the case of using two ports on the HCA simultaneously (i.e., the active-active scenario), if it is possible that concurrent flows share some ISLs or even target the same destination port, then it may make sense to share the target group between the local HCA ports.
[0598] According to an embodiment, another reason for sharing the target group between HCA ports is whether the HCA local memory bandwidth cannot sustain the full link speed of both (all) HCA ports. In this case, the target group can be set such that the total aggregated link bandwidth never exceeds the local memory bandwidth, regardless of which port on the source or destination HCA is involved.
[0599] According to an embodiment, in the case of having a fixed route towards a specific destination, any (one or more) intermediate target groups will generally represent only a single ISL at a specific stage in the path. However, when dynamic forwarding is active, both the target group and the ECN handling must take this into account. In the case where the dynamic forwarding decision will only occur to balance the traffic between parallel ISLs (e.g., the uplinks from a single leaf switch to a single spine switch) between a pair of switches, then all the handling principles are in principle very similar to when only a single ISL is used. The FECN notification will occur based on the status of all ports in the relevant group, and the signaling can be "aggressive" in that it is signaled based on congestion indications from any port, or it can be more conservative and based on the size of the shared output queue of all ports in the group. The target group configuration generally represents the aggregated bandwidth of all the links in the group, as long as the forwarding allows any packet to select the best output port at that point in time. However, if there is a strict concept of packet order preservation for each flow, then the evaluation of the bandwidth quota is more complex because some flows may "have to" use the same ISL at a certain point in time. If such a flow order scheme is based on well-defined header fields, then it is better to represent each port in the group as an independent target group. In this case, the selection of the target group on the sender-side HCA must be able to perform the same evaluation of the header fields that will be associated with the RC QP connection or address handle as the switch will perform for each packet at runtime.
[0600] According to an embodiment, by default, the initial target group rate for a new remote target can be conservatively set low. In this way, there is an inherent throttling before the target has a chance to update the relevant rate. Thus, all such rate controls are independent of the VM itself involved, but the VM will be able to request the hypervisor to update the quotas for ingress and egress traffic for different remote peers, but this can only be granted within the total constraints defined for the local and remote vHCA ports.
[0601] Associate peer nodes, paths, and target groups:
[0602] According to an embodiment, in order for a VM to identify the bandwidth limitations associated with different peer nodes and different peer node groups, there must be a way to query which target groups are associated with various communication peers (along with the associated address / path information). Based on associating a set of communication peers with various target groups and the rate limitations represented by the various target groups, the VM will be able to track what bandwidth can be achieved relative to the various communication peers. In principle, this will allow the VM to schedule communication operations in such a way that over time, the best possible bandwidth utilization is achieved by having as many concurrent transmissions as possible that do not involve conflicting target groups.
[0603] Relationship between HCA resource limit groups and target groups:
[0604] According to an embodiment, the HRLG concept and the target group concept overlap in several ways because they both represent bandwidth limitations that can be defined and shared in a flexible manner between the VM and the tenant. However, while the main concern of HRLG is to define how different portions of the local HCA / HCA port capacity are allocated to different VFs (and thereby to the VM and the tenant), the target group concept, in terms of the final destination as well as the intermediate architecture topology, concerns the traffic control constraints and bandwidth limitations that exist outside the local HCA.
[0605] According to an embodiment, in this way, it makes sense to use HRLG as a way to control which shares of the local HCA capacity can be used by the various VFs, but ensure that the granted capacity can only be used in a way that does not conflict with any architectural or remote target limitations or congestion conditions. These external constraints are then dynamically controlled and reflected via the relevant target groups.
[0606] According to an embodiment, in terms of implementation, the state of all relevant target groups will define which outstanding work requests in the flow control state the local QP will be in that allows them to generate more egress data traffic at any point in time. This state, along with the state regarding what QP actually has what to send, can then be aggregated at the VF / vHCA port level in terms of which VFs are candidates for the next transmission. The decision on which VF to schedule for the next transmission on the HCA port will be based on the state and policies of various HRLGs in the HRLG hierarchy, the set of "ready to send" VFs, and the recent history in terms of what egress traffic the VFs have generated. For the selected VF, the VF-specific arbitration policy will define which QP will be selected for data transmission.
[0607] According to an embodiment, since the set of QPs with outstanding data transmissions includes QPs with local work requests and QPs with outstanding RDMA read requests from related remote peers, the above scheduling and arbitration will be responsible for all outstanding egress data traffic.
[0608] According to an embodiment, the ingress traffic (including incoming RDMA read responses) will be controlled by the current state of all relevant target groups in the remote peer node. This (remote) state will include a dynamic flow control state based on congestion conditions and explicit updates from this HCA reflecting changes in the ingress bandwidth quota of the local VFs on this HCA. Such ingress bandwidth quotas will be based on the policies reflected by the HRLG hierarchy. In this way, various VMs can have "fine-tuned" independent bandwidth quotas for ingress and egress, and on a per-priority basis for ingress and egress.
[0609] SLA category:
[0610] According to an embodiment, the following proposal assumes that a non-blocking two-tier fat-tree topology is being used for system sizes (physical node counts) beyond the cardinality of a single leaf switch. Additionally, it is assumed that a single VM on a physical server will be able to use all the fabric bandwidth (via one or more HCA VFs). Therefore, the number of VMs per tenant per physical server is not a parameter that needs to be considered as an SLA factor.
[0611] According to an embodiment, the highest level tier (e.g., Premium Plus)
[0612] · can only use dedicated servers.
[0613] · can be allocated on the same leaf domain as much as possible, unless the number and size of the VMs (or HA policies) imply additional distance.
[0614] · If the tenant is using more than one leaf domain, then it can always have non-blocking uplink bandwidth from the local leaf.
[0615] · It is possible to use all "flow groups".
[0616] According to an embodiment, a lower level layer (e.g., premium) can be provided
[0617] · It is possible to use only dedicated servers, but the same leaf domains are not guaranteed.
[0618] · At least 50% non-blocking uplink bandwidth can be guaranteed (i.e., relative to the number of servers allocated to this tenant within the same leaf domain).
[0619] · It is possible to use all "flow groups".
[0620] According to an embodiment, a third level layer (e.g., economy plus) can be provided
[0621] · It may be possible to use shared servers, but there will be 4 dedicated "flow groups" (i.e., representing priorities such as different buffer pools and arbitration groups within the architecture).
[0622] ο These resources will be dedicated to local HCAs and switch ports, but will be shared within the architecture.
[0623] · It is possible to use all available bandwidth (egress and ingress) of local servers, but at least 50% of the total bandwidth is guaranteed.
[0624] · Each physical server is limited to one economy plus tenant.
[0625] · At least 25% non-blocking leaf uplink bandwidth can be guaranteed relative to the number of servers being used by this economy plus tenant.
[0626] According to an embodiment, a fourth layer (e.g., economy) can be provided
[0627] · It is only possible to use shared servers
[0628] · There is no dedicated priority
[0629] · It is possible to allow the use of up to 50% of the server bandwidth (egress and ingress), but this bandwidth can be shared with up to 3 other economy tenants
[0630] · It is possible to share up to 25% of the available leaf uplink bandwidth with other economy tenants within the same leaf domain.
[0631] According to an embodiment, a bottom layer (e.g., standby) that can use idle capacity without guaranteed bandwidth can be provided.
[0632] Although various embodiments of the present teachings have been described above, it should be understood that they are presented by way of example and not limitation. The embodiments were chosen and described to explain the principles of the present teachings and their practical applications. The embodiments illustrate systems and methods where the present teachings are utilized to improve the performance of systems and methods by providing new and / or improved features and / or by providing benefits such as reduced resource utilization, increased capacity, improved efficiency, and reduced latency.
[0633] In some embodiments, features of the present teachings are implemented, in whole or in part, in a computer that includes a processor, a storage medium such as a memory, and a network card for communicating with other computers. In some embodiments, features of the present teachings are implemented in a distributed computing environment where one or more computer clusters are connected by a network such as a local area network (LAN), a switched fabric network (e.g., InfiniBand), or a wide area network (WAN). The distributed computing environment can have all computers located at a single location or can have clusters of computers at different remote geographical locations connected by a WAN.
[0634] In some embodiments, features of the present teachings are implemented, in whole or in part, in the cloud, as part of or as a service of a cloud computing system that delivers shared, elastic resources to users in a self-service, metered fashion using web technologies. There are five characteristics of the cloud (defined by the National Institute of Standards and Technology: on-demand self-service; broad network access; resource pooling; rapid elasticity; and measured service). See, for example, “The NIST Definition of Cloud Computing”, Special Publication 800-145 (2011), which is incorporated herein by reference. Cloud deployment models include: public, private, and hybrid. Cloud service models include software as a service (SaaS), platform as a service (PaaS), database as a service (DBaaS), and infrastructure as a service (IaaS). As used herein, the cloud is a combination of hardware, software, networks, and web technologies that delivers shared, elastic resources to users in a self-service, metered fashion. Unless otherwise specified, as used herein, the cloud includes public cloud, private cloud, and hybrid cloud embodiments, and all cloud deployment models, including but not limited to cloud SaaS, cloud DBaaS, cloud PaaS, and cloud IaaS.
[0635] In some embodiments, the features of the present teachings are implemented using hardware, software, firmware, or combinations thereof, or with the assistance of hardware, software, firmware, or combinations thereof. In some embodiments, the features of the present teachings are implemented using a processor configured or programmed to perform one or more functions of the present teachings. In some embodiments, the processor is a single-chip or multi-chip processor, a digital signal processor (DSP), a system-on-chip (SOC), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a state machine, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. In some implementations, the features of the present teachings may be implemented by circuitry specific to a given function. In other implementations, the features may be implemented in a processor configured to execute a specific function using instructions stored on, for example, a computer-readable storage medium.
[0636] In some embodiments, the features of the present teachings are incorporated in software and / or firmware for controlling the hardware of a processing and / or networking system and for enabling a processor and / or network to interact with other systems using the features of the present teachings. Such software or firmware may include, but is not limited to, application code, device drivers, operating systems, virtual machines, hypervisors, application programming interfaces, programming languages, and execution environments / containers. Based on the teachings of the present disclosure, appropriate software code can be readily prepared by a skilled programmer, as will be apparent to those skilled in the software art.
[0637] In some embodiments, the present teachings include a computer program product, such as a computer-readable medium that conveys instructions that can be used to effectuate the present teachings. In some examples, the computer-readable medium is a storage medium or computer-readable media which conveys instructions by storing the instructions thereon / therein, which can be used to program a system such as a computer or otherwise configure it to perform any process or function of the present teachings. The storage medium or computer-readable media may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of media or device suitable for storing instructions and / or data. In certain embodiments, the storage medium or computer-readable media is a non-transitory storage medium or non-transitory computer-readable media. The computer-readable media may also include or alternatively include transitory media, such as a carrier wave or transmission signal that conveys such instructions.
[0638] Thus, from one perspective, systems and methods have been described for using multiple CE (Congestion Experienced) flags in both FECN (Forward Explicit Congestion Notification) and BECN (Backward Explicit Congestion Notification) in a high-performance computing environment. An example method can provide a first subnet that includes a plurality of switches, a plurality of host channel adapters, and a plurality of end nodes. The method can receive an ingress packet from a remote end node at an end node attached to a host channel adapter, where the ingress packet traverses at least a portion of the first subnet before being received at the end node. The method can, upon receiving the ingress packet, send a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet.
[0639] The foregoing description is not intended to be exhaustive or to limit the scope to the precise forms disclosed. Additionally, in instances where embodiments of the present teachings are described using a particular series of transactions and steps, it should be clear to those skilled in the art that the scope is not limited to the described series of transactions and steps. Further, in instances where embodiments have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of the present teachings. Additionally, while various embodiments describe particular combinations of features of the present invention, it should be understood that those skilled in the relevant art will appreciate that different combinations of features are within the scope of the present teachings such that features of one embodiment can be incorporated into another embodiment. Moreover, it will be clear to those skilled in the relevant art that various additions, subtractions, deletions, variations, and other modifications and changes can be made in form, detail, implementation, and application without departing from the spirit and scope. The present invention is intended to be defined by the proper interpretation of the following claims.
Claims
1. A system for using multiple congestion experienced (CE) flags in both forward explicit congestion notification (FECN) and backward explicit congestion notification (BECN) in a high-performance computing environment, comprising: A computer including one or more microprocessors; A first subnet provided at the computer, the first subnet including: A plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports, A plurality of host channel adapters, wherein each host channel adapter of the plurality of host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches, and A plurality of end nodes, the plurality of end nodes including a plurality of virtual machines; Wherein an ingress packet is received at an end node attached to a host channel adapter from a remote end node, wherein the ingress packet traverses at least a portion of the first subnet along a path including a plurality of path stages before being received at the end node, wherein the ingress packet includes a bit field, the bit field including a plurality of bits indicating that the ingress packet experienced congestion at a first path stage among the plurality of path stages during traversing the at least a portion of the first subnet, wherein each bit of the bit field corresponds to a different path stage among the plurality of path stages, and wherein the indication of congestion at the first path stage is set within the bit field of the ingress packet by a switch port associated with the first path stage; Wherein upon receiving the ingress packet, the end node determines different switch ports at which the ingress packet experienced congestion by associating information of the path with the ingress packet, and sends a response message to the remote end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet.
2. The system according to claim 1, wherein the response message includes a bit field of the response message including a plurality of bits.
3. The system according to claim 2, wherein each path stage includes one of an inter-switch link or a switch port.
4. The system according to claim 3, wherein each of the plurality of bits of the bit field of the response message corresponds to a different path stage among the plurality of path stages respectively.
5. The system according to claim 4, wherein a first bit corresponding to the first path stage among the plurality of path stages in the bit field of the response message indicates that the ingress packet experienced congestion at the first path stage within the response message.
6. The system according to claim 5, wherein the first bit corresponding to the first path stage among the plurality of path stages in the bit field of the response message includes a CE flag.
7. The system according to claim 6, wherein the first subnet includes an InfiniBand subnet.
8. A method for using multiple congestion experienced (CE) flags in both forward explicit congestion notification (FECN) and backward explicit congestion notification (BECN) in a high-performance computing environment, comprising: Providing a first subnet at one or more microprocessors, the first subnet comprising: A plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports, A plurality of host channel adapters, wherein each host channel adapter of the plurality of host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches, and A plurality of end nodes, the plurality of end nodes including a plurality of virtual machines; Receiving an ingress packet at an end node attached to a host channel adapter from a remote end node, wherein the ingress packet traverses at least a portion of the first subnet along a path including a plurality of path segments before being received at the end node, wherein the ingress packet includes a bit field, the bit field including a plurality of bits indicating that the ingress packet experienced congestion at a first path segment of the plurality of path segments during traversing the at least a portion of the first subnet, wherein each bit of the bit field corresponds to a different path segment of the plurality of path segments, and wherein the indication of congestion at the first path segment is set within the bit field of the ingress packet by a switch port associated with the first path segment; Upon receiving the ingress packet, determining, by the end node, different switch ports at which the ingress packet experienced congestion by associating information of the path with the ingress packet, and sending a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet.
9. The method according to claim 8, wherein the response message includes a bit field of the response message including a plurality of bits.
10. The method according to claim 9, wherein each path segment includes one of an inter-switch link or a switch port.
11. The method according to claim 10, wherein each of the plurality of bits of the bit field of the response message corresponds to a different path segment of the plurality of path segments, respectively.
12. The method according to claim 11, wherein a first bit corresponding to the first path segment of the plurality of path segments in the bit field of the response message indicates that the ingress packet experienced congestion at the first path segment within the response message.
13. The method according to claim 12, wherein the first bit corresponding to the first path segment of the plurality of path segments in the bit field of the response message includes a CE flag.
14. The method according to claim 13, wherein the first subnet includes an InfiniBand subnet.
15. A computer-readable medium having instructions thereon for using a plurality of congestion experienced (CE) flags in both forward explicit congestion notification (FECN) and backward explicit congestion notification (BECN) in a high-performance computing environment, the instructions, when read and executed, causing a computer to perform steps including the following: Provide a first subnet at one or more microprocessors, the first subnet including: A plurality of switches, the plurality of switches including at least leaf switches, wherein each switch of the plurality of switches includes a plurality of switch ports, A plurality of host channel adapters, wherein each host channel adapter of the plurality of host channel adapters includes at least one host channel adapter port, and wherein the plurality of host channel adapters are interconnected via the plurality of switches, and A plurality of end nodes, the plurality of end nodes including a plurality of virtual machines; Receive an ingress packet at an end node attached to a host channel adapter from a remote end node, wherein the ingress packet traverses at least a portion of the first subnet along a path including a plurality of path segments before being received at the end node, wherein the ingress packet includes a bit field including a plurality of bits indicating that the ingress packet experienced congestion at a first path segment of the plurality of path segments during traversing the at least a portion of the first subnet, wherein each bit of the bit field corresponds to a different path segment of the plurality of path segments, and wherein the indication of congestion at the first path segment is set within the bit field of the ingress packet by a switch port associated with the first path segment; Upon receiving the ingress packet, determine, by the end node, different switch ports at which the ingress packet experienced congestion by associating information of the path with the ingress packet, and send a response message from the end node attached to the host channel adapter to the remote end node, the response message indicating that the ingress packet experienced congestion during traversing the at least a portion of the first subnet.
16. The computer-readable medium of claim 15, wherein the response message includes a bit field of the response message including a plurality of bits.
17. The computer-readable medium of claim 16, wherein each path segment includes one of an inter-switch link or a switch port.
18. The computer-readable medium of claim 17, wherein each of the plurality of bits of the bit field of the response message corresponds to a different path segment of the plurality of path segments, respectively.
19. The computer-readable medium of claim 18, wherein a first bit corresponding to the first path segment of the plurality of path segments in the bit field of the response message indicates that the ingress packet experienced congestion at the first path segment within the response message.
20. The computer-readable medium of claim 19, wherein the first bit corresponding to the first path segment of the plurality of path segments in the bit field of the response message includes a CE flag.
Citation Information
Patent Citations
Method and apparatus for preventing activation of a congestion control process
US20070076598A1
Destination-based congestion control
US20130135999A1
System and method for supporting dual-port virtual router in a high performance computing environment
US20170257312A1