System and method for supporting fast hybrid reconfiguration in high performance computing environment

A hybrid reconfiguration scheme with custom fat-tree routing optimizes InfiniBand networks for fast partial reconfiguration, addressing performance bottlenecks and enabling efficient VM migration in large-scale cloud computing environments.

JP2025157261APending Publication Date: 2025-10-15ORACLE INT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025107101
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-08-23
Filing Date
2025-06-25
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing cloud computing environments face performance and management bottlenecks due to complex addressing and routing schemes in lossless networks, particularly when live migration of virtual machines (VMs) is required, which disrupts connectivity and impairs network performance.

Method used

A hybrid reconfiguration scheme using a custom fat-tree routing algorithm with node ordering for partial network reconfiguration, combined with a default Fat Tree routing algorithm, to manage different use cases in a single subnet, optimizing performance and resource utilization in large-scale InfiniBand networks.

Benefits of technology

Enables fast partial network reconfiguration, reducing downtime during VM migration and improving overall network performance by maintaining efficient resource utilization and scalability in high-performance computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157261000001_ABST
    Figure 2025157261000001_ABST
Patent Text Reader

Abstract

To support computer system virtualization and live migration using SR-IOV vSwitch architecture.SOLUTION: A hybrid reconfiguration scheme can allow for fast partial network reconfiguration with different routing algorithms of choice in different subparts of the network. Partial reconfigurations can be orders of magnitude faster than the initial full configuration, thus making it possible to consider performance-driven reconfigurations in lossless networks.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Copyright notice: A portion of the disclosure of this patent document contains material that is subject to copyright protection. As this patent document or patent disclosure appears in the Patent and Trademark Office patent file or records, the copyright owner has no objection to the facsimile reproduction thereof by anyone, but otherwise reserves all copyright rights whatsoever.

[0002] Field of the invention: The present invention relates generally to computer systems, and more particularly to supporting computer system virtualization and live migration using the SR-IOV vSwitch architecture. [Background technology]

[0003] background: As larger-scale cloud computing architectures are deployed, performance and management bottlenecks associated with traditional networks and storage are becoming serious issues. There is growing interest in using InfiniBand (IB) technology as the foundation for cloud computing fabrics. This is the general area that embodiments of the present invention are intended to address. Summary of the Invention [Means for solving the problem]

[0004] overview: According to one embodiment, a system and method can provide performance-driven reconfiguration in large-scale lossless networks. A hybrid reconfiguration scheme can enable fast partial network reconfiguration using different routing algorithms to select different subsections of the network. Because partial reconfiguration can be orders of magnitude faster than an initial global configuration, it may be possible to consider performance-driven reconfiguration in lossless networks. The proposed mechanism takes advantage of the fact that large HPC systems and clouds are shared by multiple tenants (e.g., different tenants on different partitions) performing isolated tasks. In such scenarios, intercommunication between tenants is impossible, and thus workload deployment and placement schedulers should attempt to avoid fragmentation to ensure efficient resource utilization. That is, most of the traffic per tenant can be contained within an aggregated subsection of the network, and the SM can reconfigure some subsections to improve overall performance. The SM can use a fat-tree topology and a fat-tree routing algorithm. Such a hybrid reconfiguration scheme can successfully reconfigure and improve performance within a subtree by using a custom fat-tree routing algorithm with the provided node ordering to reconfigure the network. If the SM wishes to reconfigure the entire network, it can use the default Fat Tree routing algorithm to effectively provide a combination of two different routing algorithms for various use cases in a single subnet.

[0005] According to one embodiment, a method for fast hybrid reconstruction in a high performance computing environment is provided. An exemplary method for supporting a network may provide a first subnet in one or more microprocessors. The first subnet includes a plurality of switches, the plurality of switches including at least leaf switches, each of the plurality of switches including a plurality of switch ports, a plurality of host channel adapters, each of the plurality of end nodes including at least one host channel adapter port, and a plurality of end nodes, each of the plurality of end nodes being associated with at least one host channel adapter of the plurality of host channel adapters. The method may arrange the plurality of switches of the first subnet in a network architecture having a plurality of levels, each of the plurality of levels including at least one switch of the plurality of switches. The method may configure the plurality of switches according to a first configuration method. The first configuration method is associated with a first ordering of the plurality of end nodes. The method may configure a subset of the plurality of switches as a sub-subnet of the first subnet. The sub-subnet of the first subnet includes several levels less than the plurality of levels of the first subnet. The method may reconfigure the sub-subnet of the first subnet according to a second configuration method. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 illustrates an example of an InfiniBand environment according to one embodiment. [Figure 2] FIG. 1 illustrates an example of a tree topology in a network environment, according to one embodiment. [Figure 3] FIG. 1 illustrates an exemplary shared port architecture according to one embodiment. [Figure 4] FIG. 1 illustrates an exemplary vSwitch architecture according to one embodiment. [Figure 5] FIG. 2 illustrates an exemplary vSwitch architecture with pre-populated LIDs, according to one embodiment. [Figure 6]FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID allocation, according to one embodiment. [Figure 7] FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs in accordance with one embodiment. [Figure 8] FIG. 2 illustrates a switch tuple according to one embodiment. [Figure 9] FIG. 1 illustrates a system for node routing stages, according to one embodiment. [Figure 10] FIG. 1 illustrates a system for node routing stages, according to one embodiment. [Figure 11] FIG. 1 illustrates a system for node routing stages, according to one embodiment. [Figure 12] FIG. 1 illustrates a system for node routing stages, according to one embodiment. [Figure 13] FIG. 1 illustrates a system including a fat tree topology with more than two levels, according to one embodiment. [Figure 14] FIG. 1 illustrates a system for fast hybrid reconstruction, according to one embodiment. [Figure 15] FIG. 1 illustrates a system for fast hybrid reconstruction, according to one embodiment. [Figure 16] 1 is a flowchart illustrating an exemplary method for supporting fast hybrid reconfiguration in a high performance computing environment, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] Detailed Description: The present invention is illustrated, by way of limitation, in the figures of the accompanying drawings in which like reference numerals refer to like elements and in which: The embodiments are described for purposes of illustration and not for purposes of illustration. Note that references in this disclosure to "an" or "one" or "several" embodiments are not necessarily to the same embodiment, and such references mean at least one. While specific implementations are described, it is understood that these specific implementations are provided for illustrative purposes only. One skilled in the art will recognize that other components and configurations can be used without departing from the scope and spirit of the invention.

[0008] Common reference numbers may be used to denote like elements throughout the drawings and detailed description, and thus a reference number used in one drawing may or may not be referenced in the detailed description specific to that drawing if the element is described elsewhere.

[0009] SUMMARY OF THE INVENTION Systems and methods for supporting fast hybrid reconfiguration in high performance computing environments are described herein.

[0010] The following description of the present invention uses an InfiniBand (IB) network as an example of a high-performance network. It will be apparent to those skilled in the art that other types of high-performance networks can be used without any limitation. The following description also uses a Fat Tree topology as an example of a fabric topology. It will be apparent to those skilled in the art that other types of fabric topologies can be used without any limitation.

[0011] In accordance with an embodiment of the present invention, virtualization can be beneficial for efficient resource utilization and elastic resource allocation in cloud computing. Live migration can optimize resource usage by moving virtual machines (VMs) between physical servers in an application-transparent manner. Because of this, virtualization can enable consolidation, on-demand provisioning of resources, and elasticity through live migration.

[0012] InfiniBand(R) InfiniBand (IB) is a registered trademark of the InfiniBand Trade Association. TM Open standard lossless network technology developed by the International Trade Association The technology is based on a serial point-to-point full-duplex interconnect that provides high-throughput and low-latency communications, especially targeted at high-performance computing (HPC) applications and data centers.

[0013] InfiniBand Architecture (IBA) is a two-tier topology. At a lower level, an IB network is referred to as a subnet, and a subnet may include a set of hosts interconnected using switches and point-to-point links. At a higher level, an IB fabric consists of one or more subnets that may be interconnected using routers.

[0014] Within a subnet, hosts may be connected using switches and point-to-point links. In addition, there may be one master management entity, a subnet manager (SM), that resides on a designated subnet device in the subnet. The subnet manager is responsible for configuring, starting, and maintaining the IB subnet. In addition, the subnet manager (SM) may be responsible for performing routing table calculations in the IB fabric. Here, for example, routing in an IB network involves all source-destination pairs in the local subnet. The goal is to achieve fair load balancing between

[0015] Through the subnet management interface, the subnet manager exchanges control packets called subnet management packets (SMPs) with the subnet management agent (SMA). An agent resides on all IB subnet devices. Using SMP, the subnet manager can discover the fabric, configure end nodes and switches, and receive notifications from the SMA.

[0016] According to one embodiment, inter- and intra-subnet routing in an IB network may be based on the LFT stored in the switch. The LFT is calculated by the SM according to the routing mechanism in use. In a subnet, Host Channel Adapter (HCA) ports on end nodes and switches are addressed using local identifiers (LIDs). Each entry in the LFT consists of a destination LID (DLID) and an output port. Only one entry per LID is supported. When a packet arrives at a switch, its output port is determined by looking up the DLID in the switch's forwarding table. Routing is deterministic because packets follow the same path in the network between a given source-destination pair (LID pair).

[0017] Generally, all other subnet managers except the master subnet manager operate in standby mode for fault tolerance. However, in the situation where the master subnet manager fails, a new master subnet manager is negotiated by the standby subnet managers. The master subnet manager also performs periodic sweeps of the subnet to detect any topology changes and migrates the subnets accordingly. Reconfigure the network.

[0018] Additionally, hosts and switches within a subnet can be addressed using a local identifier (LID), and a single subnet can be limited to 49151 unicast LIDs. In addition to the LID, which is a local address valid within the subnet, each IB device can have a 64-bit global unique identifier (GUID). The GUID can be used to form a global identifier (GID), which is an IB Layer 3 (L3) address.

[0019] The SM may calculate routing tables (i.e., connections / routes between each pair of nodes in a subnet) at network initialization time. Additionally, whenever the topology changes, the routing tables may be updated to ensure connectivity and optimal performance. During normal operation, the SM may perform periodic light sweeps of the network to check for topology changes. If a change is discovered during a light sweep, Alternatively, if the SM receives a message (trap) signaling a network change, the SM may reconfigure the network according to the discovered change.

[0020] For example, the SM may reconfigure the network when the network topology changes, such as when a link goes down, a device is added, or a link is removed. The reconfiguration step may include a step performed during network initialization. Furthermore, the reconfiguration may have a local scope that is limited to the subnet where the network change occurred. Also, segmentation of a large fabric using routers may limit the scope of the reconfiguration.

[0021] According to one embodiment, an IB network is a system that shares a network fabric. The fabric may support partitioning as a security mechanism to provide isolation of logical groups of systems. Each HCA port on a node in the fabric may be a member of one or more partitions. Partition membership is managed by a centralized partition manager, which may be part of the SM. The SM may organize the partition membership information for each port as a table of 16-bit partition keys (P-keys). The SM may also associate a LID with the partition membership information. Switches and routers can be configured with a partition enforcement table containing the assigned P-key information. Additionally, in the general case, the partition membership of a switch port may represent the set of all memberships indirectly associated with LIDs routed through the port in the egress direction (towards the link).

[0022] According to one embodiment, for communication between nodes, queue pairs (QPs) and end-to-end contexts (EECs), with the exception of the management queue pair (QP0 and QP1), can be assigned to specific partitions. P-key information can then be added to all transmitted IB transport packets. When a packet arrives at an HCA port or switch, its P-key value can be checked against a table configured by the SM. If an invalid P-key value is found, the packet is immediately discarded. In this way, communication is only allowed between ports that share a partition.

[0023] An example of an InfiniBand fabric is shown in Figure 1, which illustrates an example InfiniBand environment 100 according to one embodiment. In the example shown in Figure 1, nodes A101-E105 communicate using an InfiniBand fabric 120 via respective host channel adapters 111-115. According to one embodiment, various nodes (e.g., nodes A101-E105) may be represented by various physical devices. According to one embodiment, various nodes (e.g., nodes A101-E105) may be represented by various virtual devices, such as virtual machines.

[0024] Virtual Machines on InfiniBand Over the past decade, hardware virtualization support has virtually eliminated CPU overhead, memory overhead has been significantly reduced by virtualizing the memory management unit, storage overhead has been reduced by utilizing high-speed SAN storage or distributed network file systems, and device pass-through technologies such as Single Root Input / Output Virtualization (SR-IOV) have been introduced. The prospects for virtualized High Performance Computing (HPC) environments have improved significantly as network I / O overhead has been reduced by using high-performance interconnect solutions. Clouds now support virtual HPC (vHPC) clusters with high-performance interconnect solutions, delivering the required performance. can be provided.

[0025] However, when coupled with lossless networks such as InfiniBand (IB), some cloud features such as live migration of virtual machines (VMs) remain problematic due to the complex addressing and routing schemes used in these solutions.IB is an interconnect network technology that offers high bandwidth and low latency, making it well suited for HPC and other communication-intensive workloads.

[0026] The traditional approach to connect IB devices to VMs is by using direct-assigned SR-IOV. However, to achieve live migration of IB-assigned VMs, a Host Channel Adapter (HCA) using SR-IOV is required. CA) has proven to be challenging. Each IB-connected node has three different addresses (i.e., LID, GUID, and GID). When a live migration occurs, one or more of these addresses change. Other nodes communicating with the migrating VM (VM-in-migration) may lose connectivity. When this occurs, the IB subnet manager (SM) can reconnect by sending a Subnet Administration (SA) record route query. It can attempt to restore the lost connection by locating the new address of the virtual machine that should be connected.

[0027] IB uses three different types of addresses. The first type of address is a 16-bit local identifier (LID). At least one unique LID is assigned by the SM to each HCA port and each switch. The LID is used to route traffic within a subnet. Because the LID is 16 bits long, 65,536 unique address combinations can be configured, of which only 49,151 (0x0001-0xBFFF) can be used as unicast addresses. As a result, the number of available unicast addresses defines the maximum size of an IB subnet. The second type of address is a 64-bit globally unique identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM may assign an additional subnet-specific GUID to an HCA port, which is useful when SR-IOV is used. The third type of address is a 128-bit global identifier (GID). A GID is a valid IPv6 unicast address, at least one of which is assigned to each HCA port. The GID is formed by combining a globally unique 64-bit prefix assigned by the fabric administrator with the GUID address of each HCA port.

[0028] Fat Tree (FTree) Topology and Routing According to one embodiment, some IB-based HPC systems employ a fat-tree topology to take advantage of the useful properties that fat trees offer, including full bisection bandwidth and inherent fault tolerance due to the availability of multiple paths between each source-destination pair. The initial idea behind fat-trees was to employ thicker links between nodes with more available bandwidth as the tree approached the root of the topology. The thicker links could help avoid congestion in higher-level switches, preserving bisection bandwidth.

[0029] FIG. 2 illustrates an example of a tree topology in a network environment, according to one embodiment. As shown in FIG. 2, one or more end nodes 201-204 may be connected in a network fabric 200. Network fabric 200 may be based on a fat-tree topology including multiple leaf switches 211-214 and multiple spine or root switches 231-234. In addition, network fabric 200 may include one or more intermediate switches, such as switches 221-224.

[0030] 2, each of end nodes 201-204 may be a multi-homed node, i.e., a single node that is connected to two or more portions of network fabric 200 through multiple ports. For example, node 201 may include ports H1 and H2, node 202 may include ports H3 and H4, node 203 may include ports H5 and H6, and node 204 may include ports H7 and H8.

[0031] In addition, each switch may have multiple switch ports. For example, root switch 231 has switch ports 1-2, and root switch 232 has switch ports 3-4. The root switch 233 may have switch ports 5-6, and the root switch 234 may have switch ports 7-8.

[0032] According to one embodiment, the Fat Tree routing mechanism is one of the most popular routing algorithms for IB-based Fat Tree topologies. The Fat Tree routing mechanism is also used in OFED (Open Fabric Enterprise Data Exchange). Distribution: A standard software for building and deploying IB-based applications This is implemented in the OpenSM (Software Stack) subnet manager.

[0033] The goal of the fat-tree routing mechanism is to generate an LFT that uniformly spreads shortest-path routes across the links in the network fabric. The mechanism traverses the fabric in indexing order and assigns the target LID of the end node, and therefore the corresponding route, to each switch port. For end nodes connected to the same leaf switch, the indexing order may depend on the switch port to which the end node is connected (i.e., the port numbering sequence). For each port, the mechanism may maintain a port usage counter, and each time a new route is added, this port usage counter may be used to select the least-used port.

[0034] According to one embodiment, in a partitioned subnet, nodes that are not members of a common partition are not allowed to communicate. In practice, this means that some of the routes assigned by the Fat Tree routing algorithm will not be used for user traffic. A problem arises if the Fat Tree routing mechanism generates LFTs for those routes in the same way as other functional paths. This behavior can degrade balancing on links because nodes are routed in indexing order. Because routing is done without awareness of partitions, Fat Tree routed subnets typically result in poor isolation between partitions.

[0035] Input / Output (I / O) virtualization According to one embodiment, I / O Virtualization (IOV) can make I / O available by allowing virtual machines (VMs) to access the underlying physical resources. The combination of storage traffic and inter-server communication can place an overwhelming burden on a single server's I / O resources, resulting in backlogs and idle processors waiting for data. As the number of I / O requests increases, IOV can provide availability and improve the performance, scalability, and elasticity of (virtualized) I / O resources to rival performance levels seen in modern CPU virtualization.

[0036] According to one embodiment, IOV is desired to enable sharing of I / O resources and to allow protected access to resources from VMs. IOV separates the logical device exposed to a VM from its physical implementation. Currently, emulation, paravirtualization, direct assignment (DA), and single-root I / O There can be various types of IOV technologies, such as single root-I / O virtualization (SR-IOV).

[0037] According to one embodiment, one type of IOV technology is software emulation. Software emulation can enable a separated front-end / back-end software architecture. The front-end can be a device driver located in the VM, and the back-end can be a device driver implemented by the hypervisor to provide I / O access. The physical device sharing ratio is high and live migration of VMs is possible with only milliseconds of network downtime. However, software emulation introduces additional and undesirable computational overhead.

[0038] According to one embodiment, another type of IOV technology is direct device assignment. Direct device assignment requires that an I / O device be attached to a VM, but the device is not shared between VMs. Direct assignment, or device passthrough, offers near-unique performance with minimal overhead. The physical device bypasses the hypervisor and is directly attached to the VM. However, a drawback of such direct device assignment is that there is no sharing between virtual machines, limiting scalability, such as one physical network card being attached to one VM.

[0039] According to one embodiment, Single Root IOV (SR-IOV) is Hardware virtualization may allow a physical device to appear as multiple independent, lightweight instances of the same device. These instances can be assigned to VMs as pass-through devices and accessed as Virtual Functions (VFs). The hypervisor accesses the device through a unique (per device) fully functional Physical Function (PF). SR-IO SR-IOV mitigates the scalability issues of purely direct allocation. However, a problem presented by SR-IOV is that it can impair VM migration. Among these IOV technologies, SR-IOV extends the PCI Express (PCIe) standard with a means to allow multiple VMs to directly access a single physical device while maintaining near-inherent performance. This allows SR-IOV to offer superior performance and scalability.

[0040] SR-IOV allows a PCIe device to expose multiple virtual devices that can be shared among multiple guests by assigning one virtual device to each guest. Each SR-IOV device has at least one physical function (PF) and one or more associated virtual functions (VFs). A PF is a communication function controlled by a virtual machine monitor (VMM) or hypervisor. VFs are lightweight PCIe functions, whereas VFs are regular PCIe functions. Each VF has its own base address (BAR) and is assigned a unique requestor ID. The unique requestor ID is managed by the I / O memory management unit (I / O memory management unit). The IOMMU also applies memory and interrupt translation between PFs and VFs.

[0041] Unfortunately, direct device allocation techniques present a barrier to cloud providers in situations where transparent live migration of virtual machines is desired for data center optimization. The essence of live migration is that the memory contents of a VM are copied to a remote hypervisor. Furthermore, the VM is suspended in the source hypervisor, and the VM's operation is resumed in the destination. When using software emulation methods, network interfaces are virtual so that their internal states are stored in memory and then copied. Therefore, downtime can be reduced to a few milliseconds.

[0042] However, migration becomes more difficult when direct device assignment technologies such as SR-IOV are used. In this situation, the network interface The entire internal state of the device cannot be copied because it is tied to the hardware. Instead, the SR-IOV VF assigned to the VM is detached, a live migration is performed, and a new VF is assigned at the destination. In the case of InfiniBand and SR-IOV, this process can cause downtime on the order of several seconds. Furthermore, in the SR-IOV shared port model, the VM's address changes after migration, which adds overhead to the SM and negatively impacts the performance of the underlying network fabric.

[0043] InfiniBand SR-IOV Architecture - Shared Port There can be various types of SR-IOV models (eg, a shared port model and a virtual switch model).

[0044] 3 illustrates an exemplary shared port architecture according to one embodiment. As shown, a host 300 (e.g., a host channel adapter) may interact with a hypervisor 310. The hypervisor 310 may assign various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions may be handled by the hypervisor 310.

[0045] 3, a host (e.g., an HCA) appears as a single port to the network with a single shared LID and shared Queue Pair (QP) space between the physical function 320 and the virtual functions 330, 350, 350. However, each function (i.e., the physical function and the virtual function) may have its own GID.

[0046] 3, according to one embodiment, various GIDs can be assigned to virtual and physical functions, and a special queue pair, QP0 and QP1 (i.e., a dedicated queue pair used for InfiniBand management packets), is owned by the physical function. These QPs are exposed to VFs as well, but VFs are not allowed to use QP0 (all incoming SMPs from VFs towards QP0 are discarded), and QP1 can act as a proxy for the actual QP1 owned by the PF.

[0047] According to one embodiment, the shared port architecture may enable highly scalable data centers that are not limited by the number of VMs (attached to the network by being assigned to virtual functions) because LID space is only consumed by the physical machines and switches in the network.

[0048] However, a drawback of the shared port architecture is that it cannot provide transparent live migration, thereby hindering the potential for flexible VM placement. Because each LID is associated with a specific hypervisor and shared among all VMs residing on that hypervisor, a migrating VM (i.e., a virtual machine migrating to a destination hypervisor) must change its LID to the LID of the destination hypervisor. Furthermore, as a result of the restricted QP0 access, a subnet manager cannot be run inside a VM.

[0049] InfiniBand SR-IOV Architecture Model - Virtual Switch (vSwitch) There can be various types of SR-IOV models (eg, a shared port model and a virtual switch model).

[0050] 4 illustrates an exemplary vSwitch architecture according to one embodiment. As shown, a host 400 (e.g., a host channel adapter) can interact with a hypervisor 410, which can assign various virtual functions 430, 440, and 450 to several virtual machines. Similarly, physical functions can be handled by hypervisor 410. A virtual switch 415 can also be handled by hypervisor 401.

[0051] According to one embodiment, in the vSwitch architecture, each virtual function 430, 440, 450 is a full virtual Host Channel Adapter (vHCA), which means that in hardware, the VM assigned to the VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM, HCA 400 appears as a switch with additional nodes connected via virtual switch 415. Hypervisor 410 can use PF 420, and the VM (attached to the virtual function) uses the VF.

[0052] According to one embodiment, the vSwitch architecture provides transparent virtualization. However, because each virtual function is assigned a unique LID, the available number of LIDs is quickly consumed. Similarly, if many LID addresses are used (i.e., one for each physical function and each virtual function), more communication paths must be calculated by the SM and more subnet management packets (SMPs) must be sent to the switch to update their LFTs. For example, calculating communication paths can take several minutes in a large network. Because the LID space is limited to 49,151 unicast LIDs and each VM (through a VF) occupies one LID per physical node and switch, the number of active VMs is limited by the number of physical nodes and switches in the network, and vice versa.

[0053] InfiniBand SR-IOV Architecture Model - LID Pre-Populated vSwitch According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs.

[0054] FIG. 5 illustrates an exemplary vSwitch architecture with pre-populated LIDs, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 500 (e.g., an IB subnet). The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. Furthermore, host channel adapters 510, 520, and 530 can interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. The hypervisor 511 also assigns the virtual machine 2 551 to the virtual function 2 515 and assigns the virtual machine 3 Hypervisor 531 may assign virtual machine 552 to virtual function 3 516. Hypervisor 531 may further assign virtual machine 4 553 to virtual function 1 534. The hypervisor may access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0055] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 500.

[0056] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0057] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs. Referring to FIG. 5, LIDs are pre-populated for various physical functions 513, 523, and 533, as well as for virtual functions 514-516, 524-526, and 534-536 (even virtual functions not currently associated with active virtual machines). For example, physical function 513 is pre-populated with LID 1, and virtual function 1 534 is pre-populated with LID 10. When a network is booted, LIDs are pre-populated in an SR-IOV vSwitch-enabled subnet. Populated VFs are assigned LIDs as shown in FIG. 5, even if not all of the VFs are occupied by VMs in the network.

[0058] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0059] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, each hypervisor consumes one LID for itself via the PF and can consume one or more LIDs for each additional VF. The sum of all VFs available across all hypervisors in an IB subnet gives the maximum amount of VMs that can run in the subnet. For example, in an IB subnet with 16 virtual functions per hypervisor in the subnet, each hypervisor consumes 17 LIDs in the subnet (one LID for each of the 16 virtual functions and one LID for the physical function). In such an IB subnet, the theoretical hypervisor limit for a single subnet is defined by the number of available unicast LIDs: 2891 (49151 available LIDs divided by 17 LIDs per hypervisor), and the total number of VMs (i.e., limit) is 46256 (2891 hypervisors multiplied by 16 VFs per hypervisor). (In practice, these numbers are smaller, as each switch, router, or dedicated SM node in the IB subnet consumes LIDs as well.) Note that vSwitches do not need to occupy additional LIDs, as they can share LIDs with PFs.

[0060] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, once the network is booted, communication paths are calculated for all LIDs. When a new VM needs to be started, the system creates a new LID in the subnet. There is no need to add a new LID, which would otherwise require a complete reconfiguration of the network, including route recalculation, which is the most time-consuming operation. Instead, available ports for the VM are located in one of the hypervisors (i.e., available virtual functions), and the virtual machine is assigned to an available virtual function.

[0061] According to one embodiment, the LID pre-populated vSwitch architecture also enables the ability to compute and use different routes to reach different VMs hosted by the same hypervisor. Essentially, this allows such subnets and networks to use LID-Mask-Control-like (LMC-like) features to provide alternative routes towards one physical machine without being bound by the LMC constraint that requires LIDs to be contiguous. The freedom to use non-contiguous LIDs is particularly useful when a VM needs to migrate and its associated LID needs to be delivered to the destination.

[0062] In accordance with one embodiment, several considerations can be taken into account along with the above-described advantages of a LID pre-populated vSwitch architecture. For example, because LIDs are pre-populated in an SR-IOV vSwitch-enabled subnet when the network is booted, initial route computation (e.g., at startup) may take longer than if the LIDs were not pre-populated.

[0063] InfiniBand SR-IOV Architecture Model - vSwitch with Dynamic LID Allocation According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID allocation.

[0064] FIG. 6 illustrates an exemplary vSwitch architecture with dynamic LID assignment, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 600 (e.g., an IB subnet). The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. The host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 may additionally assign virtual machine 2 551 to virtual function 2 515 and virtual machine 3 552 to virtual function 3 516. Hypervisor 531 may further assign virtual machine 4 553 to virtual function 1 534. The hypervisor may access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0065] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 600.

[0066] According to one embodiment, virtual switches 512, 522, and 532 may be handled by respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a complete virtual host channel adapter. The HCAs 510, 520, and 530 are virtual HCAs, which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), the HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0067] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID assignment. Referring to FIG. 6 , various physical functions 513, 523, and 533 are dynamically assigned LIDs, with physical function 513 receiving LID 1, physical function 523 receiving LID 2, and physical function 533 receiving LID 3. Those virtual functions associated with active virtual machines may also receive dynamically assigned LIDs. For example, because virtual machine 1 550 is active and associated with virtual function 1 514, virtual function 514 may be assigned LID 5. Similarly, virtual function 2 515, virtual function 3 516, and virtual function 1 534 are each associated with an active virtual function. As such, these virtual functions are assigned LIDs, with LID 7 assigned to virtual function 2 515, LID 11 assigned to virtual function 3 516, and virtual function 9 assigned to virtual function 1 535. Unlike a vSwitch with pre-populated LIDs, virtual functions that are not currently associated with an active virtual machine do not receive an LID assignment.

[0068] According to one embodiment, dynamic LID assignment can substantially reduce initial route computation: When a network is booting for the first time and no VMs are present, a relatively small number of LIDs can be used for initial route computation and LFT distribution.

[0069] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0070] According to one embodiment, when a new VM is created in a system utilizing a vSwitch with dynamic LID allocation, a free VM slot is discovered and a unique, unused unicast LID is discovered as well to determine on which hypervisor the newly added VM should boot. However, there is no known route in the switch's LFT and network to handle the newly added LID. Computing a new set of routes to handle the newly added VM is undesirable in a dynamic environment where several VMs may be booted every minute. In a large IB subnet, computing a new set of routes can take several minutes, and this procedure would have to be repeated each time a new VM is booted.

[0071] Advantageously, according to one embodiment, since all VFs in a hypervisor share the same uplink with the PF, there is no need to calculate a new set of routes. All that is required is to iterate through the LFTs of all physical switches in the network, copy the forwarding ports from the LID entries belonging to the PF of the hypervisor (on which the VM is created) to the newly added LID, and send a single SMP to update the corresponding LFT block of the particular switch. This eliminates the need for the system and method to calculate a new set of routes.

[0072] According to one embodiment, the LIDs assigned in a vSwitch using a dynamic LID allocation architecture do not need to be contiguous. The LIDs assigned on the VMs on each hypervisor are compared between a vSwitch with pre-populated LIDs and a vSwitch with dynamic LID allocation. When comparing a vSwitch with a dynamic LID allocation architecture, it can be seen that the assigned LIDs are discontinuous, whereas the pre-populated LIDs are essentially continuous. Furthermore, in a vSwitch dynamic LID allocation architecture, when a new VM is created, the next available LID is used for the lifetime of the VM. Conversely, in a vSwitch with pre-populated LIDs, each VM inherits the LID already assigned to its corresponding VF, whereas in a network without live migration, VMs assigned consecutively to a given VF get the same LID.

[0073] According to one embodiment, a vSwitch using a dynamic LID allocation architecture can address the shortcomings of a vSwitch using a pre-populated LID architecture model, at the expense of some additional network and runtime SM overhead. Each time a VM is created, the LFT of the physical switch in the subnet can be updated with the newly added LID associated with the created VM. This operation requires one subnet management packet (SMP) to be sent per switch. Because each VM uses the same route as its host hypervisor, features such as LMC are also unavailable. However, there is no limit on the total number of VFs present on all hypervisors, and the number of VFs may exceed the unicast LID limit. In such a case, of course, not all VFs can be simultaneously granted on active VMs. Having more spare hypervisors and VFs adds flexibility for recovering from and optimizing fragmented network failures when operating near the unicast LID limit.

[0074] InfiniBand SR-IOV Architecture Model - Dynamic LID Allocation and Pre-Populated LID vSwitch FIG. 7 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs for a vSwitch, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 500 (e.g., an IB subnet). The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. The host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 can additionally assign virtual machine 2 551 to virtual function 2 515. Hypervisor 521 can assign virtual machine 3 552 to virtual function 3 526. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 2 535. The hypervisors can access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0075] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 700.

[0076] According to one embodiment, the virtual switches 512, 522, and 532 may be handled by respective hypervisors 511, 521, 531. In the ch architecture, each Virtual Function is a full Virtual Host Channel Adapter (vHCA), which means that in hardware, the VM assigned to the VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0077] According to one embodiment, the present disclosure provides a system and method for providing a hybrid vSwitch architecture with dynamic LID allocation and pre-populated LIDs. Referring to FIG. 7 , hypervisor 511 may be deployed with a vSwitch using a pre-populated LID architecture, while hypervisor 521 may be deployed with a vSwitch with pre-populated LIDs and dynamic LID allocation. Hypervisor 531 may be deployed with a vSwitch with dynamic LID allocation. Thus, physical function 513 and virtual functions 514-516 have their LIDs pre-populated (i.e., even virtual functions not assigned to active virtual machines are assigned LIDs). Physical function 523 and virtual function 1 524 may have their LIDs pre-populated, while virtual function 2 525 and virtual function 3 526 have their LIDs dynamically assigned (i.e., virtual function 2 525 is available for dynamic LID assignment, and virtual function 3 526 has been dynamically assigned an LID of 11 because it is attached to virtual machine 3 552). Finally, the functions (physical and virtual functions) associated with hypervisor 3 531 can have their LIDs dynamically assigned. This results in virtual function 1 534 and virtual function 3 536 being available for dynamic LID assignment, while virtual function 2 535 has been dynamically assigned an LID of 9 because it is attached to virtual machine 4 553.

[0078] 7, in which both LID pre-populated vSwitches and dynamic LID allocation vSwitches are utilized (independently or combined within any given hypervisor), the number of pre-populated LIDs per host channel adapter can be defined by a fabric administrator and can be in the range 0<=pre-populated VFs<=total VFs (per host channel adapter). The VFs available for dynamic LID allocation can be found by subtracting the number of pre-populated VFs from the total number of VFs (per host channel adapter).

[0079] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0080] Fast Hybrid Reconstruction According to one embodiment, a High Performance Computing (HPC) cluster is a massively parallel system consisting of thousands of nodes and millions of cores. Traditionally, such systems have been associated with the scientific community and can be used to perform complex, high-granularity computations. However, with the emergence of the cloud computing paradigm and big data analytics, the computer science community tends to agree that HPC and big data will converge, and that the cloud is becoming a vehicle for delivering related services to a wider audience. Large, traditional HPC clusters are typically shared environments among a diverse set of users, but with predictable workloads. However, the cloud and its more dynamic computing capabilities allow for more efficient and predictable computing. When subjected to a dynamic, prepaid model, system workload and utilization can be unpredictable, which can lead to the need to optimize performance during runtime.

[0081] According to one embodiment, one of the components that can be tuned and reconfigured to improve performance is the underlying interconnection network. Interconnection networks are a critical part of massively parallel architectures due to the intensive communication between nodes. Therefore, high-performance network technologies employing lossless Layer 2 flow control are typically used, as they offer significantly better performance. However, this performance comes at the expense of increased complexity and management costs, and network reconfiguration can be challenging. Because lossless networks do not drop packets, deadlocks can occur if routing functions allow loops to form. Subnet Manager (SM) software is responsible for managing the network. Among other tasks, the SM is responsible for computing deadlock-free communication paths between nodes in the network and distributing the corresponding Linear Forwarding Tables (LFTs) to the switches. Reconfiguration However, distributing a new LFT during the transition phase will result in a new routing function R new The old routing function R old R old and R newBoth are deadlock-free, but their combination may not be. Furthermore, path computation is the more expensive stage of the reconfiguration and can take up to several minutes depending on the topology and the selected routing function. This can introduce obstacles that turn reconfiguration into an excessive operation that should be avoided unless a major failure occurs. In the event of a failure, reconfigurations are kept to a minimum to quickly restore deadlock-free connectivity, at the expense of reduced performance.

[0082] According to one embodiment, a system and method can provide performance-driven reconfiguration in large-scale lossless networks. A hybrid reconfiguration scheme can enable fast partial network reconfiguration using different routing algorithms to select different subsections of the network. Because partial reconfiguration can be orders of magnitude faster than an initial global configuration, it may be possible to consider performance-driven reconfiguration in lossless networks. The proposed mechanism takes advantage of the fact that large HPC systems and clouds are shared by multiple tenants (e.g., different tenants on different partitions) performing isolated tasks. In such scenarios, intercommunication between tenants is impossible, and thus workload deployment and placement schedulers should attempt to avoid fragmentation to ensure efficient resource utilization. That is, most of the traffic per tenant can be contained within an aggregated subsection of the network. The SM can reconfigure some subsections to improve overall performance. The SM can use a fat-tree topology and a fat-tree routing algorithm. Such a hybrid reconfiguration scheme can successfully reconfigure and improve performance within a subtree by using a custom fat-tree routing algorithm with the provided node ordering to reconfigure the network. If the SM wishes to reconfigure the entire network, it can use the default Fat Tree routing algorithm to effectively provide a combination of two different routing algorithms for various use cases in a single subnet.

[0083] According to one embodiment, a fat tree routing algorithm (FTree) FTree is a topology-aware routing algorithm for FTree topologies. FTree first discovers the network topology, and each switch is marked with a tuple that identifies its position in the topology. Each tuple is a set of (l,a h ,...,a1), where "l" is the level at which the switch is located. Furthermore, "a h " represents the switch index within the top subtree, and the number "a1" is used recursively until "a1" represents the index of the subtree within that first subtree, etc. h-1」 For an n-level fat tree, the root-level (top or core) switch is located at level l=0, while the leaf switches (to which the nodes are connected in this case) are located at level l=n-1. The tuple allocation for an example 2-ary-4-tree is shown in Figure 8.

[0084] Figure 8 illustrates a switch tuple according to one embodiment. More specifically, the diagram illustrates the switch tuple as assigned by OpenSM's Fat Tree routing algorithm, implemented for an example Fat Tree, XGFT(4;2,2,2,2;2,2,2,2,1). Fat Tree 800 may include switches 801-808, 811-818, 921-1428, and 831-838. Because the Fat Tree has n=4 switch levels (marked as row 0 at the root level through row 3 at the leaf level), the Fat Tree consists of m=2 first-level subtrees, each with n'=n-1=3 switch levels. This is illustrated in the diagram by the two boxes defined by dashed lines surrounding the switches from level 1 to level 3, with the first-level subtrees receiving an identifier of 0 or 1. Each of these first-level subtrees is composed of m2 = 2 second-level subtrees, where each has n" = n' - 1 = 2 switch levels above the leaf switches. This is shown in the diagram by the four boxes defined by dotted lines surrounding the switches from level 2 to level 3, and each second-level subtree receives an identifier of 0 or 1. Similarly, each leaf switch can also be considered a subtree, and is shown in the diagram by the eight boxes defined by dashed lines, and each of these subtrees receives an identifier of 0 or 1.

[0085] According to one embodiment, tuples, such as a tuple of four numbers as illustrated in the figure, can be assigned to various switches, with each number in the tuple indicating a particular subtree correspondence for each value's position in the tuple. For example, switch 814 (which may be referred to as switch 1_3) can be assigned tuple 1.0.1.1, which represents its position in level 1 and the 0th first-level subtree.

[0086] According to one embodiment, once the tuples are assigned, FTree iterates through each leaf switch in ascending tuple order, and for each downward switch port to which a node is connected in ascending port order, the algorithm routes the selected node based on their LID. Figures 9-13 show the various stages of how nodes are routed according to one embodiment.

[0087] FIG. 9 illustrates a system for the node routing stage, according to one embodiment. Switches 901-912 in the diagram are labeled with numbers 1-12. Each switch may include multiple ports (not shown). For example, each switch may include 32 ports (16 downward and 16 upward). Each of switches 1, 2, 3, and 4 may also be linked to more than one node; for example, nodes A 920 and B 921 may be linked to switch 1, nodes C 922 and D 923 may be linked to switch 2, nodes E 924 and F 925 may be linked to switch 3, and nodes G 926 and H 927 may be linked to switch 4. The tree maintains port usage counters to balance routes and starts traversing the fabric upward from the least loaded port while selecting a downward route. As shown in the figure, on the first iteration, all port counters are zero, so the first available upward port is selected. With each level up, a newly arrived switch, in this case switch 5 905, is selected to route all traffic downward from the input port through which the arrived switch passed, toward the routed node (node ​​A 920). The dashed line in the figure represents the route assigned to node A.

[0088] FIG. 10 illustrates a system for the node routing stage, according to one embodiment. Switches 901-912 in the figure are labeled with numbers 1 through 12. Each switch may include multiple ports (not shown). For example, each switch may include 32 ports (16 facing downward and 16 facing upward). Each of switches 1, 2, 3, and 4 may also be linked to more than one node; for example, nodes A 920 and B 921 may be linked to switch 1, nodes C 922 and D 923 may be linked to switch 2, nodes E 924 and F 925 may be linked to switch 3, and nodes G 926 and H 927 may be linked to switch 4. After the routing step shown in FIG. 9, the FTree traverses the fabric downward and assigns routes upward to the switches in a similar manner. This is shown in the figure as a long arrow pointing from switch 5 to switch 2 and representing the routing algorithm. Route assignment then proceeds upward from switch 2 to switch 5. The dashed line in the diagram represents the route assigned to node A.

[0089] FIG. 11 illustrates a system for the node routing stage, according to one embodiment. Switches 901-912 in the figure are labeled with numbers 1 through 12. Each switch may include multiple ports (not shown). For example, each switch may include 32 ports (16 downward and 16 upward). Each of switches 1, 2, 3, and 4 may also be linked to more than one node; for example, nodes A 920 and B 921 may be linked to switch 1, nodes C 922 and D 923 may be linked to switch 2, nodes E 924 and F 925 may be linked to switch 3, and nodes G 926 and H 927 may be linked to switch 4. Recursive operations identical or similar to those described in FIGS. 9 and 10 continue until route entries for the selected nodes have been added to all of the necessary switches in the fabric. As shown in FIG. 11, the descending route is illustrated by an upward movement. As the FTree mechanism traverses the tree upwards (from switch 5 to switch 9), a route is assigned to node A between switch 9 and switch 5 (the downward route).

[0090] FIG. 12 illustrates a system for the node routing stage, according to one embodiment. Switches 901-912 in the figure are labeled with numbers 1 through 12. Each switch may include multiple ports (not shown). For example, each switch may include 32 ports (16 facing downward and 16 facing upward). Each of switches 1, 2, 3, and 4 may also be linked to two or more nodes; for example, nodes A 920 and B 921 may be linked to switch 1, nodes C 922 and D 923 may be linked to switch 2, nodes E 924 and F 925 may be linked to switch 3, and nodes G 926 and H 927 may be linked to switch 4. Recursive operations the same or similar to those described in FIGS. 9, 10, and 11 continue until route entries for the selected nodes have been added to all of the necessary switches in the fabric. As shown in FIG. 12, a route for ascending by a downward descending action exists between switch 9 and switch 7, and a route for ascending by a downward descending action exists between switch 9 and switch 7. Two routes are performed, one route between switch 7 and switch 3, and the other route between switch 7 and switch 4. The dashed line in the diagram represents the route assigned to node A. At this point, routes from all nodes to node A are defined in the system. This operation can be repeated for each node in the system, while maintaining port counters, until all routes for all nodes have been calculated.

[0091] Note that even though routing toward node A is complete, there are still several blank switches (switches 6, 8, 10, 11, and 12) that have no route toward node A. In fact, FTree can add routes even in these blank switches. When a packet destined for node A arrives at, say, switch 12, this switch knows it must forward the received packet downward toward switch 6, while switch 6 knows it must forward the packet received from switch 12 to switch 1 to reach its destination A. However, lower-level switches will not forward traffic destined for node A to switch 12 because the upward route would always push the packet toward switch 9. Note that using one root switch per destination node counter prevents the growth of wide congestion trees.

[0092] According to one embodiment, the fast hybrid reconfiguration method may be based on the concept that HPC systems and cloud environments are shared by multiple tenants performing isolated tasks (i.e., intercommunication between tenants is not permitted). To achieve better resource utilization, workload deployment or virtual machine placement schedulers attempt to avoid resource fragmentation as much as possible. As a result, per-tenant workloads are mapped onto physical machines that are close to each other in terms of physical network connections to avoid unnecessary network traffic and cross-tenant network interference. For fat-tree topologies with more than two levels, this means that per-tenant traffic can be included within subtrees of the multi-level fat-tree.

[0093] FIG. 13 illustrates a system including a fat tree topology with more than two levels, according to one embodiment. Within a fat tree topology subnet 1300 with several switch levels (three switch levels in the illustrated embodiment), subtrees 1310 (also referred to herein as sub-subnets) can be defined. In this case, traffic within subtree 1310 is self-contained; that is, traffic within subtree 1310 (i.e., between end nodes 1320 spanning from end node A to end node P) does not flow into or out of the rest of the topology. As an example, end nodes 1320 may all belong to the same partition (e.g., all nodes in 1320 share a common partition key (P_Key)). Note that, although not shown, each end node may be connected to the switching fabric via a host channel adapter (HCA).

[0094] According to one embodiment, a fast hybrid reconfiguration method can apply partial reconfiguration to locally optimize within a sub-subnet based solely on its internal traffic patterns. By applying such partial reconfiguration, the method can effectively treat the reconfiguration as a fat tree with fewer levels, thereby reducing the cost of path computation and overall reconfiguration. In practice, performance-driven reconfiguration becomes attractive even in shared, highly dynamic environments. Furthermore, when applying partial reconfiguration, the method need only modify the forwarding entries of nodes within sub-subnet 1310. Assuming that the initial routing algorithm used to route the fabric is FTree or similar, and applying a variant of upward / downward routing without using virtual lanes ensures deadlock resolution, the method can use any optimal routing algorithm to reroute a given subtree as an isolated one (hybrid reconfiguration).

[0095] According to one embodiment, once a subtree of a fat tree is reconfigured, connectivity between all end nodes is still maintained, even if they are outside the reconfigured sub-subnet. This is because the switches have LFTs that specify to which destinations traffic should be forwarded. That is, every switch S has a valid forwarding entry in its LFT for every destination x, even if no other node can actually forward packets destined for x via S. For example, after the initial routing selection within the subtree, a switch one level above the leaf switch (referred to here as switch 5) was selected to route traffic downward toward node A, while switch 6, which is at the same level as switch 5, was selected to route traffic toward node B. After the subtree reconfiguration, switch 5 is now used to route traffic toward node B, and switch 6 is used to route traffic toward node A. In this case, when nodes E and F, located within the subtree, send traffic toward node A or node B, the newly calculated path will be used, and traffic will remain entirely within the subtree. However, if a node located outside the subtree (not shown) sends traffic to nodes A and B, the old paths (i.e., those paths are not part of the reconfiguration because they are outside the subtree) will be used, and traffic destined for A and B will enter the subtree at the switch specified by the original routing for the entire subnet. This behavior outside the subtree could potentially defeat the purpose of the subtree reconfiguration, for example, by interfering with load balancing within the subtree. However, if the subtree is configured such that little or no traffic crosses the subtree boundary (e.g., if the subtree includes an entire partition), such interference will be only a minor issue.

[0096] According to one embodiment, to apply partial reconfiguration, the method may first select all nodes and switches in the subtree that must be reconfigured. The method may use switch tuples to select which subtree should be reconfigured. For partial reconfiguration, the method may select all nodes and switches in the subtree that need to be reconfigured. All nodes in the subtree need to be selected and considered. The selection process of all entities in the subtree can be done in the following steps:

[0097] 1) An administrator (or an automated solution that monitors fabric utilization) provides a list of nodes that should be involved in the reconfiguration.

[0098] 2) The leaf switch tuples of the nodes from step 1 are compared to select a common ancestor subtree.

[0099] 3) All switches that belong to the subtree selected in step 2 will be marked for reconfiguration.

[0100] 4) From the list of switches in step 3, a leaf switch is selected and all nodes connected to the selected leaf switch will be involved in the reconfiguration process.

[0101] 5) Finally, the routing algorithm must compute a new set of routes only for the nodes selected in step 4 and distribute the LFT only to the switches selected in step 3.

[0102] According to one embodiment, in a multi-stage switch topology such as a Fat Tree, the effective bisection bandwidth is typically less than the theoretical bisection bandwidth for various traffic patterns because there may be shared links in the upward direction depending on which node pairs are selected for communication. An example is shown in Figure 14.

[0103] 14 illustrates a system for fast hybrid reconfiguration according to one embodiment. Within a fat-tree topology subnet 1400 having several switch levels (three switch levels in the illustrated embodiment), a sub-tree 1410 can be defined that contains all of the traffic within the sub-subnet 1410. That is, traffic within the sub-subnet 1410 (i.e., between end nodes 1420 spanning from end node A to end node P) does not flow into or out of the rest of the topology. As an example, the end nodes 1420 may all belong to the same partition.

[0104] As shown in the figure, end nodes 1420 (end nodes A-P) can communicate within a two-level subtree (denoted as 1410) of a three-level fat-tree globally routed with the FTree routing algorithm. In the illustrated embodiment, the routing method, i.e., FTree, selected switch 5 to route downward toward nodes A, E, I, and M; switch 6 to route downward toward nodes B, F, J, and N; switch 7 to route downward toward nodes C, G, K, and O; and switch 8 to route downward toward nodes D, H, L, and P. While this subtree has a theoretical full bisection bandwidth, the effective bisection bandwidth for the illustrated communication pattern, in which nodes B, C, and D are sending traffic to nodes E, I, and M, respectively, is one-third of the full bandwidth. This is because all destination nodes are routed through the same switch (switch 5) in the second level, and the bold dashed link connecting switch 1 and switch 5 is shared by all three flows and represents a traffic bottleneck. However, there are enough free links to avoid link sharing and provide full bandwidth. To allow flexible reconfiguration that does not necessarily result in the same routing order based on port ordering, the fast hybrid reconfiguration scheme can use a fat-tree routing mechanism (also known as NoFTree), which routes the fat-tree network using a node ordering defined by the user. This can provide improvements. A simple way to determine the incoming traffic per node is to read the IB port counters. This eliminates the need for administrators to be familiar with the details of the jobs run by tenants.

[0105] According to one embodiment, NoFTree can be used in the context of a fast hybrid reconfiguration scheme to route a subtree after the switches and nodes have been selected as described above. The scheme may follow the following steps:

[0106] 1) An ordered list of nodes to be routed is provided by the user or a monitoring solution.

[0107] 2) NoFTree reorders the nodes per leaf switch. Each of the ordered nodes is then placed in the n% maximum nodes per leaf sw+1 slot to be routed in a given leaf switch, where n is the node's overall position in the reordered list for that node.

[0108] 3) The remaining nodes connected to each leaf switch but not present in the provided node ordering list fill the remaining leaf switch routing slots based on the port order to which the nodes are connected. If port ordering is not provided by the user, NoFTree can function as an FTree routing algorithm.

[0109] 4) NoFTree again iterates through each leaf switch and routes each node based on the node order established through the previous steps.

[0110] 15 illustrates a system for fast hybrid reconfiguration according to one embodiment. Within a fat-tree topology subnet 1500 having several switch levels (three switch levels in the illustrated embodiment), a sub-tree 1510 can be defined that contains all of the traffic within the sub-subnet 1510. That is, traffic within the sub-subnet 1510 (i.e., between end nodes 1520 spanning from end node A to end node P) does not flow into or out of the rest of the topology. As an example, the end nodes 1520 may all belong to the same partition.

[0111] As shown in the figure, end nodes 1520 (end nodes A-P) can communicate within a two-level subtree (shown as 1510) of a three-level fat-tree globally routed with the FTree routing algorithm. In the illustrated embodiment, the routing method involves NoFTree reconstructing the subtree of FIG. 15 using the supplied / received node ordering E, I, M, selecting switch 5 to route downward toward nodes A, E, J, and N, selecting switch 6 to route downward toward nodes B, F, I, and O, selecting switch 7 to route downward toward nodes C, G, K, and M, and selecting switch 8 to route downward toward nodes D, H, L, and P.

[0112] In this case, the provided / received node order that NoFTree uses for reconfiguration is E, I, M. Because nodes from leaf switch 1 are not specified in the node ordering, the nodes connected to switch 1 are routed based on port order. Because node E is the first node in the overall node ordering and the first node to be ordered on leaf switch 2, node E becomes the first node to be routed on switch 2 (routed downwards from switch 5). The rest of the nodes on leaf switch 2 (nodes F, G, H) are routed according to port order. The mechanism then proceeds to the third leaf switch (switch 3), where node I is connected from the provided / received node ordering. Because node I is the second node in the provided / received node ordering and the first node to be ordered on switch 3, node I is routed on switch 3 (routed downwards from switch 6). ) becomes the second node to be routed. Nodes connected to switch 4 are routed in the same manner. The remaining routing occurs as described and illustrated above. In this scenario, a 300% performance improvement can be achieved because there is no longer an upstream link shared with traffic flowing from nodes B, C, and D to nodes E, I, and M.

[0113] FIG. 16 is a flowchart illustrating an exemplary method for supporting fast hybrid reconfiguration in a high performance computing environment, according to one embodiment.

[0114] At step 1610, the method may provide, in one or more microprocessors, a first subnet including a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, a plurality of host channel adapters each including at least one host channel adapter port, and a plurality of end nodes, each of the plurality of end nodes being associated with at least one host channel adapter of the plurality of host channel adapters.

[0115] In step 1620, the method may arrange the switches of the first subnet in a network architecture having multiple levels, each of the multiple levels including at least one switch of the multiple switches.

[0116] In step 1630, the method may configure the plurality of switches according to a first configuration method, the first configuration method being associated with a first ordering of the plurality of end nodes.

[0117] In step 1640, the method may configure a subset of the plurality of switches as a sub-subnet of the first subnet, the sub-subnet of the first subnet including some levels less than the plurality of levels of the first subnet.

[0118] In step 1650, the method may reconfigure sub-subnets of the first subnet according to a second configuration method.

[0119] While various embodiments of the present invention have been described above, it should be understood that they are presented by way of example and not limitation. These embodiments were chosen and described in order to explain the principles of the invention and its practical applications. These embodiments illustrate systems and methods in which the present invention is utilized to enhance the performance of such systems and methods by providing new and / or improved features and / or by providing benefits such as reduced resource utilization, increased capacity, improved efficiency, and reduced latency.

[0120] In some embodiments, features of the present invention are implemented, in whole or in part, in a computer that includes a processor, a storage medium such as memory, and a network card for communicating with other computers. In some embodiments, features of the present invention are implemented in a computer system where one or more clusters of computers are connected to a network, such as a Local Area Network (LAN), a switched fabric network (e.g., InfiniBand), or a Wide Area Network (WAN). The distributed computing environment may have all the computers in one location, or may have a cluster of computers in various remote geographic locations connected by a WAN.

[0121] In some embodiments, features of the invention are implemented in the cloud as part of a cloud computing system or service, based in whole or in part on shared, elastic resources delivered to users in a self-service, coordinated manner using web technologies. There are five characteristics of the cloud (as defined by the National Institute of Standards and Technology): on-demand self-service, wide area network access, resource pooling, high-speed elasticity, and measured service. See, e.g., "The NIST Definition of Cloud Computing," Special Publication 800-145 (2011), which is incorporated herein by reference. Cloud deployment models include public, private, and hybrid. Cloud service models include Software as a Service. Platform as a Service (SaaS) Database as a Service (DBaaS) and and Infrastructure as a Service (Ia Cloud includes cloud services (SaaS, DBaaS, PaaS, and IaaS). As used herein, cloud is a combination of hardware, software, network, and web technologies that delivers shared, elastic resources to users in a self-service, coordinated manner. Unless otherwise specified, cloud, as used herein, encompasses public cloud, private cloud, and hybrid cloud embodiments, and all cloud deployment models, including, but not limited to, cloud SaaS, cloud DBaaS, cloud PaaS, and cloud IaaS.

[0122] In some embodiments, features of the present invention are implemented using or with the aid of hardware, software, firmware, or a combination thereof. In some embodiments, features of the present invention are implemented using a processor configured or programmed to perform one or more functions of the present invention. The processor, in some embodiments, is a single or multi-chip processor, a digital signal processor (DSP), a system on a chip (SOC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a state machine, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. In some implementations, features of the present invention may be implemented by circuitry specific to a given function. In other implementations, these features may be implemented in a processor configured to perform a particular function using instructions stored on a computer-readable storage medium, for example.

[0123] In some embodiments, features of the present invention are embodied in software and / or firmware for controlling the hardware of a processing system and / or networking system and for enabling the processor and / or network to interact with other systems that utilize features of the present invention. Such software or firmware may include, but is not limited to, application code, device drivers, operating systems, virtual machines, hypervisors, application program interfaces, programming languages, and execution environments / containers. Appropriate software coding can be readily prepared by skilled programmers based on the teachings of the present disclosure, as will be apparent to those skilled in the software arts.

[0124] In some embodiments, the present invention provides a storage medium or computer having instructions stored thereon. The present invention also includes a computer program product that is a computer-readable medium. These instructions can be used to program or otherwise configure a system, such as a computer, to perform any of the processes or functions of the present invention. The storage medium or computer-readable medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), and any type of medium or device suitable for storing instructions and / or data. In certain embodiments, the storage medium or computer-readable medium is a non-transitory storage medium or computer-readable medium.

[0125] The foregoing description is not intended to be exhaustive or to limit the invention to the precise form disclosed. In addition, while embodiments of the present invention have been described using a particular series of transactions and steps, it will be apparent to those skilled in the art that the scope of the present invention is not limited to the above-described series of transactions and steps. Furthermore, while embodiments of the present invention have been described using a particular combination of hardware and software, it will be recognized that other combinations of hardware and software are within the scope of the present invention. Furthermore, while various embodiments describe particular combinations of features of the present invention, it will be understood by those skilled in the art that various combinations of these features will be apparent to those skilled in the art as being within the scope of the present invention, such that features of one embodiment may be incorporated into another embodiment. Furthermore, it will be apparent to those skilled in the art that various additions, deductions, deletions, modifications, and other changes in form, details, implementation, and application can be made without departing from the spirit and scope of the present invention. The broader spirit and scope of the present invention is intended to be defined by the appended claims and their equivalents.

Claims

1. 1. A system for supporting fast hybrid reconfiguration in a high performance computing environment, comprising: one or more microprocessors; a first subnet, the first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each including at least one host channel adapter port; a plurality of end nodes, each of the plurality of end nodes being associated with at least one host channel adapter of the plurality of host channel adapters; the plurality of switches of the first subnet are arranged in a network architecture having a plurality of levels, each of the plurality of levels including at least one switch of the plurality of switches; the plurality of switches are initially configured according to a first configuration method, the first configuration method being associated with a first ordering of the plurality of end nodes; a subset of the plurality of switches configured as sub-subnets of the first subnet, the sub-subnets of the first subnet including some levels less than the plurality of levels of the first subnet; The system, wherein the sub-subnets of the first subnet are reconfigured according to a second configuration method.

2. The system of claim 1 , wherein the end nodes of the first subnet are interconnected via the switches.

3. a subset of the plurality of end nodes is associated with the sub-subnet of the first subnet; 3. The system of claim 2, wherein the sub-subnets of the first subnet are configured such that traffic between a subset of the plurality of end nodes is restricted to the subset of the plurality of switches configured as the sub-subnets of the first subnet.

4. 4. The system of claim 3, wherein the second reconfiguration method is associated with a second ordering of at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet.

5. 5. The system of claim 4, wherein the second ordering for the at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet is received from a system administrator.

6. 5. The system of claim 4, wherein the second ordering for at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet is received from a management entity.

7. the first subnet comprises an InfiniBand subnet; The management entity: Subnet Manager, Fabric Manager, and 7. The system of claim 6, wherein the management entity is selected from the group consisting of a global fabric manager.

8. 1. A method for supporting fast hybrid reconfiguration in a high performance computing environment, comprising: providing a first subnet in one or more microprocessors, said first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each including at least one host channel adapter port; a plurality of end nodes, each of the plurality of end nodes associated with at least one host channel adapter of the plurality of host channel adapters, the method further comprising: arranging the plurality of switches of the first subnet in a network architecture having a plurality of levels, each of the plurality of levels including at least one switch of the plurality of switches, the method further comprising: configuring the plurality of switches according to a first configuration method, the first configuration method being associated with a first ordering of the plurality of end nodes, the method further comprising: configuring a subset of the plurality of switches as sub-subnets of the first subnet, the sub-subnets of the first subnet including some levels less than the plurality of levels of the first subnet, the method further comprising: A method comprising reconfiguring the sub-subnets of the first subnet according to a second configuration method.

9. The method of claim 8 , wherein the end nodes of the first subnet are interconnected via the switches.

10. a subset of the plurality of end nodes is associated with the sub-subnet of the first subnet; 10. The method of claim 9, wherein the sub-subnets of the first subnet are configured such that traffic between a subset of the plurality of end nodes is restricted to the subset of the plurality of switches configured as the sub-subnets of the first subnet.

11. 11. The method of claim 10, wherein the second reconfiguration method is associated with a second ordering of at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet.

12. 12. The method of claim 11, wherein the second ordering for the at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet is received from a system administrator.

13. 12. The method of claim 11, wherein the second ordering for the at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet is received from a management entity.

14. the first subnet comprises an InfiniBand subnet; The management entity: Subnet Manager, Fabric Manager, and 14. The method of claim 13, wherein the management entity is selected from the group consisting of a global fabric manager.

15. 1. A non-transitory computer-readable storage medium having stored thereon instructions for supporting fast hybrid reconfiguration in a high performance computing environment, the instructions, when read and executed by one or more computers, causing the one or more computers to: performing a step of providing a first subnet in one or more microprocessors, said first subnet comprising: a plurality of switches, the plurality of switches including at least a leaf switch, each of the plurality of switches including a plurality of switch ports, the first subnet further comprising: a plurality of host channel adapters, each including at least one host channel adapter port; a plurality of end nodes, each of the plurality of end nodes being associated with at least one host channel adapter of the plurality of host channel adapters, the one or more computers further comprising: arranging the plurality of switches of the first subnet into a network architecture having a plurality of levels, each of the plurality of levels including at least one switch of the plurality of switches; and causing the one or more computers to further: configuring the plurality of switches according to a first configuration method, the first configuration method being associated with a first ordering of the plurality of end nodes; configuring a subset of the plurality of switches as sub-subnets of the first subnet, the sub-subnets of the first subnet including some levels less than the plurality of levels of the first subnet; and causing the one or more computers to further: A non-transitory computer-readable storage medium for causing a step of reconfiguring the sub-subnets of the first subnet according to a second configuration method.

16. 16. The non-transitory computer-readable storage medium of claim 15, wherein the end nodes of the first subnet are interconnected via the switches.

17. a subset of the plurality of end nodes is associated with the sub-subnet of the first subnet; 17. The non-transitory computer-readable storage medium of claim 16, wherein the sub-subnet of the first subnet is configured such that traffic between a subset of the plurality of end nodes is restricted to the subset of the plurality of switches configured as the sub-subnet of the first subnet.

18. 20. The non-transitory computer-readable storage medium of claim 17, wherein the second reconfiguration method is associated with a second ordering of at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet.

19. the plurality of end nodes associated with the sub-subnets of the first subnet; 20. The non-transitory computer-readable storage medium of claim 18, wherein the second ordering for the at least two end nodes of the subset of nodes is received from a system administrator.

20. the second ordering for the at least two end nodes of the subset of the plurality of end nodes associated with the sub-subnet of the first subnet is received from a management entity; the first subnet comprises an InfiniBand subnet; The management entity: Subnet Manager, Fabric Manager, and 20. The non-transitory computer-readable storage medium of claim 18, wherein the management entity is selected from the group consisting of: a global fabric manager.

21. A computer program comprising program instructions in a machine-readable format, said program instructions, when executed by a computer system, causing said computer system to carry out a method according to any of claims 8 to 13.

22. 22. A computer program product comprising the computer program of claim 21 stored on a non-transitory machine-readable data storage medium.

23. Apparatus comprising means for carrying out the method according to any one of claims 8 to 13.

Citation Information

Patent Citations

  • System and method for supporting sub-subnet in an infiniband (IB) network

    US20140241208A1

  • System and method for supporting partition-aware routing in a multi-tenant cluster environment

    US20160127236A1