Systems and methods for efficient virtualization in lossless networks

The vSwitch architecture with pre-populated or dynamic LID allocation addresses the challenges of live migration in InfiniBand networks, ensuring efficient virtualization and transparent live migration of virtual machines with reduced downtime and improved scalability.

JP7795575B2Active Publication Date: 2026-01-07ORACLE INT CORP

Patent Information

Application Number
JP2024063248
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-07-14
Filing Date
2024-04-10
Publication Date
2026-01-07
Estimated Expiration
2036-11-18

AI Technical Summary

Technical Problem

Existing cloud computing systems face challenges in live migration of virtual machines due to complex addressing and routing schemes in high-performance lossless networks like InfiniBand (IB), leading to connectivity issues and network overhead during migration, particularly with technologies like Single-Root I/O Virtualization (SR-IOV).

Method used

Implementing a virtual switch (vSwitch) architecture with pre-populated or dynamically allocated local identifiers (LIDs) in InfiniBand networks, allowing efficient virtualization and live migration of virtual machines by optimizing routing and reducing network reconfiguration time.

Benefits of technology

The vSwitch architecture enables transparent live migration of virtual machines with minimal downtime and improved resource utilization, enhancing scalability and flexibility in cloud computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795575000003
    Figure 0007795575000003
  • Figure 0007795575000004
    Figure 0007795575000004
  • Figure 0007795575000005
    Figure 0007795575000005
Patent Text Reader

Abstract

To provide systems and methods for supporting efficient virtualization in lossless interconnection networks.SOLUTION: In a network switched environment 600 (e.g., an InfiniBand (IB) subnet), a virtual switch (vSwitch) architecture comprises: one or more switches 501 to 504 including at least a leaf switch; a plurality of host channel adapters (HCAs) each including at least one of virtual functions 514 to 516, 524 to 526, 534 to 536, at least one virtual switch, and at least one physical function; a plurality of hypervisors; and a plurality of virtual machines 550 to 552 each associated with at least one virtual function. Physical HCAs have two or more ports, and virtual HCAs are also represented with two ports and connected via one or more virtual switches to an external IB subnet.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Copyright notice: A portion of the disclosure of this patent document contains material that is subject to copyright protection. As this patent document or patent disclosure appears in the Patent and Trademark Office patent file or records, the copyright owner has no objection to the facsimile reproduction thereof by anyone, but otherwise reserves all copyright rights whatsoever.

[0002] Field of the invention: The present invention relates generally to computer systems, and more particularly to supporting computer system virtualization and live migration using an SR-IOV vSwitch architecture. [Background technology]

[0003] background: As larger-scale cloud computing architectures are deployed, the performance and management bottlenecks associated with traditional networks and storage are becoming serious problems. There is growing interest in using high-performance, lossless interconnects, such as InfiniBand (IB) technology, as the foundation for cloud computing fabrics. This is the general area that embodiments of the present invention are intended to address. Summary of the Invention [Means for solving the problem]

[0004] overview: Systems and methods for supporting virtual machine migration in a subnet are described herein. An exemplary method can include providing one or more switches in one or more computers including one or more microprocessors, the one or more switches including at least a leaf switch, each of the one or more switches including a plurality of ports, and the method can further include providing a plurality of host channel adapters. Each of the plurality of host channel adapters includes at least one virtual function, at least one virtual switch, and at least one physical function. The plurality of host channel adapters are interconnected via the one or more switches. The method can further include providing a plurality of hypervisors, each of the plurality of hypervisors being associated with at least one host channel adapter of the plurality of host channel adapters. The method can further include providing a plurality of virtual machines, each of the plurality of virtual machines being associated with at least one virtual function. The method can further include configuring the plurality of host channel adapters with one or more of the virtual switches having a pre-populated local identifier (LID) architecture or a dynamic LID allocation architecture. The method can assign each virtual switch to a LID, the assigned LID corresponding to the LID of the associated physical function. The method can calculate one or more linear forwarding tables (LFTs) based at least on the LID assigned to each of the virtual switches. Each of the one or more LFTs can be associated with one or more switches. The switch is associated with one of the switches.

[0005] According to one embodiment, a method includes, in one or more computers including one or more microprocessors, performing a process on the one or more microprocessors and a process on the one or more computers including at least a leaf switch. The method may include one or more switches, each of the one or more switches including a plurality of ports; and may further include a plurality of host channel adapters, each of the host channel adapters including at least one virtual function, at least one virtual switch, and at least one physical function, the plurality of host channel adapters being interconnected via the one or more switches; and may further include a plurality of hypervisors, each of the plurality of hypervisors being associated with at least one host channel adapter of the plurality of host channel adapters; and may further include a plurality of virtual machines, each of the plurality of virtual machines being associated with at least one virtual function. The method may deploy the plurality of host channel adapters with one or more of the virtual switches having a pre-populated local identifier (LID) architecture or a dynamic LID assignment architecture. The method may assign one pLID of a plurality of physical LIDs (pLIDs) to each of the virtual switches, the assigned pLID corresponding to the pLID of the associated physical function. The method may also assign a plurality of virtual LIDs (vLIDs) to each of the plurality of virtual machines. , and the LID space contains multiple pLIDs and multiple vLIDs.

[0006] According to one embodiment, each pLID value can be represented using the standard SLID and DLID fields in the local route header of an InfiniBand packet. Similarly, each vLID value can be represented using a combination of the standard SLID and DLID fields, combined with two or more additional bits representing extensions. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates an example of an InfiniBand environment according to one embodiment. [Figure 2]FIG. 1 illustrates an example of a tree topology in a network environment, according to one embodiment. [Figure 3] FIG. 1 illustrates an exemplary shared port architecture according to one embodiment. [Figure 4] FIG. 1 illustrates an exemplary vSwitch architecture according to one embodiment. [Figure 5] FIG. 2 illustrates an exemplary vPort architecture according to one embodiment. [Figure 6] FIG. 2 illustrates an exemplary vSwitch architecture with pre-populated LIDs, according to one embodiment. [Figure 7] FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID allocation, according to one embodiment. [Figure 8] FIG. 2 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs in accordance with one embodiment. [Figure 9] FIG. 1 illustrates an extended local route header according to one embodiment. [Figure 10] 1A-1C illustrate two exemplary linear forwarding tables according to one embodiment. [Figure 11] FIG. 1 illustrates an example of efficient virtualization support in a lossless interconnect network, according to one embodiment. [Figure 12] FIG. 1 illustrates an example of efficient virtualization support in a lossless interconnect network, according to one embodiment. [Figure 13] FIG. 1 illustrates an example of efficient virtualization support in a lossless interconnect network, according to one embodiment. [Figure 14] FIG. 1 illustrates an example of efficient virtualization support in a lossless interconnect network, according to one embodiment. [Figure 15] FIG. 2 illustrates a potential virtual machine migration according to one embodiment. [Figure 16]FIG. 2 illustrates a switch tuple according to one embodiment. [Figure 17] FIG. 1 illustrates a reconstruction process according to one embodiment. [Figure 18] 1 is a flowchart illustrating a method for supporting efficient virtualization in a lossless interconnect network, according to one embodiment. [Figure 19] 1 is a flowchart illustrating a method for supporting efficient virtualization in a lossless interconnect network, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] Detailed Description: The present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to like elements. It should be noted that references in this disclosure to "an" or "one" or "several" embodiments are not necessarily to the same embodiment, and such references mean at least one. While specific implementations are described, it is understood that these specific implementations are provided for illustrative purposes only. Those skilled in the art will recognize that other components and configurations can be used without departing from the scope and spirit of the invention.

[0009] Common reference numbers may be used to denote like elements throughout the drawings and detailed description, and thus a reference number used in one drawing may or may not be referenced in the detailed description specific to that drawing if the element is described elsewhere.

[0010] SUMMARY Systems and methods for supporting efficient virtualization in lossless interconnect networks are described herein.

[0011] The following description of the present invention uses an InfiniBand (IB) network as an example of a high-performance network. It will be apparent to those skilled in the art that other types of high-performance networks can be used without any limitation. The following description also uses a Fat Tree topology as an example of a fabric topology. It will be apparent to those skilled in the art that other types of fabric topologies can be used without any limitation.

[0012] To meet the demands of modern clouds (e.g., the exascale era), virtual machines must be able to utilize low-overhead network communication paradigms such as Remote Direct Memory Access (RDMA). It is desirable to have RDMA bypass the OS stack and communicate directly with the hardware, allowing for pass-through technologies such as Single-Root I / O Virtualization (SR-IOV) network adapters. In accordance with one embodiment, a virtual switch (vSwitch) SR-IOV architecture can be provided for applicability in high-performance lossless interconnect networks. Since network reconfiguration time is critical to make live migration a viable option, a scalable, topology-independent dynamic reconfiguration mechanism can be provided in addition to the network architecture.

[0013] According to one embodiment, a routing strategy for a virtualized environment using a vSwitch can be further provided, and an efficient routing algorithm can be provided for a network topology (e.g., a fat-tree topology). The dynamic reconfiguration mechanism can be further tuned to minimize the overhead imposed in the fat-tree.

[0014] In accordance with an embodiment of the present invention, virtualization can be beneficial for efficient resource utilization and elastic resource allocation in cloud computing. Live migration can optimize resource usage by moving virtual machines (VMs) between physical servers in an application-transparent manner. Because of this, virtualization can enable consolidation, on-demand provisioning of resources, and elasticity through live migration.

[0015] InfiniBand(R) InfiniBand (IB) is a registered trademark of the InfiniBand Trade Association. TM It is an open-standard, lossless network technology developed by the IEEE 802.11b Trade Association. The technology is based on a serial point-to-point full-duplex interconnect that provides high-throughput and low-latency communications, especially targeted at high-performance computing (HPC) applications and data centers.

[0016] InfiniBand Architecture (IBA) is a two-layer Supports topology partitioning. At a low level, an IB network is called a subnet, and a subnet may contain a set of hosts interconnected using switches and point-to-point links. At a higher level, an IB fabric consists of one or more subnets that may be interconnected using routers.

[0017] Within a subnet, hosts may be connected using switches and point-to-point links. In addition, there may be one master management entity, a subnet manager (SM), that resides on a designated device in the subnet. The subnet manager is responsible for configuring, starting, and maintaining the IB subnet. In addition, the subnet manager (SM) may be responsible for performing routing table calculations in the IB fabric. Here, for example, routing in an IB network aims to provide fair load balancing between all source-destination pairs in the local subnet.

[0018] Through the subnet management interface, the subnet manager exchanges control packets called subnet management packets (SMP) with the subnet management agent (SMA). An agent resides on all IB subnet devices. Using SMP, the subnet manager can discover the fabric, configure end nodes and switches, and receive notifications from the SMA.

[0019] According to one embodiment, routing within a subnet in an IB network may be based on the LFT stored in the switch. The LFT is calculated by the SM according to the routing mechanism in use. In a subnet, Host Channel Adapter (HCA) ports on end nodes and switches are addressed using local identifiers (LIDs). Each entry in the LFT consists of a destination LID (DLID) and an output port. There is one LID per LID in the table. Only one entry is supported. When a packet arrives at a switch, its output port is determined by looking up the DLID in the switch's forwarding table. Routing is deterministic because packets follow the same path in the network between a given source-destination pair (LID pair).

[0020] In general, all other subnet managers except the master subnet manager It operates in standby mode for fault tolerance. However, in the situation where the master subnet manager fails, a new master subnet manager is negotiated by the standby subnet manager. The master subnet manager also performs periodic sweeps of the subnet to detect any topology changes and migrates the subnetworks accordingly. Reconfigure the network.

[0021] Additionally, hosts and switches within a subnet may be addressed using a local identifier (LID), and a single subnet may be limited to 49151 unicast LIDs. In addition to the LID, which is a local address valid within the subnet, each IB device may have a 64-bit global unique identifier (GUID). The GUID may be used to form a global identifier (GID), which is an IB Layer 3 (L3) address.

[0022] The SM may calculate routing tables (i.e., connections / routes between each pair of nodes in a subnet) at network initialization time. Additionally, whenever the topology changes, the routing tables may be updated to ensure connectivity and optimal performance. During normal operation, the SM may perform periodic light sweeps of the network to check for topology changes. If a change is discovered during a light sweep, Alternatively, if the SM receives a message (trap) signaling a network change, the SM may reconfigure the network according to the discovered change.

[0023] For example, the SM may reconfigure the network when the network topology changes, such as when a link goes down, a device is added, or a link is removed. The reconfiguration step may include a step performed during network initialization. Furthermore, the reconfiguration may have a local scope that is limited to the subnet where the network change occurred. Also, segmentation of a large fabric using routers may limit the reconfiguration scope.

[0024] According to one embodiment, an IB network may support partitioning as a security mechanism to provide isolation for logical groups of systems that share a network fabric. Each HCA port on a node in the fabric may be a member of one or more partitions. Partition membership is managed by a centralized partition manager, which may be part of the SM. The SM may configure the partition membership information for each port as a table of 16-bit partition keys (P_key). The SM may also manage the partition membership information for these ports. Switch ports and router ports can be configured with a partition enforcement table that contains P_Key information associated with end nodes that send or receive data traffic through the switch port. Additionally, in the general case, the partition membership of a switch port may represent the collection of all memberships indirectly associated with LIDs routed through the port in the egress direction (towards the link).

[0025] According to one embodiment, for communication between nodes, queue pairs (QP) and end-to-end contexts (EEC), with the exception of the management queue pair (QP0 and QP1), can be assigned to specific partitions. P_Key information can then be added to all transmitted IB transport packets. When a packet arrives at an HCA port or switch, its P_Key value can be checked against a table configured by the SM. If an invalid P_Key value is found, the packet is immediately discarded. In this way, communication is only allowed between ports that share a partition.

[0026] An example of an InfiniBand fabric is shown in Figure 1, which illustrates an example InfiniBand environment 100 according to one embodiment. In the example shown in Figure 1, nodes A101-E105 communicate using an InfiniBand fabric 120 via respective host channel adapters 111-115. According to one embodiment, various nodes (e.g., nodes A101-E105) may be represented by various physical devices. According to one embodiment, various nodes (e.g., nodes A101-E105) may be represented by various virtual devices, such as virtual machines.

[0027] Virtual Machines on InfiniBand Over the past decade, hardware virtualization support has virtually eliminated CPU overhead, memory overhead has been significantly reduced by virtualizing the memory management unit, storage overhead has been reduced by utilizing high-speed SAN storage or distributed network file systems, and device pass-through technologies such as Single Root Input / Output Virtualization (SR-IOV) have been introduced. The prospects for virtualized High Performance Computing (HPC) environments have improved significantly as network I / O overhead has been reduced by using high-performance interconnect solutions. Clouds now support virtual HPC (vHPC) clusters with high-performance interconnect solutions, delivering the required performance. can be provided.

[0028] However, when coupled with lossless networks such as InfiniBand (IB), some cloud features such as live migration of virtual machines (VMs) remain problematic due to the complex addressing and routing schemes used in these solutions.IB is an interconnect network technology that offers high bandwidth and low latency, making it well suited for HPC and other communication-intensive workloads.

[0029] The traditional approach to connecting IB devices to VMs is by using direct-assigned SR-IOV. However, achieving live migration of VMs assigned to IB host channel adapters (HCAs) using SR-IOV has proven challenging. Each IB-attached node has three different addresses (i.e., LID, GUID, and GID). When a live migration occurs, one or more of these addresses change. Other nodes communicating with the migrating VM (VM-in-migration) may lose connectivity. This When this occurs, the IB subnet manager (SM) is notified that it should reconnect by sending a Subnet Administration (SA) record route query. An attempt can be made to restore the lost connection by locating the virtual machine's new address.

[0030] IB uses three different types of addresses. The first type of address is a 16-bit local identifier (LID). At least one unique LID is assigned by the SM to each HCA port and each switch. The LID is used to route traffic within a subnet. Because the LID is 16 bits long, 65536 unique address combinations can be configured, of which only 49151 (0x0001-0xBFFF) can be used as unicast addresses. As a result, the number of available unicast addresses defines the maximum size of an IB subnet. The second type of address is a 64-bit globally unique identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM may assign additional subnet-specific GUIDs to HCA ports, which is useful when SR-IOV is used. The third type of address is a 64-bit globally unique identifier (GUID) assigned by the manufacturer to each device (e.g., HCA and switch) and each HCA port. The SM may assign additional subnet-specific GUIDs to HCA ports, which is useful when SR-IOV is used. The address of a HCA port is a 128-bit Global Identifier (GID). A GID is a valid IPv6 unicast address, and at least one is assigned to each HCA port. The GID is formed by combining a globally unique 64-bit prefix assigned by the fabric administrator with the GUID address of each HCA port.

[0031] Fat Tree (FTree) Topology and Routing According to one embodiment, some IB-based HPC systems employ a fat-tree topology to take advantage of the useful properties that fat trees offer, including full bisection bandwidth and inherent fault tolerance due to the availability of multiple paths between each source-destination pair. The initial concept behind fat-trees was to employ thicker links between nodes with more available bandwidth as the tree approached the root of the topology. The thicker links could help avoid congestion in higher-level switches, preserving bisection bandwidth.

[0032] FIG. 2 illustrates an example of a tree topology in a network environment, according to one embodiment. As shown in FIG. 2, one or more end nodes 201-204 may be connected in a network fabric 200. Network fabric 200 may be based on a fat-tree topology including multiple leaf switches 211-214 and multiple spine or root switches 231-234. In addition, network fabric 200 may include one or more intermediate switches, such as switches 221-224.

[0033] 2, each of end nodes 201-204 may be a multi-homed node, i.e., a single node that is connected to two or more portions of network fabric 200 via multiple ports. For example, node 201 may include ports H1 and H2, node 202 may include ports H3 and H4, node 203 may include ports H5 and H6, and node 204 may include ports H7 and H8.

[0034] Additionally, each switch may have multiple switch ports. For example, root switch 231 may have switch ports 1-2, root switch 232 may have switch ports 3-4, root switch 233 may have switch ports 5-6, and root switch 234 may have switch ports 7-8.

[0035] According to one embodiment, the Fat Tree routing mechanism is one of the most popular routing algorithms for IB-based Fat Tree topologies. The Fat Tree routing mechanism is also used in OFED (Open Fabric Enterprise Data Exchange). Distribution: A standard software for building and deploying IB-based applications This is implemented in the OpenSM (Software Stack) subnet manager.

[0036] The goal of a fat-tree routing mechanism is to generate an LFT that uniformly spreads shortest-path routes across links in the network fabric. The mechanism traverses the fabric in indexing order and assigns target LIDs for end nodes, and therefore corresponding routes, to each switch port. For end nodes connected to the same leaf switch, the indexing order may depend on the switch ports to which the end nodes are connected (i.e., the port numbering sequence). For each port, the mechanism may maintain a port usage counter, and each time a new route is added, the port usage counter may be used to select the least frequently used port.

[0037] According to one embodiment, in a partitioned subnet, nodes that are not members of a common partition are not allowed to communicate. In effect, this is a fat This means that some of the routes assigned by the tree routing algorithm will not be used for user traffic. Problems arise when the fat-tree routing mechanism generates LFTs for these routes in the same way as other functional paths. This behavior can degrade balancing on links because nodes are routed in indexing order. Fat-tree routed subnets generally provide poor isolation between partitions because routing is done without awareness of partitions.

[0038] According to one embodiment, a Fat-Tree is a hierarchical network topology that can scale with available network resources. Furthermore, Fat-Tree is easily constructed using commodity switches arranged at various levels of hierarchy. Furthermore, various variants of Fat-Tree are publicly available, including k-ary-n-tree, Extended Generalized Fat-Tree (XGFT), Parallel Ports Generalized Fat-Tree (PGFT), and Real Life Fat-Tree (RLFT).

[0039] Also, a k-ary-n-tree is an n-level fat tree with k n end nodes and n·k n-1 Each switch has 2k ports. Each switch has the same number of connections up and down the tree. XGFT fat trees extend k-ary-n-trees by allowing both different numbers of up and down connections for switches and different numbers of connections at each level in the tree. The PGFT definition further extends the XGFT topology to allow multiple connections between switches. A wide variety of topologies can be defined using XGFT and PGFT. However, for practical purposes, a restricted version of PGFT, RLFT, is introduced to define fat trees commonly found in modern HPC clusters. RLFT uses the same port count switches for all levels in the fat tree.

[0040] Input / Output (I / O) virtualization According to one embodiment, I / O Virtualization (IOV) can make I / O available by allowing virtual machines (VMs) access to the underlying physical resources. The combination of storage traffic and inter-server communication can place an unbearable strain on a single server's I / O resources, resulting in backlogs and idle processors waiting for data. As the number of I / O requests increases, IOV can provide availability and improve the performance, scalability, and elasticity of (virtualized) I / O resources to rival performance levels seen in modern CPU virtualization.

[0041] According to one embodiment, IOV is desired to enable sharing of I / O resources and to allow protected access to resources from VMs. IOV separates the logical device exposed to a VM from its physical implementation. Currently, emulation, paravirtualization, direct assignment (DA), and single-root I / O There can be various types of IOV technologies, such as virtualization (SR-IOV).

[0042] According to one embodiment, one type of IOV technology is software emulation. Software emulation can enable a separated front-end / back-end software architecture. The front-end can be a device driver located in a VM and communicate with a back-end implemented by a hypervisor to provide I / O access. The physical device sharing ratio is high, and live migration of VMs can be achieved with only milliseconds of network downtime. However, software emulation can also enable a separate front-end / back-end software architecture. Hardware emulation introduces additional, undesirable computational overhead.

[0043] According to one embodiment, another type of IOV technology is direct device assignment. Direct device assignment requires that an I / O device be attached to a VM, but the device is not shared between VMs. Direct assignment, or device passthrough, offers near-unique performance with minimal overhead. The physical device bypasses the hypervisor and is directly attached to the VM. However, a drawback of such direct device assignment is that there is no sharing between virtual machines, limiting scalability, such as one physical network card being attached to one VM.

[0044] According to one embodiment, Single Root IOV (SR-IOV) is Hardware virtualization may allow a physical device to appear as multiple independent, lightweight instances of the same device. These instances can be assigned to VMs as pass-through devices and accessed as Virtual Functions (VFs). The hypervisor accesses the device through a unique (per device) fully functional Physical Function (PF). SR-IO SR-IOV mitigates the scalability issues of purely direct allocation. However, a problem presented by SR-IOV is that it can impair VM migration. Among these IOV technologies, SR-IOV extends the PCI Express (PCIe) standard with a means to allow multiple VMs to directly access a single physical device while maintaining near-inherent performance. This allows SR-IOV to offer superior performance and scalability.

[0045] SR-IOV allows a PCIe device to expose multiple virtual devices that can be shared among multiple guests by assigning one virtual device to each guest. Each SR-IOV device has at least one physical function (PF) and one or more associated virtual functions (VFs). A PF is a communication function controlled by a virtual machine monitor (VMM) or hypervisor. VFs are lightweight PCIe functions, whereas VFs are regular PCIe functions. Each VF has its own base address (BAR) and is assigned a unique requestor ID. The unique requestor ID is managed by the I / O memory management unit (I / O memory management unit). The IOMMU also applies memory and interrupt translation between PFs and VFs.

[0046] Unfortunately, direct device allocation techniques present a barrier to cloud providers in situations where transparent live migration of virtual machines is desired for data center optimization. The essence of live migration is that the memory contents of a VM are copied to a remote hypervisor. Furthermore, the VM is suspended in the source hypervisor, and the VM's operation is resumed in the destination. When using software emulation methods, network interfaces are virtual so that their internal states are stored in memory and then copied. Therefore, downtime can be reduced to a few milliseconds.

[0047] However, migration becomes more difficult when direct device assignment techniques such as SR-IOV are used. In this situation, the entire internal state of the network interface cannot be copied because it is tied to the hardware. Instead, the SR-IOV VF assigned to the VM is detached and live migrated. A migration is performed and a new VF is assigned at the destination. For InfiniBand and SR-IOV, this process can cause downtime on the order of a few seconds. Furthermore, in the SR-IOV shared port model, the VM's address changes after the migration, which adds overhead to the SM and negatively impacts the performance of the underlying network fabric.

[0048] InfiniBand SR-IOV Architecture - Shared Port There can be various types of SR-IOV models (eg, a shared port model, a virtual switch model, and a virtual port model).

[0049] 3 illustrates an exemplary shared port architecture according to one embodiment. As shown, a host 300 (e.g., a host channel adapter) may interact with a hypervisor 310. The hypervisor 310 may assign various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions may be handled by the hypervisor 310.

[0050] 3, a host (e.g., an HCA) appears as a single port to the network with a single shared LID and shared Queue Pair (QP) space between the physical function 320 and the virtual functions 330, 350, 350. However, each function (i.e., the physical function and the virtual function) may have its own GID.

[0051] 3, according to one embodiment, various GIDs can be assigned to virtual and physical functions, and a special queue pair, QP0 and QP1 (i.e., a dedicated queue pair used for InfiniBand management packets), is owned by the physical function. These QPs are exposed to VFs as well, but VFs are not allowed to use QP0 (all incoming SMPs from VFs towards QP0 are discarded), and QP1 can act as a proxy for the actual QP1 owned by the PF.

[0052] According to one embodiment, the shared port architecture may enable highly scalable data centers that are not limited by the number of VMs (attached to the network by being assigned to virtual functions) because LID space is only consumed by the physical machines and switches in the network.

[0053] However, a drawback of the shared port architecture is that it cannot provide transparent live migration, thereby hindering the potential for flexible VM placement. Because each LID is associated with a specific hypervisor and shared among all VMs residing on that hypervisor, a migrating VM (i.e., a virtual machine migrating to a destination hypervisor) must change its LID to the LID of the destination hypervisor. Furthermore, as a result of the restricted QP0 access, a subnet manager cannot be run inside a VM.

[0054] InfiniBand SR-IOV Architecture Model - Virtual Switch (vSwitch) 4 illustrates an exemplary vSwitch architecture according to one embodiment. As shown, a host 400 (e.g., a host channel adapter) can interact with a hypervisor 410, which can assign various virtual functions 430, 440, and 450 to several virtual machines. Similarly, physical functions can be assigned to a host The virtual switch 415 may also be handled by the hypervisor 401.

[0055] According to one embodiment, in the vSwitch architecture, each virtual function 430, 440, 450 is a full virtual Host Channel Adapter (vHCA), which means that in hardware, the VM assigned to the VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM, HCA 400 appears as a switch with additional nodes connected via virtual switch 415. Hypervisor 410 can use PF 420, and the VM (attached to the virtual function) uses the VF.

[0056] According to one embodiment, the vSwitch architecture provides transparent virtualization. However, because each virtual function is assigned a unique LID, the available number of LIDs is quickly consumed. Similarly, if many LID addresses are used (i.e., one for each physical function and each virtual function), more communication paths must be computed by the SM and more subnet management packets (SMPs) must be sent to the switch to update their LFTs. For example, computing communication paths can take several minutes in a large network. Because the LID space is limited to 49,151 unicast LIDs and each VM (through a VF) occupies one LID per physical node and switch, the number of active VMs is limited by the number of physical nodes and switches in the network, and vice versa.

[0057] InfiniBand SR-IOV Architecture Model - Virtual Port (vPort) 5 illustrates an exemplary vPort concept according to one embodiment. As shown, a host 300 (e.g., a host channel adapter) can interact with a hypervisor 410 that can allocate various virtual functions 330, 340, and 350 to several virtual machines. Similarly, physical functions can be handled by the hypervisor 310.

[0058] According to one embodiment, the vPort concept is loosely defined to allow vendors implementation freedom (e.g., the definition does not stipulate that implementations should be SRIOV-only), and the purpose of vPort is to standardize how VMs are handled in a subnet. The vPort concept allows for the definition of both an SR-IOV shared port-like architecture and a vSwitch-like architecture, or a combination of these architectures, which may be more scalable in both the spatial and performance domains. Also, vPorts support optional LIDs, and unlike shared ports, the SM is aware of all vPorts available in a subnet, even if the vPorts do not use dedicated LIDs.

[0059] InfiniBand SR-IOV Architecture Model - LID Pre-Populated vSwitch According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs.

[0060] 6 illustrates an exemplary vSwitch architecture with pre-populated LIDs, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 600 (e.g., an IB subnet). The virtual machine 510 may include several hardware devices, such as host channel adapters 510, 520, and 530. Furthermore, host channel adapters 510, 520, and 530 may interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, may further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 may be assigned to virtual function 1 514 by hypervisor 511. Hypervisor 511 may additionally assign virtual machine 2 551 to virtual function 2 515, and virtual machine 3 552 to virtual function 3 556. Hypervisor 531 may assign virtual machine 552 to virtual function 3 516. Hypervisor 531 may further assign virtual machine 4 553 to virtual function 1 534. The hypervisor may access the host channel adapters through fully capable physical functions 513, 523, and 533 on each of the host channel adapters.

[0061] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 600.

[0062] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0063] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with pre-populated LIDs. Referring to FIG. 5, LIDs are pre-populated for various physical functions 513, 523, and 533, as well as for virtual functions 514-516, 524-526, and 534-536 (even virtual functions not currently associated with active virtual machines). For example, physical function 513 is pre-populated with LID 1, and virtual function 1 534 is pre-populated with LID 10. When a network is booted, LIDs are pre-populated in an SR-IOV vSwitch-enabled subnet. Populated VFs are assigned LIDs as shown in FIG. 5, even if not all of the VFs are occupied by VMs in the network.

[0064] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0065] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, each hypervisor consumes one LID for itself via the PF and can consume one or more LIDs for each additional VF. The sum of all VFs available across all hypervisors in an IB subnet gives the maximum amount of VMs that can run in the subnet. For example, in an IB subnet with 16 virtual functions per hypervisor in the subnet, each hypervisor consumes 17 LIDs in the subnet (one LID for each of the 16 virtual functions and one LID for the physical function). In such an IB subnet, For a single subnet, the theoretical hypervisor limit is defined by the number of available unicast LIDs: 2891 (49151 available LIDs divided by 17 LIDs per hypervisor), and the total number of VMs (i.e., limit) is 46256 (2891 hypervisors multiplied by 16 VFs per hypervisor). (In practice, these numbers are smaller, as each switch, router, or dedicated SM node in an IB subnet consumes LIDs as well.) Note that vSwitches do not need to occupy additional LIDs, as they can share LIDs with PFs.

[0066] According to one embodiment, in a vSwitch architecture with pre-populated LIDs, once the network is booted, communication paths are calculated for all LIDs. If a new VM needs to be started, the system does not need to add a new LID in the subnet. Otherwise, operations that may completely reconfigure the network, including recalculating paths, are the most time-consuming part. Instead, available ports for VMs are located in one of the hypervisors (i.e., available virtual functions), and virtual machines are assigned to available virtual functions.

[0067] According to one embodiment, the LID pre-populated vSwitch architecture also enables the ability to compute and use different routes to reach different VMs hosted by the same hypervisor. Essentially, this allows such subnets and networks to use LID-Mask-Control-like (LMC-like) features to provide alternative routes towards one physical machine without being bound by the LMC constraint that requires LIDs to be contiguous. The freedom to use non-contiguous LIDs is particularly useful when a VM needs to migrate and its associated LID needs to be delivered to the destination.

[0068] In accordance with one embodiment, several considerations can be taken into account along with the above-described advantages of a LID pre-populated vSwitch architecture. For example, because LIDs are pre-populated in an SR-IOV vSwitch-enabled subnet when the network is booted, the initial route computation (e.g., at startup) may take longer than if the LIDs were not pre-populated.

[0069] InfiniBand SR-IOV Architecture Model - vSwitch with Dynamic LID Allocation According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID allocation.

[0070] FIG. 7 illustrates an exemplary vSwitch architecture with dynamic LID assignment, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 700 (e.g., an IB subnet). The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. The host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. The hypervisor 511 also assigns virtual machine 2 551 to virtual function 2 515 and virtual machine 3 5 52 can be assigned to virtual function 3 516. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 1 534. The hypervisor can access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0071] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 700.

[0072] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, and 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0073] According to one embodiment, the present disclosure provides a system and method for providing a vSwitch architecture with dynamic LID assignment. Referring to FIG. 7 , various physical functions 513, 523, and 533 are dynamically assigned LIDs, with physical function 513 receiving LID 1, physical function 523 receiving LID 2, and physical function 533 receiving LID 3. Those virtual functions associated with active virtual machines may also receive dynamically assigned LIDs. For example, virtual machine 1 550 is active and associated with virtual function 1 514, so virtual function 514 may be assigned LID 5. Similarly, virtual function 2 515, virtual function 3 516, and virtual function 1 534 are each associated with an active virtual function. As such, these virtual functions are assigned LIDs: LID 7 is assigned to virtual function 2 515, LID 11 is assigned to virtual function 3 516, and LID 9 is assigned to virtual function 1 534. Unlike a vSwitch, which has pre-populated LIDs, virtual functions that are not currently associated with an active virtual machine do not receive an LID assignment.

[0074] According to one embodiment, dynamic LID assignment can substantially reduce initial path computation: When a network is booting for the first time and no VMs are present, a relatively small number of LIDs can be used for initial path computation and LFT distribution.

[0075] According to one embodiment, many similar physical host channel adapters can have two or more ports (with two ports shared for redundancy), and a virtual HCA can also be represented by two ports and connected to an external IB subnet via one or more virtual switches.

[0076] According to one embodiment, when a new VM is created in a system utilizing a vSwitch with dynamic LID allocation, a free VM slot is discovered and a unique, unused unicast LID is discovered as well to determine on which hypervisor the newly added VM should boot. However, there is no known route in the switch's LFT and network to handle the newly added LID. Computing a new set of routes to handle the newly added VM is undesirable in a dynamic environment where several VMs may be booted every minute. In a large IB subnet, computing a new set of routes could take several minutes, and this procedure would have to be repeated each time a new VM is booted.

[0077] Advantageously, according to one embodiment, since all VFs in a hypervisor share the same uplink with the PF, there is no need to compute a new set of routes. All that is required is to iterate through the LFTs of all physical switches in the network, copy the forwarding ports from the LID entries belonging to the PF of the hypervisor (on which the VM is created) to the newly added LID, and send a single SMP to update the corresponding LFT block of the particular switch. This eliminates the need for the system and method to compute a new set of routes.

[0078] According to one embodiment, the assigned LIDs in a vSwitch with a dynamic LID allocation architecture do not need to be contiguous. Comparing the assigned LIDs on VMs on each hypervisor between a vSwitch with pre-populated LIDs and a vSwitch with dynamic LID allocation, it can be seen that the assigned LIDs in the dynamic LID allocation architecture are discontinuous, whereas the pre-populated LIDs are essentially contiguous. Furthermore, in a vSwitch dynamic LID allocation architecture, when a new VM is created, the next available LID is used for the lifetime of the VM. Conversely, in a vSwitch with pre-populated LIDs, each VM inherits the LID already assigned to its corresponding VF, and in a network without live migration, VMs assigned consecutively to a given VF get the same LID.

[0079] According to one embodiment, a vSwitch with a dynamic LID allocation architecture can address the shortcomings of a vSwitch with a pre-populated LID architecture model at the expense of some additional network and runtime SM overhead. Each time a VM is created, the LFT of the physical switch in the subnet is updated with the newly added LID associated with the created VM. This operation requires one subnet management packet (SMP) to be sent per switch. Because each VM uses the same route as its host hypervisor, features such as LMC are also unavailable. However, there is no limit on the total number of VFs present on all hypervisors, and the number of VFs may exceed the unicast LID limit. In such a case, of course, not all VFs can be simultaneously granted on active VMs. Having more spare hypervisors and VFs adds flexibility for recovering from and optimizing fragmented network failures when operating near the unicast LID limit.

[0080] InfiniBand SR-IOV Architecture Model - Dynamic LID Allocation and Pre-Populated LID vSwitch FIG. 8 illustrates an exemplary vSwitch architecture with dynamic LID assignment and pre-populated LIDs for a vSwitch, according to one embodiment. As shown, several switches 501-504 can establish communication between members of a fabric, such as an InfiniBand fabric, within a network switching environment 800 (e.g., an IB subnet). The fabric can include several hardware devices, such as host channel adapters 510, 520, and 530. The host channel adapters 510, 520, and 530 can further interact with hypervisors 511, 521, and 531, respectively. Each hypervisor, along with the host channel adapters, can further interact with, configure, and assign to several virtual machines several virtual functions 514, 515, 516, 524, 525, 526, 534, 535, and 536. For example, virtual machine 1 550 can be assigned to virtual function 1 514 by hypervisor 511. The hypervisor 511 also runs virtual machine 2 551 can be assigned to virtual function 2 515. The hypervisor 521 can assign the virtual Machine 3 552 can be assigned to virtual function 3 526. Hypervisor 531 can further assign virtual machine 4 553 to virtual function 2 535. The hypervisor can access the host channel adapters through fully functional physical functions 513, 523, and 533 on each of the host channel adapters.

[0081] According to one embodiment, each of switches 501-504 may include several ports (not shown) that are used to configure linear forwarding tables to direct traffic within network switching environment 800.

[0082] According to one embodiment, virtual switches 512, 522, and 532 can be handled by respective hypervisors 511, 521, 531. In such a vSwitch architecture, each virtual function is a full virtual host channel adapter (vHCA), which means that in hardware, a VM assigned to a VF is assigned a set of IB addresses (e.g., GID, GUID, LID) and a dedicated QP space. To the rest of the network and SM (not shown), HCAs 510, 520, and 530 appear as switches with additional nodes connected via virtual switches.

[0083] According to one embodiment, the present disclosure provides a system and method for providing a hybrid vSwitch architecture with dynamic LID assignment and pre-populated LIDs. Referring to FIG. 7 , hypervisor 511 may be deployed with a vSwitch with a pre-populated LID architecture, while hypervisor 521 may be deployed with a vSwitch with pre-populated LIDs and dynamic LID assignment. Hypervisor 531 may be deployed with a vSwitch with dynamic LID assignment. Thus, physical function 513 and virtual functions 514-516 have their LIDs pre-populated (i.e., even virtual functions that are not assigned to active virtual machines are assigned LIDs). Physical function 523 and virtual function 1 524 may have their LIDs pre-populated, while virtual function 2 525 and virtual function 3 526 have their LIDs dynamically assigned (i.e., virtual function 2 525 is available for dynamic LID assignment, and virtual function 3 526 has been dynamically assigned an LID of 11 because it is attached to virtual machine 3 552). Finally, the functions (physical and virtual functions) associated with hypervisor 3 531 can have their LIDs dynamically assigned. This results in virtual function 1 534 and virtual function 3 536 being available for dynamic LID assignment, while virtual function 2 535 has been dynamically assigned an LID of 9 because it is attached to virtual machine 4 553.

[0084] 8, in which both LID pre-populated vSwitches and dynamic LID allocation vSwitches are utilized (independently or combined within any given hypervisor), the number of pre-populated LIDs per host channel adapter can be defined by a fabric administrator and can be in the range 0<=pre-populated VFs<=total VFs (per host channel adapter). The VFs available for dynamic LID allocation can be found by subtracting the number of pre-populated VFs from the total number of VFs (per host channel adapter).

[0085] According to one embodiment, many similar physical host channel adapters can have two or more ports (two ports shared for redundancy), and the virtual HCA is also represented by two ports and connects to the external IB subnet via one or more virtual switches. It can be connected to a

[0086] vSwitch Scalability According to one embodiment, a problem with using the vSwitch architecture is the limited LID space. To overcome the scalability issue with the LID space, the following three alternatives (each described in further detail below) can be used independently or in combination: using multiple subnets; introducing backward-compatible LID space extensions; and combining the vPort and vSwitch architectures to form a lightweight vSwitch.

[0087] According to one embodiment, multiple IB subnets can be used. The LID is a Layer 2 address and must be unique within a subnet. When an IB topology spans multiple subnets, the LID is no longer a limitation, but if a VM needs to migrate to a different subnet, its LID address can change because the address may already be in use in the new subnet. Spanning multiple subnets overcomes the LID limitation of a single-subnet topology, but it also means that Layer 3 GID addresses must be used for inter-subnet routing, which adds additional overhead and latency to the routing process because Layer 2 headers must be modified by routers at the edge of the subnets. Also, under current hardware and software implementations and loose IBA (InfiniBand Architecture) standards, SMs in individual subnets are no longer aware of the global topology to provide optimized routing paths for clusters spanning multiple subnets.

[0088] According to one embodiment, a backward-compatible LID space extension in IBA can be introduced. Problems can arise when increasing the number of LID bits, for example to 24 or 32 bits, thereby increasing the insufficient LID space. Increasing the LID space by such an amount can break backward compatibility, since the IB Local Route Header (LRH) would have to be overhauled, and legacy hardware would no longer be able to function with the new standard. According to one embodiment, the LID space can be extended in a way that maintains backward compatibility while still allowing new hardware to take advantage of the extended functionality. The LRH has 7 spare bits that are transmitted as 0 and ignored by the receiver. By utilizing two of these spare bits in the LRH for the Source LID (SLID) and two bits for the Destination LID (DLID), This allows the LID space to be expanded to 18 bits (quadruple the LID space) and allows the creation of a scheme with physical LIDs (pLIDs) assigned to physical devices and virtual LIDs (vLIDs) assigned to VMs.

[0089] According to one embodiment, if the additional two bits are sent as 0, the LID is used as currently defined in the IBA (48K unicast LID and 16K multicast LID), and the switch can look up their primary LFT for forwarding the packet. Otherwise, the LID is a vLID, and forwarding can be based on a secondary LFT with a size of 192K. Because the vLID belongs to a VM and the VM shares an uplink with the physical node with the pLID, the vLID can be excluded from the path calculation stage when configuring (e.g., initial configuration) or reconfiguring (e.g., after a topology change) the network, but the secondary LFT table in the switch can be updated as described above. When the SM boots and discovers the network, it It can identify if all of the hardware supports the extended LID space, if not, the SM can fall back to legacy compatibility mode and the VM should occupy a LID from the pLID space.

[0090] 9 illustrates an extended local route header according to one embodiment. As shown in the figure, within the local route header, a virtual lane (VL) 900 includes 4 bits, a link version (Lver) 901 includes 4 bits, a service level (SL) 902 includes 4 bits, and a LID extension flag (LID The extension flag (LEXTF) 903 contains one bit, the first reserved bit (R1) 9 04 contains 1 bit, link next header (LNH) 905 contains 2 bits, destination local ID (DLID) 906 contains 16 bits, and DLID prefix extension (DPF) 907 contains 2 bits. SLID prefix extension (SPF) 908 includes 2 bits The second reserved bit (R2) 909 comprises 1 bit, the packet length (PktLen) 910 comprises 11 bits, and the source local ID (SLID) 911 comprises 16 bits. According to one embodiment, both reserved bits 904 and 909 may be set to zero.

[0091] According to one embodiment, as described above, the LRH shown in Figure 9 utilizes four of the seven (originally) spare bits as prefix extensions for the destination local ID 906 and source local ID 908. This, when utilized, signals that the LRH is to be used in association with a vLID that can be associated with the LID extension flag and routed via a secondary LFT at the switch. Alternatively, if extensions 907 and 908 are sent as zero (and ignored by the receiver), the LID is associated with the pLID and used as currently defined in the IBA.

[0092] FIG. 10 illustrates two exemplary linear forwarding tables according to one embodiment. As shown in FIG. 10, linear forwarding table 916 is a forwarding table associated with a pLID. The LFT spans from entry 912 (entry 0, indexed by DLID=0) to entry 913 (entry 48K-1, indexed by DLID=48K-1). In this case, each entry in the LFT is indexed by a standard 16-bit DLID and contains a standard IB port number. In contrast, linear forwarding table 917 is a secondary forwarding table associated with a vLID. The LFT spans from entry 914 (entry 0, indexed by an 18-bit DPF+DLID=0) to entry 915 (entry 256K-1, indexed by an 18-bit DPF+DLID=256K-1). In this case, each entry is indexed by an extended 18-bit DPF+DLID and contains a standard IB port number.

[0093] According to one embodiment, a hybrid architecture can be used to form a lightweight vSwitch architecture. A vSwitch architecture that can migrate LIDs along with migrated VMs scales well for subnet management because there is no requirement for additional signaling to re-establish connectivity with peers after migration, as opposed to a shared LID scheme where the LID would change. On the other hand, a shared LID scheme scales well for LID space. A hybrid vSwitch+shared vPort model can be realized if the SM is aware of available SR-IOV virtual functions in a subnet, where certain VFs may receive dedicated LIDs while others are routed in a shared LID manner based on their GID. With some information about VM node roles (e.g., to calculate routes and perform load balancing in the network), the SM can easily manage the LIDs. To ensure that they are considered separately while running a migration, popular VMs with many peers (e.g., servers) can be assigned dedicated LIDs, while other VMs that do not interact with many peers or run stateless services (they do not need to be migrated and can be respawned) can share LIDs.

[0094] Routing Strategies for vSwitch-Based Subnets According to one embodiment, to achieve higher performance, the routing algorithm can take the vSwitch architecture into account when calculating routes. In a fat tree, vSwitches can be identified in the topology discovery process by their unique property that they have only one upstream link to their corresponding leaf switch. Once the vSwitches are identified, the routing function can generate LFTs for all switches so that traffic from each VM can find a route to all other VMs in the network. Each VM has its own address, so each VM can be routed independently from other VMs attached to the same vSwitch. This results in the routing function generating multiple independent routes to vSwitches in the topology, each carrying traffic for a specific VM. One drawback of this approach is that if VM distribution is uneven among vSwitches, vSwitches with more VMs are potentially allocated greater network resources. However, the single upstream link from the vSwitch to the corresponding leaf switch remains a bottleneck link shared by all VMs attached to a particular vSwitch. This can result in suboptimal network utilization. The simplest and fastest routing strategy is to generate routes between all vSwitch-vSwitch pairs and route VMs with the same routes assigned to the corresponding vSwitches. With pre-populated LID allocation schemes and dynamic LID allocation schemes, each vSwitch has a LID defined by the PF in the SR-IOV architecture. These PF LIDs for the vSwitch can be used to generate an LFT in the first stage of routing, and in the second stage, the VM's LID can be added to the generated LFT. In the pre-populated LID scheme, an entry for the VF LID can be added by copying the outgoing port of the corresponding vSwitch.Similarly, in the case of dynamic LID assignment when a new VM is booted, a new entry with the VM's LID and the corresponding vSwitch-determined egress port is added in every LFT. The problem with this strategy is that VMs belonging to different tenants sharing a vSwitch may inherently interfere with each other due to sharing the same complete route in the network. To solve this problem while maintaining high network utilization, a weighted routing scheme for virtualized subnets can be used.

[0095] According to one embodiment, a weighted routing scheme for vSwitch-based virtualized subnets can be utilized. In such a mechanism, each VM on a vSwitch is assigned a parameter weight that can be taken into account for balancing when calculating routes. The value of the weight parameter reflects the vSwitch's proportion of leaf switch link capacity allocated to VMs on that vSwitch. For example, a simple configuration may assign each VM a weight equal to 1 / num_vms, where num_vms is the number of booted VMs on the corresponding vSwitch hypervisor. Another possible implementation may be to assign a higher proportion of the vSwitch capacity to the most important VMs in order to prioritize traffic flowing towards them. However, since the cumulative weight of VMs per vSwitch may be equal across all vSwitches, the topology may be affected. Links in the network can be balanced without being affected by the actual VM distribution. At the same time, the scheme enables multiple paths where each VM can be routed independently in the network, while eliminating interference between VMs on the same vSwitch at intermediate links in the topology. This scheme can be combined with the implementation of per VM rate limits on each vSwitch to ensure that VMs do not exceed their assigned capacity. In addition, when multiple tenant groups exist in the network, techniques such as tenant-aware routing can be integrated with the proposed routing scheme to provide network-wide isolation between tenants.

[0096] According to one embodiment, the following describes weighted routing for an IB-based fat-tree topology. As a fat-tree routing algorithm, vSwitchFatTree recursively traverses the fat-tree topology to configure LFTs in all switches for the LIDs associated with each VM in the subnet. This mechanism is deterministic and supports destination-based routing, where backward computation for all routes starts from the destination node.

[0097] Weighted Fat-Tree Routing Algorithm for Virtualized Subnets

[0098]

number

[0099] According to one embodiment, the vSwitchFatTree routing mechanism works as follows: Each VM is assigned a proportional weight. This proportional weight is calculated by dividing the weight of the vSwitch node (e.g., taken as a constant 1) by the total number of VMs running on it. Various weighting schemes can also be implemented. For example, an implementation can be chosen to assign weights based on VM type. However, for the sake of brevity, this description will focus on a proportional weighting scheme. For each leaf switch, the routing mechanism sorts the connected vSwitches in decreasing order based on the number of VMs connected (line 3). This order ensures that higher weighted VMs are routed first, thus balancing the routes assigned to links. The routing mechanism recursively assigns routes to VMs in the tree by going through all leaf switches and their corresponding vSwitches, traversing the tree from each VM and calling ROUTEDOWNGOINGBYGOINGUP (line 10). The downstream port on each switch is selected based on the smallest cumulative downstream weight among all available upstream ports (ROUTEDOWNGOINGBYGOINGUP). (GBYGOINGUP; line 16). Once a downstream port is selected, the mechanism can increase the cumulative downstream weight for the corresponding port by the weight of the VM being routed (ROUTEDOWNGOINGBYGOINGUP; line 19). After the downstream port is configured, the routing mechanism can assign an upstream port for the route towards the VM on all connected downstream switches by going down the tree (updating the corresponding upward weight for the port). ) (ROUTEUPGOINGBYGOINGDOWN; line 20). The process is then repeated by going up to the next level in the tree. Once all VMs have been routed, the algorithm also routes the vSwitch's physical LID in the same way as the VMs, albeit with equal weighting to balance between vSwitch routes in the topology (not shown in the pseudocode). This is desirable to improve balancing when the minimal reconfiguration method is used in the context of live migration. Also, routing paths on the vSwitch's underlying physical LID can be used as pre-defined paths to quickly deploy new VMs without requiring reconfiguration. However, over a period of time, overall routing performance will decrease slightly while using the original vSwitchFatTree routing. To limit performance degradation, vSwitchFatTree-based reconfiguration may be performed offline when a certain performance threshold is exceeded.

[0100] According to one embodiment, the above-described routing mechanism can provide various improvements over regular / legacy routing mechanisms. Unlike the original Fat Tree routing algorithm, which does not take vSwitches or VMs into account in the topology, vSwitchFatTree marks vSwitches and routes each VM independently of other VMs connected to the vSwitch. Similarly, to achieve uneven VM distribution among vSwitches, each VM is assigned a weight corresponding to the proportion of links allocated to it on the vSwitch. The weights are used to maintain port counters to balance route distribution in the Fat Tree. The scheme also enables generalized weighted Fat Tree routing, where each VM can be assigned a weight based on its traffic profile or role priority in the network.

[0101] 11 through 14 illustrate an example of supporting efficient virtualization in a lossless interconnect network, according to one embodiment. Specifically, FIG. 11 illustrates a two-level fat-tree topology with four switches: root switches 925 and 926, leaf switches 920 and 921, and four virtual switches: VS1 931, VS2 941, VS3 951, and VS4 961. The four virtual switches, VS1 931, VS2 941, VS3 951, and VS4 961, are associated with four hosts / hypervisors 930, 940, 950, and 960, respectively. In this case, the four virtual switches support eight virtual machines: VM1 932, VM2 933, VM3 942, VM4 943, VM5 952, VM6 953, VM7 954, and VM8 955. Providing connectivity for 8,962.

[0102] To further explain vSwitch FatTree routing, consider a simple virtualized fat tree topology with four end nodes (vSwitches), as shown in Figure 11. Each of the vSwitches connected to leaf switch 920, VS1, and VS2, has two VMs running (VM1 and VM2 for VS1, and VM3 and VM4 for VS2). A second leaf switch 921 has VS3 with three VMs (VM5, VM6, and VM7), with one VM running on host vSwitch VS4. Because each leaf switch is connected to both root switches 925 and 926, there are two alternative paths available to route each VM through the root switch. Routing for the VMs connected to VS1 is shown in Figure 12, with circles indicating the selected downstream path from the root switch. VM1 is routed using 925 → 920, and VM2 is routed from 926 → 920. The corresponding downstream load counters are updated on the selected links to add 0.5 for each VM. Similarly, as shown in FIG. 13, after adding a route for VS2, VM3 and VM4 are routed via link 925 → 920 and link 926 → 920, respectively. Note that after routing all VMs connected to leaf switch 920, the total downstream load on both links is equal, even if the VMs were routed individually. The VM distribution on the vSwitch connected to leaf switch 921 is different, so the vSwitch with one VM (VS4) will be routed first. Route 925 → 921 is assigned to VM8, and all three VMs connected to VS3 are routed from 926 → 921 to maintain a balanced cumulative load on both downstream links. In the final routing shown in FIG. 14, assuming a VM distribution in the topology, the load is balanced on each link as much as possible, with independent routes to the VMs.

[0103] Minimal overhead reconfiguration on virtual machine live migration According to one embodiment, abbreviated as ItRC (Iterative Reconfiguration) The resulting dynamic reconfiguration mechanism repeats all route switching and updating as necessary when a VM is migrated. However, only a subset of switches need to actually update, depending on the existing LFTs in the subnet (i.e., the LFTs already calculated and present in each switch in the subnet).

[0104] 15 illustrates a potential virtual machine migration according to one embodiment. More specifically, FIG. 15 illustrates the special case of migration of a VM within a leaf switch where, regardless of the network topology, only the corresponding leaf switch requires an LFT update.

[0105] As shown in Figure 15, a subnet is made up of several switches, namely, Switch 1 The subnet may include switches 1301 through 1312. Some of these switches may include leaf switches, such as switch 1 1301, switch 2 1302, switch 11 1311, and switch 12 1312. The subnet may additionally include several hosts / hypervisors 1330, 1340, 1350, and 1360, and several virtual switches VS1 1331, VS2 1341, VS3 1351, and VS4 1361. The various hosts / hypervisors can host virtual machines in the subnet, such as VM1 1332, VM2 1333, VM3 1334, VM4 1342, VM5 1343, and VM6 1352, via virtual functions.

[0106] According to one embodiment, VM3 is attached to a high-performance When migrating a virtual function from hypervisor 1330 to a free virtual function in hypervisor 1340, only the LFT in leaf switch 1 1301 needs to be updated because both hypervisors are connected to the same leaf switch and local changes do not affect the rest of the network. For example, the initial routing algorithm determines that traffic from hypervisor 1360 to hypervisor 1330 follows a first path marked by a solid line (i.e., 12 → 9 → 5 → 3 → 1). Similarly, traffic from hypervisor 1360 to hypervisor 1340 follows a second path marked by a dashed line (i.e., 12 → 10 → 6 → 4 → 1). If VM3 is migrated and ItRC is used to reconfigure the network, traffic destined for VM3 will follow a first path to hypervisor 1330 before migration and a second path to hypervisor 1340 after migration. In this situation, assuming a fat-tree routing algorithm was used for the initial routing, the ItRC method would update half (6 / 12) of the total number of switches. However, only one leaf switch needs to be updated to keep the migrated VM connected.

[0107] According to one embodiment, by limiting the number of switch updates after a VM migration, the network can be reconfigured faster, reducing the time and overhead required for traditional routing updates. This is based on a topology-agnostic skyline technique, which allows for the F This can be achieved by a topology-aware fast reconfiguration method for supporting VM migration on fat trees, called TreeMinRC.

[0108] Subtrees and switch tuples in fat trees According to one embodiment, the following description utilizes the minimal overhead network reconfiguration method, FTreeMinRC, using XGFT as an exemplary fat-tree network. However, the concepts presented here are also valid for PGFT and RLFT. XGFT(n;m1,...,m n ;w1,...,w n ) is a fat tree with n+1 levels of nodes. The levels are denoted from 0 to n, with computation nodes at level n and switches at all other levels. All nodes at level i, 0≦i≦n-1, except for computation nodes with no children, have m i Similarly, except for the root switch, which has no parent, all other nodes at level i, 1≦i≦n have i It has a parent node of +1.

[0109]

number

[0110] According to one embodiment, each switch in an n+1 level XGFT has a unique n-tuple (l, x1, x2,..., x n ) where the leftmost tuple value (l) represents the level at which the tree is located, and the remaining values ​​(x1, x2, ..., x n ) represents the position of a switch in the tree relative to other switches. In particular, a switch A(l, a1,..., a l ,...,a n ) for all values ​​of a except for i=l+1. i =b i At level l+1,(l+1,b1,...,b l ,b l+1 ,...,b n ) is connected to switch B in

[0111] Figure 16 illustrates a switch tuple according to one embodiment. More specifically, the diagram illustrates a switch tuple as allocated by the OpenSM Fat Tree routing algorithm implemented for an exemplary fat tree, XGFT(4;2,2,2,2;2,2,2,2,1). Fat tree 1400 may include switches 1401-1408, 1411-1418, 1421-1428, and 1431-1438. Because the fat tree has n=4 switch levels (marked as column 0 at the root level through column 3 at the leaf level), the fat tree is composed of m1=2 first-level subtrees, each of which is n'=n-1=3 switch levels. This is illustrated in the diagram by two boxes defined by dashed lines enclosing switches at levels 1 through 3. Each first-level subtree receives an identifier of 0 or 1. Each first-level subtree is composed of m2=2 second-level subtrees, each of which is n"=n'-1=2 switch levels above a leaf switch. This is shown in the figure by the four boxes defined by dotted lines surrounding the switches from level 2 to 3. Each second-level subtree receives an identifier of 0 or 1. Similarly, each leaf switch can also be considered a subtree, shown in the figure by the eight boxes defined by dashed lines. Each of these subtrees receives an identifier of 0 or 1.

[0112] According to one embodiment, tuples, such as a tuple of four numbers as illustrated in the figure, can be assigned to various switches, with each number in the tuple indicating a particular subtree correspondence for each value's position in the tuple. For example, switch 1413 (which may be referenced as switch 1_3) can be assigned tuple 1.0.1.1, representing its position in level 1 and the 0th first-level subtree.

[0113] Fat-Tree-Aware Minimal Reconfiguration with FTreeMinRC in the Context of Live Migration According to one embodiment, the switch tuple encodes information about the location of the switch corresponding to the subtree in the topology. FTreeMinRC can use this information to enable rapid reconfiguration in the case of live VM migration. The tuple information can be used to find the skyline that minimizes the number of switches that need to be reconfigured by the SM when a VM is migrated. In particular, when a VM is migrated between two hypervisors in a fat-tree topology, the skyline representing the minimum number of switches that need to be updated is formed by all top-level switches of all subtrees involved in the migration.

[0114] According to one embodiment, when a VM is live migrated, a switch marking mechanism can be initiated from both leaf switches. In this case, the source and destination hypervisors are connected and switch tuples are compared. If the tuples match, the mechanism can determine that the VM has been migrated within the leaf switch. This marks only the corresponding leaf switch for reconfiguration. However, if the tuples do not match, the upstream links from both the source leaf switch and the destination leaf switch are traced. The switch located one level up is the highest-level switch in the immediate supertree to which the leaf-level subtree is connected, and is the only possible hop before reaching the leaf switch when traversing the tree downwards. The mechanism can then compare the source leaf switch tuple and destination leaf switch tuple with the newly traced switch, adjusting the tuple values ​​to reflect the current level, and then wildcarding the values ​​corresponding to the current subtree. Furthermore, the traced switch (which is the highest-level switch for the corresponding subtree) is marked for update, and if the comparisons from both the source switch tuple and the destination switch tuple match the tuples of all traced switches, the trace is stopped. Otherwise, the same procedure is repeated until the mechanism identifies a common ancestor switch from both ends. In the worst case, the mechanism can stop after reaching the root switch of the fat-tree topology. Since all upstream path tracing starts from the leaf level and marks the skyline switches of successive subtrees, when the mechanism reaches the top subtree affected by the migration, it has already selected all switches along the way that are potential traffic gateways towards lower level switches, as well as hypervisors involved in the live migration. Thus, the mechanism has marked all switches that form the skyline of the part of the network affected by the live migration.

[0115] According to one embodiment, the switch marking mechanism finds the minimum number of switches that need to be updated in terms of physical connectivity. However, it is possible that not all of these switches contain active routes calculated by the routing algorithm for the LIDs affected by the reconfiguration. Therefore, switches that contain active routes are prioritized in the update procedure, while the remaining switches with secondary routes can be updated later.

[0116] According to one embodiment, the fat-tree routing mechanism always routes traffic to a given destination through the same root switch. Because only a single path exists between the root switch and an end node in the topology, once the root switch selected to represent a given end node is located, the intermediate switches used to route traffic to the end node can be found. To discover the active route, the path can be traced from the source LID of the participating hypervisors to the destination LID, or vice versa. Switches that are a subset of the switches already selected for reconfiguration can be marked, and their LFT updates can be prioritized. The remaining selected switches can then be updated to keep all LFTs valid.

[0117] Figure 17 illustrates a reconfiguration process according to one embodiment. Fat tree 1400 may include switches 1401-1408, 1411-1418, 1421-1428, and 1431-1438. Because the Fat tree has n=4 switch levels (marked as column 0 at the root level through column 3 at the leaf level), the Fat tree is composed of m1=2 first-level subtrees, each with n'=n-1=3 switch levels. Each of these first-level subtrees is composed of m2=2 second-level subtrees, each with n"=n'-1=2 switch levels above the leaf switches. Similarly, each of the leaf switches can also be considered a subtree.

[0118] According to one embodiment, Figure 17 illustrates a situation in which a VM is being migrated between two hypervisors connected to leaf switches with tuples 3.0.0.0 and 3.0.1.1. These two tuples are used as a basis for comparison when the mechanism traces the path upward from the selected leaf switch. In this example, a common ancestor switch is found on level 1. Level 0 is the root level, and level 3 is the leaf level. The links between switches with the displayed tuple information are links that can be traced throughout the execution of the mechanism, and all of those same switches can be marked for update. The five highlighted switches (switches 1431, 1421, 1411, 1423, and 1434) and the links between them represent active routes, and their LFT updates can be prioritized.

[0119] According to one embodiment, to provide fast connectivity with minimal overhead in a virtualized data center that supports live migration, FTreeMinRC minimizes the number of LFT updates that need to be sent to the switch.

[0120] 18 is a flowchart of a method for supporting efficient virtualization in a lossless interconnect network according to one embodiment. In step 1810, the method includes, in one or more computers including one or more microprocessors, providing one or more switches including at least a leaf switch, each of the one or more switches including a plurality of ports, and providing a plurality of host channel adapters, each of the host channel adapters including at least one virtual function, at least one virtual switch, and at least one physical function, the plurality of host channel adapters being interconnected via the one or more switches, providing a plurality of hypervisors, each of the plurality of hypervisors being associated with at least one host channel adapter of the plurality of host channel adapters, and providing a plurality of virtual machines, each of the plurality of virtual machines being associated with at least one virtual function.

[0121] In step 1820, the method can deploy multiple host channel adapters with one or more of a virtual switch with a pre-populated local identifier (LID) architecture or a virtual switch with a dynamic LID assignment architecture.

[0122] In step 1830, the method can assign a LID to each virtual switch, the assigned LID corresponding to the LID of the associated physical function.

[0123] In step 1840, the method can calculate one or more linear forwarding tables based at least on the LIDs assigned to each of the virtual switches, each of the one or more LFTs being associated with a switch of the one or more switches.

[0124] 19 is a flowchart of a method for supporting efficient virtualization in a lossless interconnect network, according to one embodiment. In step 1910, the method includes, in one or more computers including one or more microprocessors, providing one or more microprocessors and one or more switches, the one or more switches may include at least leaf switches, each of the one or more switches including a plurality of ports, and providing a plurality of host channel adapters, each of the host channel adapters including at least one virtual function, at least one virtual switch, and at least one physical function, the plurality of host channel adapters communicating with each other via the one or more switches. The system is connected to the network, and further, a plurality of hypervisors can be provided, each of the plurality of hypervisors being associated with at least one host channel adapter among the plurality of host channel adapters, and further, a plurality of virtual machines can be provided, each of the plurality of virtual machines being associated with at least one virtual function.

[0125] In step 1920, the method can deploy multiple host channel adapters with one or more of a virtual switch with a pre-populated local identifier (LID) architecture or a virtual switch with a dynamic LID assignment architecture.

[0126] In step 1930, the method can assign each of the virtual switches a pLID from the plurality of pLIDs, the assigned pLID corresponding to the pLID of the associated physical function.

[0127] In step 1940, the method may assign a vLID from the plurality of vLIDs to each of the plurality of virtual machines, the LID space including the plurality of pLIDs and the plurality of vLIDs.

[0128] Many features of the present invention can be implemented in, using, or with the aid of hardware, software, firmware, or a combination thereof. Thus, features of the present invention can be implemented using a processing system (e.g., including one or more processors).

[0129] Features of this invention may be implemented in, using, or with the aid of a computer program product, which is a storage medium or computer-readable medium storing instructions usable to program a processing system to perform any of the features presented herein. The storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0130] Features of this invention may be incorporated into software and / or firmware, stored on any machine-readable medium, for controlling the hardware of a processing system and for enabling the processing system to interact with other mechanisms that utilize the results of this invention. Such software or firmware may include, but is not limited to, application code, device drivers, operating systems, and execution environments / containers.

[0131] Aspects of the present invention may also be implemented in hardware using hardware components such as, for example, application specific integrated circuits (ASICs). Implementing a hardware state machine to perform the functions described herein will be apparent to those skilled in the relevant art.

[0132] Additionally, the present invention may be conveniently implemented using one or more conventional general-purpose or specialized digital computers, computing devices, machines, or microprocessors, including one or more processors, memory, and / or computer-readable storage media, programmed according to the teachings of this disclosure. As will be apparent to those skilled in the software art, appropriate software coding, based on the teachings of this disclosure, may be readily implemented by skilled artisans. It can be easily prepared by the programmer.

[0133] While various embodiments of the present invention have been described above, it should be understood that they have been presented by way of example and not by way of limitation. It will be apparent to those skilled in the relevant art that various changes in form and detail can be made therein without departing from the spirit and scope of the invention.

[0134] The present invention has been described above with the help of functional building blocks illustrating the performance of specified functions and relationships thereof. For convenience of explanation, the boundaries of these functional building blocks have often been arbitrarily defined in this specification. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Thus, any such alternative boundaries are within the scope and spirit of the present invention.

[0135] The foregoing description of the invention has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. The breadth and scope of the invention should not be limited by any of the exemplary embodiments described above. Many modifications and variations will be apparent to those skilled in the art. These modifications and variations include any relevant combination of the disclosed features. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, so as to enable those skilled in the art to understand the invention in its various modifications and in various embodiments suited to the particular uses contemplated. It is intended that the scope of the invention be defined by the claims and their equivalents.

Claims

1. 1. A system for supporting efficient virtualization in a lossless interconnect network, comprising: one or more microprocessors; a subnet including a plurality of switches, the plurality of switches providing point-to-point connectivity between a plurality of host channel adapters, each host channel adapter running one of a plurality of hypervisors; a virtual machine performs a migration from a first hypervisor of the plurality of hypervisors to a second hypervisor of the plurality of hypervisors; as a result of the virtual machine migration, a subnet manager updates only one linear forwarding table in one switch of the plurality of switches within the subnet; The system, wherein the one linear forwarding table is determined by the subnet manager based on a skyline determination between the first hypervisor and the second hypervisor.

2. 2. The system of claim 1, wherein the subnet manager updates only the one linear forwarding table within the subnet by sending instructions to one of the plurality of switches via subnet management packets.

3. The system of claim 2 , wherein the one switch of the plurality of switches comprises a leaf switch.

4. 4. The system of claim 3, wherein the first hypervisor runs on a first host channel adapter and the second hypervisor runs on a second host channel adapter.

5. The system of claim 4 , wherein the first host channel adapter and the second host channel adapter are connected to the leaf switch.

6. the determined skyline includes the one switch; The system of any one of claims 1 to 5, wherein only one linear forwarding table is located in one switch.

7. Each of the plurality of switches is assigned one of a plurality of switch tuples, and the switch tuple encodes information about the position of the switch; The system of any one of claims 1 to 6, wherein the switch tuples are used to discover the skyline.

8. 1. A method for supporting efficient virtualization in a lossless interconnect network, comprising: providing a computer including one or more microprocessors; providing a subnet including a plurality of switches, the plurality of switches providing point-to-point connectivity between a plurality of host channel adapters, each host channel adapter running one of a plurality of hypervisors, the method further comprising: migrating a virtual machine from a first hypervisor of the plurality of hypervisors to a second hypervisor of the plurality of hypervisors; as a result of the virtual machine migration, a subnet manager updates only one linear forwarding table in one switch of the plurality of switches within the subnet; The method, wherein the one linear forwarding table is determined by the subnet manager based on a skyline determination between the first hypervisor and the second hypervisor.

9. 9. The method of claim 8, wherein the subnet manager updates only the one linear forwarding table within the subnet by sending instructions to one of the switches via subnet management packets.

10. The method of claim 9 , wherein the one switch of the plurality of switches comprises a leaf switch.

11. 11. The method of claim 10, wherein the first hypervisor runs on a first host channel adapter and the second hypervisor runs on a second host channel adapter.

12. The method of claim 11 , wherein the first host channel adapter and the second host channel adapter are connected to the leaf switch.

13. the determined skyline includes the one switch; The method of any one of claims 8 to 12, wherein only one linear forwarding table is located in one switch.

14. Each of the plurality of switches is assigned one of a plurality of switch tuples, the switch tuple encoding information about the position of the switch; The method of any one of claims 8 to 13, wherein the switch tuples are used to discover the skyline.

15. A computer readable program for causing one or more computers to carry out the method of any one of claims 8 to 14.

Citation Information

Patent Citations

  • Processing program for information processing apparatus, processing method for information processing apparatus, and information processing apparatus

    JP2010220103A

  • A system and method for supporting live migration of virtual machines in a virtualized environment.

    JP2015514271A

  • System and method for supporting multi-homed fat-tree routing in a middleware machine environment

    US20150030034A1

Cited By

  • Imaging lens and imaging device provided therewith

    WO2023176593A1