Determining the possibility of moving a virtual volume in a storage node of a storage cluster based on specified virtual machine boot conditions

By monitoring the historical boot time of the virtual machine and analyzing the boot conditions, identifying potential boot storm points, and migrating the virtual volumes to storage nodes with light loads, the resource overload problem caused by storage nodes due to large amounts of virtual machine boots is solved, and system stability and performance is improved.

CN115617256BActive Publication Date: 2025-07-29DELL PROD LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110783640.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-07-29
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively avoid the guidance storm caused by the storage nodes caused by the booting of a large number of virtual machines in a short time, resulting in resource flooding and performance degradation.

Method used

By monitoring the historical boot time of the virtual machine, analyzing the possibilities of boot conditions, identifying potential boot storm points, and migrating virtual volumes from high-risk storage nodes to lighter load storage nodes to avoid boot storms.

Benefits of technology

It effectively avoids resource overload of storage nodes, improves system stability and performance, and prevents service interruptions caused by boot storms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115617256B_ABST
    Figure CN115617256B_ABST
Patent Text Reader

Abstract

An apparatus, the apparatus including a processing device configured to obtain information characterizing a historical boot time of a virtual machine associated with a virtual volume hosted on a storage cluster including a plurality of storage nodes, and to determine, at least in part based on the obtained information, whether any of the storage nodes has at least a threshold likelihood of experiencing a specified virtual machine boot condition during a given time period. The processing device is further configured to, in response to determining that a first storage node among the storage nodes has at least the threshold likelihood of experiencing the specified virtual machine boot condition during the given time period, identify a subset of the virtual machines associated with a subset of the virtual volume hosted on the first storage node and move at least one of the subset of the virtual volume to a second storage node among the storage nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This field generally relates to information processing, and more particularly to techniques for managing an information processing system. Background Art

[0002] Information processing systems increasingly utilize reconfigurable virtual resources to meet changing user needs in an efficient, flexible, and cost-effective manner. For example, cloud computing environments implemented using various types of virtualization technologies are known. These illustratively include operating system-level virtualization technologies such as Linux containers. Such containers can be used to provide at least a portion of the cloud infrastructure of a given information processing system. Other types of virtualization, such as virtual machines implemented using a hypervisor, can be used additionally or alternatively. Summary of the Invention

[0003] Exemplary embodiments of the present disclosure provide techniques for moving virtual volumes in storage nodes of a storage cluster based at least in part on a determined likelihood of specified virtual machine boot conditions.

[0004] In one embodiment, a device includes at least one processing device including a processor coupled to a memory. The at least one processing device is configured to perform the following steps: obtain information characterizing historical boot times of a plurality of virtual machines associated with a plurality of virtual volumes hosted on a storage cluster including a plurality of storage nodes. The at least one processing device is further configured to perform the following steps: determine, at least in part based on the obtained information characterizing the historical boot times of the plurality of virtual machines, whether any of the plurality of storage nodes has at least a threshold likelihood of experiencing a specified virtual machine boot condition during a given period. The at least one processing device is further configured to perform the following steps: in response to determining that a first storage node of the plurality of storage nodes has at least the threshold likelihood of experiencing the specified virtual machine boot condition during the given period, identify a subset of the plurality of virtual machines associated with a subset of the plurality of virtual volumes hosted on the first storage node; and move at least one of the subset of the plurality of virtual volumes associated with at least one of the subset of the plurality of virtual machines from the first storage node to a second storage node of the plurality of storage nodes.

[0005] These and other exemplary embodiments include, but are not limited to, methods, devices, networks, systems, and processor-readable storage media. Brief Description of the Drawings

[0006] Figure 1FIG. 0 is a block diagram of an information processing system configured to move virtual volumes among storage nodes of a storage cluster at least in part based on a determined likelihood of specified virtual machine boot conditions.

[0007] Figure 2 FIG. 4 is a flow chart of an exemplary process configured to move virtual volumes among storage nodes of a storage cluster at least in part based on a determined likelihood of specified virtual machine boot conditions.

[0008] Figure 3 FIG. 8 illustrates a virtualized environment utilizing virtual volume storage in an illustrative embodiment.

[0009] Figure 4 FIG. 12 illustrates a process flow for relocating a virtual volume to avoid a boot storm condition on a storage node of a storage cluster in an illustrative embodiment.

[0010] Figure 5 and Figure 6 FIG. 18 illustrates an example of a processing platform that may be used to implement at least a portion of an information processing system in an illustrative embodiment. DETAILED DESCRIPTION

[0011] Exemplary information processing systems and associated computers, servers, storage devices, and other processing devices will be described herein. However, it should be understood that the embodiments are not limited to use with the specific illustrative system and device configurations shown. Thus, the term "information processing system" as used herein is intended to be broadly construed to cover, for example, processing systems including cloud computing and storage systems, and other types of processing systems including various combinations of physical and virtual processing resources. Thus, an information processing system may include, for example, at least one data center or other type of cloud-based system that includes one or more clouds hosting tenants accessing cloud resources.

[0012] Figure 1An information processing system 100 configured according to an illustrative embodiment is shown that provides functionality for moving virtual volumes among storage nodes of a storage cluster based at least in part on a determined likelihood of specified virtual machine (VM) boot conditions (e.g., to prevent or avoid having too many VM boots on a given storage node over a period of time). The information processing system 100 includes one or more host devices 102-1, 102-2, ... 102-H (collectively host devices 102) that communicate with a virtual desktop infrastructure (VDI) environment 112 via a network 104. The VDI environment 112 includes a virtualization infrastructure 114 for providing secure virtual desktop services in the form of VMs to multiple users in one or more enterprises (e.g., users of the host devices 102). User data for VMs provided using the virtualization infrastructure 114 can be stored on virtual volumes in one or more data stores. Each of the data stores can host multiple virtual volumes for one or more VMs. One or more storage arrays 106-1, 106-2, ... 106-S (collectively storage arrays 106) are also coupled to the network 104 and provide the underlying physical storage used by the data stores in the VDI environment 112. The storage arrays 106 can represent respective storage nodes of a storage cluster that hosts virtual volumes for VMs provided using the virtualization infrastructure 114. The network 104 can include a storage area network (SAN).

[0013] Storage array 106-1 (as Figure 1 shown) includes a plurality of storage devices 108, each of which stores data utilized by one or more of the applications running on the host device 102 (e.g., where such applications can include one or more applications running in a virtual desktop or VM in the VDI environment 112, possibly including the VM itself). The storage devices 108 are illustratively arranged in one or more storage pools. Storage array 106-1 also includes one or more storage controllers 110 that facilitate IO processing of the storage devices 108. Storage array 106-1 and its associated storage devices 108 are examples of what is more generally referred to herein as a "storage system". Such a storage system in this embodiment is shared by the host devices 102 and is thus also referred to herein as a "shared storage system". In an embodiment where there is only a single host device 102, the host device 102 can be configured to use the storage system exclusively.

[0014] The host device 102 and the virtualization infrastructure 114 of the VDI environment 112 illustratively include respective computers, servers, or other types of processing devices capable of communicating with the storage array 106 via the network 104. For example, the virtualization infrastructure 114 of the VDI environment 112 may implement respective VMs of a computing service platform or other type of processing platform. Similarly, at least a subset of the host device 102 may be implemented as respective VMs of a computing service platform or other type of processing platform. In such an arrangement, the virtualization infrastructure 114 of the host device 102 and / or the VDI environment 112 illustratively provides computing services, such as executing one or more applications on behalf of each of one or more users (associated with the respective host device and / or VDI environment in the host device 102 and / or the VDI environment 112).

[0015] The term "user" herein is intended to be interpreted broadly to cover numerous arrangements of people, hardware, software, or firmware entities, and combinations of such entities.

[0016] Computing and / or storage services may be provided to users according to a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, and / or a function as a service (FaaS) model, but it will be understood that many other cloud infrastructure arrangements may be used. Moreover, illustrative implementations may be implemented outside of a cloud infrastructure context, such as in the case of stand-alone computing and storage systems implemented within a given enterprise.

[0017] The storage device 108 of the storage array 106-1 may implement logical unit numbers (LUNs) configured to store objects of users associated with the host device 102 (e.g., virtual desktops or VMs in the VDI environment 112 for utilization by users of the host device 102). These objects may include files, blocks, or other types of objects. The host device 102 interacts with the storage array 106-1 using read and write commands and other types of commands transmitted via the network 104. In some embodiments, such commands more specifically include Small Computer System Interface (SCSI) commands, but in other embodiments other types of commands may be used. The term a given IO operation as used broadly herein illustratively includes one or more such commands. References to terms such as "input-output" and "IO" herein should be understood to refer to input and / or output. Thus, an IO operation involves at least one of input and output.

[0018] Moreover, as used herein, the term "storage device" is intended to be broadly interpreted to cover, for example, logical storage devices such as LUNs or other logical storage volumes. A logical storage device may be defined in storage array 106-1 to include different portions of one or more physical storage devices. Thus, storage device 108 may be considered to include the corresponding LUN or other logical storage volume.

[0019] The VDI environment 112, as described above, includes a virtualization infrastructure 114 for providing secure virtual desktop services in the form of VMs to multiple users in one or more enterprises (e.g., users of host device 102). Examples of processing platforms that may be used to provide the virtualization infrastructure 114 will be described in more detail below in connection with Figure 5 and Figure 6 The VDI environment 112 also includes a VM boot condition detection module 116 and a virtual volume relocation module 118. The VM boot condition detection module 116 is configured to detect one or more specified VM boot conditions. At least one of such specified VM boot conditions includes at least a threshold number of VMs booting within a specified time range (e.g., a VM "boot storm" that will be described in further detail below). The virtual volume relocation module 118 is configured to relocate virtual volumes associated with one or more VMs between different storage nodes (e.g., different storage arrays in storage array 106) to avoid overwhelming the resources of any one particular storage node in the storage nodes due to concurrently booting VMs on the storage nodes within the specified time range.

[0020] At least portions of the functionality of the VM boot condition detection module 116 and the virtual volume relocation module 118 may be implemented at least in part in the form of software stored in a memory and executed by a processor.

[0021] Although shown external to host device 102 and storage array 106 in the Figure 1 embodiment, it should be understood that in other embodiments, the VDI environment 112 may be implemented at least in part inside one or more of host device 102 and / or one or more of storage array 106 (e.g., on storage controller 110 of storage array 106-1). For example, one or more of host device 102 and / or storage array 106 may provide at least a portion of the virtualization infrastructure 114 that supports virtual desktops, VMs, and a data store (e.g., a virtual volume) for user data for the virtual desktops and VMs.

[0022] Figure 1In the embodiments, the host device 102, the storage array 106, and the VDI environment 112 are assumed to be implemented using at least one processing platform, where each processing platform includes one or more processing devices, and each processing device has a processor coupled to a memory. Such processing devices can illustratively include a specific arrangement of computing, storage, and network resources. For example, in some embodiments, the processing devices are implemented at least in part using virtual resources (such as VMs or Linux containers (LXC) or a combination of both), as in an arrangement where Docker containers or other types of LXC are configured to run on a VM.

[0023] Although the host device 102, the storage array 106, and the VDI environment 112 can be implemented on corresponding different processing platforms, numerous other arrangements are possible. For example, in some embodiments, at least portions of one or more of the host device 102, the storage array 106, and the VDI environment 112 are implemented on the same processing platform. One or more of the VDI environment 112, the storage array 106, or a combination thereof can thus be implemented at least in part within at least one processing platform that implements at least a subset of the host device 102.

[0024] The network 104 can be implemented using a variety of different types of networks to interconnect the storage system components. For example, the network 104 can include a SAN that is part of a global computer network such as the Internet, but other types of networks can be part of the SAN, including wide area networks (WANs), local area networks (LANs), satellite networks, telephone or cable networks, cellular networks, wireless networks (such as WiFi or WiMAX networks), or various parts or combinations of these and other types of networks. Thus, in some embodiments, the network 104 includes a combination of a variety of different types of networks, each network including processing devices configured to communicate using Internet Protocol (IP) or other related communication protocols.

[0025] As a more specific example, some embodiments can utilize one or more high-speed local area networks, where the associated processing devices communicate with each other using the peripheral component interconnect high-speed (PCIe) cards of those devices and networking protocols such as InfiniBand, Gigabit Ethernet, or Fibre Channel. As those skilled in the art will appreciate, numerous alternative networking arrangements are possible in a given embodiment.

[0026] Although in some embodiments, certain commands used by host device 102 to communicate with storage array 106 illustratively include SCSI commands, in other embodiments other types of commands and command formats may be used. For example, some embodiments may utilize command features and functionality associated with Non-Volatile Memory Express (NVMe), such as described in the May 2017 revision 1.3 of the NVMe specification, which is incorporated herein by reference, to implement IO operations. Other storage protocols of this type that may be utilized in the illustrative embodiments disclosed herein include: fabric-based NVMe, also known as NVMeoF; and Transmission Control Protocol (TCP)-based NVMe, also known as NVMe / TCP.

[0027] Assume that in this embodiment, storage array 106-1 includes persistent memory implemented using flash memory or other types of non-volatile memory of storage array 106-1. More specific examples include NAND-based flash memory or other types of non-volatile memory, such as resistive RAM, phase change memory, spin transfer torque magnetoresistive RAM (STT-MRAM), and Intel Optane based on 3D XPoint TM memory. TM device. Further assume that the persistent memory is separate from storage devices 108 of storage array 106-1, but in other embodiments, the persistent memory may be implemented as one or more designated portions of one or more of storage devices 108. For example, in some embodiments, such as in embodiments involving all-flash storage arrays, storage devices 108 may include flash-based storage devices, or may be implemented in whole or in part using other types of non-volatile memory.

[0028] As mentioned above, communication between host device 102 and storage array 106 may utilize a PCIe connection or other type of connection implemented over one or more networks. For example, illustrative embodiments may use interfaces such as Internet Small Computer System Interface (iSCSI), Serial Attached SCSI (SAS), and Serial ATA (SATA). In other embodiments, numerous other interfaces and associated communication protocols may be used.

[0029] In some embodiments, storage array 106 and other parts of system 100, such as VDI environment 112, may be implemented as part of a cloud-based system.

[0030] The storage device 108 of the storage array 106-1 may be implemented using a solid state drive (SSD). Such SSDs are implemented using non-volatile memory (NVM) devices such as flash memory. Other types of NVM devices that may be used to implement at least a portion of the storage device 108 include non-volatile random access memory (NVRAM), phase change RAM (PC-RAM), and magnetic RAM (MRAM). These and various combinations of multiple different types of NVM devices or other storage devices may also be used. For example, a hard disk drive (HDD) may be used in combination with or in place of an SSD or other type of NVM device. Thus, numerous other types of electronic or magnetic media may be used to implement at least a subset of the storage device 108.

[0031] The storage array 106 may alternatively or additionally be configured to implement multiple different storage layers of a multi-layer storage system. For example, a given multi-layer storage system may include a fast or performance layer implemented using flash storage devices or other types of SSDs and a capacity layer implemented using HDDs, where one or more such layers may be server-based. It will be apparent to those skilled in the art that numerous other types of storage devices and multi-layer storage systems may be used in other embodiments. The particular storage devices used in a given storage layer may vary according to the specific needs of a given embodiment, and multiple different storage device types may be used in a single storage layer. As previously indicated, the term "storage device" as used herein is intended to be interpreted broadly and thus may encompass, for example, SSDs, HDDs, flash drives, hybrid drives, or other types of storage products and devices or portions thereof, and illustratively includes logical storage devices such as LUNs.

[0032] As another example, the storage array 106 may be used to implement one or more storage nodes in a clustered storage system that includes multiple storage nodes interconnected by one or more networks.

[0033] Thus, it should be apparent that the term "storage array" as used herein is intended to be interpreted broadly and may encompass multiple different instances of commercially available storage arrays.

[0034] Other types of storage products that may be used to implement a given storage system in an illustrative embodiment include software-defined storage, cloud storage, object-based storage, and scale-out storage. In an illustrative embodiment, combinations of multiple of these and other storage types may also be used to implement a given storage system.

[0035] In some embodiments, the storage system includes a first storage array and a second storage array arranged in an active-active configuration. For example, such an arrangement can be used to ensure that data stored in one of the storage arrays is replicated to the other storage array using a synchronous replication process. Such data replication between multiple storage arrays can be used to facilitate fault recovery in system 100. Thus, one of the storage arrays can act as a production storage array relative to the other storage array that acts as a backup or recovery storage array.

[0036] However, it should be understood that the embodiments disclosed herein are not limited to an active-active configuration or any other particular storage system arrangement. Thus, the illustrative embodiments herein can be configured using a variety of other arrangements, including, for example, active-passive arrangements, active-active asymmetric logical unit access (ALUA) arrangements, and other types of ALUA arrangements.

[0037] These and other storage systems can be part of what is more generally referred to herein as a processing platform, which includes one or more processing devices, each including a processor coupled to a memory. A given such processing device can correspond to one or more virtual machines or other types of virtualization infrastructure, such as Docker containers or other types of LXC. As indicated above, communication between such elements of system 100 can occur over one or more networks.

[0038] As used herein, the term "processing platform" is intended to be interpreted broadly to include, for example but not limited to, multiple sets of processing devices configured to communicate over one or more networks and one or more associated storage systems. For example, a distributed implementation of host device 102 is possible, where some of the host devices in host device 102 reside in one data center at a first geographic location, while other host devices in host device 102 reside in one or more other data centers at one or more other geographic locations that may be remote from the first geographic location. Storage array 106 and VDI environment 112 can be implemented at least in part at the first geographic location, the second geographic location, and one or more other geographic locations. Thus, in some implementations of system 100, different host devices in host device 102, storage array 106, and VDI environment 112 can reside in different data centers.

[0039] Numerous other distributed implementations of host device 102, storage array 106, and VDI environment 112 are possible. Thus, host device 102, storage array 106, and VDI environment 112 can also be implemented in a distributed manner across multiple data centers.

[0040] Additional examples of processing platforms for implementing the various parts of system 100 in the illustrative embodiments will be described below in connection with Figure 5 and Figure 6 in more detail.

[0041] It should be understood that Figure 1 the set of specific elements shown in for moving virtual volumes in the storage nodes of a storage cluster based at least in part on the determined likelihood of specified VM boot conditions is presented by way of illustrative example only, and additional or alternative elements may be used in other embodiments. Thus, another embodiment may include additional or alternative systems, devices, and other network entities, as well as different arrangements of modules and other components.

[0042] It should be understood that these and other features of the illustrative embodiments are presented by way of example only and should not be construed as limiting in any way.

[0043] An exemplary process for moving virtual volumes in the storage nodes of a storage cluster based at least in part on the determined likelihood of specified VM boot conditions will now be described in more detail with reference to the flowchart of Figure 2 It should be understood that this particular process is merely an example, and additional or alternative processes for moving virtual volumes in the storage nodes of a storage cluster based at least in part on the determined likelihood of specified VM boot conditions may be used in other embodiments.

[0044] In this embodiment, the process includes steps 200 through 206. It is assumed that these steps are performed by the VDI environment 112 using the VM boot condition detection module 116 and the virtual volume relocation module 118. The process begins at step 200: obtaining information characterizing the historical boot times of a plurality of VMs, the plurality of VMs being associated with a plurality of virtual volumes hosted on a storage cluster including a plurality of storage nodes (e.g., storage array 106). Step 200 may include monitoring the creation times of one or more specified types of the plurality of virtual volumes. The one or more specified types of the plurality of virtual volumes may include at least one type of virtual volume created when a VM is powered on. The one or more specified types of the plurality of virtual volumes may also or alternatively include at least one type of virtual volume containing a copy of a VM memory page not retained in the memory of the plurality of VMs. The one or more specified types of the plurality of virtual volumes include swap virtual volumes (e.g., virtual volumes containing swap files for a VM).

[0045] In step 202, determine whether any one of a plurality of storage nodes has at least a threshold likelihood of experiencing a specified VM boot condition during a given time period (based on the information obtained in step 200). In step 204, in response to determining that a first storage node among the plurality of storage nodes has at least a threshold likelihood of experiencing a specified VM boot condition during a given time period, identify a subset of a plurality of VMs associated with a subset of a plurality of virtual volumes hosted on the first storage node. The specified VM boot condition may include at least a threshold likelihood of booting more than at least a threshold number of a plurality of VMs on the first storage node during a given time period. The threshold number of a plurality of VMs may be selected at least in part based on resources available at the first storage node.

[0046] In step 206, move at least one of a subset of a plurality of virtual volumes associated with at least one of the subset of a plurality of VMs from the first storage node to a second storage node among the plurality of storage nodes. In some embodiments, step 206 includes moving all virtual volumes associated with a given one of the subset of a plurality of VMs from the first storage node to the second storage node.

[0047] Step 206 may include: generating a probability density function of the boot time of each of the subset of a plurality of VMs; using the probability density function of the boot time of the subset of a plurality of VMs to determine the likelihood that each of the subset of a plurality of VMs boots during a given time period; and selecting at least one of the subset of a plurality of VMs at least in part based on the determined likelihood that each of the subset of a plurality of VMs boots during a given time period. Selecting at least one of the subset of a plurality of VMs at least in part based on the determined likelihood that each of the subset of a plurality of VMs boots during a given time period includes selecting a given one of the subset of a plurality of VMs having the highest boot probability of the probability density function during the given time period.

[0048] Step 206 may further include: identifying a subset of a plurality of storage nodes as candidate destination storage nodes, each of the candidate destination storage nodes currently hosting fewer than a specified number of a plurality of virtual volumes; and selecting the second storage node from the candidate destination storage nodes. The specified number of a plurality of virtual volumes may be determined at least in part based on the average number of virtual volumes hosted on each of the plurality of storage nodes. Selecting the second storage node from the candidate destination storage nodes may be at least in part based on determining the likelihood that each of the candidate storage nodes experiences a specified VM boot condition during a given time period. Selecting the second storage node from the candidate destination storage nodes may include selecting the candidate storage node having the lowest likelihood of experiencing a specified VM boot condition during a given time period.

[0049] "Boot storm" is a term used to describe service degradation that occurs when a large number of VMs are booted within a very narrow time frame. A boot storm can flood the network with data requests and can overwhelm system storage. A boot storm can severely disrupt a VDI environment by reducing performance and hindering productivity, and thus it is necessary to prevent a boot storm from occurring. Illustrative embodiments provide techniques for avoiding a boot storm or more generally avoiding VM boot conditions that can potentially have a negative impact on a VDI environment or other virtualization and IT infrastructure on which VMs operate. To that end, some embodiments appropriately construct and balance VM placement between servers and storage nodes that provide the virtualization and IT infrastructure on which VMs operate. A suitable architecture can be determined by measuring storage requirements and allocating storage resources to meet peak and average storage requirements.

[0050] Boot storm conditions can be common in various use cases. For example, consider a typical office workday: employees or other end users log in to the system around 8:30 am and log out around 5:00 pm. The servers utilized by the office may be able to handle the usage during the entire workday duration, but problems can occur if too many VMs are booted within a short time frame (e.g., 8:30 am to 9:00 am). This typically inevitable synchronous startup can overwhelm system resources and storage, preventing users from fully accessing the system and its VMs until sufficient resources are available.

[0051] In some data centers or other IT infrastructures that include virtualization infrastructure (e.g., a VDI environment that includes multiple VMs), SAN and NAS arrays may be virtualized. For example, VMware Virtual Volume (VVol) integration and management frameworks can be used to virtualize SAN and NAS arrays, enabling a more efficient operating model that is optimized for virtualized environments and is application - rather than infrastructure - centric.

[0052] VVol is an object type that corresponds to a VM disk (VMDK) (e.g., object type). On a storage system, a VVol resides in a VVol data store, which is also referred to as a storage container. A VVol data store is a type of A data store that allows VVols to be directly mapped to a storage system at a finer granularity level than VM File System (VMFS) and Network File System (NFS) data stores. Although VMFS and NFS data stores are managed and configured at the LUN or file system level, VVols allow VMs or virtual disks to be managed independently. For example, an end user can create a VVol data store based on an underlying storage pool and allocate a specific portion of one or more storage pools for the VVol data store and its VVols. A hypervisor (such as VMware ESXi TM ) can use NAS and SCSI protocol endpoints (PEs) as access points for IO communication between a VM and its VVol data store on the storage system.

[0053] VVols and VMware vSAN TM share a common storage operation model, namely, a leading hyper-converged infrastructure (HCI) solution. VMware vSAN TM is a software-defined enterprise storage solution that supports HCI systems. vSAN is fully integrated within VMware as a distributed software layer within the ESXi hypervisor . vSAN is configured to aggregate local or directly attached storage to create a single storage pool shared among all hosts in a vSAN cluster.

[0054] Both VVols and vSAN utilize Storage Policy-Based Management (SBPM) to eliminate storage configuration and use descriptive policies that can be applied or changed quickly (e.g., within minutes) at the VM or VMDK level. SBPM accelerates storage operations and reduces the need for storage infrastructure expertise. Advantageously, VVols make it easier to provision and enable the correct storage service level according to the specific requirements of individual VMs. By having better control of storage resources and data services at the VM level, administrators of virtualized infrastructure environments can create exact combinations and precisely provision storage service levels.

[0055] In a virtualized infrastructure environment that utilizes VVols, a swap VVol is created when a VM is powered on. The swap VVol contains copies of the VM memory pages that are not retained in memory. Advantageously, the VM boot time can be traced and monitored from the swap VVol.

[0056] If many VMs boot in a short period of time and the associated VVols of those VMs are located at the same storage node, this can lead to a boot storm that may cripple traditional system storage. Illustrative embodiments provide techniques for avoiding such boot storm conditions by learning and analyzing VM boot patterns from VM boot time history data learned by monitoring the swap VVols of VMs. VM boot "hot spots" are sought and boot storms under such boot hot spots are avoided by implementing balancing of VVols to different storage nodes. In some embodiments, a VM boot probability distribution is learned from VM boot history data to identify potential boot storm points. Then potential boot storm points are avoided by implementing VVol balancing among different storage nodes in a storage cluster.

[0057] Figure 3 A framework of a storage environment 300 utilizing VVol technology is shown. The storage environment 300 includes a VM management environment 301 (e.g., a VMware environment) coupled via a PE 303 to a virtual volume or VVol-enabled storage cluster 305. The VM management environment 301 includes a set of server nodes 310-A, 310-B, 310-C, ... 310-M (collectively referred to as server nodes 310). The VM management environment 301 also includes a virtual volume (VVol) data store 312 for a set of VMs 314-1, 314-2, 314-3, 314-4, ... 314-V (collectively referred to as VMs 314). The VVol-enabled storage cluster 305 includes a set of storage nodes 350-A, 350-B, 350-C, ... 350-N (collectively referred to as storage nodes 350) and a storage container 352 that includes VVols 354-1, 354-2, 354-3, 354-4, ... 354-O (collectively referred to as VVols 354). It should be noted that the number M of server nodes 310 may be the same as or different from the number N of storage nodes 350.

[0058] The VVols 354 are exported to the VM management environment 301 via the PE 303 (e.g., which may include server nodes 310 that implement corresponding ESXi hosts). The PE 303 is part of a physical storage fabric and establishes data paths from the VMs 314 to their corresponding VVols 354 as needed. The storage nodes 350 enable data services on the VVols 354. The storage container 352 may provide a storage capacity pool logically grouped into the VVols 354. The VVols 354 are created inside the storage container 352. The storage container 352 may be presented to the server nodes 310 (e.g., ESXi hosts) of the VM management environment 301 in the form of the VVol data store 312.

[0059] There are various different types of virtual volumes or VVols 354 that provide specific functions based on their role in a VM management environment 301 (e.g., VMware and / or VMware environment). Different types of VVols include: Config VVol; Data VVol; Swap VVol; Memory VVol; and Other VVols. Each VM has a Config VVol. The Config VVol holds information previously in the VM directory (e.g.,.vmx file, VM logs, etc.). Each virtual data disk has a Data VVol. The Data VVol holds the system data of the VM and is similar to a VMDK file. Each swap file has a Swap VVol. As described above, the Swap VVol is created when the VM is powered on and contains a copy of the VM memory pages not retained in memory. Each snapshot has a Memory VVol. The Memory VVol of a given VM contains a complete copy of the VM memory as part of a memory VM snapshot of the given VM. Other VVols are VMware specific solution types.

[0060] Since the Swap VVol is only created when the VM is powered on, the illustrative embodiment monitors the status of the Swap VVol associated with different VMs to check when the VMs are powered on and off. Thus, the illustrative embodiment is able to collect the boot history information of all VMs in the system by monitoring the associated Swap VVol of the VM. By learning the history of VM boot activities (e.g., by monitoring the Swap VVol associated with the VM), the probability distribution function of each VM's boot event at a time point t can be determined. The probability that VMi boots at time point t (where t is measured in minutes) is represented as:

[0061] P_bootup{VM i,t}

[0062] Then the probability density function of the VM boot event is determined according to the following equation:

[0063]

[0064] P_bootup{VM i,(t - 10)≤t≤(t + 10)} represents the probability that VM i boots during the time period {t - 10, t + 10}. Here, the 10-minute floating time is selected or defined based on the assumption that a complex VM may take 10 to 20 minutes to complete booting. The specific floating time value (e.g., "10" in the above equation) can be adjusted by the end user as needed for a specific implementation (e.g., based on the expected time to boot a VM in the end user's environment).

[0065] Similarly, by learning the historical data of the boot storm activities of each storage cluster node, the probability distribution function of the boot storm event of each storage cluster node j at time point t is expressed as:

[0066] P_bootup_storm{Node j,t}

[0067] The event probability of the VM boot storm on storage node j at time point t is defined as the probability when the concurrent VM boot count reaches the threshold M (for example, such as 30 VMs). The end user can adjust the threshold M as needed. Generally, the greater the power of the server or storage node, the more concurrent VM boot operations it can support.

[0068] The process flow for VM boot balancing among storage nodes (e.g., storage node 350) in a storage cluster enabled with virtual volumes (e.g., VVol-enabled storage cluster 305) will now be described with respect to Figure 4 is described. Figure 4 The process utilizes the historical VM boot activities learned by monitoring the swap VVols associated with the VMs to predict the possible upcoming boot storm times and adjusts the VM location distribution among the storage nodes in the cluster to avoid such possible upcoming boot storms. VM location balancing advantageously helps to avoid concurrent VM boot activities exceeding the specified threshold. Figure 4 The process begins at step 401. In step 403, the VM boot historical data is collected by monitoring whether there is an associated Swap VVol for a given VM. Based on the data collected in step 403, the probability distribution function (e.g., P_bootup_storm{Node j,t}) of the boot storm event of each storage cluster node is evaluated in step 405.

[0069] In step 407, it is determined whether the value of P_bootup_storm{Node j,t} of any of the storage cluster nodes exceeds the acceptable threshold Φ. If the determination result of step 407 is no (e.g., corresponding to the case where no boot storm is predicted for any of the storage cluster nodes), then Figure 4 the process flow ends at step 417. If the determination result of step 417 is yes (e.g., for any of the storage nodes in the storage cluster), then Figure 4The process flow proceeds to step 409. In step 409, for each VM residing on storage node j (e.g., having an associated P_bootup_storm{node j,t}>Φ), P_bootup{VM i,(t - 10)≤t≤(t - 10)} of the VM is evaluated. Then the VM with the maximum value of P_bootup{VM i,(t - 10)≤t≤(t - 10)} is selected as the current target VM.

[0070] For each other storage cluster node k that meets a specified condition (e.g., having a number of VMs less than a threshold Θ), P_bootup_storm{node k,t} is evaluated in step 411. The storage node with the minimum P_bootup_storm{node k,t} value is selected as the current target destination storage node. Here, the threshold Θ is used to ensure that the workload of each storage cluster node is not overly heavy after VM relocation. The user can select the value of the threshold Θ. In some embodiments, the threshold Θ can be defined according to the following formula:

[0071]

[0072] In step 413, the relevant VVol of the target VM is moved from storage cluster node j to the target destination storage node. Then the target VM is removed from the historical data of storage cluster node j, and in step 415, the probability distribution function P_bootup_storm{node j,t} of storage cluster node j is re-evaluated. Figure 4 The process flow then returns to step 407 to see if the re-evaluated P_bootup_srorm{node j,t} value still exceeds the acceptable threshold Φ. The VM balancing operation then continues (in another iteration of steps 409 to 413) or the balancing algorithm ends for node j. Then step 407 is repeated for each storage node (e.g., to determine if any other storage node j has P_bootup_storm{node j,t}>Φ). When there are no remaining storage nodes j that meet P_bootup_storm{node j,t}>Φ, Figure 4 The process flow ends at step 417.

[0073] It should be understood that the specific advantages described above and elsewhere in this document are associated with specific illustrative embodiments and need not exist in other embodiments. Moreover, the specific types of information processing system features and functionality shown in the figures and described above are merely exemplary, and numerous other arrangements may be used in other embodiments.

[0074] An illustrative implementation of a processing platform for implementing functionality for moving virtual volumes in storage nodes of a storage cluster based at least in part on a determined likelihood of specified VM boot conditions will now be described with reference to FIGS. 5 and Figure 6 in more detail. Although described in the context of system 100, in other implementations, these platforms can also be used to implement at least a portion of other information processing systems.

[0075] Figure 5 An exemplary processing platform including cloud infrastructure 500 is shown. Cloud infrastructure 500 includes a combination of physical and virtual processing resources that can be used to implement at least a portion of Figure 1 information processing system 100 therein. Cloud infrastructure 500 includes a plurality of virtual machines (VMs) and / or sets of containers 502-1, 502-2, ... 502-L that are implemented using virtualization infrastructure 504. Virtualization infrastructure 504 runs on physical infrastructure 505 and illustratively includes one or more hypervisors and / or operating system-level virtualization infrastructure. Operating system-level virtualization infrastructure illustratively includes kernel control groups of a Linux operating system or other types of operating systems.

[0076] Cloud infrastructure 500 also includes multiple sets of applications 510-1, 510-2, ... 510-L that run under the control of virtualization infrastructure 504 on corresponding ones of the VM / container sets 502-1, 502-2, ... 502-L. The VM / container sets 502 can include corresponding VMs, one or more sets of corresponding containers, or one or more sets of corresponding containers running within a VM.

[0077] In Figure 5 some implementations of the implementation of, the VM / container sets 502 include corresponding VMs implemented using virtualization infrastructure 504 that includes at least one hypervisor. A hypervisor platform can be used to implement a hypervisor within virtualization infrastructure 504, where the hypervisor platform has an associated virtual infrastructure management system. The underlying physical machine can include one or more distributed processing platforms that include one or more storage systems.

[0078] In Figure 5 other implementations of the implementation of, the VM / container sets 502 include corresponding containers implemented using virtualization infrastructure 504 that provides operating system-level virtualization functionality such as supporting Docker containers running on bare metal hosts or Docker containers running on VMs. Containers are illustratively implemented using corresponding kernel control groups of an operating system.

[0079] As is apparent from the above, one or more of the processing modules or other components of system 100 may each operate on a computer, server, storage device, or other processing platform element. A given such element may be considered an example of what is more generally referred to herein as a "processing device." Figure 5 The illustrated cloud infrastructure 500 may represent at least a portion of a processing platform. Another example of such a processing platform is Figure 6 the illustrated processing platform 600.

[0080] In this embodiment, processing platform 600 includes a portion of system 100 and includes a plurality of processing devices represented as 602-1, 602-2, 602-3,... 602-K that communicate with each other via network 604.

[0081] Network 604 may include any type of network, such as including a global computer network (such as the Internet), WAN, LAN, satellite network, telephone or cable network, cellular network, wireless network (such as a WiFi or WiMAX network), or various portions or combinations of these and other types of networks.

[0082] Processing device 602-1 in processing platform 600 includes a processor 610 coupled to a memory 612.

[0083] Processor 610 may include a microprocessor, microcontroller, application specific integrated circuit (ASIC), field programmable gate array (FPGA), central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), video processing unit (VPU), or other type of processing circuitry, as well as portions or combinations of such circuit elements.

[0084] Memory 612 may include random access memory (RAM), read only memory (ROM), flash memory, or other types of memory in any combination. Memory 612 and other memories disclosed herein should be considered illustrative examples of what is more generally referred to as a "processor-readable storage medium" that stores executable program code for one or more software programs.

[0085] An article of manufacture that includes such a processor-readable storage medium is considered an illustrative embodiment. A given such article of manufacture may include, for example, a storage array, a storage disk, or an integrated circuit that includes RAM, ROM, flash memory, or other electronic memory, or any one of a number of other types of computer program products. As used herein, the term "article of manufacture" should be understood to exclude transient propagated signals. Numerous other types of computer program products that include a processor-readable storage medium may be used.

[0086] The processing device 602-1 further includes a network interface circuit 614, which is used to interface the processing device with the network 604 and other system components, and may include a conventional transceiver.

[0087] The other processing devices 602 of the processing platform 600 are assumed to be configured in a manner similar to that shown for the processing device 602-1 in the figure.

[0088] Moreover, the specific processing platform 600 shown in the figure is presented only by way of example, and the system 100 may include additional or alternative processing platforms, and may include numerous different processing platforms in any combination, where each such platform includes one or more computers, servers, storage devices, or other processing devices.

[0089] For example, other processing platforms for implementing the illustrative embodiments may include converged infrastructure.

[0090] Therefore, it should be understood that in other embodiments, different arrangements of additional or alternative elements may be used. At least a subset of these elements may be implemented together on a common processing platform, or each such element may be implemented on a separate processing platform.

[0091] As previously indicated, the components of the information processing system as disclosed herein may be implemented at least in part in the form of one or more software programs stored in a memory and executed by a processor of a processing device. For example, at least some parts of the functionality for moving virtual volumes in the storage nodes of a storage cluster as disclosed herein are illustratively implemented in the form of software running on one or more processing devices, at least partially based on the determined likelihood of specified VM boot conditions.

[0092] It should be emphasized again that the above embodiments are presented for illustrative purposes only. Many variations and other alternative embodiments may be used. For example, the disclosed techniques may be applicable to many other types of information processing systems, storage systems, virtualization infrastructures, etc. Moreover, the specific configurations of the system and device elements illustratively shown in the figures and the associated processing operations may vary in other embodiments. In addition, the various assumptions made above during the description of the illustrative embodiments should also be considered exemplary, rather than requirements or limitations of the present disclosure. Numerous other alternative embodiments within the scope of the appended claims will be apparent to those skilled in the art.

Claims

1. A device, comprising: at least one processing device, the at least one processing device including a processor coupled to a memory; the at least one processing device is configured to perform the following steps: obtain information characterizing the historical boot times of a plurality of virtual machines, the plurality of virtual machines being associated with a plurality of virtual volumes hosted on a storage cluster including a plurality of storage nodes; determine, at least in part based on the obtained information characterizing the historical boot times of the plurality of virtual machines, whether any one of the plurality of storage nodes has at least a threshold likelihood of experiencing respective specified virtual machine boot conditions during a given time period, a first specified virtual machine boot condition for a first storage node among the plurality of storage nodes including: booting at least a first threshold number of the plurality of virtual machines simultaneously on the first storage node, a second specified virtual machine boot condition for a second storage node among the plurality of storage nodes including: booting at least a second threshold number of the plurality of virtual machines simultaneously on the second storage node, the first and second threshold numbers of the plurality of virtual machines respectively including at least two of the plurality of virtual machines, the first and second threshold numbers being at least in part based on the available resources of the first and second storage nodes, the first threshold number being different from the second threshold number, the length of the given time period being selected at least in part based on the expected time to complete booting each of the plurality of virtual machines; in response to determining that the first storage node has at least the threshold likelihood of experiencing the first specified virtual machine boot condition during the given time period, identify a subset of the plurality of virtual machines associated with a subset of the plurality of virtual volumes hosted on the first storage node; and move at least one of the subset of the plurality of virtual volumes associated with at least one of the subset of the plurality of virtual machines from the first storage node to another storage node among the plurality of storage nodes; wherein obtaining the information characterizing the historical boot times of the plurality of virtual machines includes: monitoring the creation times of one or more specified types of virtual volumes among the plurality of virtual volumes, the one or more specified types of virtual volumes among the plurality of virtual volumes including a first type of virtual volume associated with a specific file corresponding to a respective virtual machine among the plurality of virtual machines, the first type of virtual volume of a given virtual machine among the plurality of virtual machines being created when the given virtual machine is powered on.

2. The device according to claim 1, wherein the one or more specified types of virtual volumes among the plurality of virtual volumes include at least one type of virtual volume containing a copy of a virtual machine memory page not retained in the memory of the plurality of virtual machines.

3. The device according to claim 1, wherein the first type of virtual volume includes a swap virtual volume, and wherein the specific file includes a swap file.

4. The apparatus according to claim 1, wherein moving at least one of the subsets of the plurality of virtual volumes associated with at least one of the subsets of the plurality of virtual machines from the first storage node to the other storage node comprises: generating a probability density function of the boot time of each of the subsets of the plurality of virtual machines; using the probability density function of the boot time of the subsets of the plurality of virtual machines to determine the likelihood that each of the subsets of the plurality of virtual machines boots during the given time period; and selecting at least one of the subsets of the plurality of virtual machines based at least in part on the determined likelihood that each of the subsets of the plurality of virtual machines boots during the given time period.

5. The apparatus according to claim 4, wherein selecting at least one of the subsets of the plurality of virtual machines based at least in part on the determined likelihood that each of the subsets of the plurality of virtual machines boots during the given time period comprises selecting a given one of the subsets of the plurality of virtual machines having the highest boot probability during the given time period in the probability density function.

6. The apparatus according to claim 4, wherein moving at least one of the subsets of the plurality of virtual volumes associated with at least one of the subsets of the plurality of virtual machines from the first storage node to the other storage node further comprises: identifying a subset of the plurality of storage nodes as candidate destination storage nodes, each of the candidate destination storage nodes currently hosting fewer than a specified number of the plurality of virtual volumes; and selecting the other storage node from the candidate destination storage nodes.

7. The apparatus according to claim 6, wherein the specified number of the plurality of virtual volumes is determined based at least in part on the average number of virtual volumes hosted on each of the plurality of storage nodes.

8. The apparatus according to claim 6, wherein selecting the other storage node from the candidate destination storage nodes is based at least in part on determining the likelihood that each of the candidate destination storage nodes experiences the specified virtual machine boot condition during the given time period.

9. The apparatus according to claim 8, wherein selecting the other storage node from the candidate destination storage nodes comprises selecting a candidate storage node having the lowest likelihood of experiencing the specified virtual machine boot condition during the given time period.

10. The apparatus according to claim 1, wherein moving at least one of the subsets of the plurality of virtual volumes associated with at least one of the subsets of the plurality of virtual machines from the first storage node to the other storage node comprises moving all of the virtual volumes associated with a given one of the subsets of the plurality of virtual machines from the first storage node to the other storage node.

11. A computer program product comprising a non-transitory processor-readable storage medium storing program code of one or more software programs, wherein the program code, when executed by at least one processing device, causes the at least one processing device to perform the following steps: Obtaining information characterizing historical boot times of a plurality of virtual machines, the plurality of virtual machines being associated with a plurality of virtual volumes hosted on a storage cluster including a plurality of storage nodes; Determining, at least in part based on the obtained information characterizing the historical boot times of the plurality of virtual machines, whether any one of the plurality of storage nodes has at least a threshold likelihood of experiencing a specified virtual machine boot condition during a given time period, a first specified virtual machine boot condition for a first storage node among the plurality of storage nodes including: Booting at least a first threshold number of the plurality of virtual machines simultaneously on the first storage node, a second specified virtual machine boot condition for a second storage node among the plurality of storage nodes including: booting at least a second threshold number of the plurality of virtual machines simultaneously on the second storage node, the first and second threshold numbers of the plurality of virtual machines respectively including at least two of the plurality of virtual machines, the first and second threshold numbers being at least in part based on available resources of the first and second storage nodes, the first threshold number being different from the second threshold number, and the length of the given time period being selected at least in part based on an expected time to complete booting each of the plurality of virtual machines; In response to determining that the first storage node has at least the threshold likelihood of experiencing the first specified virtual machine boot condition during the given time period, identifying a subset of the plurality of virtual machines associated with a subset of the plurality of virtual volumes hosted on the first storage node; and Moving at least one of the subset of the plurality of virtual volumes associated with at least one of the subset of the plurality of virtual machines from the first storage node to another storage node among the plurality of storage nodes; wherein obtaining the information characterizing the historical boot times of the plurality of virtual machines includes: monitoring creation times of one or more specified types of virtual volumes among the plurality of virtual volumes, the one or more specified types of virtual volumes among the plurality of virtual volumes including a first type of virtual volume associated with a specific file corresponding to a respective virtual machine among the plurality of virtual machines, the first type of virtual volume of a given virtual machine among the plurality of virtual machines being created when the given virtual machine is powered on.

12. The computer program product according to claim 11, wherein the first type of virtual volume includes a swap virtual volume, and wherein the specific file includes a swap file.

13. A method comprising: Obtaining information characterizing historical boot times of a plurality of virtual machines, the plurality of virtual machines being associated with a plurality of virtual volumes hosted on a storage cluster including a plurality of storage nodes; Determine whether any one of the multiple storage nodes has at least a threshold likelihood of experiencing a specified virtual machine boot condition during a given time period, at least in part based on the obtained information characterizing the historical boot times of the multiple virtual machines. The first specified virtual machine boot condition for a first storage node among the multiple storage nodes includes: booting at least a first threshold number of the multiple virtual machines simultaneously on the first storage node. The second specified virtual machine boot condition for a second storage node among the multiple storage nodes includes: booting at least a second threshold number of the multiple virtual machines simultaneously on the second storage node. The first and second threshold numbers of the multiple virtual machines respectively include at least two of the multiple virtual machines. The first and second threshold numbers are at least in part based on the available resources of the first and second storage nodes. The first threshold number is different from the second threshold number. The length of the given time period is selected at least in part based on the expected time to complete booting each of the multiple virtual machines; In response to determining that the first storage node has at least the threshold likelihood of experiencing the first specified virtual machine boot condition during the given time period, identify a subset of the multiple virtual machines associated with a subset of the multiple virtual volumes hosted on the first storage node; and Move at least one of the subset of the multiple virtual volumes associated with at least one of the subset of the multiple virtual machines from the first storage node to another storage node among the multiple storage nodes; Wherein, obtaining the information characterizing the historical boot times of the multiple virtual machines includes: monitoring the creation times of one or more specified types of virtual volumes among the multiple virtual volumes. The one or more specified types of virtual volumes among the multiple virtual volumes include first-type virtual volumes associated with specific files corresponding to respective virtual machines among the multiple virtual machines. The first-type virtual volumes of a given virtual machine among the multiple virtual machines are created when the given virtual machine is powered on; Wherein the method is executed by at least one processing device, and the processing device includes a processor coupled to a memory.

14. The method according to claim 13, wherein the first-type virtual volume includes a swap virtual volume, and wherein the specific file includes a swap file.

Citation Information

Patent Citations

  • Synchronously replicating datasets and other managed objects to cloud-based storage systems

    CN110392876A

  • Providing an instance availability estimate

    US9256452B1