Pipeline rolling update

By adopting a cross-fault domain migration strategy in the HCI cluster, allowing multiple hosts to be upgraded simultaneously within each fault domain, the problem of time-consuming and downtime risks of online upgrade of the HCI cluster is solved, and a more efficient upgrade process is achieved.

CN113805907BActive Publication Date: 2025-07-01DELL PROD LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010541459.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-15
Publication Date
2025-07-01
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

In hyperconverged infrastructure (HCI) clusters, the online upgrade process can be time-consuming, affecting system load capacity and increasing the risk of downtime, mainly due to the tight coupling of virtualized computing and software definition storage on a single node, resulting in only one node or up to two nodes at a time.

Method used

By adopting a cross-domain migration strategy in an HCI cluster, multiple hosts are divided into independent update units using the fault domain concept, allowing multiple hosts to be upgraded simultaneously within each fault domain. Specific steps include receiving fault domain information, bringing the host into protection mode and maintenance mode, and performing an upgrade in these modes.

Benefits of technology

This method significantly reduces the time complexity of the entire cluster upgrade, from O (number of hosts) to O (number of fault domains), thereby shortening the upgrade time, reducing the risk of system downtime, and improving the efficiency of the upgrade process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113805907B_ABST
    Figure CN113805907B_ABST
Patent Text Reader

Abstract

The present invention relates to rolling updates in a pipeline. The present invention provides an information processing system, comprising: at least one processor; and a non-transitory memory coupled to the at least one processor. The information processing system is configured to upgrade a plurality of hosts in an information processing system cluster by: receiving information about a fault domain of the cluster such that each of the plurality of hosts is a member of exactly one fault domain; and for each fault domain: causing all hosts of the fault domain to enter a protection mode in which no new virtual machines can be created or no new virtual machines can be accepted for migration; causing the hosts to enter a maintenance mode in which any existing virtual machines are migrated away from the hosts; and causing the hosts to perform an upgrade, wherein the plurality of hosts are configured to perform the upgrade simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to information processing systems, and more particularly, to accelerating update events in a clustered environment such as a hyper-converged infrastructure (HCI) cluster. Background Art

[0002] As the value and use of information continue to increase, both individuals and enterprises are seeking alternative ways to process and store information. One option available to users is an information processing system. An information processing system generally processes, compiles, stores, and / or communicates information or data for commercial, personal, or other purposes, thereby allowing users to leverage the value of the information. Since technology and information processing needs and requirements vary among different users or applications, information processing systems can also vary in terms of what information is processed, how it is processed, how much information is processed, stored, or communicated, and how quickly and efficiently the information can be processed, stored, or communicated. Variations in information processing systems allow the information processing system to be general-purpose or configured for a particular user or particular use (such as financial transaction processing, airline reservation booking, enterprise data storage, or global communication). Additionally, an information processing system can include a variety of hardware and software components that can be configured to process, store, and communicate information, and can include one or more computer systems, data storage systems, and networking systems.

[0003] Hyper-converged infrastructure (HCI) is an IT framework that combines storage, computing, and networking into a single system in an attempt to reduce data center complexity and improve scalability. Hyper-converged platforms can include a hypervisor for virtualized computing, software-defined storage, and virtualized networking, and they typically run on standard off-the-shelf servers.

[0004] Using rolling updates allows for an online upgrade of an HCI cluster without interrupting services. During such an upgrade, each node of the cluster can typically be upgraded sequentially (e.g., using upgrade components such as firmware, drivers, application software, etc.) until the entire cluster reaches the same system version.

[0005] Since virtualized computing and software-defined storage are tightly coupled on a single node, typically such an upgrade can only be performed on one node (or at most two nodes) at a time to ensure that the number of hosts being maintained does not exceed the allowable failure tolerance level (FTT, typically FTT = 1). Therefore, an online upgrade of a cluster can become a very time-consuming task. It can also affect the system load capacity and increase the risk of downtime.

[0006] Accordingly, the present disclosure provides techniques for accelerating update speeds in an HCI cluster by leveraging several cross-fault-domain migration strategies. Using this approach, the time complexity of updating the entire cluster in big O notation can be reduced from O(number of hosts) to O(number of fault domains).

[0007] It should be noted that the discussion of techniques in the background section of the present disclosure does not constitute an admission of the state of the prior art. No such admission is made herein unless clearly and expressly acknowledged. SUMMARY OF THE INVENTION

[0008] In accordance with the teachings of the present disclosure, disadvantages and problems associated with upgrading a cluster of information processing systems may be reduced or eliminated.

[0009] According to an embodiment of the present disclosure, an information processing system may include: at least one processor; and a non-transitory memory coupled to the at least one processor. The information processing system may be configured to upgrade a plurality of hosts of an information processing system cluster by: receiving information about a fault domain of the cluster such that each of the plurality of hosts is a member of exactly one fault domain; and for each fault domain: putting all hosts of the fault domain into a protection mode in which no new virtual machines can be created or accepted for migration; putting a host into a maintenance mode in which any existing virtual machines are migrated away from the host; and causing the host to perform the upgrade, wherein the plurality of hosts are configured to perform the upgrade simultaneously.

[0010] According to these and other embodiments of the present disclosure, a method may include: a management information processing system receiving information about a fault domain of an information processing system cluster including a plurality of hosts such that each of the plurality of hosts is a member of exactly one fault domain; and for each fault domain, the management information processing system: putting all hosts of the fault domain into a protection mode in which no new virtual machines can be created or accepted for migration; putting a host into a maintenance mode in which any existing virtual machines are migrated away from the host; and causing the host to perform an upgrade, wherein the plurality of hosts are configured to perform the upgrade simultaneously.

[0011] According to these and other embodiments of the present disclosure, an article may include a non-transitory computer-readable medium having computer-executable code thereon that is executable by a processor of an information processing system to manage an upgrade of a cluster of information processing systems including a plurality of host systems by: receiving information about fault domains of the cluster such that each of the plurality of hosts is a member of exactly one fault domain; and for each fault domain: putting all hosts of the fault domain into a protection mode in which no new virtual machines can be created or no new virtual machines can be accepted for migration; putting a host into a maintenance mode in which any existing virtual machines are migrated away from the host; and causing the host to perform the upgrade, wherein the plurality of hosts are configured to perform the upgrade simultaneously.

[0012] Based on the accompanying drawings, the description, and the claims included herein, the technical advantages of the present disclosure may be apparent to those skilled in the art. The objectives and advantages of the embodiments will be realized and attained at least by the elements, features, and combinations particularly pointed out in the claims.

[0013] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims set forth in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] A more complete understanding of the present embodiments and their advantages may be obtained by reference to the following description taken in conjunction with the accompanying drawings, in which like reference numerals indicate like features and in which:

[0015] Figure 1 A block diagram of an exemplary information processing system according to an embodiment of the present disclosure is shown;

[0016] Figure 2 A block diagram of a fault domain according to an embodiment of the present disclosure is shown;

[0017] Figure 3 An exemplary method according to an embodiment of the present disclosure is shown; and

[0018] Figure 4 An exemplary method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0019] By reference Figures 1 to 3 The preferred embodiments and their advantages are best understood, where like reference numerals are used to indicate identical and corresponding parts.

[0020] For purposes of this disclosure, the term "information processing system" may include any tool or collection of tools capable of operating to compute, classify, process, transmit, receive, retrieve, originate, switch, store, display, manifest, detect, record, reproduce, handle, or utilize any form of information, intelligence, or data for commercial, scientific, control, entertainment, or other purposes. For example, an information processing system may be a personal computer, a personal digital assistant (PDA), a consumer electronic device, a network storage device, or any other suitable device, and may vary in size, shape, performance, functionality, and price. An information processing system may include a memory, one or more processing resources such as a central processing unit ("CPU") or hardware or software control logic. Additional components of an information processing system may include one or more storage devices, one or more communication ports for communicating with external devices, and various input / output ("I / O") devices (such as a keyboard, a mouse, and a video display). An information processing system may also include one or more buses capable of operating to transfer communications between the various hardware components.

[0021] For purposes of this disclosure, when two or more elements are referred to as being "coupled" to each other, such term indicates that such two or more elements communicate electronically or mechanically, as applicable, whether they are directly connected or indirectly connected, with or without intermediate elements.

[0022] When two or more elements are referred to as being "couplable" to each other, such term indicates that they are capable of being coupled together.

[0023] For purposes of this disclosure, the term "computer-readable medium" (e.g., a transient or non-transient computer-readable medium) may include any tool or collection of tools capable of retaining data and / or instructions for a period of time. A computer-readable medium may include, but is not limited to: storage media such as direct access storage devices (e.g., a hard disk drive or a floppy disk), sequential access storage devices (e.g., a magnetic tape drive), optical disks, CD-ROMs, DVDs, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and / or flash memory; communication media such as wires, optical fibers, microwaves, radio waves, and other electromagnetic and / or optical carriers; and / or any combination of the foregoing.

[0024] For purposes of this disclosure, the term "information processing resource" may broadly refer to any subsystem, device, or equipment of an information processing system, including but not limited to a processor, a service processor, a basic input / output system, a bus, a memory, an I / O device and / or interface, a storage resource, a network interface, a motherboard, and / or any other component and / or element of the information processing system.

[0025] For purposes of this disclosure, the term "management controller" can broadly refer to an information processing system that provides management functions (typically out-of-band management functions) to one or more other information processing systems. In some embodiments, the management controller can be a service processor, a baseboard management controller (BMC), a chassis management controller (CMC), or a remote access controller (e.g., a Dell Remote Access Controller (DRAC) or an Integrated Dell Remote Access Controller (iDRAC)) (or can be an integral part thereof).

[0026] Figure 1 FIG. shows a block diagram of an exemplary information processing system 102 in accordance with an embodiment of the present disclosure. In some embodiments, the information processing system 102 can include a server chassis configured to house a plurality of servers or "blades". In other embodiments, the information processing system 102 can include a personal computer (e.g., a desktop computer, a laptop computer, a mobile computer, and / or a notebook computer). In still other embodiments, the information processing system 102 can include a storage cabinet configured to house a plurality of physical disk drives and / or other computer-readable media for storing data (which can generally be referred to as "physical storage resources"). As Figure 1 shown, the information processing system 102 can include a processor 103, a memory 104 communicatively coupled to the processor 103, a BIOS 105 (e.g., a UEFI BIOS) communicatively coupled to the processor 103, a network interface 108 communicatively coupled to the processor 103, and a management controller 112 communicatively coupled to the processor 103.

[0027] In operation, the processor 103, the memory 104, the BIOS 105, and the network interface 108 can form at least a portion of the host system 98 of the information processing system 102. In addition to the elements explicitly shown and described, the information processing system 102 can also include one or more other information processing resources.

[0028] The processor 103 can include any system, device, or apparatus configured to interpret and / or execute program instructions and / or process data, and can include, but is not limited to: a microprocessor, a microcontroller, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or any other digital or analog circuit configured to interpret and / or execute program instructions and / or process data. In some embodiments, the processor 103 can interpret and / or execute program instructions and / or process data stored in the memory 104 and / or another component of the information processing system 102.

[0029] The memory 104 may be communicatively coupled to the processor 103 and may include any system, device, or apparatus (e.g., a computer-readable medium) configured to retain program instructions and / or data for a period of time. The memory 104 may include RAM, EEPROM, PCMCIA cards, flash memory, magnetic storage devices, magneto-optical storage devices, or any suitable selection and / or array of volatile or non-volatile memory that retains data after power to the information processing system 102 is cut off.

[0030] As Figure 1 shown, the operating system 106 may be stored on the memory 104. The operating system 106 may include any program (or collection of programs) of executable instructions configured to manage and / or control the allocation and use of hardware resources (such as memory, processor time, disk space, and input and output devices) and to provide an interface between such hardware resources and the application programs hosted by the operating system 106. Additionally, the operating system 106 may include all or a portion of a network stack for network communication via a network interface (e.g., the network interface 108 for communicating over a data network). Although the operating system 106 is shown in Figure 1 as being stored in the memory 104, in some embodiments, the operating system 106 may be stored on a storage medium accessible to the processor 103, and the active portions of the operating system 106 may be transferred from such storage medium to the memory 104 for execution by the processor 103.

[0031] The network interface 108 may include one or more suitable systems, devices, or apparatuses capable of operating to serve as an interface between the information processing system 102 and one or more other information processing systems via an in-band network. The network interface 108 may enable the information processing system 102 to communicate using any suitable transport protocol and / or standard. In these and other embodiments, the network interface 108 may include a network interface card or “NIC”. In these and other embodiments, the network interface 108 may be capable of being enabled as a local area network (LAN)-on-motherboard (LOM) card on the motherboard.

[0032] The management controller 112 may be configured to provide management functions for managing the information processing system 102. Such management may also be performed by the management controller 112 even if the information processing system 102 and / or the host system 98 are powered off or powered to a standby state. The management controller 112 may include a processor 113, memory, and a network interface 118 separate and physically isolated from the network interface 108.

[0033] As Figure 1As shown, the processor 113 of the management controller 112 is communicatively coupled to the processor 103. Such coupling may be via a Universal Serial Bus (USB), a System Management Bus (SMBus), and / or one or more other communication channels.

[0034] The network interface 118 may be coupled to a management network, which may be separate from and physically isolated from the data network, as shown. The network interface 118 of the management controller 112 may include any suitable system, device, or apparatus capable of operating to serve as an interface between the management controller 112 and one or more other information processing systems via an out-of-band management network. The network interface 118 may enable the management controller 112 to communicate using any suitable transport protocol and / or standard. In these and other embodiments, the network interface 118 may include a network interface card or "NIC". The network interface 118 may be the same type of device as the network interface 108, or in other embodiments, it may be a different type of device.

[0035] As discussed above, it will be desirable to reduce the amount of time required to perform an online rolling upgrade of a cluster of information processing systems. Accordingly, "fault domains" may be determined within the cluster to allow for more efficient pipelining of update procedures. A fault domain is a group of hosts that share a single point of failure. For example, if the failure of a single power cord or a single network connection could cause a group of hosts to fail, they may be considered to be in the same fault domain. In this way, the cluster may be protected from a certain level of failures, such as power or connection outages.

[0036] Based on this principle, multiple hosts may be updated simultaneously within a fault domain. A typical HCI cluster may generally include at least three fault domains to support FTT = 1, each fault domain consisting of one or more hosts. The fault domain definition may identify physical hardware structures that may represent potential failure zones, such as separate computer rack cabinets.

[0037] As an example, Figure 2 A block diagram of nine hosts 202 is shown, which have been partitioned into three fault domains.

[0038] Currently, a rolling update procedure may require a host to enter "maintenance mode" before being updated. Any virtual machines (VMs) running on the host entering maintenance mode are either migrated to another host or shut down. For the purposes of this disclosure, a "protection mode" may also be implemented to allow for the upgrade of multiple hosts in a pipelined manner. The protection mode may allow for the migration of virtual machines in a more efficient manner.

[0039] Specifically, in maintenance mode, both the hypervisor service and the storage service on the host are stopped. In contrast, in protection mode, both the hypervisor service and the storage service can remain running, but creating new virtual machines on the host or migrating new virtual machines to the host is not allowed. Thus, protection mode can be regarded as a "lighter weight" version of the maintenance mode restrictions. Table 1 below provides details of these two modes.

[0040] Maintenance mode Protection mode Hypervisor service Stop Run Storage service Stop Run Does it trigger VM migration? Yes No Does it trigger data migration? Conditional No VM migration source Not applicable Allow VM migration destination Reject Reject

[0041] Table 1.

[0042] When the update process starts, all hosts in the first failure domain can be placed in protection mode. Then, the host can enter maintenance mode to migrate all virtual machines running on it to one or more other failure domains. When a particular host enters maintenance mode, the update task for that host and the entry into maintenance mode for the next host can be performed simultaneously. After the update process is completed on the host, it can exit maintenance mode and protection mode. This operation can be repeated on the remaining hosts in the failure domain, and then the entire process can be repeated on the remaining failure domains to update the entire cluster to a higher version.

[0043] Now turning to Figure 3 , a flowchart of an exemplary method 300 for updating a cluster of an information processing system according to some embodiments is shown. As Figure 2 shown, hosts 1, 4, and 7 are in failure domain 1. Hosts 2, 5, and 8 are in failure domain 2. Hosts 3, 6, and 9 are in failure domain 3.

[0044] At step 302, the update process for failure domain 1 starts. All three hosts in failure domain 1 first enter protection mode. The hypervisor service and the storage service can continue to run, but creating new virtual machines on these three hosts or migrating new virtual machines to these three hosts is not allowed.

[0045] Next, host 1 enters maintenance mode. Both the hypervisor service and the storage service on host 1 are stopped, and its virtual machines are migrated away. Specifically, they may be migrated to hosts within some other failure domain because the other hosts in failure domain 1 are already in protection mode and thus cannot accept such migrations.

[0046] Then, host 1 starts its upgrade task, and host 4 enters maintenance mode. Similarly, the virtual machines are migrated from host 4 to other failure domains.

[0047] Then, host 4 starts its upgrade task, and host 7 enters maintenance mode. The virtual machines are migrated from host 7 to some other failure domain.

[0048] When each host has completed its upgrade task, it exits the protection mode and the maintenance mode respectively. At the end of step 302, hosts 1, 4, and 7 are upgraded and available to receive the migrated virtual machines.

[0049] At step 304, a similar process occurs within fault domain 2. At step 306, a similar process occurs within fault domain 3.

[0050] As can be seen from Figure 3 a large number of upgrade tasks can be executed in parallel instead of being forced to execute serially.

[0051] Those of ordinary skill in the art who benefit from this disclosure will understand that Figure 3 the preferred starting point of the method shown and the order of the steps that make up the method may depend on the implementation chosen. In these and other embodiments, the method can be implemented as hardware, firmware, software, applications, functions, libraries, or other instructions. Additionally, although Figure 3 a specific number of steps to be performed in the disclosed method are disclosed, the method can be performed using more or fewer steps than those depicted. The method can be implemented using any of the various components disclosed herein (such as Figure 1 components) and / or any other system capable of operating to implement the method.

[0052] As Figure 3 depicted, during the upgrade process, each host may go through four stages. In the first stage, it enters the protection mode. In the second stage, it enters the maintenance mode. In the third stage, it runs the upgrade task. And in the fourth stage, it exits the maintenance mode and the protection mode.

[0053] In some embodiments, entering the maintenance mode can be considered the critical path of the upgrade process because only one node in the entire cluster can execute "enter the maintenance mode" simultaneously. If there are not enough resources to migrate all VMs to other fault domains, the upgrade pipeline may stall until one of the previous hosts completes stage 4 (exiting the maintenance mode and the protection mode).

[0054] Figure 4 depicts a part of an example similar to Figure 3 but in this example, initially there are not enough resources available to migrate the VMs from host 7. Therefore, Figure 4 shows an example of how the entire workflow may stall in this case. Specifically, Figure 4 shows a barrier 452, which is used (as discussed below) to ensure that all hosts are in the protection mode. Additionally, due to the stall, host 7 has to wait until time 454 to enter the maintenance mode.

[0055] The implementation of some embodiments may rely on three synchronization primitives to control the upgrade workflow: pm_barrier is a barrier object used to ensure that all hosts are in the protected mode; mm_mutex is a mutex lock that facilitates exclusive operations for "entering the maintenance mode"; and exit_cv is a condition variable that can be used to wait for the exit operation on the previous host. An example listing of the pseudocode for each node is shown below:

[0056] do_stage_1();

[0057] pm_barrier.await();

[0058] while (not enough resources for migration) {

[0059] exit_cv.wait();

[0060] }

[0061] mm_mutex.lock();

[0062] do_stage_2();

[0063] mm_mutex.unlock();

[0064] do_stage_3();

[0065] do_stage_4();

[0066] exit_cv.notify();

[0067] ---------------------------------------------

[0068] Therefore, using the "protected mode" described in the present disclosure allows for a more efficient way to migrate virtual machines. This can provide a mechanism to update hosts in parallel with the multi-phase pipeline model, as Figure 3 shown. In addition, a mechanism is provided for scheduling rolling update workflows in multiple phases by means of a failure domain (e.g., a storage failure domain).

[0069] These features can provide significant time savings in cluster upgrades, especially for large-scale clusters. The maintenance window can be correspondingly shortened, and the risk of downtime is reduced. In addition, unnecessary virtual machine migration overhead during upgrades can be reduced.

[0070] The scheduling and management of the upgrade program can be managed via a management system. In some embodiments, the management system can be an element of a cluster, while in other embodiments, it can be a separate system. In some embodiments, the management system can be a management controller, such as the management controller 112 discussed above. In other embodiments, the management system can be implemented on the host system rather than on the management controller.

[0071] Although various possible advantages of the embodiments of the present disclosure have been described, those of ordinary skill in the art who benefit from the present disclosure should understand that in any particular embodiment, not all such advantages may be applicable. In any particular embodiment, some of the listed advantages may be applicable, all may be applicable, or even none may be applicable.

[0072] The present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that would be understood by those of ordinary skill in the art. Similarly, where appropriate, the appended claims encompass all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that would be understood by those of ordinary skill in the art. Additionally, the recitation in the appended claims of a device or system or a component of a device or system that is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses the device, system, or component, whether or not the particular function is activated, turned on, or unlocked, so long as the device, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative.

[0073] Furthermore, the recitation in the appended claims of a structure being “configured to” or “operable to” perform one or more tasks is not intended to invoke 35 U.S.C. § 112(f) with respect to the claim element. Accordingly, no claim in the present application being filed is intended to be construed as having a means-plus-function element. If the applicant wishes to invoke § 112(f) during examination, the applicant will use the “means for [performing the function]” structure to recite the claim element.

[0074] All of the examples and conditional language recited herein are intended for pedagogical purposes to aid the reader in understanding the invention and the concepts contributed by the inventor to further the art, and are to be construed as not being limited to such specifically recited examples and conditions. Although the embodiments of the invention have been described in detail, it should be understood that various changes, substitutions, and alterations can be made without departing from the spirit and scope of the present disclosure.

Claims

1. An information processing system, comprising: At least one processor; And A non-transitory memory coupled to the at least one processor; Wherein the information processing system is configured to upgrade a plurality of hosts of an information processing system cluster by: Receiving information about a fault domain of the cluster such that each of the plurality of hosts is a member of exactly one fault domain; And For each fault domain: Put all hosts of the fault domain into a protection mode, in which new virtual machines cannot be created or new virtual machines cannot be accepted for migration; Put the host into a maintenance mode, in which any existing virtual machines are migrated away from the host; And Cause the host to perform the upgrade, wherein the plurality of hosts within the fault domain are configured to perform the upgrade simultaneously.

2. The information processing system according to claim 1, wherein the information processing system is further configured to cause each host to exit the protection mode and the maintenance mode after completing the upgrade.

3. The information processing system according to claim 1, wherein the cluster is a hyper-converged infrastructure (HCI) cluster.

4. The information processing system according to claim 1, wherein each of the fault domains includes an equal number of hosts.

5. The information processing system according to claim 1, wherein the information processing system is a host of the cluster.

6. The information processing system according to claim 1, wherein the information processing system is a management controller.

7. The information processing system according to claim 1, wherein at a specific time, a first host is configured to start the upgrade and a second host is configured to enter the maintenance mode.

8. A method, comprising: A management information processing system receives information about a fault domain of an information processing system cluster including a plurality of hosts such that each of the plurality of hosts is a member of exactly one fault domain; And For each fault domain, the management information processing system: Put all hosts of the fault domain into a protection mode, in which new virtual machines cannot be created or new virtual machines cannot be accepted for migration; Put the host into a maintenance mode, in which any existing virtual machines are migrated away from the host; And Cause the host to perform an upgrade, wherein the plurality of hosts within the fault domain are configured to perform the upgrade simultaneously.

9. The method according to claim 8, the method further comprising causing each host to exit the protection mode and the maintenance mode after completing the upgrade.

10. The method according to claim 8, wherein each fault domain involves a single point of failure.

11. The method according to claim 8, wherein each of the fault domains includes an equal number of hosts.

12. The method according to claim 8, wherein the management information processing system is a host of the cluster.

13. The method according to claim 8, wherein at a specific time, a first host starts the upgrade and a second host enters the maintenance mode.

14. A program product includes a non-transitory computer-readable medium having computer-executable code thereon that is executable by a processor of an information processing system to manage an upgrade of a cluster of information processing systems including a plurality of host systems by: Receiving information about the fault domains of the cluster such that each of the plurality of hosts is a member of exactly one fault domain; and for each failure domain: put all hosts in the failure domain into a protection mode in which no new virtual machines can be created or no new virtual machines can be accepted for migration; put the hosts into a maintenance mode in which any existing virtual machines are migrated away from the hosts; and cause the hosts to perform the upgrade, wherein a plurality of hosts within the failure domain are configured to perform the upgrade simultaneously.

15. The program product of claim 14, wherein the code is further executable to: cause each host to exit the protection mode and the maintenance mode after completion of the upgrade.

16. The program product of claim 14, wherein the cluster is a hyper-converged infrastructure (HCI) cluster.

17. The program product of claim 14, wherein the failure domains each include an equal number of hosts.

18. The program product of claim 14, wherein the information processing system is a host of the cluster.

19. The program product of claim 14, wherein the information processing system is a management controller.

20. The program product of claim 14, wherein at a particular time, a first host is configured to start the upgrade and a second host is configured to enter the maintenance mode.

Citation Information

Patent Citations

  • Virtual machine (vm) migration between processor architectures

    CN101382906A

  • Controlled automatic healing of data-center services

    CN102385541A