High availability multi-single tenant service
Through dynamic adjustment and pooling auxiliary virtual machines, the problem of low resource usage and allocation efficiency in multi-tenant service environments is solved, efficient resource utilization and cost reduction are achieved, and high service availability is ensured.
Patent Information
- Application Number
- CN202510105550.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2018-03-01
- Publication Date
- 2025-06-06
AI Technical Summary
In a multi-tenant service environment, it is difficult for prior art to use and allocate resources efficiently, especially to reduce resource waste and costs while ensuring failover capabilities.
By dynamically determining secondary VM capacity requirements, pool secondary VM instances to ensure failover capability. The data processing hardware dynamically adjusts the number of secondary VM instances based on the number of primary VM instances and unavailability rates, and gradually grows the resource level of secondary VM instances during failover until the target resource level is reached.
Significantly reduces the number of secondary VM instances allocated to each primary VM instance, frees up compute resources, reduces costs, and ensures high availability of service instances.
Smart Images

Figure CN120104310A_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 201880090608.4, filed on March 1, 2018. Technical Field
[0003] The present disclosure relates to multi-tenant services with high availability. Background Art
[0004] A multi-single-tenant (MST) service executes software / service instances on virtual machines. In a single-tenant, each instance executes on a separate virtual machine. When a virtual machine fails, or is unavailable for update or maintenance for a period of time, the service executing on that virtual machine may be transferred to and executed on a secondary virtual machine. When capacity is lacking in a computing environment, delays may occur in generating secondary virtual machines. As a result, it is known to assign one secondary virtual machine to each primary virtual machine so that the MST service can responsively fail over to the secondary virtual machine without delaying the execution of software instances associated with a primary virtual machine that becomes unavailable due to failure, maintenance / update, or other reasons. However, MST services that use a large number of virtual machines utilize a large amount of resources to allocate and maintain the availability of the corresponding virtual machines, which are typically idle most of the time. Summary of the invention
[0005] One aspect of the present disclosure provides a method for maintaining the availability of service instances on a distributed system. The method includes executing a pool of primary virtual machine (VM) instances through data processing hardware of the distributed system, each primary VM instance executing a corresponding single service instance and including an unavailability rate. The method also includes determining, by the data processing hardware, the number of secondary VM instances required to maintain the availability of a single service instance when one or more primary VM instances are unavailable based on the number of primary VM instances in the pool of primary VM instances and the unavailability rate. The method also includes instantiating, by the data processing hardware, a pool of secondary VM instances based on the number of secondary VM instances required to maintain the availability of a single service instance.
[0006] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the method further includes identifying, by data processing hardware, the unavailability of one of the primary VM instances in the pool of primary VM instances, and causing, by the data processing hardware, the unavailable primary VM instance to fail over to one of the secondary VM instances in the secondary VM instance pool to begin executing a single service instance associated with the unavailable primary VM instance. In these implementations, the method may also include determining, by the data processing hardware, that the secondary VM instance includes a corresponding resource level that is less than the target resource level associated with the corresponding single service instance; and increasing, by the data processing hardware, the corresponding resource level of the secondary VM instance during the execution of the single service instance by the secondary VM instance until the target resource level associated with the single service instance is met. In some examples, the number of secondary VM instances in the pool of secondary VM instances is less than the number of primary VM instances in the pool of primary VM instances. Optionally, each primary VM instance may execute multiple primary containers, and each primary container executes a corresponding single service instance in a secure execution environment isolated from other primary containers.
[0007] In some examples, the method further includes: when the number of primary VM instances executing in the pool changes, updating, by the data processing hardware, the number of secondary VM instances required to maintain the availability of the single service instance. The unavailability of the primary instance can be based on at least one of a failure of the primary VM instance, a delay in recreating the primary VM instance, or a scheduled maintenance period of the primary VM instance. In addition, the unavailability rate can include at least one of an unavailability frequency or an unavailability period.
[0008] In some embodiments, the method further includes determining, by data processing hardware, an unavailability rate for each primary VM instance in the pool of primary VM instances based on a mean time to failure (MTTF) and an expected length of time to recreate the corresponding VM instance. Instantiating the pool of secondary VM instances may include determining a corresponding VM type for each primary VM instance in the pool of primary VM instances, and for each different VM type in the pool of primary VM instances, instantiating at least one secondary VM instance of the same VM type. In some examples, the corresponding VM type for each primary VM instance indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, or a local storage requirement for the VM instance.
[0009] In some examples, the method also includes receiving a planned failover message at the data processing hardware, the planned failover message indicating the number of primary VM instances in the pool of primary VM instances that will be unavailable during the planned maintenance period. In these examples, instantiating the pool of secondary VM instances is further based on the number of primary VM instances that will be unavailable during the planned maintenance period. For example, instantiating the pool of secondary VM instances can include instantiating a number of secondary VM instances equal to the larger of the following: the number of secondary VM instances required to maintain the availability of a single service instance; or the number of primary VM instances that will be unavailable during the planned maintenance period. The pool of secondary VM instances can be instantiated for use by a single customer of the distributed system. Optionally, the pool of secondary VM instances can be instantiated for use by multiple customers of the distributed system.
[0010] The above aspects can be applied to the above-mentioned multi-tenant services. In this case, each single service instance can be executed on a separate virtual machine instance. Similarly, when each main service instance is running in a container, the above aspects can be applied so that the main service instance fails over to a secondary container instead of the entire virtual machine instance. Multiple containers can run in a single VM. A method based on this aspect can provide a method for maintaining the availability of service instances on a distributed system, the method comprising: executing a primary virtual machine (VM) instance by data processing hardware of a distributed system, the primary VM instance executing a pool of primary containers. Each primary container executes a corresponding single service instance and includes an unavailability rate. The method also includes: determining by the data processing hardware the number of primary containers in the pool of primary containers and the corresponding unavailability rate for maintaining the availability of a single service instance in a secondary VM instance when one or more primary containers are unavailable; and instantiating by the data processing hardware a pool of secondary containers on the secondary VM instance based on the number of required secondary containers.
[0011] In addition, the above aspects may be applied to bare metal machines (e.g., bare metal servers) rather than VM instances. For example, single-tenant, bare metal servers may provide different options, such as data security and privacy controls, for some organizations subject to regulatory measures. A method based on this aspect may provide a method for maintaining service instances on a distributed system, the method comprising executing a pool of master bare metal machines, each master bare metal machine executing a corresponding single service instance and including an unavailability rate. The method also includes determining, by data processing hardware, the number of secondary bare metal machines required to maintain the availability of a single service instance when one or more master bare metal machines are unavailable based on the number of master bare metal machines in the master bare metal machine pool and the unavailability rate. The method also includes instantiating, by data processing hardware, a secondary bare metal machine pool based on the number of secondary bare metal machines required to maintain the availability of a single service instance.
[0012] Another aspect of the present disclosure provides a system for maintaining the availability of service instances on a distributed system. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including executing a pool of primary virtual machine (VM) instances. Each primary VM instance executes a corresponding single service instance and includes an unavailability rate. The operations also include determining the number of secondary VM instances required to maintain the availability of a single service instance when one or more primary VM instances are unavailable based on the number of primary VM instances in the primary VM instance pool and the unavailability rate. The operations also include instantiating a pool of secondary VM instances based on the number of secondary VM instances required to maintain the availability of a single service instance.
[0013] This aspect may include one or more of the following optional features. For example, the operation may optionally include identifying the unavailability of one of the primary VM instances in the pool of primary VM instances; and causing the unavailable primary VM instance to fail over to one of the secondary VM instances in the pool of secondary VM instances to begin executing a single service instance associated with the unavailable primary VM instance. In addition, the operation may further include determining that the secondary VM instance includes a corresponding resource level that is less than the target resource level associated with the corresponding single service instance; and increasing the corresponding resource level of the secondary virtual machine instance during the execution of the single service instance by the secondary virtual machine instance until the target resource level associated with the single service instance is met. In some examples, the number of secondary VM instances in the pool of secondary VM instances is less than the number of primary VM instances in the pool of primary VM instances. Optionally, each primary VM instance can execute multiple primary containers, and each primary container executes a corresponding single service instance in a secure execution environment isolated from other primary containers.
[0014] In some examples, the operation further includes: updating the number of secondary VM instances required to maintain the availability of a single service instance when the number of primary VM instances executing in the pool changes. The unavailability of the primary instance can be based on at least one of a failure of the primary VM instance, a delay in recreating the primary VM instance, or a scheduled maintenance period of the primary VM instance. Moreover, the unavailability rate can include at least one of an unavailability frequency or an unavailability period.
[0015] In some embodiments, the operations further include determining an unavailability rate for each primary VM instance in the pool of primary VM instances based on a mean time to failure (MTTF) and an expected length of time to recreate the corresponding VM instance. Instantiating the pool of secondary VM instances may include determining a corresponding VM type for each primary VM instance in the pool of primary VM instances, and for each different VM type in the pool of primary VM instances, instantiating at least one secondary VM instance of the same VM type. In some examples, the corresponding VM type for each primary VM instance indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, or a local storage requirement of the corresponding VM instance.
[0016] In some examples, the operations also include receiving a planned failover message indicating the number of primary VM instances in a pool of primary VM instances that will be unavailable during a planned maintenance period. In these examples, instantiating the pool of secondary VM instances is further based on the number of primary VM instances that will be unavailable during the planned maintenance period. For example, instantiating the pool of secondary VM instances can include instantiating a plurality of secondary VM instances equal to the larger of: the number of secondary VM instances required to maintain the availability of a single service instance; or the number of primary VM instances that will be unavailable during the planned maintenance period. The pool of secondary VM instances can be instantiated for use by a single customer of a distributed system. Optionally, the pool of secondary VM instances can be instantiated for use by multiple customers of a distributed system.
[0017] The details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of an example system including a user device that communicates with a distributed system via a network.
[0019] Figure 2 is a schematic diagram of an example distributed system that implements a virtual computing environment.
[0020] Figure 3 is a diagram of an example virtual computing environment with a primary virtual machine pool and a secondary virtual machine pool.
[0021] Figure 4 is a flow diagram of an example arrangement of operations of a method for maintaining availability of a virtual machine instance.
[0022] Figure 5 is an example computing device.
[0023] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0024] Users of distributed systems, such as customers developing projects on distributed systems, can execute one or more virtual machine (VM) instances to execute software applications or services by creating a virtual computing environment that typically simulates a physical computing environment that may involve specialized hardware, software, or a combination thereof. The VM instance itself is an application running on a host machine (e.g., a computer server) that executes a host operating system (OS) for managing VM instances running on the host machine. The host machine can run any number of VMs at the same time, such as a portion of a VM (e.g., a VM distributed on multiple host machines), one VM, and multiple VMs. Each VM can be executed on the host machine in a separate environment isolated from other VMs running on the host machine. Such a separate environment can be called a security sandbox or container. In addition, a single VM can run multiple containers, each of which runs in a separate environment isolated from other containers running on the VM.
[0025] A distributed system may provide a multi-tenant (MST) service that executes software / service instances on VMs. In a single tenant, each instance executes on a separate VM instance. An example MST service may include a structured query language (SQL) relational database. Thus, a distributed system may execute a pool of VM instances for an MST service, wherein each VM instance in the pool uses virtualized hardware provided by the VM instance to run / execute a corresponding single service instance of a service (e.g., a software application). Optionally, a distributed system may execute a pool of bare metal machines (e.g., bare metal machine servers) for an MST service, wherein each bare metal machine executed in the pool runs / executes a corresponding single service instance of a service.
[0026] Embodiments of the present invention are directed to an adaptive and dynamic method for maintaining the availability of service instances in a computing environment (e.g., a distributed system) when a primary virtual machine instance executing a corresponding single service instance becomes unavailable. Based on the knowledge that unplanned virtual machine failures are relatively rare (e.g., approximately once every two years) and there is little or no correlation between them, embodiments of the present invention dynamically adjust the number of secondary VMs to maintain availability based on the unavailability rate (or planned maintenance time period) of each primary VM instance and the number of primary VM instances. The unavailability rate can indicate the statistical frequency of failures and the length of time to recreate the corresponding primary VM instance. As a result, the total number of secondary VM instances instantiated to maintain the availability of the service instance is significantly reduced compared to conventional techniques that allocate one secondary VM instance to each primary VM instance. By reducing the total number of secondary VM instances, other computing resource requirements in the computing environment (e.g., a distributed system) are released, which would otherwise be reserved for maintaining a separate secondary VM instance for each primary VM instance. In addition to consuming fewer computing resources, the reduction in the total number of secondary VM instances also reduces costs due to the reduction in the number of physical resources required to create secondary VM instances.
[0027] The technical problems to be solved include how to better use and allocate resources in a multi-tenant (MST) service environment while still ensuring failover capabilities. The technical solution includes dynamically determining the capacity requirements of secondary virtual machines, and in particular, pooling secondary VM instances to ensure failover capabilities. Here, the pooling of secondary VM instances allows for a reduction in resources (e.g., processing time, memory, etc.) reserved for secondary VM instances, thereby freeing up additional resources for the computing environment to use for other purposes. In an optional embodiment in which each primary bare metal machine in the primary bare metal machine pool executes / runs a corresponding service instance, the embodiments of this article can similarly ensure failover capabilities by dynamically instantiating a secondary bare metal machine pool based on the number of secondary bare metal machines required to maintain the availability of the service instance when one or more primary bare metal machines fail.
[0028] Embodiments further reduce resources used for secondary VM instances by optionally running secondary VM instances with lower resource levels until called from the corresponding primary VM instance upon failover, and then increasing / growing resources to match the resource requirements (target level) of the failed / unavailable primary VM instance. Here, the target level of the unavailable VM instance corresponds to the target resource level specified by the corresponding service instance that failed over to the secondary VM instance to execute thereon. Additional embodiments allow for the use of a reduced number of high-resource secondary VM instances, as low-resource primary VM instances can fail over to high-resource secondary VM instances without negative impact. Here, the secondary VM instance can be reduced to an appropriate size after failover.
[0029] Figure 1An exemplary system 100 is depicted that includes a distributed system 200 configured to run a software application 360 (e.g., a service) executed on a pool of primary VM instances 350, 350P in a virtual computing environment 300. A user device 120 (e.g., a computer) associated with a user 130 (customer) communicates with the distributed system 200 via a network 140 to provide commands 150 for deploying, removing, or modifying primary VM instances 350P running in the virtual computing environment 300. Thus, the number of primary VM instances 350P in the pool of primary VM instances 350P can be dynamically changed based on the commands 150 received from the user device 120.
[0030] In some examples, the software application 360 is associated with an MST service, and each primary VM instance 350P is configured to execute a corresponding single service instance 362 of the software application 360 (e.g., a single tenant of the MST service). In the event that one or more primary VM instances 350P become unavailable, the distributed system 200 executes a computing device 112 configured to instantiate a pool of secondary VM instances 350, 350S in order to maintain the availability of the one or more single service instances 362 associated with the unavailable primary VM instances 350P. For example, in response to the primary VM instance 350P executing the corresponding single service instance 362 becoming unavailable, the computing device 112 (e.g., data processing hardware) can cause the single service instance 362 to fail over to one of the secondary VM instances 350S in the pool of secondary VM instances 350S so that the single service instance 362 does not become unavailable for an indeterminate period of time. Thus, the distributed system 200 can dynamically fail over to one of the instantiated secondary VM instances 350, rather than experiencing downtime to recreate the unavailable primary VM 350P, thereby not disrupting the existing single service instance 362 for the user 130. The primary VM instance 350P may become unavailable due to unplanned and unexpected failures, delays in recreating the primary VM instance 350P, and / or due to a planned maintenance period for the primary VM instance 350P, such as an update of a critical security patch to the kernel of the primary VM instance 350P.
[0031] While techniques for creating separate secondary VM instances 350S as backups for each primary VM instance 350P advantageously provide high availability, a disadvantage of these techniques is that most of these passive secondary VM instances 350S remain idle indefinitely because primary VM instances 350P rarely fail, resulting in a large number of resources not being used. In addition, undue costs are incurred based on the resource requirements required to create a secondary VM instance 350S for each primary VM instance 350P. In order to maintain the availability of a single service instance 362 when one or more primary VM instances 350P are unavailable, but without incurring the inefficiencies and high costs associated with the above-mentioned high availability techniques, embodiments of the present invention are directed to instantiating a pool of secondary VM instances 350S by reducing the number of secondary VM instances 350S relative to the number of primary VM instances 350P in the pool of primary VM instances 350P. Here, the computing device 112 may determine the number of secondary VM instances required to maintain availability when one or more primary VM instances are unavailable based on the number of primary VM instances 350P in the pool of primary VM instances 350P and the unavailability rate for each primary VM instance 350P.
[0032] The unavailability rate of the primary VM instance 350P may include at least one of the unavailability frequency or the unavailability period. For example, each primary VM instance 350P may include a corresponding mean time to failure (MTTF), which indicates how long (e.g., days) the primary VM instance 350P is expected to run before a failure occurs. The MTTF value may be 365 days (e.g., 1 year) or 720 days (e.g., 2 years). The unavailability rate of each primary VM instance 350P may further include the expected length of time to recreate the corresponding primary VM instance (e.g., a stock-out value). For example, when the distributed system 200 waits for resources (i.e., processing resources and / or memory resources) to become available to recreate the VM instance 350, the VM instance 350 may be associated with a stock-out value. The MTTF and the expected length of time to recreate each primary VM instance 350P may be obtained by observing the statistical analysis and / or machine learning techniques of the execution of VM instances 350 with the same or similar VM types (i.e., processing resources, memory resources, storage resources, network configuration).
[0033] As the number of primary VM instances 350P executed in the virtual environment 300 continuously changes due to users 130 adding / removing primary VM instances 350P and / or primary VM instances 350 becoming unavailable, the distributed system 200 (i.e., via the computing device 112) is configured to dynamically update the number of secondary VM instances 350S instantiated in the pool of secondary VM instances 350S. In some examples, the pool of primary VM instances 350P is associated with a single user / customer 130, and the pool of secondary VM instances 350S is instantiated only for use by the single user / customer 130. In other examples, the pool of primary VM instances 350P includes multiple sub-pools of primary VM instances 350P, each sub-pool being associated with a different user / customer 130 and isolated from the other sub-pools. In these examples, the pool of secondary VM instances 350S is shared between multiple different users / customers 130 in the event that one or more primary VM instances 350P in any sub-pool are unavailable.
[0034] In some embodiments, the virtual computing environment 300 is overlaid on the resources 110, 110a-n of the distributed system 200. The resources 110 may include hardware resources 110 and software resources 110. The hardware resources 110 may include a computing device 112 (also referred to as a data processing device and data processing hardware) or a non-transitory memory 114 (also referred to as memory hardware). The software resources 110 may include software applications, software services, application programming interfaces (APIs), etc. The software resources 110 may reside in the hardware resources 110. For example, the software resources 110 may be stored in the memory hardware 114 or the hardware resources 110 (e.g., computing device 112) may execute the software resources 110.
[0035] The network 140 may include various types of networks, such as a local area network (LAN), a wide area network (WAN), and / or the Internet. Although the network 140 may represent a long-range network (e.g., the Internet or a WAN), in some embodiments, the network 140 includes a short-range network, such as a local area network (LAN). In some embodiments, the network 140 uses standard communication technologies and / or protocols. Therefore, the network 140 may include links using technologies such as Ethernet, wireless fidelity (WiFi) (e.g., 802.11), Worldwide Interoperability for Microwave Access (WiMAX), 3G, Long Term Evolution (LTE), Digital Subscriber Line (DSL), Asynchronous Transfer Mode (ATM), InfiniBand, PCI Express Advanced Switching, Bluetooth, Bluetooth Low Energy (BLE), etc. Similarly, the networking protocols used on the network 132 may include Multi-Protocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged over the network 140 may be represented using techniques and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. In addition, all or certain links may be encrypted using conventional encryption techniques such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. In other examples, the network 140 uses custom and / or proprietary data communication techniques instead of or in addition to the above techniques.
[0036] exist Figure 2 In the example shown, the distributed system 200 includes a collection 220 of resources 110 (e.g., hardware resources 110) executing a virtual computing environment 300. The virtual computing environment 300 includes a virtual machine manager (VMM) 320 and a virtual machine (VM) layer 340 that runs one or more virtual machines (VMs) 350, 350a-n, which are configured to execute instances 362a, 362a-n of one or more software applications 360. Each hardware resource 110 may include one or more physical central processing units (pCPUs) 112 ("data processing hardware 112") and memory hardware 114. Although each hardware resource 110 is shown with a single physical processor 112, any hardware resource 110 may include multiple physical processors 112. A host operating system (OS) 312 may execute on the collection 220 of resources 110.
[0037] In some examples, the VMM 320 corresponds to a virtual machine monitor 320 (e.g., a computing engine) including at least one of software, firmware, or hardware configured to create, instantiate / deploy, and execute VM 350. A computer (i.e., data processing hardware 112) associated with the VMM 320 that executes one or more VMs 350 may be referred to as a host 310, and each VM 350 may be referred to as a client. Here, the VMM 320 or virtual machine monitor is configured to provide each VM 350 with a corresponding client operating system (OS) 354, 354a-n having a virtual operating platform, and manage the execution of the corresponding client OS 354 on the VM 350. As used herein, each VM 350 may be referred to as an "instance" or "VM instance". In some examples, multiple instances of various operating systems may share virtualized resources. For example, a first VM 350 of a Linux® operating system, a second VM 350 of a Windows® operating system, and a third VM 350 of an OS X® operating system may all run on a single physical x86 machine.
[0038] The VM layer 340 includes one or more virtual machines 350. The distributed system 200 enables the user 130 to start the VM 350 on demand, that is, by sending a command 150 ( Figure 1 ). For example, the command 150 may include an image or snapshot associated with the corresponding operating system 312, and the distributed system 200 may use the image or snapshot to create a root resource 110 for the corresponding VM 350. Here, the image or snapshot within the command 150 may include a boot loader, a corresponding operating system 312, and a root file system. In response to receiving the command 150, the distributed system 200 may instantiate the corresponding VM 350 and automatically start the VM 350 when instantiated. The VM 350 simulates a real computer system (e.g., the host 310) and operates based on the computer architecture and functions of the real computer system or a hypothetical computer system, which may involve dedicated hardware, software, or a combination thereof. In some examples, the distributed system 200 authorizes and authenticates the user 130 before starting one or more VMs 350. An instance 362 of a software application 360, or simply an instance, refers to a VM 350 hosted on (executed on) the processing hardware 112 of the distributed system 200.
[0039] The host OS 312 virtualizes the underlying host hardware and manages the concurrent execution of one or more VM instances 350. For example, the host OS 312 can manage VM instances 350a-n, and each VM instance 350 can include an emulated version of the underlying host hardware or a different computer architecture. The emulated version of the hardware associated with each VM instance 350, 350a-n is referred to as virtual hardware 352, 352a-n. The virtual hardware 352 can include one or more virtual central processing units (vCPUs) (“virtual processors”) that emulate one or more physical processors 112 of the host 310 ( Figure 3 ). The virtual processors may be interchangeably referred to as “computing resources” associated with a VM instance 350 . The computing resources may include a target computing resource level required to execute a corresponding single service instance 362 .
[0040] The virtual hardware 352 may further include a virtual memory that communicates with the virtual processor and stores client instructions (e.g., client software) that can be executed by the virtual processor to perform operations. For example, the virtual processor may execute instructions from the virtual memory that cause the virtual processor to execute a corresponding single service instance 362 of the software application 360. Here, the single service instance 362 may be referred to as a client instance that cannot determine whether it is being executed by the virtual hardware 352 or the physical data processing hardware 112. If a client service instance 362 or the VM instance 350 itself executed on the corresponding VM instance 350 fails or terminates, other VM instances executing the corresponding single service instance 362 will not be affected. The host's microprocessor may include a processor-level mechanism to enable the virtual hardware 352 to effectively execute the software instance 362 of the application 360 by allowing the client software instructions to be executed directly on the host's microprocessor without code rewriting, recompiling, or instruction emulation. The virtual memory may be interchangeably referred to as a "memory resource" associated with the VM instance 350. The memory resource may include a target memory resource level required to execute the corresponding single service instance 362.
[0041] The virtual hardware 352 may further include at least one virtual storage device that provides storage capacity for services on the physical storage hardware 114. The at least one virtual storage device may be referred to as a storage resource associated with the VM instance 350. The storage resource may include a target storage resource level required to execute a corresponding single service instance 362. The client software executed on each VM instance 350 may further be assigned a network boundary (e.g., assigned a network address) through which the respective client software may communicate with the internal network 330 ( Figure 3 )、External network 140( Figure 1 ) or other process communications achieved by both. A network boundary may be referred to as a network resource associated with a VM instance 350.
[0042] The guest OS 354 executing on each VM 350 includes software that controls the execution of a corresponding single service instance 362, 362a-n of an application 360 by the VM instance 350. The guest OS 354, 354a-n executing on the VM instance 350, 350a-n can be the same as or different from other guest OS 354 executing on other VM instances 350. In some embodiments, the VM instance 350 does not require a guest OS 354 in order to execute a single service instance 362. The host OS 312 can further include virtual memory reserved for a kernel 316 of the host OS 312. The kernel 316 can include kernel extensions and device drivers, and can perform certain privileged operations that are prohibited for processes running in the user process space of the host OS 312. Examples of privileged operations include access to different address spaces, access to special function processor units (such as a memory management unit) in the host 310, and the like. The communication process 314 running on the host OS 312 may provide a portion of the VM network communication functionality and may execute in a user process space or in a kernel process space associated with the kernel 316 .
[0043] refer to Figure 3 In some embodiments, a virtual computing environment 300 running on a distributed system 200 includes multiple hosts 310, 310a-n (e.g., one or more data processing devices, such as rack-mounted servers or different computing devices) that may be located in different physical locations and may have different functions and computer architectures. The hosts 310 may communicate with each other through an internal data communication network 330 (internal network). The internal network 330 may include, for example, one or more wired (e.g., Ethernet) or wireless (e.g., Wi-Fi) networks. In some embodiments, the internal network 330 is an intranet. Optionally, the hosts 310 may also communicate with devices on an external network 140, such as the Internet. Other types of external networks are also possible.
[0044] In the illustrated example, each host 310 executes a respective host operating system (OS) 312, 312a-n, which virtualizes the underlying hardware (i.e., data processing hardware 112 and memory hardware 114) of the host 310 and manages the parallel execution of multiple VM instances 350. For example, the host operating systems 312a-312n-1 each manage the parallel execution of multiple primary VM instances 350P to collectively provide a pool of primary VMs 350P, while the host operating system 312n executing on the host 310n manages the execution of a pool of secondary VM instances 350S. Here, a dedicated host (e.g., host 310n) hosts the entire pool of secondary VM instances 350S, thereby ensuring that in the event of a failover, sufficient resources are available for use by the secondary VM instances 350S (without requiring the failover secondary VM instances 350S to migrate to a different host 310 with sufficient resources). However, in other examples, one or more secondary VM instances 350S may be instantiated across multiple hosts 310 that may also execute one or more primary VM instances 350P.
[0045] In some embodiments, the virtual machine manager 320 uses the master VM manager 322 to create and deploy each master VM instance 350P in the pool of master VM instances 350 to execute on the designated host 310. The VMM 320 can create each master VM instance 350 by allocating the computing resource level, memory resource level, network specification, and / or storage resource level required for executing the corresponding single service instance 362. Therefore, each master VM instance 350P in the pool of master VM instances 350P may include a corresponding VM type 380, which indicates at least one of the memory resource requirement, computing resource requirement, network specification requirement, or storage resource requirement for the corresponding master VM instance 350. In the example shown, all master VM instances 350P in the pool of master VM instances 350P have a VM type 380 of type A or type B. Therefore, the VM type 380 of type A may include at least one of the computing resource level, memory resource level, network specification, or storage resource level that is different from the VM type 380 of type B.
[0046] The master VM manager 322 at the VMM 320 can maintain an activity log for each VM instance 350P deployed into the pool of master VM instances 350P, the VM type 380 of each VM instance 350P, and the corresponding single service instance 362 executed on each master VM instance 350P. The log can be updated when the master VM instance 350P is deployed into or removed from the pool of master VM instances 350P. In addition, the pool of master VM instances 350P can be further divided into sub-pools based on distributing the master VM instances 350P in various fault domains such as buildings, regions, or districts. In some embodiments, the single service instance 362 is executed in a corresponding container running on a single master VM instance 350P with multiple other containers. Therefore, the log can indicate a list of containers running on each master VM instance 350P, and the corresponding service instance 362 executed in each container.
[0047] The master VM manager 322 further obtains an unavailability rate for each master VM instance 350P. In some examples, all master VM instances 350P in a pool of master VM instances 350P include the same unavailability rate. In other examples, the unavailability rate of the master VM instance 350P associated with the type A VM type 380 is different from the availability rate of the master VM instance 350P associated with the type B VM type 380. As described above, each master VM instance 350P can include a corresponding MTTF value indicating how long (e.g., days) the master VM instance 350P is expected to run before a failure occurs and a stock-out value indicating the expected length of time to recreate the master VM instance 350P. The MTTF value and the stock-out value can be derived from the observed monitoring data and a machine learning algorithm that observes the execution of similar VM instances 350 over time.
[0048] For example, each primary VM instance 350P may include a corresponding mean time to failure (MTTF) indicating how long (e.g., number of days) the primary VM instance 350P is expected to run before a failure occurs. The MTTF value may be 365 days (e.g., 1 year) or 720 days (e.g., 2 years). The unavailability rate of each primary VM instance 350P may further include the expected length of time to recreate the corresponding primary VM instance (e.g., a stock-out value). For example, when the distributed system 200 waits for resources (i.e., processing resources and / or memory resources) to become available to recreate the VM instance 350, the VM instance 350 may be associated with a stock-out value. The MTTF and the expected length of time to recreate each primary VM instance 350P may be obtained by observing the execution of VM instances 350 having the same or similar VM types (i.e., processing resources, memory resources, storage resources, network configurations) through statistical analysis and / or machine learning techniques.
[0049] The VMM 320 may further maintain a service instance repository 324 indicating each single service instance 362 of the software application 360 executed on the corresponding primary VM instance 350P in the pool of primary VM instances 350P and a target resource level required for executing the corresponding single service instance 362. In some examples, each single service instance 362 in the repository 324 may specify whether the corresponding service instance 362 allows reduced performance requirements for a temporary period after failover. In these examples, the service instance 362 that allows reduced performance requirements allows the service instance 362 to fail over to a secondary VM instance 362 having a corresponding resource level (e.g., a processor resource level, a memory resource level, or a storage resource level) that is less than the corresponding target resource level associated with the corresponding single service instance 362, and then dynamically grows / increases the corresponding resource level of the secondary VM instance until the target resource level is met. Therefore, the secondary VM instance 350S may initially execute the single service instance 362 with reduced performance until the corresponding resource level grows to meet the corresponding target resource level associated with the single service instance 362. In these examples, the secondary VM instances 350S may be allocated fewer resources when idle, and may be resized / grown as needed once required to execute a single service instance 362 during a failover.
[0050] In some examples, the VMM 320 includes a maintenance scheduler 326 that identifies a maintenance time period when one or more primary VM instances 350P in a pool of primary VM instances 350P will be unavailable for maintenance / updates performed offline. For example, the maintenance scheduler 326 may indicate the number of primary VM instances 350P that will be unavailable during a scheduled maintenance time period for performing maintenance / updates. In one example, the distributed system 200 periodically pushes out kernel updates at a deployment rate of two percent (2%) (or other percentage / value) such that two percent of the primary VM instances 350P in a pool of primary VM instances 350P will be unavailable during the scheduled maintenance time period for completing the updates. The kernel update may include a fix security patch in the kernel 216 associated with the VM instance 350. In some examples, the VMM 320 receives a planned failover message 302 from a computing device 304 indicating the number (or percentage) of primary VM instances 350P that will be unavailable during a scheduled maintenance time period for performing maintenance / updates. The computing device 304 may belong to an administrator of the distributed system 200. Optionally, when the user 130 wants to update one or more primary VM instances 350P in the pool of primary VM instances 350P, the user device 120 may provide a planned failover message 302 via the external network 140 .
[0051] In some embodiments, the VMM 320 includes a secondary VM instantiator 328 in communication with the primary VM manager 322, a service instance repository 324, and a maintenance scheduler 326 for instantiating a pool of secondary VM instances 350S. In these embodiments, the secondary VM instantiator 328 determines the number of secondary VM instances 350S required to maintain the availability of a single service instance 362 when one or more primary VM instances 350P are unavailable based on the number of primary VM instances 350P in the pool of primary VM instances 350P and the unavailability rate of each primary VM instance 350P. In some examples, the secondary VM instantiator 328 determines the number of secondary VM instances by calculating the following formula:
[0052]
[0053] Among them, VM secondary_N is the number of secondary VM instances 350S required to maintain the availability of a single service instance 362, VM Primary_Pool is the number of primary VM instances 350P in the pool of primary VM instances 350P, SO is the highest out-of-stock value among the primary VM instances 350P indicating the number of days to recreate the corresponding primary VM instance 350P, and MTTF is the lowest mean time to failure value in days among the primary VM instances 350P.
[0054] In one example, when the number of primary VM instances 350P is equal to one thousand (1,000), and the MTTF value of each VM instance 350P is equal to one year (365 days), and the stock-out value is equal to five (5) days, the number of secondary VM instances 350S (VM secondary_N ) is equal to fourteen (14). Unless the maintenance scheduler 326 identifies that more than fourteen primary VM instances will be unavailable during the planned maintenance time period, the secondary VM instantiator 328 will instantiate a number of secondary VM instances 350S equal to the number of secondary VM instances 350S required to maintain the availability of the single service instance 362 (e.g., 14). Otherwise, when the maintenance scheduler 326 identifies that more (e.g., more than 14) primary VM instances 350P will be unavailable during the planned maintenance time period, the secondary VM instantiator 328 will instantiate a number of secondary VM instances 350S equal to the number of primary VM instances 350P that will be unavailable during the planned maintenance time period. For instance, in the example above, when a kernel update is rolled out at a deployment rate of two percent (2%) (or other deployment rate), the secondary VM instantiator 328 will need to ensure that twenty (20) primary VM instances (e.g., 2% of 1,000 primary VM instances) are instantiated into a pool of secondary VM instances 350S to provide failover coverage for planned failover events.
[0055] In the case when a kernel update (or other maintenance / update process that makes the primary VM instance 350P unavailable during a planned time period) occurs after initially instantiating a pool of secondary VM instances 350 to maintain availability for unplanned failover (e.g., in the above example, instantiating 14 secondary VM instances 350S), the secondary VM instantiator 328 can simply instantiate additional secondary VM instances 350S to provide failover coverage during the planned maintenance time period. In some examples, once the planned maintenance time period expires, the pool of secondary VM instances 350S is updated by removing one or more secondary VM instances 350S (i.e., deallocating resources).
[0056] Instantiating a plurality of secondary VM instances 350S that are less than the number of primary VM instances 350P in the pool of primary VM instances 350P alleviates the need to provide one secondary VM instance 350S for each primary VM instance 350P (and keep them idle). Each secondary VM instance 350S in the pool of secondary VM instances 350S is idle and does not execute any workload (e.g., service instance) unless one of the primary VM instances 350P of the corresponding VM type 380 becomes unavailable (e.g., fails), thereby causing the unavailable primary VM instance 350P to fail over to the idle / passive secondary VM instance 350S to begin executing the single service instance 362 associated therewith. Since the secondary VM instances 350S are used during failover, the secondary VM instantiator 328 dynamically adjusts the pool of secondary VM instances 350S as needed to ensure that there are always enough secondary VM instances 350S to maintain the availability of the single service instance 362 of the software application 360. During a failover, the secondary VM instance 350S may be reconfigured with appropriate resources (eg, network configuration settings and / or memory resources) and execute boot applications associated with the unavailable primary VM instance 350P.
[0057] In some embodiments, when a customer / user 130 deploys a large number of primary VM instances 350P and has specific networking or isolation requirements that prevent sharing the pool of secondary VM instances 350S with other users / customers of the distributed system 200, the pool of secondary VM instances 350S is per customer / user 130, rather than global. In other embodiments, the pool of secondary VM instances 350S is shared between all individual service instances 362 among all customers / users of the distributed system 200.
[0058] In some examples, the primary VM manager 322 determines a corresponding VM type 380 for each primary VM instance 350P in the pool of primary VM instances 350P, and the secondary VM instantiator 328 instantiates at least one secondary VM instance 350S having the same (or corresponding) VM type 380 for each different VM type 380 in the pool of primary VM instances 350P. Thus, the secondary VM instantiator 328 can ensure that sufficient secondary VM instances 350S are available for each VM type 380 in the pool of primary VM instances 350P. In some configurations (not shown), the pool of secondary VM instances 350S is divided into multiple sub-pools based on the VM type 380. For example, each sub-pool of secondary VM instances 350S will include one or more secondary VM instances 350S having a respective VM type 380.
[0059] In some examples, secondary VM instantiator 328 instantiates one or more secondary VM instances 350S from a pool of secondary VM instances 350S having corresponding resource levels that are less than corresponding target resource levels associated with a single service instance 362 (i.e., when service instance repository 324 indicates that a single service instance 362 allows for reduced performance requirements). For example, Figure 3 A secondary VM instance 350S having a corresponding VM type 380 of type B' that is related to but not identical to a VM type 380 of type B in a pool of primary VM instances 350P is shown. Here, the secondary VM instance 350S having a VM type 380 of type B' includes at least one corresponding resource level (such as a computing resource level and / or a memory resource level) that is less than the corresponding resource level associated with the VM type 380 of type B. However, during failover to the secondary VM instance 350S (of type B'), at least one corresponding resource level (e.g., a computing resource level and / or a memory resource level) may dynamically grow / increase until at least one corresponding target resource level (e.g., defined by type B) is met. In these examples, the required resources of the pool of secondary VM instances 350S are reduced while the secondary VM instances 350S are passive and in an idle state, thereby reducing the cost of maintaining unused resource levels that would otherwise be incurred unless a failover occurs.
[0060] In some embodiments, the secondary VM instantiator 328 instantiates one or more secondary VM instances 350S in the pool of secondary VM instances 350S, and the corresponding resource level of the one or more secondary VM instances 350S is greater than the corresponding resource level of the primary VM instances 350P in the pool of primary VM instances 350P. For example, a secondary VM instance 350S having a corresponding VM type 380 of type B' may instead indicate that the secondary VM instance 350S is associated with a corresponding resource level that is greater than the corresponding resource level associated with the VM type 380 of type B, while the secondary VM instance 350S is in an idle state. However, during failover to the secondary VM instance 350S (of type B'), the corresponding resource level (e.g., computing resource level, memory resource level, and / or storage resource level) may be dynamically lowered / reduced until the corresponding target resource level (e.g., defined by type B) is met. By providing a secondary VM instance 350S that is larger (a larger resource level) than the primary VM instances 350P in the pool of primary VM instances 350P, the size of the pool of secondary VM instances 350S (the number of secondary VM instances 350S) may be reduced.
[0061] In an embodiment when the primary workload (e.g., service instance 362) is executed in a primary container, the secondary VM instantiator 328 can instantiate multiple secondary containers in a single secondary VM instance. Therefore, the primary container running on the primary VM instance 350P that becomes unavailable (e.g., fails) can fail over to a secondary container in the pool of secondary VM instances 350S. In this case, the secondary VM instance 350S running the secondary container can de-allocate the remaining resources not used by the secondary container for use by other VM instances 350, 350P, 350S in the virtual environment 300.
[0062] In some cases, the VMM 320 (or the host 310) identifies the unavailability of one of the primary VM instances 350P in the pool of primary VM instances 350P. For example, each primary VM instance 350P may employ an agent to collect an operating state 370 indicating whether the primary VM instance 350P is operating or unavailable due to a failure. The hosts 310 may communicate the operating state 370 of the VM instances 350 to the VMM 320 in addition to each other. As used herein, the term "agent" is a broad term that encompasses its ordinary and common meaning, including but not limited to a portion of code deployed inside a VM instance 350 (as part of a client OS 354 and / or as an application running on a client OS 354) to identify the operating state 370 of a VM instance 350. Thus, VMM 320 and / or host 310 may receive an operational status 370 indicating unavailability of one of primary VM instances 350 and failover the unavailable primary VM instance 350P to one of secondary VM instances 350S to begin executing a single service instance 362 associated with the unavailable primary VM instance 350P. Figure 3 In the example shown, the operating state 370 indicates the unavailability (e.g., due to a failure) of one of the primary VM instances 350P executing on the host 310n-1 and having a VM type 380 of type B, thereby causing the primary VM instance 350P to fail over to a secondary VM instance 350S having a VM type 380 of type B' to begin executing a single service instance 362 associated with the unavailable primary VM instance 350P having a VM type 380 of type B.
[0063] In some examples, a VM type 380 of type B' indicates that a corresponding secondary VM instance 350S includes at least one corresponding resource level (e.g., a computing resource level and / or a memory resource level) that is less than a target resource level associated with a corresponding single service instance 362 (as defined by the VM type 380 of type B). During execution of the single service instance 362 by the corresponding secondary VM instance 350S after a failover, the secondary VM instance 350S is configured to increase the at least one corresponding resource level until the target resource level associated with the single service instance 362 is met. In other examples, a VM type 380 of type B' indicates that the corresponding secondary VM instance 350S is greater than (e.g., includes a greater resource level) than the unavailable primary VM instance 350P of the VM type 380 of type B. In these examples, the secondary VM instance 350S may reduce / lower its resource level to meet the target resource level associated with the single service instance 362.
[0064] Figure 4An example arrangement of operations of a method 400 for maintaining the availability of a service instance 362 on a distributed system 200 is provided. The service instance 362 may correspond to a single service instance 362 in a multi-tenant service 360. At block 402, the method 400 includes executing a pool of master virtual machine (VM) instances 350P by the data processing hardware 112 of the distributed system 200. Each master VM instance 350P executes a corresponding single service instance 362 (associated with a service / software application 360) and includes an unavailability rate. Optionally, each master VM instance 350P may execute multiple master containers, and each master container executes a corresponding single service instance 362 in a secure execution environment isolated from other master containers. The unavailability rate may include at least one of an unavailability frequency or an unavailability period. In some embodiments, the method 400 also includes determining, by the data processing hardware 112, an unavailability rate of each master VM instance 350P in the pool of master VM instances 350P based on a mean time to failure (MTTF) and an expected length of time to recreate the corresponding master VM instance (e.g., a stock-out value). As used herein, MTTF indicates how often (eg, days) a primary VM instance 350P is expected to be operational before a failure occurs. MTTF values and stock-out values can be derived from observed monitoring data and machine learning algorithms that observe the execution of similar VM instances 350 over time.
[0065] At block 404, the method 400 also includes determining, by the data processing hardware 112, the number of secondary VM instances required to maintain the availability of the single service instance 362 when one or more primary VM instances 350P are unavailable based on the number of primary VM instances 350P in the pool of primary VM instances 350P and the unavailability rate. As used herein, the unavailability of the primary VM instance 350P is based on at least one of a failure of the primary VM instance 350P, a delay in recreating the primary VM instance 350P, or a scheduled maintenance period of the primary VM instance 350P. In some examples, the method uses Equation 1 to calculate the number of secondary VM instances required to maintain the availability of the single service instance 362.
[0066] At box 406, the method 400 also includes instantiating, by the data processing hardware 112, a pool of secondary VM instances 350S based on the number of secondary VM instances 350S required to maintain the availability of the single service instance 362. The number of secondary VM instances instantiated into the pool can further take into account the number of primary VM instances 350 that may be unavailable during the planned maintenance time period. For example, the planned failover message 302 can indicate that the primary VM instance 350P will be unavailable to perform a deployment rate for maintenance / updates during the planned maintenance time period. Here, the method 400 can include instantiating a number of secondary VM instances equal to the larger of the following: the number of secondary VM instances 350S required to maintain the availability of the single service instance 362; or the number of primary VM instances 350P that will be unavailable during the planned maintenance time period.
[0067] Each secondary VM instance 350S in the pool of secondary VM instances may be passive and idle (i.e., not executing a workload) unless a failover causes the corresponding secondary VM instance 350S to begin executing a single service instance 362 associated with the unavailable primary VM instance 350P. In some examples, the method 400 further determines a corresponding VM type 380 for each primary VM instance 350P in the pool of primary VM instances 350P. In these examples, instantiating the pool of secondary VM instances 350S includes instantiating at least one secondary VM instance 350S of the same VM type 380 for each different VM type 380 in the pool of primary VM instances 350P.
[0068] Optionally, when the primary VM instance 350P runs multiple primary containers that each execute a corresponding single service instance, the method may optionally include instantiating multiple secondary containers to maintain the availability of the single service instance when one or more primary containers are unavailable (e.g., fail). Here, each secondary VM instance 350S instantiated into the pool of secondary VM instances 350S is configured to run multiple secondary containers.
[0069] Software applications (i.e., software resources 110) may refer to computer software that causes a computing device to perform tasks. In some examples, software applications may be referred to as "applications," "apps," or "programs." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0070] Non-transitory memory (e.g., memory hardware 114) can be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) on a temporary or permanent basis for use by a computing device (e.g., data processing hardware 112). Non-volatile memory 114 can be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory can include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0071] Figure 5 is a schematic diagram of an example computing device 500 that can be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not meant to limit implementations of the inventions described and / or claimed in this document.
[0072] The computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connected to the memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 connected to a low-speed bus 570 and the storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or in other appropriate ways. The processor 510 is capable of processing instructions for execution within the computing device 500, including instructions stored in the memory 520 or stored on the storage device 530, to display graphical information for a graphical user interface (GUI) on an external input / output device such as a display 580 coupled to the high-speed interface 540. In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and memory types, can be used appropriately. Moreover, multiple computing devices 500 can be connected, each of which provides a portion of the necessary operations (for example, as a server group, a blade server group, or a multi-processor system).
[0073] Memory 520 stores information non-temporarily within computing device 500. Memory 520 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-temporary memory 520 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0074] The storage device 530 can provide mass storage for the computing device 500. In some embodiments, the storage device 530 is a computer readable medium. In various embodiments, the storage device 530 can be a floppy disk device, a hard disk device, an optical disk device or a tape device, a flash memory or other similar solid state storage device, or a device array, including a device in a storage area network or other configuration. In other embodiments, the computer program product is tangibly embodied as an information carrier. The computer program product contains instructions for performing one or more methods, such as those described above, when executed. The information carrier is a computer or machine readable medium, such as a memory 520, a storage device 530, or a memory on a processor 510.
[0075] The high-speed controller 540 manages bandwidth-intensive operations of the computing device 500, while the low-speed controller 560 manages less bandwidth-intensive operations. This division of responsibilities is illustrative only. In some embodiments, the high-speed controller 540 is coupled to a memory 520, a display 580 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some embodiments, the low-speed controller 560 is coupled to a storage device 530 and a low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device, such as a switch or a router, for example, through a network adapter.
[0076] As shown, computing device 500 may be implemented in a variety of different forms. For example, computing device 500 may be implemented as a standard server 500a or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0077] Various implementations of the systems and techniques described herein can be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to these devices.
[0078] These computer programs (also referred to as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0079] The processes and logic flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by dedicated logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits). For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include one or more large-capacity storage devices such as magnetic disks, magneto-optical disks, or optical disks for storing data, or be operably coupled to a large-capacity storage device to receive data from it or transfer data to it, or both. However, a computer does not have to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0080] To provide interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, to display information to the user, and an optional keyboard and pointing device, such as a mouse and trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0081] Many embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Therefore, other embodiments are also within the scope of the appended claims.
Claims
1. A computer-implemented method executed on data processing hardware, the method causing the data processing hardware to perform an operation, the operation include: A pool of master virtual machine (VM) instances is executed, each of which executes a single service instance. Instantiate a shared auxiliary VM instance; Identifying unavailability of a particular primary VM instance in the pool of primary VM instances; causing the corresponding single service instance executing on the specific primary VM instance to fail over to the shared secondary VM instance to begin executing the corresponding single service instance; and After the corresponding single service instance fails over to the shared secondary VM instance: determining a difference between a current resource level of the shared secondary VM instance and a target resource level associated with the corresponding single service instance; and The current resource level of the shared secondary VM instance is adjusted based on the difference.
2. The computer-implemented method of claim 1, in, The target resource level indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
3. The computer-implemented method of claim 1, in, Identifying the unavailability of the particular primary VM instance is based on at least one of a failure of the particular primary VM instance, a delay in recreating the particular primary VM instance, and a planned maintenance time period for the particular primary VM instance.
4. The computer-implemented method of claim 1, in: The operations further include selecting the shared secondary VM instance from a shared pool of secondary VM instances; and The number of secondary VM instances in the shared pool of secondary VM instances is less than the number of primary VM instances in the pool of primary VM instances.
5. The computer-implemented method of claim 4, in, The operations further include changing the number of secondary VM instances in the shared pool of secondary VM instances when the number of the primary VM instances changes.
6. The computer-implemented method of claim 4, in, The operations further include changing a number of secondary VM instances in the shared pool of secondary VM instances based on at least one of an unavailability frequency or an unavailability period.
7. The computer-implemented method of claim 4, in, The operations further include: determining a corresponding VM type for each primary VM instance in the pool of primary VM instances; and For each specific VM type in the pool of primary VM instances, at least one secondary VM instance of the specific VM type is instantiated.
8. The computer-implemented method of claim 7, in, The specific VM type indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
9. The computer-implemented method of claim 4, in, The shared pool of secondary VM instances is shared among multiple clients of the distributed system.
10. The computer-implemented method of claim 1, in, The secondary VM instance is passive and idle until a failover causes the secondary VM instance to begin executing a corresponding single service instance associated with the particular primary VM instance that is unavailable.
11. A system, include: Data processing hardware; Storage hardware, the storage hardware is in communication with the data processing hardware, the storage hardware stores instructions, the instructions, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations comprising: A pool of master virtual machine (VM) instances is executed, each of which executes a single service instance. Instantiate a shared auxiliary VM instance; Identifying unavailability of a particular primary VM instance in the pool of primary VM instances; causing the corresponding single service instance executing on the specific primary VM instance to fail over to the shared secondary VM instance to begin executing the corresponding single service instance; and After the corresponding single service instance fails over to the shared secondary VM instance: determining a difference between a current resource level of the shared secondary VM instance and a target resource level associated with the corresponding single service instance; and A resource level of the shared secondary VM instance is adjusted based on the difference.
12. The system according to claim 11, in, The target resource level indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
13. The system according to claim 11, in, Identifying the unavailability of the particular primary VM instance is based on at least one of a failure of the particular primary VM instance, a delay in recreating the particular primary VM instance, and a planned maintenance time period for the particular primary VM instance.
14. The system according to claim 11, in: The operations further include selecting the shared secondary VM instance from a shared pool of secondary VM instances; and The number of secondary VM instances in the shared pool of secondary VM instances is less than the number of primary VM instances in the pool of primary VM instances.
15. The system according to claim 14, in, The operations further include changing the number of secondary VM instances in the shared pool of secondary VM instances when the number of the primary VM instances changes.
16. The system according to claim 14, in, The operations further include changing a number of secondary VM instances in the shared pool of secondary VM instances based on at least one of an unavailability frequency or an unavailability period.
17. The system according to claim 14, in, The operations further include: determining a corresponding VM type for each primary VM instance in the pool of primary VM instances; and For each specific VM type in the pool of primary VM instances, at least one secondary VM instance of the specific VM type is instantiated.
18. The system according to claim 17, in, The specific VM type indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
19. The system according to claim 14, in, The shared pool of secondary VM instances is shared among multiple clients of the distributed system.
20. The system according to claim 11, in, The secondary VM instance is passive and idle until a failover causes the secondary VM instance to begin executing a corresponding single service instance associated with the particular primary VM instance that is unavailable.
21. A computer-implemented method executed by data processing hardware, the method causing the data processing hardware to perform an operation, the operation include: executing a master pool of master virtual machine (VM) instances, the master pool of master VM instances comprising a first number of master VM instances, each master VM instance executing a corresponding single service instance; determining, based on the first number of primary VM instances in the primary pool of primary VM instances, a first number of secondary VM instances required to maintain availability of each single service instance when one or more of the primary VM instances are unavailable; instantiating a secondary pool of secondary VM instances based on the first number of secondary VM instances required to maintain availability of each single service instance; After instantiating the secondary pool of secondary VM instances, determining that the primary pool of primary VM instances includes a second number of primary VM instances that is different from the first number of primary VM instances; determining, based on the second number of primary VM instances in the primary pool of primary VM instances, a second number of secondary VM instances required to maintain availability of each single service instance when one or more of the primary VM instances are unavailable; and The secondary pool of secondary VM instances is updated based on the second number of secondary VM instances required to maintain availability of each single service instance.
22. The computer-implemented method of claim 21, in, The secondary pool of secondary VM instances is passive and idle until a failover causes the secondary pool of secondary VM instances to begin executing a corresponding single service instance associated with the particular primary VM instance that is unavailable.
23. The computer-implemented method of claim 21, in, The secondary pool of secondary VM instances is shared among multiple clients of the distributed system.
24. The computer-implemented method of claim 21, in, The first number of secondary VM instances is less than the first number of primary VM instances.
25. The computer-implemented method of claim 21, in, The second number of secondary VM instances is less than the second number of primary VM instances.
26. The computer-implemented method of claim 21, in, The operations further include: determining a corresponding VM type for each primary VM instance in the primary pool of primary VM instances; and For each specific VM type in the primary pool of primary VM instances, at least one secondary VM instance of the specific VM type is instantiated.
27. The computer-implemented method of claim 26, in, The specific VM type indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
28. The computer-implemented method of claim 21, in, The operations further include identifying unavailability of a particular primary VM instance in the pool of primary VM instances.
29. The computer-implemented method of claim 28, in, The operations further include causing the corresponding single service instance executing on the particular primary VM instance to fail over to one of the secondary VM instances to begin executing the corresponding single service instance.
30. The computer-implemented method of claim 21, in, Each master VM instance further executes the corresponding container.
31. A system, include: Data processing hardware; Storage hardware, the storage hardware is in communication with the data processing hardware, the storage hardware stores instructions, the instructions, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations comprising: executing a master pool of master virtual machine (VM) instances, the master pool of master VM instances comprising a first number of master VM instances, each master VM instance executing a corresponding single service instance; determining, based on the first number of primary VM instances in the primary pool of primary VM instances, a first number of secondary VM instances required to maintain availability of each single service instance when one or more of the primary VM instances are unavailable; instantiating a secondary pool of secondary VM instances based on the first number of secondary VM instances required to maintain availability of each single service instance; After instantiating the secondary pool of secondary VM instances, determining that the primary pool of primary VM instances includes a second number of primary VM instances that is different from the first number of primary VM instances; determining, based on the second number of primary VM instances in the primary pool of primary VM instances, a second number of secondary VM instances required to maintain availability of each single service instance when one or more of the primary VM instances are unavailable; and The secondary pool of secondary VM instances is updated based on the second number of secondary VM instances required to maintain availability of each single service instance.
32. The system according to claim 31, in, The secondary pool of secondary VM instances is passive and idle until a failover causes the secondary pool of secondary VM instances to begin executing a corresponding single service instance associated with the particular primary VM instance that is unavailable.
33. The system according to claim 31, in, The secondary pool of secondary VM instances is shared among multiple clients of the distributed system.
34. The system according to claim 31, in, The first number of secondary VM instances is less than the first number of primary VM instances.
35. The computer-implemented method of claim 31 , in, The second number of secondary VM instances is less than the second number of primary VM instances.
36. The system according to claim 31, in, The operations further include: determining a corresponding VM type for each primary VM instance in the primary pool of primary VM instances; and For each specific VM type in the primary pool of primary VM instances, at least one secondary VM instance of the specific VM type is instantiated.
37. The system according to claim 36, in, The specific VM type indicates at least one of a memory resource requirement, a computing resource requirement, a network specification requirement, and a storage resource requirement.
38. The system according to claim 31, in, The operations further include identifying unavailability of a particular primary VM instance in the pool of primary VM instances.
39. The system according to claim 38, in, The operations further include causing the corresponding single service instance executing on the particular primary VM instance to fail over to one of the secondary VM instances to begin executing the corresponding single service instance.
40. The system according to claim 31, in, Each master VM instance further executes the corresponding container.