Customer initiated virtual machine resource allocation sharing

By adopting single-slot oversubscription technology on edge servers of cloud computing systems, allowing replicas of virtual machines to be run on the same slot, solving the downtime problem caused by limited edge server resources, and achieving efficient resource utilization and rapid failover.

CN119968620APending Publication Date: 2025-05-09AMAZON TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069982.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-26
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In cloud computing systems, edge servers have limited resource capacity, resulting in the need to terminate existing virtual machines for failover, deployment, or other purposes, resulting in downtime or interruption of customer applications.

Method used

Using single-slot oversubscription technology, customers allow replicas of virtual machines to be run on the same slot, enabling failover and updates to virtual machines without using two separate slots for both running and standby replicas.

Benefits of technology

Efficiently utilize the limited capacity of edge servers, reduce downtime, and avoid the time required to redeploy virtual machines, improving the availability and efficiency of cloud services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968620A_ABST
    Figure CN119968620A_ABST
Patent Text Reader

Abstract

Techniques for customer initiated virtual machine resource allocation sharing are described. A hardware virtualization service of a cloud provider network receives a request to launch a first virtual machine, where the first virtual machine belongs to a first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type. The hardware virtualization service causes startup of the first virtual machine on a host computer system of the cloud provider network. The host computer system shares an allocation of the amount of resources from the host computer system for corresponding resources between the first virtual machine and a second virtual machine, wherein the second virtual machine belongs to the first virtual machine type.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cloud computing systems typically provide on-demand, managed computing resources to customers. Such computing resources (e.g., computing and storage capacity) are typically provided by large pools of capacity installed in data centers. Customers can request computing resources from the "cloud," and the cloud can provide computing resources to those customers. Technologies such as virtual machines and containers are often used to allow customers to securely share the capacity of a computer system. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Various examples according to the present disclosure will be described with reference to the accompanying drawings.

[0003] Figure 1 An environment for customer-initiated sharing of virtual machine resource allocation according to some examples is illustrated.

[0004] Figure 2 An exemplary system including a cloud provider network and also including various edge locations of the cloud provider network according to some examples is illustrated.

[0005] Figure 3 An exemplary cloud provider network including geographically dispersed edge locations is illustrated according to some examples.

[0006] Figure 4 An exemplary system is illustrated in which a cloud provider network edge location is deployed within a communications service provider network according to some examples.

[0007] Figure 5 Exemplary components of edge locations within a cloud provider network and a communications service provider network and connectivity therebetween according to some examples are illustrated in greater detail.

[0008] Figure 6 An environment for sharing processing resources among resource-sharing virtual machines according to some examples is illustrated.

[0009] Figure 7 An environment for sharing memory resources among resource-sharing virtual machines according to some examples is illustrated.

[0010] FIG. 8A to FIG. 8D An environment for sharing networking resources among resource-sharing virtual machines according to some examples is illustrated.

[0011] Fig. 9 An environment for health-based resource sharing virtual machine replacement according to some examples is illustrated.

[0012] Fig.10 An environment for transferring a virtual machine image according to some examples is illustrated.

[0013] Fig.11Operation of a method for customer-initiated sharing of virtual machine resource allocations according to some examples is illustrated.

[0014] Fig.12 An example provider network environment according to some examples is illustrated.

[0015] Fig.13 An example provider network that provides storage services and hardware virtualization services to customers according to some examples is illustrated.

[0016] Fig.14 An example computer system is illustrated that may be used for some of the examples. DETAILED DESCRIPTION

[0017] The present disclosure relates to methods, devices, systems, and non-transitory computer-readable storage media for customer-initiated virtual machine resource allocation sharing. More specifically, embodiments of the present disclosure relate to "single slot oversubscription" in which a customer can request to have a duplicate copy of an instance run on the same "slot". A slot refers to a set of physical hardware resources (e.g., CPU, memory) allocated for use by a specific virtual machine instance (VM), or in the case of the present disclosure, a VM and its copies. The duplicate VMs may be a running VM and a standby copy of the VM, or may be a running VM and an updated version of the VM, which can replace the old version of the VM when ready. The disclosed slot oversubscription technology for customer-initiated standby VM copies beneficially enables the customer's application to fail over to the standby VM in a scenario where the running copy of the VM encounters a problem, without using two separate slots for the running copy and the standby copy. This is particularly beneficial for workloads running on edge servers that are more limited in capacity, enabling efficient use of limited capacity while also preventing customer application downtime or interruption caused by the time spent re-provisioning virtual machines from scratch.

[0018] A network of cloud providers offers a variety of computing products and services to their customers. Virtualization technology is an important foundation for these products, allowing customers to access their own virtualized computing environments, and the underlying hardware resources that support virtualized computing environments are usually shared among many virtualized computing environments. Since virtualization separates the relationship between physical hardware and virtualized computing environments, virtual machines are usually described by the amount and / or performance level of different resources (e.g., computing, memory, network throughput, storage, accelerators, etc.) that they can use (or use at most) from the underlying host computer system. Virtual machines (also referred to as VMs or instances) with different resource allocations may be referred to as different virtual machine types. For example, one virtual machine type may have two virtual processors, 8GB of memory, and 10GB / sec of network throughput, while another virtual machine type may have four virtual processors, 16GB of memory, 15GB / sec of network throughput, and an attached acceleration card (e.g., graphics processor, signal processor, etc.).

[0019] Although the resources of a host computer system (or "host" for short) may be shared among many different virtual machines, the resources allocated to a particular virtual machine are typically dedicated to that virtual machine. The examples described herein relate to customer-initiated sharing of resources allocated to a particular virtual machine among multiple virtual machines. Each of the virtual machines that share the underlying virtual machine resource allocation is referred to as a "resource sharing" virtual machine. Sharing may also be referred to as "oversubscription" because the declared amount and / or performance level associated with a virtual machine type reflects the total amount and / or performance level available to each resource sharing virtual machine.

[0020] In some examples, a customer may request to start one or more resource-sharing virtual machines from a virtualization service of a cloud provider network by identifying an existing virtual machine that shares resources with a newly requested virtual machine. In this case, the virtualization service may start the requested virtual machine on the same host that executes the existing virtual machine, and configure the host to share the resource allocation of the existing virtual machine with the newly started virtual machine. In other examples, a customer may request to start multiple resource-sharing virtual machines of a specific type. In this case, the virtualization service may start the virtual machine on the host, and configure the host to share the resource allocation that is usually associated with a single virtual machine of that type among the virtual machines. In either case, resources that are traditionally allocated to a single virtual machine of a given type are shared among multiple resource-sharing virtual machines of that type. Of course, if all resource-sharing virtual machines except one resource-sharing virtual machine are terminated, the remaining virtual machines will no longer share resources, and the full allocation and / or performance level of the virtual machine will be made available to the virtual machine customer.

[0021] In many cases, it can be advantageous to allow a virtual machine's resource allocation to be shared among virtual machines. For example, some customers may use a rolling deployment strategy to periodically update their applications. Under a rolling deployment strategy, virtual machines running an older version of a customer's application are slowly replaced by virtual machines running a newer version of the customer's application. Rather than launching virtual machines that require separate resource allocations, customers can launch resource-sharing virtual machines along with existing virtual machines executing the older version of the application. Once the virtual machines executing the newer version of the application begin running, the virtual machines executing the older version of the application can be terminated.

[0022] As another example of a beneficial use of resource-sharing VMs, some customers may want to have a backup VM ready in the event that a primary VM fails. Such a backup may be referred to as a "shadow" VM. While the primary VM is operating normally, the backup VM may consume very few resources. Doing so eliminates the costs associated with a second VM with separate resource allocations while also reducing downtime in the event of a failure (e.g., there may be delays in rerouting network traffic from the primary VM to the backup VM if it is hosted on another host).

[0023] The benefits of resource-sharing VMs are further amplified in the context of cloud provider network edge locations. For present purposes, edge locations typically extend the management infrastructure experience typically associated with cloud provider networks to other environments (e.g., third-party networks, customer networks, etc.). Due to being deployed outside the typical confines of a cloud provider network, edge locations have relatively limited physical computing resources available to virtualize customer environments and other services compared to the cloud. Given the smaller resource capacity of edge locations relative to the cloud, edge locations may lack sufficient available resources to allocate to new parallel VM launches for failover, deployment, or other purposes. As a result, customers may need to terminate existing VMs before launching replacement VMs, which may result in significant downtime for the customer's applications. Downtime may be exacerbated if any dependencies need to be transferred from the cloud provider network to the edge location before launching new VMs, and if there are limited communication channels between the cloud provider network and the edge location. Therefore, resource-sharing VMs can play an important role in minimizing downtime in situations where physical computing resources are limited.

[0024] As an example use case, the disclosed single-slot oversubscription technology can be implemented on a radio access network (RAN) edge server that runs network functions corresponding to distributed units (DUs) and / or centralized units (CUs) of a wireless communication network (such as a 5G network). The RAN also includes a radio unit (RU) and one or more antennas. These units can be geographically distributed and provided in different proportions, but generally at least the RU will be located close to the antenna. Multiple RUs can be connected to the DU, and multiple DUs can also be connected to the CU. The CU of a 5G network may be located away from the antenna and in a more centralized location. In general, the radio unit (RU), distributed unit (DU), and central unit (CU) convert analog radio signals received from the antenna into digital packets that can be routed through the network, and similarly, they also convert digital packets into radio signals that can be transmitted through the antenna. This signal conversion is performed by a series of network functions that can be distributed between the RU, DU, and CU in various ways to achieve different balances of latency, throughput, and network performance. These are called "functional splits" of the RAN.

[0025] The network functions implemented in the RAN correspond to the lowest three network layers in the seven-layer OSI model of computer networks. The physical layer (PHY) or layer 1 (L1) is the first and lowest layer in the OSI model. In a radio-based network 103, the PHY is the layer that sends and receives radio signals. This can be split into two parts: "high PHY" and "low PHY". Each of these can be considered a network function. The high-end PHY converts binary bits into electrical pulses representing the binary data, and the low-end PHY converts these electrical pulses into radio waves for wireless transmission through an antenna. The PHY similarly converts received radio waves into digital signals. This layer can be implemented by a dedicated PHY chip.

[0026] The PHY interfaces with the data link layer, layer 2 (L2) in the OSI model. The main task of L2 is to provide an interface between higher transport layers and the PHY. 5G L2 has three sublayers: media access control (MAC), radio link control (RLC), and packet data convergence protocol (PDCP). Each of these can be considered a network function. PDCP provides security for radio resource control (RRC) traffic and signaling data, sequence numbering and in-order delivery of RRC messages and IP packets, and IP packet header compression. The RLC protocol provides control of the radio link. The MAC protocol maps information between logical channels and transport channels.

[0027] The data link layer interfaces with Layer 3 (L3) (Network Layer) in the OSI model. 5G L3 is also called the Radio Resource Control (RRC) layer and is responsible for functions such as packet forwarding, quality of service management, and the establishment, maintenance, and release of the RRC connection between the UE and the RAN.

[0028] Various functional splits can be selected for the RAN. The functional split defines different sets of L1 and L2 functions running on the RU and on the CU and DU. L3 also runs on the CU. For example, in a RAN architecture that follows Split 7, the functionality of the baseband unit (BBU) used in previous generations of wireless networks is split into two functional units: the DU, which is responsible for real-time L1 and L2 scheduling functions, and the CU, which is responsible for non-real-time, higher-level L2 and L3 functions. In contrast, for example, in a RAN architecture that follows Split 2, only the PDCP from L2 is handled by the DU and CU, while the RLC, MAC, PHY, and radio frequency signals (RF) are handled by the RU. For example, in Split 5, the DU and CU handle PDCP, RLC, and part of the MAC functions, while the RU handles part of the MAC as well as PHY and RF. For example, in Split 6, the DU and CU handle PDCP, RLC, MAC, while the RU only handles PHY and RF. For example, in Split 8, the DU and CU handle PDCP, RLC, MAC, and PHY, while the RU only handles RF.

[0029] An outage in any of these network functions may cause the operation of the entire network to fail, thereby affecting any UE connected to the network. Therefore, a customer may request that the disclosed single-slot oversubscription technique be implemented for any or all network functions running on a RAN edge server to provide a more resilient and available network for UEs. The RAN edge server may be a cloud provider underlying extension as described herein, which may further include a dedicated PHY chip for L1 processing as described above. It is understood that 5G RAN is just one use case for such an edge server, and customers may deploy cloud-managed edge servers within their premises to handle any desired type of workload, such as latency-sensitive workloads that provide better performance when placed near other local customer workloads (e.g., control systems for manufacturing operations, real-time data analytics workloads, real-time machine learning (ML) inference workloads). In addition, although this article presents an example of using single-slot oversubscription on an edge server with limited capacity, it is understood that this technique is also applicable to servers with larger capacity, such as servers running in a data center in a cloud provider environment, to provide higher workload availability while making the most efficient use of the underlying hardware.

[0030] Figure 1An example of a shared environment for virtual machine resource allocation initiated by a customer according to some examples. As shown, a cloud provider network 100 includes a hardware virtualization service 110 and a host computer system 150, which is typically one of many systems that form a pool of physical hardware resources that can be used for virtualization. Edge location 199 includes a host computer system 160. As described above, cloud provider network 100 typically relies on virtualization technology. Virtualization technology can provide users with the ability to control or use virtualized computing environments, wherein one or more virtualized computing environments are hosted by the same host. Therefore, users can directly use computing resources hosted by the provider network (e.g., provided by hardware virtualization services) to perform a variety of computing tasks. An exemplary virtualized computing environment includes virtual machines and containers. A virtual machine is typically a process executed within a host operating system managed by a hypervisor or a virtual machine manager (VMM). A virtual machine typically operates a guest operating system in which other applications are executed. In some examples, an offload card including a dedicated processor can execute a hypervisor or VMM and other virtualization management components, thereby releasing other host system resources.

[0031] The hardware virtualization service 110 (referred to as an elastic computing service, a virtual machine service, a computing cloud service, a computing engine, or a cloud computing service in various embodiments) enables users of the provider network 100 to configure and manage virtualized computing environments, such as virtual machines. The startup of a virtual machine is typically performed as follows. The hardware virtualization service 110 receives a startup request including one or more parameters. One such parameter is an indication of the type of virtual machine to be started. The virtual machine type typically defines the various amounts and / or levels of resources to be provided to the virtual machine from the physical hardware resources of the underlying host. Other parameters may include an identification of the software environment of the virtual machine (e.g., an identification of a guest operating system) or an identification of a machine image (typically a snapshot of a virtual machine), the machine image including various preloaded software from which the virtual machine is to be started.

[0032] The hardware virtualization service 110 may identify a host with sufficient resources to launch the requested virtual machine, and then cause or otherwise direct an agent (e.g., hypervisor, VMM) on the identified host to configure and launch the virtual machine using a particular machine image (whether specified in the request or associated with the identified software environment). If the machine image is not already stored locally on the host, the agent may retrieve the machine image from a data store, launch a virtual machine process from the machine image, and allocate some portion of the resources of the underlying host system to the process according to the corresponding virtual machine type.

[0033] keep going, Figure 2An exemplary system including a cloud provider network 100 and also including various edge locations 240, 242, 244 of the cloud provider network is illustrated according to some examples. A cloud provider network (sometimes simply referred to as a "cloud") refers to a pool of network-accessible computing resources (such as computing, storage, and networking resources; applications, and services), which may be virtualized or bare metal. A cloud may provide convenient, on-demand network access to a shared pool of configurable computing resources that may be programmatically provisioned and released in response to user commands. These resources may be dynamically provisioned and reconfigured to adjust to variable loads. Cloud computing may therefore be viewed as both applications delivered as services over publicly accessible networks (e.g., the Internet, cellular communication networks) and the hardware and software in cloud provider data centers that provide these services.

[0034] The cloud provider network provides users with the ability to use one or more of a variety of types of computing-related resources, such as computing resources (e.g., executing virtual machines and / or containers, executing batch jobs, executing code without provisioning servers), data / storage resources (e.g., object storage, block-level storage, data archive storage, databases and database tables, etc.), network-related resources (e.g., configuring virtual networks (including multiple groups of computing resources), content delivery networks (CDNs), domain name services (DNS)), application resources (e.g., databases, application build / deployment services), access policies or roles, identity policies or roles, machine images, routers, and other data processing resources, etc. These and other computing resources can be provided as services, such as hardware virtualization services that can manage virtual machines, storage services that can store data objects, etc. Users (or "customers") of the provider network 100 can use one or more user accounts associated with a customer account, but these items can be used interchangeably to a certain extent depending on the usage scenario. Users can also refer to other services or applications executing within the cloud provider network (e.g., a service or application executing on one virtual machine requests the launch of another virtual machine).

[0035] Users can connect to and interact with cloud provider network resources and services using various interfaces, typically application programming interfaces ("APIs"). Communications between users and the cloud provider network typically go through one or more intermediate networks (e.g., the public Internet). For example, user 238 of electronic device 234 can interact with cloud provider network 100 via intermediate network 236. The interaction can be performed via interface 204, such as using an API or command line, web-based, or other interface.

[0036] API refers to an interface and / or communication protocol between a client and a server such that if a client issues a request in a predefined format, the client should receive a response in a specific format or cause a defined action to be initiated. In the context of a cloud provider network, API provides a gateway to enable users to access the cloud infrastructure by allowing them to obtain data from the cloud provider network or cause actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted in the cloud provider network. APIs can also enable different services of the cloud provider network to exchange data with each other. Users may choose to deploy their virtual computing systems to provide network-based services for their own use and / or for use by their users or clients.

[0037] Cloud provider network 100 may include a physical network (e.g., sheet metal boxes, cables, rack hardware) called a substrate. The substrate can be thought of as a network fabric that contains the physical hardware that runs the provider network's services. The substrate can be isolated from the rest of cloud provider network 100, e.g., it may not be possible to route from a substrate network address to an address in a production network running a cloud provider service, or to a user network hosting user resources.

[0038] The cloud provider network 100 may also include various overlay networks of virtualized computing resources running on the bottom layer. In at least some examples, a hypervisor or other device or process on the network bottom layer may use encapsulation protocol technology to encapsulate and route network packets (e.g., client IP packets) between client resource instances on different hosts within the provider network through the network bottom layer. Encapsulation protocol technology may be used on the network bottom layer to route encapsulated packets (also referred to as network bottom layer packets) between endpoints on the network bottom layer via an overlay network path or route. Encapsulation protocol technology may be viewed as providing a virtual network topology overlaid on the network bottom layer. Thus, network packets may be routed along the bottom layer network according to the construction in the overlay network (e.g., a virtual network that may be referred to as a virtual private cloud (VPC), a port / protocol firewall configuration that may be referred to as a security group). A mapping service (not shown) may coordinate the routing of these network packets. A mapping service may be a regional distributed lookup service that maps a combination of an overlay Internet Protocol (IP) and a network identifier to an underlying IP so that a distributed underlying computing device can find where to send the packet.

[0039] For example, each physical computer system may have an IP address in the underlying network. Hardware virtualization technology can enable multiple virtual computing environments to run simultaneously on a host computer, for example as a virtual machine (VM) on a computing server. The hypervisor or virtual machine manager (VMM) on the host allocates the hardware resources of the host among the various VMs executed on the host and monitors the execution of the VM. Each VM may be provided with one or more IP addresses in one or more overlay networks, and the VMM on the host may know the IP address of the VM on the host. The VMM (and / or other devices or processes on the network bottom layer) can use encapsulation protocol technology to encapsulate network packets (e.g., client IP packets) and route the network packets between virtualized resources on different hosts within the cloud provider network 100 through the network bottom layer. Encapsulation protocol technology can be used on the network bottom layer to route encapsulated packets between endpoints on the network bottom layer via an overlay network path or route. Encapsulation protocol technology can be viewed as providing a virtual network topology overlaid on the network bottom layer. In some examples, the encapsulation protocol technology includes a mapping service that maintains a mapping directory that maps IP overlay addresses (e.g., IP addresses visible to a user) to underlying IP addresses (IP addresses not visible to a user) that are accessible by various processes on the cloud provider network for routing packets between endpoints.

[0040] As shown, in various examples, the traffic and operations underlying the cloud provider network can be broadly divided into two categories: control plane traffic carried on the logical control plane 214 and data plane operations carried on the logical data plane 216. The data plane 216 represents the movement of user data through the distributed computing system, while the control plane 214 represents the movement of control signals through the distributed computing system. The control plane 214 typically includes one or more control plane components or services distributed on and implemented by one or more control servers 212. Control plane traffic typically includes management operations, such as establishing an isolated virtual network (or "virtual private cloud") for the user, monitoring resource usage and health, identifying a specific host on which to start a virtual machine, configuring additional hardware as needed, etc. The data plane 216 includes user resources (e.g., virtual machines, containers, block storage volumes, databases, file storage, etc.) implemented on the cloud provider network. Data plane traffic typically includes non-management operations, such as transmitting data to and from resources.

[0041] As shown, the data plane 216 may include one or more host computer systems 206, which may be bare metal (e.g., a single tenant) or may be virtualized by a hypervisor or VMM to run multiple VMs or microVMs for a user. These host computer systems 206 may support virtualized computing services of a cloud provider network, such as hardware virtualization services 110. In some examples, the virtualized computing services are part of the control plane 214, allowing users to issue commands via an interface (e.g., interface 204) to launch and manage computing instances of their applications.

[0042] Edge locations 202 provide the resources and services of cloud provider network 100 within a separate network, thereby extending the functionality of cloud provider network 100 to new locations (e.g., for reasons related to latency of communications with user devices, legal compliance, security, etc.). As indicated, such edge locations 202 may include cloud provider network managed edge locations 240 (e.g., formed by servers located in cloud provider managed facilities separate from those associated with cloud provider network 100), communications service provider edge locations 242 (e.g., formed by servers associated with communications service provider facilities), user managed edge locations 244 (e.g., formed by servers located locally at user or partner facilities), as well as other possible types of underlying extensions.

[0043] As shown in the example edge location 240, the edge location 202 may similarly include a logical separation between a control plane 218 and a data plane 220, which extend the control plane 214 and the data plane 216 of the cloud provider network 100, respectively. In some examples, the edge location 202 is pre-configured, for example, by the cloud provider network operator with an appropriate combination of hardware and software and / or firmware elements to support various types of computing-related resources, and to do so in a manner that reflects the experience of using the cloud provider network. For example, one or more edge location servers may be configured by the cloud provider to be deployed within the edge location 202. As described above, in some examples, the cloud provider network 100 provides a set of predefined instance types, each instance type having a different type and amount of underlying hardware resources. Various sizes of each instance type may also be provided. In order to enable users to continue to use the same instance types and sizes that they use in the region in the edge 202, the servers may be heterogeneous servers. Heterogeneous servers may support multiple instance sizes of the same type at the same time, and may also be reconfigured to host any instance type supported by its underlying hardware resources. The reconfiguration of heterogeneous servers can occur instantly using the available capacity of the servers, i.e., while other VMs are still running and consuming other capacity of the edge location servers. This can improve the utilization of computing resources within the edge location by allowing better packaging of running instances on the servers, and also provide a seamless experience with regard to the use of instances on the cloud provider network 100 and the cloud provider network edge locations.

[0044] As shown, the edge server may host one or more computing instances 222. The computing instance 222 may be a VM, or a container that packages code and dependencies, so that the application can run quickly and reliably across computing environments (e.g., including VMs). In addition, if the user needs it, the server may host one or more data volumes 224. In the region of the cloud provider network 100, such volumes may be hosted on dedicated block storage servers. However, due to the possibility of having significantly smaller capacity at the edge location 202 than in the region, if the edge location includes such a dedicated block storage server, it may not provide the best utilization experience. Therefore, the block storage service may be virtualized in the edge location 202, so that one of the VMs runs the block storage software and stores the data of the volume 224. Similar to the operation of the block storage service in the region of the cloud provider network 100, the volume 224 in the edge location 202 can be replicated to achieve persistence and availability. The volume can be configured in its own isolated virtual network in the edge location 202. The computing instance 222 and any volume 224 together constitute the data plane extension 220 of the provider network data plane 216 in the edge location 202.

[0045] In some embodiments, servers within edge locations 202 may host certain local control plane components 226, such as components that enable edge locations 202 to continue to operate in the event of a loss of connectivity back to cloud provider network 100. Examples of these components include a migration manager that can move compute instances 222 between edge location servers if necessary to maintain availability, and a key value data store that indicates where volume replicas are located. However, control plane 218 functionality for edge locations will typically remain in cloud provider network 100 to allow users to use as much of the resource capacity of the edge location as possible.

[0046] In some examples, server software running at edge location 202 may be designed by a cloud provider to run on a cloud provider underlay network, and the software may be able to run unmodified in edge location 202 by using a local network manager 228 to create a dedicated replica of the underlay network (“shadow underlay”) within the edge location. The local network manager 228 may run on edge location 202 servers and bridge the shadow underlay with edge location 202 network, for example, by acting as a VPN endpoint or endpoint between edge location 202 and proxies 230, 232 in cloud provider network 100 and by implementing a mapping service (for traffic encapsulation and decapsulation) to correlate data plane traffic (from data plane proxies) and control plane traffic (from control plane proxies) with the appropriate servers. By implementing a local version of the provider network's underlay-to-overlay mapping service, the local network manager 228 allows resources in edge location 202 to communicate seamlessly with resources in cloud provider network 100. In some embodiments, a single local network manager may perform these actions for all servers hosting compute instances 222 in edge location 202. In other embodiments, each of the servers hosting the compute instances 222 has a dedicated local network manager. In a multi-rack edge location, inter-rack communications can go through the local network managers, where the local network managers maintain open tunnels between each other.

[0047] The edge location may utilize secure networking tunnels through the edge location 202 network to the cloud provider network 100, for example, to maintain the security of user data while traversing the edge location 202 network and any other intermediate networks (possibly including the public Internet). Within the cloud provider network 100, these tunnels consist of virtual infrastructure components including isolated virtual networks (e.g., in an overlay network), control plane agents 230, data plane agents 232, and underlying network interfaces. In some examples, such agents may be implemented as containers running on computing instances. In some examples, each server in the edge location 202 hosting a computing instance may utilize at least two tunnels: one tunnel for control plane traffic (e.g., Constrained Application Protocol (CoAP) traffic), and one tunnel for encapsulated data plane traffic. A connectivity manager (not shown) within the cloud provider network manages their cloud provider network-side lifecycles, for example, by automatically provisioning these tunnels and their components when needed and maintaining them in a healthy operating state. In some examples, a direct connection between the edge location 202 location and the cloud provider network 100 may be used to control the data plane and conduct data plane communications. Compared to VPNs that go through other networks, direct connections provide constant bandwidth and more consistent network performance because its network path is relatively fixed and stable.

[0048] A control plane (CP) agent 230 may be configured in the cloud provider network 100 to represent a specific host in an edge location. The CP agent is an intermediary between the control plane 214 in the cloud provider network 100 and the control plane target in the control plane 218 of the edge location 202. That is, the CP agent 230 provides an infrastructure for tunneling management API traffic destined for edge location servers from the regional bottom layer to the edge location 202. For example, a virtualized computing service of the cloud provider network 100 may issue a command to the VMM of a server at the edge location 202 to start a computing instance 222. The CP agent maintains a tunnel (e.g., VPN) to the local network manager 228 of the edge location. The software implemented within the CP agent ensures that only well-formed API traffic leaves and returns to the bottom layer. The CP agent provides a mechanism to expose remote servers on the cloud provider bottom layer while still protecting the bottom layer security materials (e.g., encryption keys, security tokens) from leaving the cloud provider network 100. The unidirectional control plane traffic tunnel imposed by the CP agent also prevents any (potentially compromised) device from calling back to the bottom layer. The CP agent may be instantiated one-to-one with a server at the edge location 202, or may manage control plane traffic for multiple servers at the same edge location.

[0049] Data plane (DP) agents 232 may also be configured in the cloud provider network 100 to represent a specific server in the edge location 202. The DP agent 232 acts as a shadow or anchor point for the server and may be used by services within the cloud provider network 100 to monitor the health of the host (including its availability, used / idle computing and capacity, used / idle storage and capacity, and network bandwidth usage / availability). The DP agent 232 also allows isolated virtual networks to span the edge location 202 and the cloud provider network 100 by acting as an agent for servers in the cloud provider network 100. Each DP agent 232 may be implemented as a packet forwarding computing instance or container. As shown, each DP agent 232 may maintain a VPN tunnel with a local network manager 228, which manages traffic to the server represented by the DP agent 232. The tunnel may be used to send data plane traffic between edge location servers and the cloud provider network 100. Data plane traffic flowing between the edge location 202 and the cloud provider network 100 may pass through the DP agent 232 associated with the edge location. For data plane traffic flowing from edge location 202 to cloud provider network 100, DP proxy 232 can receive the encapsulated data plane traffic, verify its correctness, and allow it to enter cloud provider network 100. DP proxy 232 can forward the encapsulated traffic from cloud provider network 100 directly to edge location 202.

[0050] The local network manager 228 may provide secure network connectivity with the proxies 230, 232 established in the cloud provider network 100. After the connection between the local network manager 228 and the proxies has been established, the user may issue commands via the interface 204 to instantiate (and / or perform other operations using) computing instances using edge location resources in a manner similar to the manner in which such commands would be issued for computing instances hosted within the cloud provider network 100. From the user's perspective, the user may now seamlessly use local resources within the edge location (as well as resources located in the cloud provider network 100, if desired). The computing instances provided on the servers at the edge location 202 may communicate with electronic devices located in the same network, as well as other resources provided in the cloud provider network 100 as needed. The local gateway 246 may be implemented to provide network connectivity between the edge location 202 and a network associated with the extension (e.g., a communications service provider network in the example of the edge location 242).

[0051] Figure 3An exemplary cloud provider network including geographically dispersed edge locations according to some examples is illustrated. The cloud provider network may be formed into multiple regions, where a region is a geographic area in which a cloud provider clusters data centers. Each region may include multiple (e.g., two or more) availability zones (AZs) connected to each other via a private high-speed network (e.g., a fiber optic communication connection). Thus, an AZ (also referred to as an availability domain) is a type of "deployment zone" that provides an isolated fault domain including one or more data center facilities that have separate power, separate networking, and separate cooling relative to a data center facility in another AZ. A data center refers to a physical building or enclosure that houses the servers of a cloud provider network and provides power and cooling to the servers of the cloud provider network. Preferably, AZs within a region are located far enough apart from each other so that a natural disaster (or other event that causes a failure) does not affect more than one AZ simultaneously or take more than one AZ offline. As Figure 3 As shown, a cloud provider network 300 (e.g., cloud provider networks 100, 200) is formed into multiple regions 312, where a region is a separate geographic area in which a cloud provider has one or more data centers 304. Each region 312 may include two or more AZs (not shown) connected to each other via a private high-speed network such as, for example, a fiber optic communication connection. Of course, the cloud provider network can be extended to a global scale outside the United States.

[0052] Users can connect to AZs of a cloud provider network via publicly accessible networks (e.g., the Internet, cellular communication networks), for example, through a switching center (TC). TCs are the main backbone locations that link users to the cloud provider network and can be co-located at other network provider facilities (e.g., Internet service providers (ISPs), telecommunications providers) and securely connected to AZs (e.g., via virtual private network (VPN) tunnels or direct connections). Each region can operate two or more TCs for redundancy. The regions are connected to a global network that includes a private networking infrastructure (e.g., a fiber connection controlled by a cloud provider) that connects each region to at least one other region. The cloud provider network can deliver content from access points (or "POPs") located outside of these regions but networked to these regions through edge locations and regional edge cache servers. This partitioning and geographic distribution of computing hardware enables the cloud provider network to provide users with low-latency resource access with a high degree of fault tolerance and stability on a global scale.

[0053] Compared to the number of regional data centers or AZs, the number of edge locations 316 can be much higher. This widespread deployment of edge locations 316 can provide low-latency connectivity to the cloud for a much larger group of end-user devices (compared to those end-user devices that happen to be very close to regional data centers). In some examples, each edge location 316 may be peered or "belong" to a portion of the cloud provider network 300 (e.g., a parent availability zone or regional data center). This peering allows various components operating in the cloud provider network to manage computing resources at edge locations. In some cases, multiple edge locations are located or installed in the same facility (e.g., a separate rack of a computer system) and are managed by different zones or data centers to provide additional redundancy. It should be noted that although edge locations are generally depicted herein as being within a communications service provider ("CSP") network, in some cases, such as when a cloud provider network facility is relatively close to a communications service provider facility, the edge location may remain within the physical premises of the cloud provider network while being connected to the communications service provider network via optical fiber or other network links.

[0054] The parenting of a given edge location as an AZ or region of a cloud provider network can be based on many factors. One such parenting factor is data sovereignty. For example, in order to keep data originating from a CSP network in a certain country in the same country, an edge location deployed in the CSP network can be a parent of an AZ or region within the country. Another factor can be the availability of services. For example, some edge locations may have different hardware configurations, such as whether there are components, such as local non-volatile storage devices (e.g., solid-state drives), graphics accelerators, etc. for user data. Some AZs or regions may lack services that utilize these additional resources, so edge locations can be parents of AZs or regions that support the use of those resources. Another factor can be the latency between an AZ or region and an edge location. Although the deployment of edge locations within a CSP network has latency benefits, those benefits can be offset by using edge locations as parents of distant AZs or regions, which introduces significant latency to edge location to regional traffic. Therefore, edge locations are typically parents of nearby (in terms of network latency) AZs or regions.

[0055] Figure 4An exemplary system in which a cloud provider network edge location is deployed within a communication service provider network according to some examples is illustrated. The CSP network 400 typically includes a downstream interface to an end-user electronic device and an upstream interface to other networks (e.g., the Internet). In this example, the CSP network 400 is a wireless "cellular" CSP network, which includes a radio access network (RAN) 402, 404, an aggregation site (AS) 406, 408, and a core network (CN) 410. The RAN 402, 404 includes base stations (e.g., NodeB, eNodeB, gNodeB) that provide wireless connectivity to electronic devices 412. The core network 410 typically includes functionality related to the management of the CSP network (e.g., billing, mobility management, etc.) and transmission functionality that relays traffic between the CSP network and other networks. Aggregation sites 406, 408 can be used to integrate traffic from many different radio access networks into the core network and direct traffic originating from the core network to various radio access networks.

[0056] End user electronic device 412 Figure 4 4 shows a base station (or radio base station) 414 wirelessly connected to the radio access network 402 from left to right in the figure. Such electronic devices 412 are sometimes referred to as user equipment (UE) or customer premises equipment (CPE). Data traffic is typically routed to the core network 410 through a fiber optic transmission network consisting of multiple hops of layer 3 routers (e.g., at aggregation sites). The core network 410 is typically housed in one or more data centers. For data traffic to locations outside the communication network 400, the network components 422-426 typically include firewalls through which traffic can enter or leave the CSP network 400 to reach external networks, such as the Internet or cloud provider network 100. Note that in some examples, the CSP network 400 may include facilities that permit traffic to enter or leave from sites further downstream in the core network 410 (e.g., at an aggregation site or RAN).

[0057] Edge locations 416-420 (or "wavelength zones") include computing resources that are managed as part of the cloud provider network but are installed or located within various points of the CSP network (e.g., locally in space owned or leased by the CSP). Computing resources generally provide a certain amount of computing and memory capacity that the cloud provider can allocate for use by its users. Computing resources may also include storage and accelerator capacity (e.g., solid-state drives, graphics accelerators, etc.). Here, edge locations 416, 418, and 420 are in communication with the cloud provider network 100.

[0058] Typically, the farther the edge location is from the cloud provider network 100 (or closer to the electronic device 412), for example, in terms of network hops and / or distance, the lower the network latency between the computing resources within the edge location and the electronic device 412. However, physical site constraints typically limit the amount of edge location computing capacity that can be installed at various points within the CSP, or determine whether computing capacity can be installed at all at various points. For example, an edge location located within the core network 410 can typically have a much larger footprint (in terms of physical space, power requirements, cooling requirements, etc.) than an edge location located within the RAN 402, 404.

[0059] The installation or location of edge locations within a CSP network may vary depending on the specific network topology or architecture of the CSP network. Figure 4 As indicated, edge locations are generally connectable to any location where a CSP network can interrupt packet-based traffic (e.g., IP-based traffic). Additionally, communications between a given edge location and cloud provider network 100 are generally securely transited through at least a portion of CSP network 400 (e.g., via a secure tunnel, a virtual private network, a direct connection, etc.). In the example shown, network component 422 facilitates routing of data traffic to and from edge location 416 integrated with RAN 402, network component 424 facilitates routing of data traffic to and from edge location 418 integrated with AS 406, and network component 426 facilitates routing of data traffic to and from edge location 420 integrated with CN 410. Network components 422-426 may include routers, gateways, or firewalls. To facilitate routing, the CSP may assign one or more IP addresses from the CSP network address space to each of the edge locations.

[0060] In the 5G wireless network development work, edge locations can be considered as a possible implementation of multi-access edge computing (MEC). Such edge locations can be connected to various points in the CSP 5G network, which provide interruptions for data traffic as part of the user plane function (UPF). Older wireless networks can also contain edge locations. For example, in 3G wireless networks, edge locations can be connected to the packet-switched network portion of the CSP network, such as to the serving general packet radio service support node (SGSN) or to the gateway general packet radio service support node (GGSN). In 4G wireless networks, edge locations can be connected to a serving gateway (SGW) or a packet data network gateway (PGW) as part of the core network or evolved packet core (EPC).

[0061] In some examples, traffic between edge locations 428 and cloud provider network 100 may escape CSP network 400 without being routed through core network 410. For example, network element 430 of RAN 404 may be configured to route traffic between edge location 416 of RAN 404 and cloud provider network 100 without traversing aggregation site or core network 410. As another example, network element 431 of aggregation site 408 may be configured to route traffic between edge location 432 of aggregation site 408 and cloud provider network 100 without traversing core network 410. Network elements 430, 431 may include gateways or routers having routing data to direct traffic from edge locations destined for cloud provider network 100 to cloud provider network 100 (e.g., via direct connection or intermediate network 434) and to direct traffic from cloud provider network 100 destined for edge locations to edge locations.

[0062] In some examples, an edge location may be connected to more than one CSP network. For example, when two CSPs share or route traffic through a common point, the edge location may be connected to two CSP networks. For example, each CSP may allocate some portion of its network address space to the edge location, and the edge location may include a router or gateway that may differentiate between traffic exchanged with each of the CSP networks. For example, traffic from one CSP network destined for an edge location may have a different destination IP address, source IP address, and / or virtual local area network (VLAN) tag than traffic received from another CSP network. Traffic originating from the edge location to a destination on one of the CSP networks may be similarly encapsulated with the appropriate VLAN tag, source IP address (e.g., from a pool allocated to the edge location from the destination CSP network address space), and destination IP address.

[0063] Note that although Figure 4 An exemplary CSP network architecture includes a radio access network, an aggregation site, and a core network, but the naming and structure of the architecture of the CSP network may vary between generations of wireless technologies, between different CSPs, and between wireless CSP networks and fixed-line CSP networks. Figure 4 Several locations where edge locations may be located within a CSP network are illustrated, but other locations are possible (eg, at a base station).

[0064] Figure 5Exemplary components of edge locations within a cloud provider network and a CSP network and connectivity therebetween are illustrated in greater detail according to some examples. Edge locations 500 provide resources and services of the cloud provider network within a CSP network 502, thereby extending the functionality of the cloud provider network 100 closer to end-user devices 504 connected to the CSP network.

[0065] Similar to the case of edge location 202, edge location 500 similarly includes a logical separation between a control plane 506 and a data plane 508, which respectively extend the control plane 214 and the data plane 216 of the cloud provider network 100. Edge location 500 can be pre-configured, for example, by a cloud provider network operator with an appropriate combination of hardware and software and / or firmware elements to support various types of computing-related resources, and to do so in a manner that reflects the experience of using a cloud provider network. The computer system of edge location 500 can host a control plane component 514, a local network manager 518, volumes 524, and computing instances 512 (e.g., virtual machines).

[0066] A local gateway 516 may be implemented to provide network connectivity between the edge location 300 and the CSP network 502. The cloud provider may configure the local gateway 516 with an IP address on the CSP network 502 and exchange routing data (e.g., via Border Gateway Protocol (BGP)) with the CSP network components 520. The local gateway 516 may include one or more routing tables that control the routing of inbound traffic to the edge location 500 and outbound traffic away from the edge location 500. The local gateway 516 may also support multiple VLANs in situations where the CSP network 502 uses separate VLANs for different portions of the CSP network 502 (e.g., one VLAN tag for wireless networks and another VLAN tag for fixed networks).

[0067] In some examples of edge locations 500, the extension includes one or more switches, sometimes referred to as top-of-rack (ToR) switches (e.g., in rack-based examples). The ToR switches are connected to CSP network routers (e.g., CSP network components 520), such as provider edge (PE) or software-defined wide area network (SD-WAN) routers. Each ToR switch may include an uplink link aggregation (LAG) interface to the CSP network router, each LAG supporting multiple physical links (e.g., 1G / 10G / 40G / 100G). The links may run Link Aggregation Control Protocol (LACP) and be configured as IEEE802.1q trunks to enable multiple VLANs on the same interface. This LACP-LAG configuration allows the edge location management entity of the control plane of the cloud provider network 100 to add more peer links to the edge location without adjusting routing. Each of the ToR switches may establish an eBGP session with a carrier PE or SD-WAN router. The CSP may provide a dedicated autonomous system number (ASN) for the edge location and the ASN of the CSP network 502 to facilitate the exchange of routing data.

[0068] Data plane traffic originating from edge location 500 may have multiple different destinations. For example, traffic addressed to a destination in data plane 216 of cloud provider network 100 may be routed via a data plane connection between edge location 500 and cloud provider network 100. Local network manager 518 may receive a packet from compute instance 512 that is addressed to, for example, another compute instance in cloud provider network 100, and encapsulate the packet with a destination that is the underlying IP address of a server that hosts the other compute instance, and then send it to cloud provider network 100 (e.g., via a direct connection or a tunnel). For traffic from compute instance 512 that is addressed to another compute instance hosted in another edge location 522, local network manager 518 may encapsulate the packet with a destination that is the IP address assigned to the other edge location 522, thereby allowing CSP network component 520 to handle routing of the packet. Alternatively, if the CSP network component 520 does not support inter-edge location traffic, the local network manager 518 may address the packet to a repeater in the cloud provider network 100, which may send the packet to another edge location 522 via its data plane connection (not shown) to the cloud provider network 100. Similarly, for traffic from a compute instance 512 address to a location outside the CSP network 502 or the cloud provider network 100 (e.g., on the Internet), if the CSP network component 520 allows routing to the Internet, the local network manager 518 may encapsulate the packet with a source IP address corresponding to an IP address in the operator address space assigned to the compute instance 512. Otherwise, the local network manager 518 may send the packet to an Internet gateway in the cloud provider network 100, which may provide Internet connectivity for the compute instance 512. For traffic from a compute instance 512 addressed to an electronic device 504, the local gateway 516 may use network address translation (NAT) to change the source IP address of the packet from an address in the address space of the cloud provider network to an address in the operator network space.

[0069] The local gateway 516, local network manager 518, and other local control plane components 514 may run on the same server that hosts the compute instance 512, may run on a dedicated processor integrated with an edge location server (e.g., on an offload card), or may be executed by servers separate from those hosting user resources.

[0070] Back to Figure 1 , hardware virtualization service 110 manages the hosting of computing instances (such as virtual machines) by host computer systems within a cloud provider network (e.g., host computer system 150) or at a cloud provider network edge location (e.g., host computer system 160).

[0071] As shown, host computer system 150 has some set of hardware resources 152, and host computer system 160 has some set of hardware resources 162. Hardware resources may include one or more processors or central processing units (CPUs), memory (e.g., system memory, storage devices), network adapters, and other hardware (such as graphics accelerators, signal processors, etc.). Given that the number of hosts across a cloud provider network and its edge locations may be very large, and many of them have different hardware configurations, various techniques may be employed to track the availability of hardware resources.

[0072] To recap, a cloud provider network can offer virtual machines characterized by instance "types." Various instance types can offer different levels of CPU, memory, and network capacity, and can include other features such as storage capacity, special hardware access, etc. For example, one virtual machine type might have two virtual processors, 8GB of memory, 10GB / sec of network throughput, while another virtual machine type might have four virtual processors, 16GB of memory, 15GB / sec of network throughput, and an attached accelerator card (e.g., graphics processor, signal processor, etc.).

[0073] In parallel, each host computer system may logically divide its associated hardware resources into multiple slots, where each slot is associated with one or more instance types and represents an amount of the host resources that will be allocated or assigned to the instance that "fills" the slot. In some examples, each host has an associated template (typically tracked in the control plane) that defines the slots on that host. Different templates can apportion the resources of the host in different ways. By way of example, assume that there are two instance types, small and large, and that the resource allocation of the large instance type is twice that of the small instance type. If a particular host computer system has the hardware resources to support two large instance types, various templates may include slots for two large instances, one large instance and two small instances, and four small instances.

[0074] Hardware virtualization service 110 may use host resource allocation data 111 to track the availability of host hardware resources that may be allocated to virtual machines. For each host computer system, host resource allocation data 111 may have identification of multiple slots (e.g., such as, for example, slots defined by a template associated with the host computer system) and an associated state identifier for each slot. In some examples, the state identifier may be an indication of whether the slot is used or available (e.g., "1" and "0"), and host resource allocation data 111 may also include identification of instances to which resources associated with the slot have been allocated (if necessary). For example, assume that each of host computer systems 150, 160 is associated with a template that divides them into four equally sized slots for use with instances of a given type. As shown in FIG. Figure 1 , the slots are based on a template in which the populated slots correspond to the occupied instances. Before any additional instances are started, the initial state of host computer system 150 is a single occupied slot, while the initial state of host computer system 160 is all four occupied slots. Hardware virtualization service 110 may store host resource allocation data 111 in a database or other data storage area, the host resource allocation data being as follows:

[0075]

[0076] In the above example, host identifier "1234" corresponds to host computer system 150, while host identifier "5678" corresponds to host computer system 160. An instance identifier corresponds to an instance of an allocated portion of a host system resource represented by a corresponding slot.

[0077] In some examples, the empty state identifier may indicate that the slot is available, while the non-empty state identifier may include an identification of an instance that has allocated a resource associated with the slot. Using the same scenario as before, host resource allocation data 111 may be stored in a database or other data storage area, the host resource allocation data is as follows:

[0078]

[0079] Two exemplary resource-sharing virtual machine startup workflows are now described. In the first resource-sharing virtual machine startup workflow, the hardware virtualization service 110 receives a startup request from an electronic device 101, as indicated by circle (1), which can be operated by a user 108. In this example workflow, the startup request includes an indication of an existing virtual machine with which the requested virtual machine will share resources. For example, the request may include an identification of a virtual machine 168-1, which is an independent virtual machine (e.g., operating with all resource allocations available to virtual machines of that type) prior to starting the requested virtual machine.

[0080] The hardware virtualization service 110 can determine the identity (e.g., IP address) of the host computer system 160 that hosts the virtual machine identified in the instance tracking data (not shown), which provides a lookup of the virtual machine identifier of its current host computer system. Here, the hardware virtualization service 110 identifies the host computer system 160, which currently hosts four virtual machines, as indicated by the host resource allocation data 111.

[0081] Once the host computer system is identified, the hardware virtualization service 110 may send a request to the VMM 166 (or other hypervisor, agent) of the host computer system 160 that manages the hosted virtual machine to cause the resource-sharing VM to be started, as indicated by circle (2). The request may include an identification of the instance with which the resource-sharing VM will share resources. The request may also include an identification of a machine image used to start the virtual machine and / or networking configuration data (such as whether to attach an existing or new elastic network interface, as described below, the addressing configuration of the new elastic network interface, etc.). In the event that the identified host computer system is part of an edge location, the hardware virtualization service 110 may send the request to the edge location 199 via a secure tunnel through one or more intermediate networks (not shown). In some examples, the request is sent via a tunnel dedicated to control plane traffic. Reference Figures 2 to 5 Provides additional details about the connectivity between the cloud provider network and the edge locations.

[0082] VMM 166 may then start the resource-sharing VM as a new VM (e.g., process) 168-2 within host operating system 164. VMM 166 may then configure host operating system 164 to share the resources originally allocated to VM 168-1 among VMs 168-1 and 168-2 (both of which are now resource-sharing virtual machines). At a high level, the host operating system may include one or more software applications or tools that generally ensure that the two resource-sharing VMs compete for the same processing resources, memory resources, networking resources, etc., without interfering with the operation of any other VMs 169 hosted by host computer system 160. Figure 6 Figures 8 through 8 provide additional details regarding various resource sharing techniques.

[0083] Typically, VMM 166 will provide a positive response to the request at circle (2) to inform hardware virtualization service 110 that the requested resource sharing VM has been successfully started. At circle (3), hardware virtualization service 110 may update host resource allocation data 111 to reflect the resource sharing instance. Figure 3 As graphically illustrated in FIG, the leftmost slot of the host computer system 160 now has two virtual machines associated therewith. Using the above data structure, the hardware virtualization service 110 may update the host resource allocation data 111 such as described in the above example, the host resource allocation data being as follows:

[0084]

[0085]

[0086] In the above updated host resource allocation data 111 example, slot 0 corresponds to the slot originally assigned to the virtual machine with identifier BCDE, which is now a resource-sharing virtual machine together with the virtual machine with identifier FABC.

[0087] In the second resource sharing virtual machine startup workflow, hardware virtualization service 110 receives a startup request from electronic device 101, as indicated by circle (4), which is operable by user 108. In this example workflow, the startup request is for two resource sharing virtual machines that have not yet been started and include one type of virtual machine.

[0088] The hardware virtualization service 110 may determine one or more candidate host systems on which to launch the requested instance. Since the request is for a resource-sharing VM, the hardware virtualization service 110 may identify one or more host computer systems with one or more available slots, which represent available host system resources for an independent virtual machine of the requested type. The hardware virtualization service 110 may then select one of the identified host computer systems. In this example, the hardware virtualization service 110 identifies and selects the host computer system 150, which currently hosts a single virtual machine (as indicated by the rightmost slot of the host resource allocation data 111).

[0089] The hardware virtualization service 110 may then send a request to the VMM 156 (or other hypervisor, agent) of the host computer system 150 that manages the hosted virtual machines to cause the resource-sharing VM to be started, as indicated by circle (5). The request at circle (5) may include hardware configuration data, such as the number of processor cores allocated based on the virtual machine type, the amount of memory allocated or otherwise limited based on the virtual machine type, whether any special hardware is attached, etc. The VMM 156 may use the hardware configuration data to allocate resources associated with one virtual machine of the identified virtual machine type to the two resource-sharing virtual machines of the requested virtual machine type.

[0090] The request at circle (5) may also include the identification of the machine image used to start the two resource-sharing virtual machines or the identification of the two machine images (if different). The request at circle (5) may also include networking configuration data (such as whether to attach an existing or new elastic network interface, as described below, the addressing configuration of the new elastic network interface, etc.).

[0091] VMM 156 may then launch the resource-sharing VMs as two new VMs (e.g., processes) 158-1 and 158-2 within host operating system 164. VMM 156 may then configure host operating system 154 to share resources associated with a single virtual machine of the specified type among resource-sharing VMs 158-1 and 158-2 of that type.

[0092] Typically, VMM 156 will provide a positive response to the request at circle (5) to inform hardware virtualization service 110 that the requested resource sharing VM has been successfully started. At circle (6), hardware virtualization service 110 may update host resource allocation data 111 to reflect the resource sharing instance. Figure 3 As graphically illustrated in FIG, the leftmost slot of the host computer system 150 now has two virtual machines associated therewith. Using the above data structure, the hardware virtualization service 110 may update the host resource allocation data 111 such as described in the above example, which is as follows:

[0093]

[0094] In the above updated host resource allocation data 111 example, slot 0 corresponds to the leftmost originally empty slot of the host computer system 1234 and now indicates the allocation of host system resources between two resource sharing virtual machines (identifiers AB12 and AB34).

[0095] More generally, a launch request, such as indicated by circles (1) and (4), may originate from within or outside of cloud provider network 100 (e.g., from electronic device 101, as shown, or from another virtual machine or service in or associated with the cloud provider network). The request may be sent on behalf of a user or customer, such as at the direction of a customer controlling the source of the request or by another service (such as a hosted service) performing operations for the customer.

[0096] Although the above examples consider a pair of resource-sharing virtual machines (such as VM 158 and VM 168), additional resource-sharing VMs using the same resource allocation may be launched using the techniques described herein, thereby forming a resource-sharing group of two or more virtual machines.

[0097] In some examples, the process of allocating resources to a virtual machine includes starting a virtual machine process and then limiting the virtual machine process (or multiple resource-sharing virtual machine processes) to a portion of the total amount of resources of the host computer system based on the type of virtual machine. Similarly, starting a virtual machine that shares resources with an existing virtual machine includes starting a new virtual machine and then limiting the new virtual machine to the same portion of the total amount of resources of the host computer system that the original virtual machine was allowed to use.

[0098] In some examples, the request at circles (1) and (4) may include an indication of whether to place a started virtual machine (e.g., one of the newly started VM 168-2, VM 158-1, and 158-2) in a paused or suspended state after startup. Such an indication is useful in use cases such as where one of the resource-sharing virtual machines is used as a failover backup, but it is not needed until the primary virtual machine fails. The hardware virtualization service 110 may pass such an indication to the VMM of the host computer system to cause the VMM to suspend the virtual machine after startup. At circles (2) and (5), the VMM is caused to place the newly started virtual machine in a paused state. Later, the hardware virtualization service 110 may receive a request to resume the suspended virtual machine, such as via the electronic device 101 from the user 108, from another client virtual machine hosted within the provider network that is monitoring the primary virtual machine, from a health monitoring service of the provider network, etc.

[0099] The hardware virtualization service 110 may also support resource sharing virtual machine termination. For example, the hardware virtualization service 110 may receive a request to terminate a resource sharing virtual machine, the request including an identifier of the virtual machine to be terminated. The hardware virtualization service 110 may determine the identifier of the host computer system that has identified the virtual machine in the hosting instance tracking data, as described above. The hardware virtualization service 110 may send a request to cause the VMM of the associated host computer system to terminate the identified virtual machine.

[0100] Figure 6 8 illustrate various resource sharing techniques. Various techniques can be used to share resources among resource sharing VMs and to limit resource usage of resource sharing VMs according to associated types. For example, Linux cgroups can be used to group VM processes that share resources.

[0101] Figure 6 An environment for sharing processing resources among resource sharing virtual machines according to some examples is illustrated. As shown, a host computer system 600 includes one or more processors 602, such as a CPU, a graphics accelerator, or other device. Processor 602 is a multi-core processor including a core 604. Host 600 executes a host operating system 610 including a VMM 611 and two resource sharing VMs 612-1, 612-2. Resource sharing VMs of associated types may be allocated a certain amount of processing resources. In this example, the VM type has a single core allocation. Therefore, a pair of resource sharing VMs share the single core.

[0102] After receiving a start request from the hardware virtualization service 110 (eg, Figure 1), VMM 611 may link or otherwise associate processes associated with the started VMs (e.g., whether one of VM 612 is started with an existing VM 612, or the two VMs 612 are started as a pair) with core 604-1. For example, VMM 611 may configure scheduler 614 of host operating system 610 to share core 604-1 between the two VMs 612. Scheduler 614 may then divide computing time on core 604-1 between the resource-sharing VMs 612.

[0103] Figure 7 An environment for sharing memory resources among resource-sharing virtual machines according to some examples is illustrated. As shown, a host computer system 700 includes a memory 702, such as a system memory. The host 700 executes a host operating system 710 including a VMM 711 and two resource-sharing VMs 712-1, 712-2. Resource-sharing VMs of associated types may allocate a certain amount of memory (e.g., 8 GB) for sharing.

[0104] After receiving a start request from the hardware virtualization service 110 (eg, Figure 1 , (2), (5) in FIG. 1 , the VMM 711 may configure the memory manager 714 of the host operating system 710 so that the total memory allocation to the resource-sharing VMs 712-1 and 712-2 is limited to a maximum limit.

[0105] In operation, VM 712-1 may request and release memory allocations from host memory 702 via memory manager 714 (memory allocations for VM 712-1 are indicated by diagonal fills). Similarly, VM 712-2 may request and release memory allocations from host memory 702 via memory manager 714 (memory allocations for VM 712-2 are indicated by hash fills). Memory manager 714 may track the total memory allocated to the resource-sharing VMs—as additional memory is added, memory manager 714 increases the total allocation by the amount of the increase; as allocated memory is released, memory manager 714 decreases the total allocation by the amount of the decrease. If one of the resource-sharing VMs requests an allocation that exceeds the difference between the maximum memory quota for the virtual machine type and the amount of memory currently allocated to the resource-sharing VM, memory manager 714 may return an insufficient memory or insufficient memory error.

[0106] In some examples, memory manager 714 can have a communication channel to a local memory manager agent 716 operating within the environment of VM 712. Memory manager 714 can communicate with agent 716 (e.g., a process within a guest operating system) to bias the proportional sharing of the maximum memory quota between resource-sharing virtual machines. For example, when launching a resource-sharing virtual machine, a user can include in the request an indication of which virtual machine to consider as the primary virtual machine or which virtual machine to consider as the secondary virtual machine, as well as an optional memory biasing factor.

[0107] The memory manager 714 may communicate with the agent 716 of the backup resource sharing virtual machine to cause the agent to request an amount of memory predefined for the backup VM or an amount of memory based on a memory bias factor within the associated VM environment. However, the agent does not use the requested allocation, so the memory manager 714 does not need to allocate a corresponding amount of memory from memory 702. Instead, the requested allocation reduces the amount of memory that other processes within the VM environment can access. For example, if both resource sharing VMs are configured with 8GB of memory (based on the associated instance type), the agent 716-1 in the VM environment of VM 712-1 may request 5GB from the memory manager 714, resulting in the VM environment of VM 712-1 having 3GB of remaining capacity. Since the agent does not use the allocation, it can be treated as "reserved" for the other VM environments of VM 712-2, thereby establishing a 5GB bottom line for VM 712-2.

[0108] In some examples, if VMM 711 receives a request from hardware virtualization service 110 to start a resource-sharing VM and another VM has already been allocated the maximum memory quota for the virtual machine type (or if the base memory footprint of the additional VM plus the quota of the existing VM would exceed the maximum), the VMM may deny the start request.

[0109] FIG. 8A to FIG. 8D An environment for sharing networking resources among resource sharing virtual machines 812-1 and 812-2 according to some examples is illustrated. To provide isolation between virtual machines or groups of virtual machines, the virtual machines communicate within a virtual network (such as a VPC). A virtual network interface (VNI) attaches the virtual machine to the VPC. The VNI has an associated configuration including, for example, a network address (e.g., an IP address) on the VPC.

[0110] like FIG. 8A to FIG. 8DAs shown by circle (1) in each example of , a network manager 814 (typically a component of a host operating system or VMM) limits the aggregate network throughput of the resource-sharing VMs 812 to the throughput limit associated with a single instance of the VM type. For example, the network manager 814 may use a token bucket algorithm that fills a bucket associated with the resource-sharing VMs at a rate associated with the throughput limit of a single instance of the VM type. When traffic is sent from any one of the resource-sharing VMs, the network manager 814 depletes the token bucket relative to the traffic sent. Thus, over time, if the rate at which two VMs 812 transmit over the network exceeds the typical throughput of a single VM of the corresponding type, that throughput will be split between the VMs. Those skilled in the art will be aware of other techniques for limiting and sharing network capacity.

[0111] Fig. 8A An example is shown in which a VNI is attached to each resource-sharing VM. VM 812-1 is connected to VNI 802-1, and VM 812-2 is connected to VNI 802-2. As shown, both VNIs connect their respective VMs to the same VPC, although in other examples, the VNIs may be connected to different VPCs.

[0112] Figure 8B A second example is illustrated in which a VNI is detached from one of the resource-sharing VMs and attached to another resource-sharing VM. Here, the network manager 814 has attached VNI 822-1 to VM 812-1 (VM 812-2 may have another attached VNI (not shown)). This example may be used in the case of a backup resource-sharing VM. As indicated by circle (2), the network manager 814 receives a switching signal. The switching signal typically originates from an external entity, such as an application owned by another customer, a hardware virtualization service, a health monitoring service, etc., and may be routed through another local entity (such as, a VMM). Regardless of its source, the switching signal changes the "backup" or "shadow" virtual machine to the primary virtual machine. In this example, the switching signal causes the network manager 814 to change the connection of VNI 822 from VM 812-1 to 812-2, as indicated by circle (3). The configuration of VNI 822 remains unchanged, so from the perspective of other entities on the VPC or other entities that previously communicated with VM 812-1, VM 812-2 now appears to be VM 812-1 (e.g., the network address of VM 812-1 moves with the VNI to VM 812-2).

[0113] Figure 8CA third example is illustrated in which a VNI is attached to each resource-sharing VM, where inbound traffic (from the VPC) is mirrored to the resource-sharing VM. Similarly, this example can be used in the case of backing up resource-sharing VMs. Here, the network manager 814 has attached VNI 842 to VM 812-1 and VM 812-2, mirroring inbound traffic to each VM 812, and packaging outbound traffic as coming from the same source. In this scenario, the backup VM is typically placed in an idle state (e.g., process suspended) by VMM 811. As indicated by circle (2), VMM 811 receives a switching signal, causing VMM 811 to change the process state of the VM, as indicated by circle (3). Specifically, VMM 811 can resume idle VMs and suspend running VMs.

[0114] Fig.8D A fourth example is illustrated in which resource-sharing VMs each have an attached VNI, and the configuration (e.g., IP address) of the VNI attached to the primary VM is switched to the configuration of the VNI attached to the backup VM. Here, the network manager 814 has attached VNI 862-1 to VM 812-1, and VNI 862-2 to VM 812-2. Similarly, this example can be used in the case of backing up resource-sharing VMs. As indicated by circle (2), the network manager 814 receives a switching signal. In this example, the switching signal causes the network manager 814 to update the configuration of VNI 862-2 to match one or more configuration parameters of VNI 862-1, such as a network address. Additionally, the network manager 814 can change the configuration of VNI 862-1 (e.g., change its network address). Based on the configuration update, VM 812-2 now appears to be VM 812-1 from the perspective of other entities on the VPC or other entities that previously communicated with VM 812-1 (eg, the network address of VM 812-1 moves to VM 812-2 with the configuration change).

[0115] Fig. 9 An environment based on healthy resource sharing virtual machine replacement according to some examples is illustrated. In this example, a host computer system 900 hosts two resource sharing VMs 912-1 and 912-2 that are usually started from the same machine image. The VMM of the host computer system 900 has suspended VM 912-2 (a shadow or backup VM), while VM 912-1 remains active (a primary VM).

[0116] The health monitoring service 950 of the cloud provider network can monitor the status of the VM. In order to monitor the status of the VM, the health monitoring service 950 can perform one or more checks. Such checks may include environmental checks (e.g., querying the VMM of the host system to understand the status of the monitored VM), activity checks (e.g., performing an echo check on the monitored VM to determine whether it is responsive), and custom checks (e.g., executing custom code to interact with the monitored VM in a prescribed manner). For each check, the health monitoring service 950 may have one or more rules to determine whether the check result indicates that there is a problem with the monitored VM.

[0117] A customer may request that a VM be monitored by health monitoring service 950. A health monitoring request typically includes an identification of the VM to be monitored (whether via an instance identifier, network address, etc.). In some examples, the request may also identify which checks are to be performed, including whether any custom checks are to be performed. If custom checks are requested, the customer may also provide or identify the code to be used to perform the custom checks.

[0118] As indicated by circle (1), health monitoring service 950 may perform one or more checks on the health of VM 912-1. Such checks may include querying VMM 902 to check whether VM 912-1 is not damaged and querying VM 912-1 itself to check activities / perform custom checks. At a certain point in time, based on the check responses, health monitoring service 950 determines that VM 912-1 is no longer healthy, as indicated by circle (2). Health monitoring service 950 may send a switch signal to VMM 902 to cause a failover to backup VM 912-2. As indicated by circle (4), VMM 902 may change the state of VM 912, including resuming execution of VM 912-2 and pausing or terminating VM 912-1. Failover may also include various network configuration changes, such as illustrated and described with reference to FIG. 8.

[0119] Go to Fig. 9 In the lower part of FIG. 1 , after the switch, the health monitoring service 950 may begin to obtain health status data from at least one of the VM 912-2 and the VMM 902 to monitor the health status of the now primary VM 912-2. In some examples, the original launch request (e.g., Figure 1912-2) may include parameters to enable persistent backups. In such cases, hardware virtualization service 110 may set corresponding parameters associated with the resource sharing VM in VMM 902 to cause VMM to start a new backup (e.g., 912-3) as a virtual machine that shares resources with the previous backup, now primary VM (e.g., VM 912-2), as indicated by circle (5). VMM 902 may start the new backup resource sharing VM from the same machine image used to start the previous backup resource sharing VM.

[0120] Fig.10 An environment for transmitting virtual machine images according to some examples is illustrated. Here, the host computer system 1052 hosts multiple VMs, including VM 1058-1. As indicated by the host resource allocation data, the host computer system 1052 does not have an available slot. In this scenario, the VMM 1054 of the host computer system typically rejects any additional launch requests. However, as indicated by circles (1) and (2), the hardware virtualization service 110 receives a request to launch a resource-sharing VM, the request including an identification of VM 1058-1. The hardware virtualization service 110 may identify the host computer system 1052 as a target for launch, and send a request to launch a resource-sharing VM to the VMM 1054, the request including an identification of a machine image from which to launch the requested resource-sharing virtual machine. As indicated by circle (3), the VMM 1054 may retrieve the identified machine image 1056 from the VM machine image (MI) storage device 1000 of the cloud provider network, and store the machine image 1056 in the local storage device 1054. As indicated at circle (4), VMM 1054 may launch resource-sharing VM 1058 - 2 from machine image 1056 .

[0121] refer to Fig.10 The illustrated and described scenarios can be used to facilitate rolling deployments, particularly at edge locations that may lack free capacity to launch additional virtual machines. For example, VM 1058-1 can remain executing while VMM 1054 launches VM 1058-2 with an updated version of an application. Thus, rather than terminating VM 1058-1 to free up capacity to launch VM 1058-2, the time associated with launching VM 1058-2 (including the time to transfer machine image 1056) can elapse while VM 1058-1 remains active. Once VMM 1054 has launched VM 1058-2, VMM 1054 can terminate VM 1058-1.

[0122] Fig.11Operation 1100 of a method for client-initiated virtual machine resource allocation sharing according to some examples is illustrated. Some or all of operation 1100 (or other processes described herein, or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors. The code is stored on a computer-readable storage medium, for example, in the form of a computer program including instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) operations in operation 1100 are performed by at least one of the host computer systems of the hardware virtualization service and / or other figures.

[0123] Operation 1100 includes, at block 1102, receiving, by a hardware virtualization service of a cloud provider network, a request to launch a first virtual machine, wherein the first virtual machine is of a first virtual machine type, the first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type. Operation 1100 also includes, at block 1104, causing, by the hardware virtualization service, launching of the first virtual machine on a host computer system of the cloud provider network. Operation 1100 also includes, at block 1106, sharing, by the host computer system, an allocation of the amount of resources in corresponding resources of the host computer system between the first virtual machine and the second virtual machine, wherein the second virtual machine is of the first virtual machine type. Various other operations of one or more entities illustrated and described herein may be performed, including the operations set forth in the claims.

[0124] Fig.12 An example provider network environment according to some examples is illustrated. Provider network 1200 may provide resource virtualization to customers via one or more virtualization services 1210, which allow customers to purchase, rent, or otherwise obtain instances 1212 of virtualized resources (including, but not limited to, computing resources and storage resources) implemented on devices within one or more provider networks in one or more data centers. A local Internet Protocol (IP) address 1216 may be associated with resource instance 1212; the local IP address is the internal network address of the resource instance 1212 on provider network 1200. In some examples, provider network 1200 may also provide public IP addresses 1214 and / or public IP address ranges (e.g., Internet Protocol version 4 (IPv4) or Internet Protocol version 6 (IPv6) addresses) that customers can obtain from provider 1200.

[0125] Conventionally, the provider network 1200 may allow a customer of the service provider (e.g., a customer operating one or more customer networks 1250A-1250C (or "client networks") including one or more customer devices 1252) via the virtualization service 1210 to dynamically associate at least some of the public IP addresses 1214 assigned or allocated to the customer with specific resource instances 1212 assigned to the customer. The provider network 1200 may also allow the customer to remap a public IP address 1214 previously mapped to one virtualized computing resource instance 1212 allocated to the customer to another virtualized computing resource instance 1212 also allocated to the customer. For example, a customer of a service provider (such as an operator of the customer networks 1250A-1250C) may use the virtualized computing resource instances 1212 and public IP addresses 1214 provided by the service provider to implement customer-specific applications and present the customer's applications over an intermediate network 1240 such as the Internet. Other network entities 1220 on the intermediate network 1240 may then generate traffic to the destination public IP address 1214 published by the customer networks 1250A to 1250C; the traffic is routed to the service provider data center and there, via the network underlay, to the local IP address 1216 of the virtualized computing resource instance 1212 that is currently mapped to the destination public IP address 1214. Similarly, response traffic from the virtualized computing resource instance 1212 may be routed back onto the intermediate network 1240 via the network underlay to reach the source entity 1220.

[0126] As used herein, a local IP address refers to an internal or "private" network address of a resource instance, for example, in a provider network. A local IP address may be within an address block reserved by Request for Comments (RFC) 1918 of the Internet Engineering Task Force (IETF) and / or have an address format specified by IETF RFC 4193, and may be variable within a provider network. Network traffic originating outside the provider network is not routed directly to a local IP address; instead, traffic uses a public IP address that is mapped to the local IP address of the resource instance. The provider network may include a networking device or apparatus that provides network address translation (NAT) or similar functionality to perform mappings from public IP addresses to local IP addresses and from local IP addresses to public IP addresses.

[0127] A public IP address is an Internet-mutable network address assigned to a resource instance by a service provider or customer. Traffic routed to a public IP address is translated, for example, via 1:1 NAT, and forwarded to the corresponding local IP address of the resource instance.

[0128] Some public IP addresses may be assigned to specific resource instances by the provider network infrastructure; these public IP addresses may be referred to as standard public IP addresses, or simply standard IP addresses. In some examples, the mapping of standard IP addresses to local IP addresses of resource instances is the default launch configuration for all resource instance types.

[0129] At least some public IP addresses may be assigned to or obtained by customers of the provider network 1200; then, the customer may assign its assigned public IP address to a specific resource instance assigned to the customer. These public IP addresses may be referred to as customer public IP addresses, or simply customer IP addresses. Rather than being assigned to a resource instance by the provider network 1200 as in the case of a standard IP address, a customer IP address may be assigned to a resource instance by a customer, for example, via an API provided by a service provider. Unlike a standard IP address, a customer IP address is assigned to a customer account and may be remapped to other resource instances by the corresponding customer as needed or desired. A customer IP address is associated with a customer account rather than a specific resource instance, and the customer controls the IP address until the customer chooses to release the IP address. Unlike conventional static IP addresses, a customer IP address allows a customer to mask a resource instance or availability zone failure by remapping the customer's public IP address to any resource instance associated with the customer account. For example, a customer IP address enables a customer to solve a problem with a customer's resource instance or software by remapping the customer IP address to an alternative resource instance.

[0130] Fig.13 An example provider network that provides storage services and hardware virtualization services to customers according to some examples is illustrated. Hardware virtualization service 1320 provides multiple computing resources 1324 (e.g., computing instances 1325 such as VMs) to customers. Computing resources 1324 can be provided as a service to customers of provider network 1300 (e.g., customers implementing customer network 1350), for example. Each computing resource 1324 can be provided with one or more local IP addresses. Provider network 1300 can be configured to route packets from the local IP addresses of computing resources 1324 to public Internet destinations and to route packets from public Internet sources to the local IP addresses of computing resources 1324.

[0131] The provider network 1300 may provide a customer network 1350 coupled to the intermediate network 1340, for example, via a local network 1356, with the ability to implement a virtual computing system 1392 via a hardware virtualization service 1320 coupled to the intermediate network 1340 and the provider network 1300. In some examples, the hardware virtualization service 1320 may provide one or more APIs 1302, such as web service interfaces, via which the customer network 1350 may access the functionality provided by the hardware virtualization service 1320, for example, via a console 1394 (e.g., a web-based application, a standalone application, a mobile application, etc.) of a customer device 1390. In some examples, at the provider network 1300, each virtual computing system 1392 at the customer network 1350 may correspond to a computing resource 1324 that is leased, rented, or otherwise provided to the customer network 1350.

[0132] A customer may access the functionality of the storage service 1310, for example, from an instance of a virtual computing system 1392 and / or another customer device 1390 (e.g., via a console 1394), via one or more APIs 1302, to access and store data from storage resources 1318A through 1318N of a virtual data store 1316 (e.g., folders or “buckets,” virtualized volumes, databases, etc.) provided by the provider network 1300. In some examples, a virtualized data storage gateway (not shown) may be provided at the customer network 1350, which may cache at least some data (e.g., frequently accessed data or critical data) locally and may communicate with the storage service 1310 via one or more communication channels to upload new or modified data from the local cache so that a primary data store (the virtualized data store 1316) is maintained. In some examples, a user via a virtual computing system 1392 and / or another client device 1390 may mount and access volumes of a virtual data repository 1316 via a storage service 1310 acting as a storage virtualization service, and these volumes may appear to the user as local (virtualized) storage 1398 .

[0133] Although Fig.13 Although not shown, the virtualization service may also be accessed from a resource instance within the provider network 1300 via API 1302. For example, a customer, device service provider, or other entity may access the virtualization service via API 1302 from within a corresponding virtual network on the provider network 1300 to request allocation of one or more resource instances within the virtual network or within another virtual network.

[0134] Descriptive System

[0135] In some examples, a system implementing part or all of the techniques described herein may include a general purpose computer system (such as Fig.14 The computer system 1400 shown in the figure is a general purpose computer system that includes or is configured to access one or more computer accessible media. In the illustrated example, the computer system 1400 includes one or more processors 1410 coupled to a system memory 1420 via an input / output (I / O) interface 1430. The computer system 1400 also includes a network interface 1440 coupled to the I / O interface 1430. Although Fig.14 Computer system 1400 is shown as a single computing device, but in various examples, computer system 1400 may include one computing device or any number of computing devices configured to work together as a single computer system 1400 .

[0136] In various examples, computer system 1400 may be a uniprocessor system including one processor 1410 or a multiprocessor system including several processors 1410 (e.g., two, four, eight, or another suitable number). Processor 1410 may be any suitable processor capable of executing instructions. For example, in various examples, processor 1410 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as, for example, an x86, ARM, PowerPC, SPARC, or MIPS ISAs or any other suitable ISAs. In a multiprocessor system, each of processors 1410 may typically, but not necessarily, implement the same ISA.

[0137] The system memory 1420 may store instructions and data that may be accessed by the processor 1410. In various examples, the system memory 1420 may be implemented using any suitable memory technology, such as random access memory (RAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash-type memory, or any other type of memory. In the illustrated example, program instructions and data that implement one or more desired functions (such as the methods, techniques, and data described above) are shown as being stored within the system memory 1420 as code 1425 (e.g., executable to implement, in whole or in part, the hardware virtualization service 110; an agent of the host computer system, such as a VMM or hypervisor; agent components, such as a scheduler, a memory manager, and a network manager; a health monitoring service; and other components depicted and described with reference to the above figures) and data 1426.

[0138] In some examples, I / O interface 1430 may be configured to coordinate I / O traffic between processor 1410, system memory 1420, and any peripheral devices in the device, including network interface 1440 and / or other peripheral interfaces (not shown). In some examples, I / O interface 1430 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., system memory 1420) into a format suitable for use by another component (e.g., processor 1410). In some examples, for example, I / O interface 1430 may include support for devices attached via various types of peripheral buses (such as, for example, a variant of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard). In some examples, for example, the functionality of I / O interface 1430 may be split into two or more separate components, such as a north bridge and a south bridge. In addition, in some examples, some or all of the functionality of I / O interface 1430 (such as an interface to system memory 1420) may be directly incorporated into processor 1410.

[0139] For example, the network interface 1440 may be configured to allow the computer system 1400 to communicate with other devices 1460 attached to one or more networks 1450 (such as, for example, Figure 1 In various examples, network interface 1440 may support communication via any suitable wired or wireless general purpose data network, such as various types of Ethernet networks, for example. Additionally, network interface 1440 may support communication via a telecommunications / telephone network, such as an analog voice network or a digital fiber optic communications network, via a storage area network (SAN), such as a Fibre Channel SAN, and / or via any other suitable type of network and / or protocol.

[0140] In some examples, computer system 1400 includes one or more offload cards 1470A or 1470B (including one or more processors 1475 and possibly one or more network interfaces 1440) connected using an I / O interface 1430 (e.g., a bus implementing a version of the Peripheral Component Interconnect-Express (PCI-E) standard or another interconnect such as Quick Path Interconnect (QPI) or Ultra Path Interconnect (UPI)). For example, in some examples, computer system 1400 may act as a host electronic device hosting computing resources such as computing instances (e.g., operating as part of a hardware virtualization service), and one or more offload cards 1470A or 1470B execute a virtualization manager that can manage computing instances executed on the host electronic device. As an example, in some examples, offload cards 1470A or 1470B may perform computing instance management operations, such as pausing and / or unpausing computing instances, starting and / or terminating computing instances, performing memory transfer / copy operations, etc. In some examples, these management operations may be performed by offload card 1470A or 1470B in cooperation with hypervisors (e.g., based on requests from the hypervisors) executed by other processors 1410A through 1410N of computer system 1400. However, in some examples, the virtualization manager implemented by offload card 1470A or 1470B may adapt to requests from other entities (e.g., from the computing instances themselves) and may not cooperate with (or serve) any individual hypervisor.

[0141] In some examples, system memory 1420 may be an example of a computer-accessible medium configured to store program instructions and data as described above. However, in other examples, program instructions and / or data may be received, sent, or stored on different types of computer-accessible media. In general, a computer-accessible medium may include any non-transitory storage medium or memory medium, such as a magnetic or optical medium, such as a disk or DVD / CD coupled to the computer system 1400 via an I / O interface 1430. A non-transitory computer-accessible storage medium may also include any volatile or non-volatile medium that may be included in some examples of the computer system 1400 as system memory 1420 or another type of memory, such as RAM (e.g., SDRAM, double data rate (DDR) SDRAM, SRAM, etc.), read-only memory (ROM), etc. In addition, a computer-accessible medium may include a transmission medium or a signal transmitted via a communication medium (such as a network and / or a wireless link), such as an electrical signal, an electromagnetic signal, or a digital signal, the communication medium being implemented, for example, via a network interface 1440.

[0142] The various examples discussed or proposed herein can be implemented in a variety of operating environments, in some cases, the operating environment may include one or more user computers, computing devices or processing devices that can be used to operate any of many applications. User equipment or client devices may include any of many general personal computers, such as desktop computers or laptop computers running standard operating systems, and cellular devices, wireless devices and handheld devices that run mobile software and can support many networking and messaging protocols. Such a system may also include many workstations, which run various commercially available operating systems and any of other known applications for purposes such as development and database management. These devices may also include other electronic devices, such as virtual terminals, thin clients, game systems and / or other devices that can communicate via a network.

[0143] Most examples use at least one network familiar to those skilled in the art to support communications using any of a variety of widely available protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), File Transfer Protocol (FTP), Universal Plug and Play (UPnP), Network File System (NFS), Common Internet File System (CIFS), Extensible Messaging and Presence Protocol (XMPP), AppleTalk, etc. The network may include, for example, a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), the Internet, an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network, and any combination thereof.

[0144] In the example of using a web server, the web server can run any of a variety of server or middle-tier applications, including HTTP servers, file transfer protocol (FTP) servers, common gateway interface (CGI) servers, data servers, Java servers, business application servers, etc. The server may also be capable of executing programs or scripts in response to requests from user devices, such as by executing a program that can be implemented in any programming language (such as The server may include one or more web applications written in one or more scripts or programs in C, C#, or C++, or any scripting language such as Perl, Python, PHP, or TCL, and combinations thereof. The server may also include a database server, including but not limited to commercially available database servers from Oracle (R), Microsoft (R), Sybase (R), IBM (R), etc. The database server may be relational or non-relational (e.g., "NoSQL"), distributed or non-distributed, etc.

[0145] The environment disclosed herein may include various data storage areas and other memories and storage media as discussed above. These may reside in a variety of locations, such as on a storage medium that resides locally (and / or resides therein) on one or more computers, or on a storage medium that resides away from any or all computers on a network. In a specific example set, information may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing functions belonging to a computer, server or other network device may be stored locally and / or remotely as appropriate. In the case where the system includes a computerized device, each such device may include a hardware element that can be electrically coupled via a bus, the element including, for example, at least one central processing unit (CPU), at least one input device (e.g., a mouse, keyboard, controller, touch screen or keypad) and / or at least one output device (e.g., a display device, printer or speaker). Such a system may also include one or more storage devices, such as a disk drive, an optical storage device and a solid-state storage device (such as a random access memory (RAM) or a read-only memory (ROM)), as well as a removable media device, a memory card, a flash memory card, etc.

[0146] Such devices may also include a computer-readable storage medium reader, a communication device (e.g., a modem, a network card (wireless or wired), an infrared communication device, etc.), and a working memory as described above. A computer-readable storage medium reader may be connected to a computer-readable storage medium or configured to receive a computer-readable storage medium, which represents a remote, local, fixed and / or removable storage device and a storage medium for temporarily and / or longer accommodating, storing, transmitting and retrieving computer-readable information. The system and various devices will also typically include many software applications, modules, services or other elements located in at least one working memory device, including an operating system and applications, such as a client application or a web browser. It should be understood that alternative examples may have many variations different from the above description. For example, custom hardware may also be used, and / or specific elements may be implemented in hardware, software (including portable software, such as applets), or both. In addition, connections with other computing devices (such as network input / output devices) may be employed.

[0147] Storage media and computer-readable media for containing code or code portions may include any suitable media known or used in the art, including storage media and communication media, such as but not limited to volatile and non-volatile media, removable and non-removable media implemented in any method or technology to store and / or transmit information (such as computer-readable instructions, data structures, program modules or other data), including RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk-read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other media that can be used to store the desired information and can be accessed by system devices. Based on the present disclosure and the teachings provided herein, a person of ordinary skill in the art will understand other ways and / or methods of implementing various examples.

[0148] In the foregoing description, various examples are described. For explanation purposes, specific configurations and details are set forth in order to provide a thorough understanding of the examples. However, it will be apparent to those skilled in the art that the examples may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified in order not to obscure the described examples.

[0149] Parenthesized text and boxes with dashed borders (e.g., large dashes, small dashes, dot dashes, and dots) are used herein to illustrate optional aspects that add additional features to some examples. However, this notation should not be taken to mean that these are the only options or optional operations, and / or that in some examples, boxes with solid borders are not optional.

[0150] Reference numerals with suffix letters (e.g., 1318A-1318N) may be used to indicate that there may be one or more instances of the referenced entity in various examples, and when there are multiple instances, each instance need not be identical, but may share some general characteristics or function in a common manner. Additionally, the particular suffix used is not meant to imply that there is a particular number of entities unless specifically indicated to the contrary. Thus, in various examples, two entities using the same or different suffix letters may or may not have the same number of instances.

[0151] Reference to "one example," "an example," etc. indicates that the example being described may include a particular feature, structure, or characteristic, but each example may not necessarily include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same example. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example, it should be considered that it is within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in conjunction with other examples, whether or not explicitly described.

[0152] Furthermore, in the various examples above, unless specifically noted otherwise, disjunctive language such as the phrase "at least one of A, B, or C" is intended to be understood to mean A, B, or C, or any combination thereof (e.g., A, B, and / or C). Similarly, language such as "at least one or more of A, B, and C" (or "one or more of A, B, and C") is intended to be understood to mean A, B, or C, or any combination thereof (e.g., A, B, and / or C). Thus, disjunctive language is neither intended nor should be understood to imply that a given example requires that at least one of A, at least one of B, and at least one of C each be present.

[0153] As used herein, the term "based on" (or similar) is an open-ended term used to describe one or more factors that influence a determination or other action. It should be understood that this term does not exclude additional factors that may influence the determination or action. For example, a determination may be based only on the factors listed or on the factors and one or more additional factors. Thus, if action A is "based on" B, it should be understood that B is a factor that influences action A, but this does not exclude that the action is also based on one or more other factors, such as factor C. However, in some cases, action A may be based entirely on B.

[0154] Unless expressly stated otherwise, articles such as "a" or "an" should generally be interpreted as including one or more of the described items. Thus, phrases such as "a device configured to..." or "computing device" are intended to include one or more of the described devices. The one or more described devices may be collectively configured to perform the stated operations. For example, "a processor configured to perform operations A, B, and C" may include a first processor configured to perform operation A working together with a second processor configured to perform operations B and C.

[0155] In addition, the words "may" or "can" are used in a permissive sense (i.e., meaning to have a certain possibility), rather than in a mandatory sense (i.e., meaning to have to). The words "include", "including" and "includes" are used to indicate an open relationship, and therefore mean including but not limited to. Similarly, the words "have", "having" and "has" also indicate an open relationship, and therefore mean having but not limited to. The terms "first", "second", "third" etc. used in this article are used as labels for the nouns following them, and do not imply any type of ordering (e.g., space, time, logic, etc.), unless such ordering is clearly indicated in addition. Similarly, the values ​​of such digital labels are generally not used to indicate the required amount of a specific noun in the claims cited herein, and therefore the "fifth" element generally does not mean that there are other four elements, unless these elements are clearly included in the claims or their existence is clearly indicated in other ways.

[0156] At least some embodiments of the disclosed technology may be described in terms of the following:

[0157] 1. A computer-implemented method, the method comprising:

[0158] hosting a first virtual machine of a first virtual machine type on a host computer system of a cloud provider network, the first virtual machine type having an amount of resources allocated to a virtual machine of the first virtual machine type, and wherein the amount of resources is allocated to the first virtual machine from corresponding physical resources of the host computer system;

[0159] Receiving, by a hardware virtualization service of the cloud provider network, a request from a customer of the cloud provider network, the request being used to start a second virtual machine to share resources with the first virtual machine, the request including an identifier of the first virtual machine;

[0160] determining an identification of the host computer system based at least in part on the identification of the first virtual machine;

[0161] The hardware virtualization service enables the agent of the host computer system to start the second virtual machine in a standby state on the host computer system;

[0162] sharing, by the agent, the amount of resources allocated to the first virtual machine between the first virtual machine and the second virtual machine; and

[0163] Tracking data of the host computer system is updated, by the hardware virtualization service, to associate the first virtual machine and the second virtual machine with a socket, wherein the socket is associated with the first virtual machine type.

[0164] 2. The computer-implemented method of clause 1, wherein causing the agent of the host computer system to start the second virtual machine on the host computer system comprises:

[0165] A launch request is sent to the agent via a secure tunnel for control plane traffic, wherein the host computer system is part of an edge location of the cloud provider network within a third-party network.

[0166] 3. A computer-implemented method as described in any of clauses 1 to 2, wherein the host computer system is a radio access network (RAN) edge server and the first virtual machine performs RAN network functions.

[0167] 4. A computer-implemented method, the method comprising:

[0168] Receiving, by a hardware virtualization service of a cloud provider network, a request to start a first virtual machine, wherein the first virtual machine is of a first virtual machine type, the first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type;

[0169] causing, by the hardware virtualization service, launching of the first virtual machine on a host computer system of the cloud provider network; and

[0170] The allocation of the amount of a corresponding resource from the host computer system is shared by the host computer system between the first virtual machine and a second virtual machine, wherein the second virtual machine is of the first virtual machine type.

[0171] 5. The computer-implemented method of clause 4, wherein the hardware virtualization service maintains tracking data to track virtual machines hosted on the host computer system, the tracking data for the host computer system comprising a plurality of slots, each slot being associated with a virtual machine type, the method further comprising:

[0172] The tracking data is updated, by the hardware virtualization service, with an indication that a first slot of the plurality of slots is used by the first virtual machine and the second virtual machine.

[0173] 6. The computer-implemented method of any of clauses 4 to 5, wherein the amount of resources is computing capacity, and wherein sharing the allocation of the amount of resources comprises:

[0174] associating a first process with a first core of a multi-core processor of the host computer system, wherein the first process corresponds to the first virtual machine;

[0175] associating a second process with the first core, wherein the second process corresponds to the second virtual machine; and

[0176] The first process and the second process are scheduled on the first core of the multi-core processor.

[0177] 7. The computer-implemented method of any of clauses 4 to 5, wherein the amount of resource is memory capacity, and wherein sharing the allocation of the amount of resource comprises:

[0178] allocating a first amount of memory of the host computer system to a first process, wherein the first process corresponds to the first virtual machine;

[0179] allocating a second amount of memory of the host computer system to a second process, wherein the second process corresponds to the second virtual machine; and

[0180] A sum of the first amount of memory and the second amount of memory is limited to the memory capacity.

[0181] 8. The computer-implemented method of any of clauses 4 to 5, wherein the amount of resources is network throughput capacity, and wherein sharing the allocation of the amount of resources comprises:

[0182] A sum of first network traffic sent from the first virtual machine during a first time period and second network traffic sent from the second virtual machine during the first time period is limited by the network throughput capacity.

[0183] 9. The computer-implemented method of any of clauses 4 to 8, wherein the first virtual machine is connected to the cloud provider network via a virtual network interface having a first Internet Protocol (IP) address, wherein the second virtual machine is connected to the cloud provider network via the virtual network interface, the method further comprising:

[0184] Traffic received at the virtual network interface is sent to both the first virtual machine and the second virtual machine.

[0185] 10. The computer-implemented method of any of clauses 4 to 9, wherein the second virtual machine is executing an application on the host computer system prior to receiving the request, and wherein causing the launching of the first virtual machine on the host computer system of the cloud provider network comprises:

[0186] sending a second request to an agent of the host computer system, the request including an identification of a machine image from which to launch the first virtual machine, wherein the machine image includes an updated version of the application;

[0187] retrieving, by the agent, the machine image from a machine image data store of the cloud provider network; and

[0188] The first virtual machine is launched from the machine image by the agent on a host computer system of the cloud provider network.

[0189] 11. The computer-implemented method of any of clauses 4 to 10, further comprising causing, by the hardware virtualization service, termination of the second virtual machine on the host computer system of the cloud provider network.

[0190] 12. The computer-implemented method of any one of clauses 4 to 11, wherein the first virtual machine and the second virtual machine are booted from a machine image, wherein the first virtual machine executes a backup application and is in a paused state, wherein the second virtual machine executes a primary application, the method further comprising:

[0191] detecting, by a health monitoring service of the cloud provider network, a failure of the primary application based at least in part on metrics obtained from at least one of the second virtual machine or an agent of the host computer system; and

[0192] The health monitoring service causes the first virtual machine to resume execution.

[0193] 13. The computer-implemented method of clause 12, further comprising:

[0194] The hardware virtualization service causes the second virtual machine to terminate;

[0195] causing, by the hardware virtualization service, the launching of a third virtual machine on the host computer system of the cloud provider network, wherein the third virtual machine is launched from the machine image in a suspended state; and

[0196] The allocation of the amount of a corresponding resource from the host computer system is shared by the host computer system between the first virtual machine and the third virtual machine, wherein the third virtual machine is of the first virtual machine type.

[0197] 14. The computer-implemented method of any of clauses 4 to 13, wherein causing the first virtual machine to be launched on a host computer system of the cloud provider network comprises:

[0198] A launch request is sent to the host computer via a secure tunnel for control plane traffic, wherein the host computer system is part of an edge location of the cloud provider network within a third-party network.

[0199] 15. A system, comprising:

[0200] The first one or more electronic devices are used to implement a hardware virtualization service of a cloud provider network, wherein the hardware virtualization service includes instructions that, when executed, cause the hardware virtualization service to:

[0201] receiving a request to start a first virtual machine, wherein the first virtual machine is of a first virtual machine type, the first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type;

[0202] causing the first virtual machine to be launched on a host computer system of the cloud provider network; and

[0203] a second one or more electronic devices, the second one or more electronic devices being used to implement the host computer system of the cloud provider network, the host computer system comprising instructions that, when executed, cause the host computer system to:

[0204] An allocation of the amount of a corresponding resource from the host computer system is shared between the first virtual machine and a second virtual machine, wherein the second virtual machine is of the first virtual machine type.

[0205] 16. The system of clause 15, wherein the hardware virtualization service comprises further instructions that, when executed, cause the hardware virtualization service to:

[0206] maintaining tracking data to track virtual machines hosted on the host computer system, the tracking data for the host computer system comprising a plurality of slots, each slot being associated with a virtual machine type; and

[0207] The tracking data is updated with an indication that a first slot of the plurality of slots is used by the first virtual machine and the second virtual machine.

[0208] 17. The system of any of clauses 15 to 16, wherein the amount of resource is computing capacity, and wherein the instructions that when executed cause the host computer system to share an allocation of the amount of resource further comprise instructions that when executed cause the host computer system to:

[0209] associating a first process with a first core of a multi-core processor of the host computer system, wherein the first process corresponds to the first virtual machine;

[0210] associating a second process with the first core, wherein the second process corresponds to the second virtual machine; and

[0211] The first process and the second process are scheduled on the first core of the multi-core processor.

[0212] 18. The system of any of clauses 15 to 17, wherein the second virtual machine is executing an application on the host computer system prior to receiving the request, wherein the instructions that, when executed, cause the hardware virtualization service to cause the launch of the first virtual machine on the host computer system of the cloud provider network further include instructions that, when executed, cause the hardware virtualization service to:

[0213] sending a second request to the host computer system, the request including an identification of a machine image from which to launch the first virtual machine, wherein the machine image includes an updated version of the application; and

[0214] wherein the host computer system comprises further instructions which, when executed, cause the host computer system to:

[0215] retrieving the machine image from a machine image data store of the cloud provider network; and

[0216] The first virtual machine is launched from the machine image on a host computer system of the cloud provider network.

[0217] 19. The system of any one of clauses 15 to 18, wherein the first virtual machine and the second virtual machine are started from a machine image, wherein the first virtual machine executes a backup application and is in a paused state, wherein the second virtual machine executes a primary application, the system further comprising:

[0218] a third one or more electronic devices, the third one or more electronic devices implementing a health monitoring service of the cloud provider network, the health monitoring service comprising instructions that, when executed, cause the health monitoring service to:

[0219] detecting a failure of the primary application based at least in part on metrics obtained from at least one of the second virtual machine or an agent of the host computer system; and

[0220] The first virtual machine is caused to resume execution.

[0221] 20. The system of any of clauses 15 to 19, wherein the instructions that, when executed, cause the hardware virtualization service to cause the launch of the first virtual machine on a host computer system of the cloud provider network further comprise instructions that, when executed, cause the hardware virtualization service to:

[0222] A launch request is sent to the host computer via a secure tunnel for control plane traffic, wherein the host computer system is part of an edge location of the cloud provider network within a third-party network.

[0223] The specification and drawings are accordingly to be regarded in an illustrative rather than a restrictive sense.It will, however, be evident that various modifications and changes may be made thereto without departing from the broader scope of the disclosure as set forth in the claims.

Claims

1. A computer-implemented method, the method comprising: Receiving, by a hardware virtualization service of a cloud provider network, a request to start a first virtual machine, wherein the first virtual machine is of a first virtual machine type, the first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type; causing, by the hardware virtualization service, the first virtual machine to be launched on a host computer system of the cloud provider network; as well as The allocation of the amount of a corresponding resource from the host computer system is shared by the host computer system between the first virtual machine and a second virtual machine, wherein the second virtual machine is of the first virtual machine type.

2. The computer-implemented method of claim 1 , wherein the hardware virtualization service maintains tracking data to track virtual machines hosted on the host computer system, the tracking data for the host computer system comprising a plurality of slots, each slot being associated with a virtual machine type, the method further comprising: The tracking data is updated, by the hardware virtualization service, with an indication that a first slot of the plurality of slots is used by the first virtual machine and the second virtual machine.

3. The computer-implemented method of any one of claims 1 to 2, wherein the amount of resources is computing capacity, and wherein sharing the allocation of the amount of resources comprises: associating a first process with a first core of a multi-core processor of the host computer system, wherein the first process corresponds to the first virtual machine; associating a second process with the first core, wherein the second process corresponds to the second virtual machine; as well as The first process and the second process are scheduled on the first core of the multi-core processor.

4. The computer-implemented method of any one of claims 1 to 2, wherein the amount of resource is memory capacity, and wherein sharing the allocation of the amount of resource comprises: allocating a first amount of memory of the host computer system to a first process, wherein the first process corresponds to the first virtual machine; allocating a second amount of memory of the host computer system to a second process, wherein the second process corresponds to the second virtual machine; as well as A sum of the first amount of memory and the second amount of memory is limited to the memory capacity.

5. The computer-implemented method of any one of claims 1 to 2, wherein the amount of resources is network throughput capacity, and wherein sharing the allocation of the amount of resources comprises: A sum of first network traffic sent from the first virtual machine during a first time period and second network traffic sent from the second virtual machine during the first time period is limited by the network throughput capacity.

6. The computer-implemented method of any one of claims 1 to 5, wherein the first virtual machine is connected to the cloud provider network via a virtual network interface having a first Internet Protocol (IP) address, wherein the second virtual machine is connected to the cloud provider network via the virtual network interface, the method further comprising: Traffic received at the virtual network interface is sent to both the first virtual machine and the second virtual machine.

7. The computer-implemented method of any one of claims 1 to 6, wherein the second virtual machine is executing an application on the host computer system prior to receiving the request, and wherein causing the launching of the first virtual machine on the host computer system of the cloud provider network comprises: sending a second request to an agent of the host computer system, the request including an identification of a machine image from which to launch the first virtual machine, wherein the machine image includes an updated version of the application; retrieving, by the agent, the machine image from a machine image data store of the cloud provider network; as well as The first virtual machine is launched from the machine image by the agent on a host computer system of the cloud provider network.

8. The computer-implemented method of any one of claims 1 to 7, further comprising: Termination of the second virtual machine on the host computer system of the cloud provider network is caused by the hardware virtualization service.

9. The computer-implemented method of any one of claims 1 to 8, wherein the first virtual machine and the second virtual machine are started from a machine image, wherein the first virtual machine executes a backup application and is in a paused state, wherein the second virtual machine executes a primary application, the method further comprising: detecting, by a health monitoring service of the cloud provider network, a failure of the primary application based at least in part on metrics obtained from at least one of the second virtual machine or an agent of the host computer system; as well as The health monitoring service causes the first virtual machine to resume execution.

10. The computer-implemented method of claim 9, further comprising: The hardware virtualization service causes the second virtual machine to terminate; causing, by the hardware virtualization service, the launching of a third virtual machine on the host computer system of the cloud provider network, wherein the third virtual machine is launched from the machine image in a suspended state; and The allocation of the amount of a corresponding resource from the host computer system is shared by the host computer system between the first virtual machine and the third virtual machine, wherein the third virtual machine is of the first virtual machine type.

11. The computer-implemented method of any one of claims 1 to 10, wherein causing the first virtual machine to be launched on a host computer system of the cloud provider network comprises: A launch request is sent to the host computer via a secure tunnel for control plane traffic, wherein the host computer system is part of an edge location of the cloud provider network within a third-party network.

12. A system, comprising: The first one or more electronic devices are used to implement a hardware virtualization service of a cloud provider network, wherein the hardware virtualization service includes instructions that, when executed, cause the hardware virtualization service to: receiving a request to start a first virtual machine, wherein the first virtual machine is of a first virtual machine type, the first virtual machine type having an amount of resources allocated to virtual machines of the first virtual machine type; causing the first virtual machine to be launched on a host computer system of the cloud provider network; as well as a second one or more electronic devices, the second one or more electronic devices being used to implement the host computer system of the cloud provider network, the host computer system comprising instructions that, when executed, cause the host computer system to: An allocation of the amount of a corresponding resource from the host computer system is shared between the first virtual machine and a second virtual machine, wherein the second virtual machine is of the first virtual machine type.

13. The system of claim 12, wherein the hardware virtualization service comprises further instructions that, when executed, cause the hardware virtualization service to: maintaining tracking data to track virtual machines hosted on the host computer system, the tracking data for the host computer system comprising a plurality of slots, each slot being associated with a virtual machine type; and The tracking data is updated with an indication that a first slot of the plurality of slots is used by the first virtual machine and the second virtual machine.

14. The system of any one of claims 12 to 13, wherein the amount of resource is computing capacity, and wherein the instructions that when executed cause the host computer system to share an allocation of the amount of resource further comprise instructions that when executed cause the host computer system to: associating a first process with a first core of a multi-core processor of the host computer system, wherein the first process corresponds to the first virtual machine; associating a second process with the first core, wherein the second process corresponds to the second virtual machine; and The first process and the second process are scheduled on the first core of the multi-core processor.

15. The system of any one of claims 12 to 14, wherein the second virtual machine is executing an application on the host computer system prior to receiving the request, wherein the instructions that when executed cause the hardware virtualization service to cause the launch of the first virtual machine on the host computer system of the cloud provider network further include instructions that when executed cause the hardware virtualization service to: sending a second request to the host computer system, the request including an identification of a machine image from which to launch the first virtual machine, wherein the machine image includes an updated version of the application; and wherein the host computer system comprises further instructions which, when executed, cause the host computer system to: retrieving the machine image from a machine image data store of the cloud provider network; and The first virtual machine is launched from the machine image on a host computer system of the cloud provider network.