Mobility of cloud computing instances hosted within a communication service provider network

By deploying cloud provider networks at the edge within the communication service provider network, and utilizing encapsulation protocols and local network managers, the high latency issue between end users and computing resources in cloud computing platforms is resolved, enabling low-latency access to computing resources and improving the responsiveness of related applications.

CN118102380BActive Publication Date: 2025-11-25AMAZON TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410183540.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-29
Filing Date
2020-10-30
Publication Date
2025-11-25
Estimated Expiration
2040-10-30

AI Technical Summary

Technical Problem

The centralized deployment of computing resources on existing cloud computing platforms limits low-latency access due to the network distance and number of network hops between end-user devices and computing resources, making it impossible to meet the needs of applications such as game streaming, virtual reality, real-time rendering, and autonomous vehicles.

Method used

Deploying cloud provider networks at the edge within a communication service provider network enables low-latency access to computing resources through provider-level extensions, encapsulation protocol technology, and local network managers. This supports virtualized computing and storage services, and resource management and communication are handled through control plane and data plane proxies.

Benefits of technology

It significantly reduces access latency to end-user devices, supports low-latency computing resource access, and improves the responsiveness of applications such as game streaming, virtual reality, real-time rendering, and autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118102380B_ABST
    Figure CN118102380B_ABST
Patent Text Reader

Abstract

This application discloses mobility of cloud computing instances hosted within a communication service provider network. Techniques for managing communication latency between a compute instance and a mobile device are described. A message is received that includes an indication of a mobility event associated with a mobile device of a communication service provider network. The mobility event indicates a change in a point of connection of the mobile device to the communication service provider network from a first access point to a second access point. It is determined that a communication latency of at least a portion of a network path between the mobile device and a compute instance via the second access point does not satisfy a latency constraint. A second provider substrate extension of the cloud provider network that satisfies the latency constraint for communicating with the mobile device via the second access point is identified, and a message is sent to the second provider substrate extension to cause another compute instance to be launched.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application entitled "Mobility of Cloud Computing Instances Hosted Within a Telecommunication Service Provider Network", which has PCT international application number PCT / US2020 / 058173, international application date of October 30, 2020, and Chinese national phase application number 202080094442.0. Background Technology

[0002] Cloud computing platforms typically provide customers with on-demand, managed computing resources. These resources (e.g., compute and storage capacity) are usually provided by large pools of capacity located in data centers. Customers can request computing resources from the “cloud,” and the cloud can provide those resources to those customers. Technologies such as virtual machines and containers are often used to allow customers to securely share the capacity of a computing system. Attached Figure Description

[0003] Various embodiments according to this disclosure will be described with reference to the following figures.

[0004] Figure 1 An exemplary system is shown that includes a cloud provider network and various provider underlying extensions of the cloud provider network according to some implementation schemes.

[0005] Figure 2 An exemplary system is shown, according to some implementation schemes, in which a cloud provider network underlying extension is deployed within a communications service provider network.

[0006] Figure 3 Exemplary components of provider underlying extensions within cloud provider networks and communications service provider networks, and the connectivity between them, are shown in more detail according to some implementation schemes.

[0007] Figure 4 An exemplary cloud provider network, including geographically dispersed provider underlying extensions (or “edge locations”), is shown according to some implementation schemes.

[0008] Figure 5 An exemplary environment is shown, in which a computing instance is launched at the edge location of a cloud provider's network, according to some implementation schemes.

[0009] Figure 6 This illustrates another exemplary environment in which a computing instance is launched at the edge location of a cloud provider's network, according to some implementation schemes.

[0010] Figure 7 This illustrates another exemplary environment in which a computing instance is launched at the edge location of a cloud provider's network, according to some implementation schemes.

[0011] Figure 8This is an exemplary environment in which a computing instance is initiated due to the mobility of an electronic device, according to some implementation schemes.

[0012] Figure 9 This is a flowchart illustrating the operation of a method for launching a computing instance at the edge location of a cloud provider's network, according to some implementation schemes.

[0013] Figure 10 This is a flowchart illustrating the operation of another method for launching compute instances at the edge of a cloud provider's network, according to some implementation schemes.

[0014] Figure 11 This is a flowchart illustrating the operation of a method for initiating a computing instance due to the mobility of an electronic device, according to some implementation schemes.

[0015] Figure 12 An exemplary provider network environment according to some implementation schemes is shown.

[0016] Figure 13 This is a block diagram of an exemplary provider network that provides storage services and hardware virtualization services to customers according to some implementation schemes.

[0017] Figure 14 This is a block diagram illustrating an exemplary computer system that can be used in some implementations. Detailed Implementation

[0018] This disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media for providing cloud provider network computing resources within a communications service provider network. According to some embodiments, computing resources managed by a cloud provider are deployed at the edge of the cloud provider network integrated within a communications service provider (CSP) network. CSPs typically include companies that have deployed networks through which end users obtain network connectivity. For example, CSPs may include mobile or cellular network providers (e.g., operating 3G, 4G, and / or 5G networks), wired internet service providers (e.g., cable, digital subscriber lines, fiber optics, etc.), and WiFi providers (e.g., at locations such as hotels, coffee shops, airports, etc.). While traditional deployments of computing resources in data centers offer various benefits due to centralization, physical constraints between end-user devices and those computing resources (such as network distance and the number of network hops) can prevent the achievement of very low latency. By installing or deploying capacity within a CSP network, cloud provider network operators can provide significantly reduced access latency to end-user devices for computing resources, in some cases providing single-digit millisecond latency. This low-latency access to computing resources is a major driver for improving the responsiveness of existing cloud-based applications and enabling next-generation applications in game streaming, virtual reality, real-time rendering, industrial automation, and autonomous vehicles.

[0019] A cloud provider network (or “cloud”) refers to a large pool of network-accessible computing resources (such as compute, storage, and networking resources, applications, and services). The cloud provides convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically configured and deployed in response to client commands. Therefore, cloud computing can be viewed as applications delivered as a service over a publicly accessible network (e.g., the Internet, cellular networks) along with the hardware and software in the cloud provider’s data centers that provide those services. Some customers may wish to use the resources and services of such a cloud provider network but, for various reasons (e.g., communication latency with their device, legal compliance, security, or others), prefer to configure these resources and services within their own network (e.g., locally), on a separate network managed by the cloud provider, within the network of a communications service provider, or in another independent network.

[0020] In some implementations, fragments of the cloud provider network (referred herein to as “Provider Underlying Extensions (PSEs)” or “Edge Locations (ELs)”) may be configured within a network separate from the cloud provider network. For example, a cloud provider network typically comprises a physical network (e.g., metal enclosures, cabling, rack hardware) referred to as the underlying layer. An underlying layer can be viewed as a network structure containing the physical hardware that runs the services of the provider network. In some implementations, a provider underlying “extension” can be an extension of the cloud provider network underlying layer consisting of one or more servers located locally in a customer or partner facility, a separate cloud provider-managed facility, a communications service provider facility, or any other type of facility including servers, where such servers communicate with a nearby availability zone or region of the cloud provider network via a network (e.g., a publicly accessible network such as the Internet). Customers can access the provider underlying extension via the cloud provider underlying layer or another network and can use the same application programming interface (API) to create and manage resources within the provider underlying extension as they would use to create and manage resources within the region of the cloud provider network.

[0021] As indicated above, one exemplary type of provider underlying extension is a provider underlying extension formed by servers located on-premises in a customer's or partner's facility. This type of underlying extension, located outside the cloud provider network's data center, can be referred to as an "outpost" of the cloud provider network. Another exemplary type of provider underlying extension is a provider underlying extension formed by servers located in a facility managed by the cloud provider, but including data plane capacity that is at least partially controlled by a separate control plane of the cloud provider network.

[0022] In some implementations, another example of provider-level extension is a network deployed within a communications service provider (CSM) network. CSMs typically include companies that have deployed networks through which end users obtain network connectivity. For example, CSMs can include mobile or cellular network providers (e.g., operating 3G, 4G, and / or 5G networks), wired internet service providers (e.g., cable, digital subscriber lines, fiber optics, etc.), and WiFi providers (e.g., in locations such as hotels, coffee shops, airports, etc.). While the traditional deployment of computing resources in data centers offers various benefits due to centralization, physical constraints between end-user devices and those computing resources (such as network distance and the number of network hops) can prevent the achievement of very low latency. By installing or deploying capacity within a CSM network, cloud provider network operators can provide significantly reduced access latency to end-user devices for computing resources, in some cases single-digit millisecond latency. This low-latency access to computing resources is a significant driver for improving the responsiveness of existing cloud-based applications and enabling next-generation applications such as game streaming, virtual reality, real-time rendering, industrial automation, and autonomous vehicles.

[0023] As used herein, computing resources within a cloud provider network (or possibly another network) installed within a communications service provider network are sometimes referred to as “cloud provider network edge locations” or simply “edge locations” because they are closer to the “edge” where end users connect to the network than computing resources in a centralized data center. Such edge locations may include one or more networked computing systems that provide computing resources to customers of the cloud provider network to serve end users with lower latency than that achievable with compute instances hosted in data center sites. Provider underlying extensions deployed within a communications service provider network may also be referred to as “wavelength zones.”

[0024] Figure 1 Exemplary systems including a cloud provider network and various provider underlying extensions of the cloud provider network are illustrated according to some implementation schemes. The cloud provider network 100 (sometimes simply referred to as the "cloud") refers to a network-accessible pool of computing resources (such as computing, storage, and networking resources, applications, and services), which may be virtualized or bare metal. The cloud can provide convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically configured and released in response to client commands. These resources can be dynamically provisioned and reconfigured to adapt to variable loads. Therefore, cloud computing can be viewed as applications delivered as a service over a publicly accessible network (e.g., the Internet, cellular communication networks) and the hardware and software in the cloud provider's data center providing those services.

[0025] Cloud provider network 100 can provide users with on-demand, scalable computing platforms via a network, for example, allowing users to have scalable “virtual computing devices” available for their use via computing servers (which provide computing instances via one or both of a central processing unit (CPU) and a graphics processing unit (GPU) optionally used with local storage devices) and block storage servers (which provide virtualized persistent block storage for specified computing instances). These virtual computing devices have the attributes of a personal computing device, including hardware (various types of processors, local memory, random access memory (RAM), hard disk and / or solid-state drive (“SSD”) storage devices), operating system options, networking capabilities, and pre-loaded application software. Each virtual computing device can also virtualize its console input and output (e.g., keyboard, monitor, and mouse). This virtualization allows users to connect to their virtual computing devices using computer applications such as browsers, application programming interfaces (APIs), and software development kits (SDKs) to configure and use their virtual computing devices as if they were personal computing devices. Unlike personal computing devices that have a fixed amount of hardware resources available to the user, the hardware associated with a virtual computing device can be scaled up or down depending on the resources the user needs.

[0026] As indicated above, a user (e.g., user 138) can connect to virtualized computing devices and other cloud provider network 100 resources and services via intermediate network 136 using various interfaces 104 (e.g., APIs). An API refers to an interface and / or communication protocol between a client (e.g., electronic device 134) and a server, such that if a client issues a request in a predefined format, the client should receive a response in a specific format or cause a defined action to be initiated. In the context of a cloud provider network, an API provides a gateway for customers to access cloud infrastructure by allowing customers to obtain data from the cloud provider network or to cause actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted in the cloud provider network. APIs can also enable different services within the cloud provider network to exchange data with each other. Users may choose to deploy their virtual computing systems to provide network-based services for their own use and / or for use by their customers or clients.

[0027] The cloud provider network 100 may include a physical network referred to as the underlying layer (e.g., metal enclosures, cables, rack hardware). The underlying layer can be viewed as a network structure containing the physical hardware running the provider network's services. The underlying layer may be isolated from the rest of the cloud provider network 100; for example, it may not be possible to route from the underlying network address to addresses in the production network running the cloud provider services, or to customer networks hosting customer resources.

[0028] The cloud provider network 100 may also include an overlay network of virtualized computing resources running on the underlying layer. In at least some embodiments, a hypervisor or other device or process on the network layer may use encapsulation protocol technology to encapsulate and route network packets (e.g., client IP packets) between client resource instances on different hosts within the provider network via the network layer. Encapsulation protocol technology can be used on the network layer to route encapsulated packets (also referred to as network layer packets) between endpoints on the network layer via overlay network paths or routes. Encapsulation protocol technology can be viewed as providing a virtual network topology overlay on the network layer. Thus, network packets can be routed along the underlying network based on the construction in the overlay network (e.g., a virtual network that may be referred to as a Virtual Private Cloud (VPC), a port / protocol firewall configuration that may be referred to as a security group). A mapping service (not shown) may coordinate the routing of these network packets. The mapping service may be a regionally distributed lookup service that maps a combination of overlay Internet Protocol (IP) and network identifiers to the underlying IP, enabling distributed underlying computing devices to find out where to send packets.

[0029] For illustration, each physical host device (e.g., compute server 106, block storage server 108, object storage server 110, control server 112) may have an IP address in the underlying network. Hardware virtualization technology enables multiple operating systems to run simultaneously on the host computer, for example, as virtual machines (VMs) on compute server 106. A hypervisor or virtual machine monitor (VMM) on the host allocates the host's hardware resources among the various VMs on the host and monitors the execution of the VMs. Each VM may be configured with one or more IP addresses in the overlay network, and the VMM on the host may be aware of the IP addresses of the VMs on the host. The VMM (and / or other devices or processes at the network layer) may use encapsulation protocol technology to encapsulate network packets (e.g., client IP packets) and route said network packets between virtualized resources on different hosts within the cloud provider network 100 via the network layer. Encapsulation protocol technology can be used at the network layer to route encapsulated packets between endpoints at the network layer via overlay network paths or routes. Encapsulation protocol technology can be viewed as providing a virtual network topology overlayed at the network layer. Encapsulation protocol technology may include a mapping service that maintains a mapping directory that maps IP overlay addresses (e.g., IP addresses visible to clients) to underlying IP addresses (IP addresses not visible to clients), which can be accessed by various processes on the cloud provider's network for routing packets between endpoints.

[0030] As shown in the figure, in various implementations, the traffic and operations at the underlying layer of the cloud provider network can be broadly subdivided into two categories: control plane traffic carried on the logical control plane 114A and data plane operations carried on the logical data plane 116A. While the data plane 116A represents the movement of user data through the distributed computing system, the control plane 114A represents the movement of control signals through the distributed computing system. The control plane 114A typically includes one or more control plane components or services distributed across and implemented by one or more control servers 112. Control plane traffic typically includes administrative operations such as establishing isolated virtual networks for various customers, monitoring resource utilization and health, identifying specific hosts or servers to initiate compute instances, provisioning additional hardware as needed, and so on. The data plane 116A includes customer resources implemented on the cloud provider network (e.g., compute instances, containers, block storage volumes, databases, file storage). Data plane traffic typically includes non-administrative operations such as transferring data to and from customer resources.

[0031] Control plane components are typically implemented on a different set of servers than data plane servers, and control plane traffic and data plane traffic can be sent over separate / different networks. In some implementations, control plane traffic and data plane traffic may be supported by different protocols. In some implementations, messages (e.g., packets) sent over the cloud provider network 100 include flags indicating whether the traffic is control plane traffic or data plane traffic. In some implementations, the payload of the traffic can be examined to determine its type (e.g., control plane or data plane). Other techniques for distinguishing traffic types are possible.

[0032] As shown in the figure, data plane 116A may include one or more compute servers 106, which may be bare metal (e.g., a single tenant) or may be virtualized by a hypervisor to run multiple VMs (sometimes referred to as "instances") or microVMs for one or more clients. These compute servers 106 may support virtualized computing services (or "hardware virtualization services") of a cloud provider network. Virtualized computing services may be part of control plane 114A, allowing clients to issue commands via interface 104 (e.g., an API) to launch and manage compute instances (e.g., VMs, containers) of their applications. Virtualized computing services may provide virtual compute instances with different compute and / or memory resources. In one embodiment, each of the virtual compute instances may correspond to one of several instance types. Instance types may be characterized by their hardware type, compute resources (e.g., the number, type, and configuration of CPUs or CPU cores), memory resources (e.g., the capacity, type, and configuration of local memory), storage resources (e.g., the capacity, type, and configuration of locally accessible storage devices), network resources (e.g., the characteristics of their network interfaces and / or network capabilities), and / or other suitable descriptive characteristics. Using instance type selection functionality, instance types can be selected for customers, for example (at least in part) based on input from the customer. For instance, a customer can choose an instance type from a set of predefined instance types. As another example, a customer can specify the desired resources for the instance type and / or the workload requirements for the instance, and the instance type selection functionality can select the instance type based on such a specification.

[0033] Data plane 116A may also include one or more block storage servers 108, which may include persistent storage devices for storing customer data volumes and software for managing these volumes. These block storage servers 108 may support managed block storage services on a cloud provider network. The managed block storage service may be part of control plane 114A, allowing customers to issue commands via interface 104 (e.g., API) to create and manage volumes of applications running on their compute instances. Block storage servers 108 include one or more servers on which data is stored as blocks. A block is a sequence of bytes or bits, typically containing a certain integer number of records, with a maximum length of block size. Block data is typically stored in a data buffer and read or written to the entire block at a time. Typically, a volume may correspond to a logical collection of data, such as a set of data maintained on behalf of a user. A user volume (which may be considered, for example, an individual hard drive ranging in size from 1 GB to 1 terabyte (TB) or larger) consists of one or more blocks stored on a block storage server. While considered as an individual hard drive, it should be understood that a volume may be stored as one or more virtualized devices implemented on one or more underlying physical host devices. A volume can be partitioned several times (e.g., up to 16 times), with each partition hosted by a different host. Volume data can be replicated across multiple devices within a cloud provider's network to provide multiple copies of the volume (where such copies can collectively represent the volume on the computing system). Volume replicas in a distributed computing system can beneficially provide automatic failover and recovery, for example, by allowing users access to a primary copy of the volume or secondary copies of the volume synchronized with the primary copy at the block level, so that failure of the primary or secondary copy does not prevent access to volume information. The primary copy's role can be to facilitate reads and writes on the volume (sometimes referred to as "input / output operations" or simply "I / O operations") and propagate any writes to the secondary copy (preferably synchronously along the I / O path, but asynchronous replication can also be used). The secondary copy can be updated synchronously with the primary copy and provide a seamless transition during failover operations, whereby the secondary copy assumes the role of the primary copy, and the former primary copy is designated as the secondary copy or a new replacement secondary copy is provisioned. While some examples in this document discuss primary and secondary copies, it should be understood that a logical volume can include multiple secondary copies. Compute instances can virtualize their I / O as volumes via clients. The client represents instructions that enable compute instances to connect to remote data volumes (e.g., data volumes stored on physically separate compute devices accessible over a network) and perform I / O operations at the remote data volume. The client can be implemented on an offload card of a server that includes the processing units (e.g., CPUs or GPUs) of the compute instances.

[0034] Data plane 116A may also include one or more object storage servers 110, which represent another type of storage device within a cloud provider network. Object storage servers 110 include one or more servers on which data is stored as objects within resources called buckets and can be used to support managed object storage services within the cloud provider network. Each object typically includes the stored data, variable metadata enabling the object storage server to analyze various aspects of the stored object, and a globally unique identifier or key that can be used to retrieve the object. Each bucket is associated with a given user account. Users can store as many objects as desired in their buckets, can write, read, and delete objects in their buckets, and can control access to their buckets and the objects contained within them. Furthermore, in implementations with multiple different object storage service servers distributed across different regions in the aforementioned areas, users can select the region (or multiple regions) of the storage bucket, for example, to optimize latency. Clients can use buckets to store various types of objects, including machine images that can be used to boot VMs, and snapshots representing point-in-time views of volume data.

[0035] Provider underlying extension 102 (“PSE”) provides the resources and services of cloud provider network 100 within a separate network, thereby extending the functionality of cloud provider network 100 to new locations (e.g., for reasons related to latency in communication with customer devices, legal compliance, security, etc.). As indicated, such provider underlying extension 102 may include cloud provider network managed provider underlying extension 140 (e.g., formed by servers located in cloud provider managed facilities separate from those facilities associated with cloud provider network 100), communication service provider underlying extension 142 (e.g., formed by servers associated with communication service provider facilities), customer managed provider underlying extension 144 (e.g., formed by servers located on-premises at customer or partner facilities), and other possible types of underlying extensions.

[0036] As illustrated in the exemplary provider underlying extension 140, provider underlying extension 102 can similarly include a logical separation between a control plane 118B and a data plane 120B, which extend the control plane 114A and data plane 116A of the cloud provider network 100, respectively. Provider underlying extension 102 can be pre-configured by the cloud provider network operator with appropriate combinations of hardware and software and / or firmware elements to support various types of compute-related resources and in a manner that reflects the experience of using the cloud provider network. For example, one or more provider underlying extension location servers can be provisioned by the cloud provider to be deployed within provider underlying extension 102. As described above, cloud provider network 100 can provide a set of predefined instance types, each with different types and quantities of underlying hardware resources. Each instance type can also be provided in various sizes. To enable customers to continue using the same instance types and sizes they use in the region within provider underlying extension 102, the servers can be heterogeneous servers. Heterogeneous servers can simultaneously support multiple instance sizes of the same type and can also be reconfigured to host any instance type supported by their underlying hardware resources. Reconfiguration of heterogeneous servers can occur instantly using available server capacity—that is, while other VMs are still running and consuming additional capacity on the provider's underlying extended location servers. This can improve the utilization of compute resources in edge locations by allowing for better packaging of running instances on servers and provides a seamless experience regarding instance usage on both the provider's network 100 and the provider's underlying extended network.

[0037] As shown in the figure, the provider underlying extension server can host one or more compute instances 122. Compute instances 122 can be VMs, or containers that package code and all its dependencies, allowing applications to run quickly and reliably across compute environments, such as VMs. Additionally, the server can host one or more data volumes 124 if needed by customers. In the regions of the cloud provider network 100, such volumes can be hosted on dedicated block storage servers. However, due to the possibility of significantly smaller capacity at the provider underlying extension 102 compared to the regions described, optimal utilization may not be provided if the provider underlying extension includes such dedicated block storage servers. Therefore, block storage services can be virtualized within the provider underlying extension 102, allowing one of the VMs to run block storage software and store the data in volume 124. Similar to the operation of block storage services in the regions of the cloud provider network 100, volume 124 within the provider underlying extension 102 can be replicated for persistence and availability. The volume can be provisioned in its own isolated virtual network within the provider underlying extension 102. Computation instance 122 and any volume 124 together constitute the provider network data plane 116A and the data plane extension 120B within the provider underlying extension 102.

[0038] In some implementations, servers within the provider underlying extension 102 may host certain local control plane components 126, such as components that enable the provider underlying extension 102 to continue operating in the event of an interruption in the connection back to the cloud provider network 100. Examples of such components include: a migration manager that can move compute instances 122 between provider underlying extension servers if availability needs to be maintained; and a key value data store that indicates the location of volume copies. However, the control plane 118B functionality typically used for the provider underlying extension will remain in the cloud provider network 100 to allow customers to utilize as much of the provider underlying extension's resource capacity as possible.

[0039] A migration manager may have a centralized coordination component running in a region, and a local controller running on a PSE server (and servers in the cloud provider's data center). When a migration is triggered, the centralized coordination component can identify the target edge location and / or target host, while the local controller can coordinate data transfer between the source host and the target host. The described resource movement between hosts in different locations can take one of several forms of migration. Migration refers to moving virtual machine instances (and / or other resources) between hosts in a cloud computing network or between hosts outside the cloud computing network and hosts within the cloud. Different types of migration exist, including live migration and restart migration. During a restart migration, the customer experiences an interruption and effective power cycling of their virtual machine instances. For example, a control plane service can coordinate a restart migration workflow that involves tearing down the current domain on the source host and then creating a new domain for the virtual machine instances on the new host. The instances are restarted by shutting down on the source host and then restarting on the new host.

[0040] Live migration refers to the process of moving a running virtual machine or application between different physical machines without significantly disrupting the availability of the virtual machine (e.g., the end user will not notice the downtime). When the control plane performs a live migration workflow, it can create a new "inactive" domain associated with the instance, while the instance's original domain continues to run as the "active" domain. The virtual machine's memory (including any state in memory of the running application), storage, and network connectivity are transferred from the original host with the active domain to the target host with the inactive domain. The virtual machine may be briefly paused to prevent state changes while transferring memory contents to the target host. The control plane can transition the inactive domain to become an active domain and degrade the original active domain to an inactive domain (sometimes called "flipping"), after which the inactive domain can be discarded.

[0041] The technologies used for various types of migration involve managing a critical phase: the time when virtual machine instances are unavailable to clients, which should be kept as short as possible. This can be particularly challenging in currently disclosed migration technologies because resources are being moved between hosts in geographically separated locations that can be connected via one or more intermediate networks. For live migration, the disclosed technologies can dynamically determine, for example, the amount of memory state data to be pre-copied (e.g., while the instance is still running on the source host) and post-copied (e.g., after the instance begins running on the target host) based on latency between locations, network bandwidth / usage patterns, and / or based on which memory pages the instance most frequently uses. Furthermore, the specific time for transferring memory state data can be dynamically determined based on network conditions between locations. This analysis can be performed by a migration management component in the region or by a migration management component running locally in the source edge location. If the instance already has access to virtualized storage, both the source and target domains can be attached to the storage device simultaneously to enable uninterrupted access to its data during migration and in the event of a rollback to the source domain.

[0042] Server software running on Provider Underlying Extension 102 may be designed by the cloud provider to run on the cloud provider's underlying network, and this software may be able to create a private copy of the underlying network ("shadow underlying") within the edge location by running unmodified in Provider Underlying Extension 102 using a local network manager 128. The local network manager 128 may run on the Provider Underlying Extension 102 server and bridge the shadow underlying network with the Provider Underlying Extension 102 network, for example, by acting as a Virtual Private Network (VPN) endpoint or an endpoint between Provider Underlying Extension 102 and proxies 130, 132 in the cloud provider network 100, and by implementing a mapping service (for traffic encapsulation and decapsulation) to correlate data plane traffic (from data plane proxies) and control plane traffic (from control plane proxies) with appropriate servers. By implementing a local version of the provider network's underlying-overlay mapping service, the local network manager 128 allows resources in Provider Underlying Extension 102 to communicate seamlessly with resources in the cloud provider network 100. In some implementations, a single local network manager may perform these actions for all servers hosting compute instances 122 in Provider Underlying Extension 102. In other implementations, each of the servers hosting compute instance 122 may have a dedicated local network manager. In a multi-rack edge location, inter-rack communication can be achieved through the local network manager, which maintains open tunnels between them.

[0043] The provider's underlying extension location can utilize secure networking tunnels through the provider's underlying extension 102 network to the cloud provider network 100, for example, to maintain the security of customer data when traversing the provider's underlying extension 102 network and any other intermediate networks (potentially including the public internet). Within the cloud provider network 100, these tunnels consist of virtual infrastructure components including isolated virtual networks (e.g., in an overlay network), control plane agent 130, data plane agent 132, and underlying network interfaces. Such agents can be implemented as containers running on compute instances. In some implementations, each server in the provider's underlying extension 102 location hosting compute instances can utilize at least two tunnels: one tunnel for control plane traffic (e.g., Constrained Application Protocol (CoAP) traffic) and one tunnel for encapsulated data plane traffic. A connectivity manager (not shown) within the cloud provider network manages their cloud provider network-side lifecycle, for example, by automatically provisioning these tunnels and their components as needed and maintaining them in a healthy operational state. In some implementations, a direct connection between the provider's underlying extension 102 location and the cloud provider network 100 can be used for control data plane and data plane communication. Compared to VPNs that go through other networks, direct connections can provide constant bandwidth and more consistent network performance because their network path is relatively fixed and stable.

[0044] A control plane (CP) agent 130 can be provisioned in the cloud provider network 100 to represent a specific host in an edge location. The CP agent acts as an intermediary between the control plane 114A in the cloud provider network 100 and the control plane 118B in the provider underlying extension 102. Specifically, the CP agent 130 provides the infrastructure for tunneling management API traffic destined for the provider underlying extension server from the regional layer to the provider underlying extension 102. For example, the virtualized compute service of the cloud provider network 100 can issue commands to the VMM of the server in the provider underlying extension 102 to start compute instance 122. The CP agent maintains a tunnel (e.g., VPN) to the local network manager 128 of the provider underlying extension. Software implemented in the CP agent ensures that only well-formed API traffic leaves and returns to the underlying layer. The CP agent provides a mechanism to expose remote servers on the cloud provider underlying layer while still protecting underlying security materials (e.g., encryption keys, security tokens) from leaving the cloud provider network 100. The unidirectional control plane traffic tunnel imposed by the CP agent also prevents any (potentially compromised) devices from calling back to the underlying layer. CP proxies can be instantiated one-to-one with servers at provider underlying extension 102, or may be able to manage control plane traffic for multiple servers in the same provider underlying extension.

[0045] Data plane (DP) proxy 132 can also be provisioned in cloud provider network 100 to represent a specific server in provider underlying extension 102. DP proxy 132 acts as a shadow or anchor point for the server and can be used by services within cloud provider network 100 to monitor the health of the host (including its availability, used / free compute and capacity, used / free storage and capacity, and network bandwidth utilization / availability). DP proxy 132 also allows isolated virtual networks to cross provider underlying extension 102 and cloud provider network 100 by acting as proxies for servers in cloud provider network 100. Each DP proxy 132 can be implemented as a packet-forwarding compute instance or container. As shown, each DP proxy 132 can maintain a VPN tunnel with a local network manager 128 that manages traffic to the server represented by the DP proxy 132. This tunnel can be used to send data plane traffic between the provider underlying extension server and cloud provider network 100. Data plane traffic flowing between provider underlying extension 102 and cloud provider network 100 can be transmitted through the DP proxy 132 associated with that provider underlying extension. For data plane traffic flowing from provider underlying extension 102 to cloud provider network 100, DP proxy 132 can receive the encapsulated data plane traffic, verify its correctness, and allow it to enter cloud provider network 100. DP proxy 132 can then forward the encapsulated traffic directly from cloud provider network 100 to provider underlying extension 102.

[0046] Local network manager 128 can provide secure network connectivity with agents 130, 132 established within cloud provider network 100. Once a connection is established between local network manager 128 and the agents, the client can issue commands via interface 104 to instantiate (and / or perform other operations using) compute instances in a manner similar to how such commands would be issued to compute instances hosted within cloud provider network 100. From the client's perspective, the client can now seamlessly use local resources within the provider underlying extension (and resources located within cloud provider network 100, if needed). Compute instances hosted on servers at provider underlying extension 102 can communicate with electronic devices located on the same network, and, as needed, with other resources hosted within cloud provider network 100. Local gateway 146 can be implemented to provide network connectivity between provider underlying extension 102 and networks associated with said extension (e.g., the communication service provider network in the example of provider underlying extension 142).

[0047] There may be situations where data needs to be transferred between the object storage service and the provider's underlying extension 102. For example, the object storage service may store machine images used to launch VMs, as well as snapshots representing point-in-time backups of volumes. The object gateway may be provided on a PSE server or dedicated storage device and provide customers with configurable per-bucket caching of the contents of object storage buckets within their PSE to minimize the impact of PSE region latency on customer workloads. The object gateway may also temporarily store snapshot data from snapshots of volumes in the PSE and then synchronize it with the object server in the region, where possible. The object gateway may also store machine images specified by the customer for use within the PSE or at the customer's premises. In some implementations, data within the PSE may be encrypted with a unique key, and for security reasons, the cloud provider may restrict the sharing of the key from the region to the PSE. Therefore, data exchanged between the object storage server and the object gateway may utilize encryption, decryption, and / or re-encryption to maintain secure boundaries regarding encryption keys or other sensitive data. A transformation intermediary can perform these operations and can (on the object storage server) create PSE buckets using the PSE encryption key to store snapshot and machine image data.

[0048] In this manner, PSE 102 forms an edge location because it provides the resources and services of the cloud provider's network outside of and closer to the customer's equipment, outside of the traditional cloud provider's data center. Edge locations, as referred to herein, can be structured in various ways. In some implementations, an edge location can be an extension of the underlying cloud provider network, including a limited amount of capacity provided outside of an Availability Zone (e.g., in a small data center of the cloud provider or in other facilities located near the customer's workload and potentially far from any Availability Zone). Such edge locations can be referred to as "far zones" (due to their distance from other Availability Zones) or "near zones" (due to their proximity to the customer's workload). Near zones can be connected to publicly accessible networks such as the Internet in various ways, such as directly, via another network, or via a dedicated connection to the region. While near zones typically have more limited capacity than a region, in some cases, near zones can have considerable capacity, such as thousands or more racks.

[0049] In some implementations, an edge location can be an extension of the underlying layer of a cloud provider network, consisting of one or more servers located locally at a customer's or partner's facility, where such servers communicate with nearby availability zones or areas of the cloud provider network via a network (e.g., a publicly accessible network, such as the Internet). This type of underlying extension located outside the cloud provider network's data center can be referred to as an "outpost" of the cloud provider network. Some outposts can be integrated into a communications network, for example, as multi-access edge computing (MEC) sites, whose physical infrastructure is distributed across telecommunications data centers, telecommunications aggregation sites, and / or telecommunications base stations within the telecommunications network. In a local example, the limited capacity of an outpost may be available only to the customer who owns the site (and any other accounts permitted by the customer). In a telecommunications example, the limited capacity of an outpost can be shared among multiple applications (e.g., games, virtual reality applications, healthcare applications) sending data to users on the telecommunications network.

[0050] Edge locations may include data plane capacity that is at least partially controlled by the control plane of a nearby availability zone within the provider network. Thus, an availability zone group may include a “parent” availability zone and any “child” edge locations belonging to the parent availability zone (e.g., at least partially controlled by its control plane). Certain limited control plane functionality (e.g., features requiring low-latency communication with customer resources, and / or features enabling the edge location to continue operating when disconnected from the parent availability zone) may also exist in some edge locations. Therefore, in the example above, an edge location refers to an extension of at least the data plane capacity located at the edge of the cloud provider network, close to customer devices and / or workloads.

[0051] Figure 2 An exemplary system is illustrated, according to some implementations, in which a cloud provider network edge location is deployed within a communications service provider network. The communications service provider (CSP) network 200 typically includes downstream interfaces to end-user electronic devices and upstream interfaces to other networks (e.g., the Internet). In this example, the CSP network 200 is a wireless “cellular” CSP network, which includes radio access networks (RAN) 202, 204, aggregation sites (AS) 206, 208, and a core network (CN) 210. RAN 202, 204 include base stations (e.g., NodeB, eNodeB, gNodeB) providing wireless connectivity to electronic devices 212. The core network 210 typically includes functionalities related to the management of the CSP network (e.g., billing, mobility management, etc.) and relay traffic between the CSP network and other networks. Aggregation sites 206, 208 can be used to consolidate traffic from many different radio access networks into the core network and to route traffic originating from the core network to various radio access networks.

[0052] End-user electronic device 212 Figure 2 From left to right, a base station (or radio base station) 214 is wirelessly connected to radio access network 202. Such electronic devices 212 are sometimes referred to as user equipment (UE) or customer premises equipment (CPE). Data traffic is typically routed to core network 210 via a fiber optic transmission network consisting of multiple hops of Layer 3 routers (e.g., at aggregation sites). Core network 210 is typically housed in one or more data centers. For data traffic destined for locations outside CSP network 200, network components 222 to 226 typically include firewalls through which traffic can enter or leave CSP network 200 to reach external networks, such as the Internet or cloud provider network 100. Note that in some implementations, CSP network 200 may include facilities that allow traffic to enter or leave from sites further downstream of core network 210 (e.g., at aggregation sites or RANs).

[0053] Provider underlying extensions 216 to 220 include computing resources that are managed as part of the cloud provider network but are installed or located at various points within the CSP network (e.g., locally in space owned or leased by the CSP). These computing resources typically provide a certain amount of compute and storage capacity that the cloud provider can allocate to its customers. Computing resources may also include storage and accelerator capacity (e.g., solid-state drives, graphics accelerators, etc.). Here, provider underlying extensions 216, 218, and 220 communicate with the cloud provider network 100.

[0054] Typically, for example, in terms of network hops and / or distance, the farther the provider's underlying extension is from the cloud provider network 100 (or closer to the device 212), the lower the network latency between the computing resources within the provider's underlying extension and the device 212. However, physical site constraints typically limit the amount of computing capacity that provider's underlying extensions can be installed at various points within the CSP, or determine whether computing capacity can be installed at all points. For example, provider's underlying extensions located within the core network 210 can typically have a much larger footprint (in terms of physical space, power requirements, cooling requirements, etc.) compared to provider's underlying extensions located within RANs 202 and 204.

[0055] The installation or location of provider-level extensions within a CSP network may vary depending on the specific network topology or architecture of the CSP network. For example... Figure 2The provider underlying extension is typically connected to any location on the CSP network where packet-based traffic (e.g., IP-based traffic) can be interrupted. Additionally, communication between a given provider underlying extension and cloud provider network 100 typically securely relays at least a portion of CSP network 200 (e.g., via secure tunnels, VPNs, direct connections, etc.). In the example shown, network component 222 facilitates data traffic routing to and from provider underlying extension 216 integrated with RAN 202, network component 224 facilitates data traffic routing to and from provider underlying extension 218 integrated with AS 206, and network component 226 facilitates data traffic routing to and from provider underlying extension 220 integrated with CN 210. Network components 222 through 226 may include routers, gateways, or firewalls. To facilitate routing, the CSP may assign one or more IP addresses from the CSP network address space to each of the edge locations.

[0056] In 5G wireless network development, edge locations can be considered a possible implementation of Multi-Access Edge Computing (MEC). Such edge locations can connect to various points within a CSP 5G network, which provide interruption for data traffic as part of the User Plane Function (UPF). Older wireless networks can also include edge locations. For example, in 3G wireless networks, edge locations can connect to the packet-switched network portion of the CSP network, such as connecting to the Serving General Packet Radio Service Support Node (SGSN) or the Gateway General Packet Radio Service Support Node (GGSN). In 4G wireless networks, edge locations can connect to the Serving Gateway (SGW) or Packet Data Network Gateway (PGW) as part of the core network or Evolved Packet Core (EPC).

[0057] In some implementations, traffic between Provider Underlay Extension 228 and Cloud Provider Network 100 can be interrupted from CSP Network 200 without routing through Core Network 210. For example, Network Component 230 of RAN 204 can be configured to route traffic between Provider Underlay Extension 216 of RAN 204 and Cloud Provider Network 100 without traversing Aggregator Site or Core Network 210. As another example, Network Component 231 of Aggregator Site 208 can be configured to route traffic between Provider Underlay Extension 232 of Aggregator Site 208 and Cloud Provider Network 100 without traversing Core Network 210. Network Components 230, 231 may include gateways or routers with routing data to direct traffic from edge locations to Cloud Provider Network 100 (e.g., via direct connection or intermediate network 234) and to Provider Underlay Extensions from Cloud Provider Network 100.

[0058] In some implementations, the provider underlying extension can connect to more than one CSP network. For example, when two CSPs share or route traffic through a common point, the provider underlying extension can connect to two CSP networks. For example, each CSP may allocate a portion of its network address space to the provider underlying extension, and the provider underlying extension may include a router or gateway that can distinguish traffic exchanged with each of the CSP networks. For example, traffic from one CSP network destined for the provider underlying extension may have different destination IP addresses, source IP addresses, and / or VLAN tags compared to traffic received from another CSP network. Traffic originating from the provider underlying extension destined for one of the CSP networks can similarly be encapsulated with appropriate VLAN tags, source IP addresses (e.g., from a pool allocated from the destination CSP network's address space to the provider underlying extension), and destination IP addresses.

[0059] It should be noted that, although Figure 2 An exemplary CSP network architecture includes a radio access network, aggregation sites, and a core network; however, the naming and structure of a CSP network architecture may differ between generations of wireless technologies, between different CSPs, and between wireless CSP networks and fixed-line CSP networks. Furthermore, although... Figure 2 Several locations within a CSP network are shown for edge locations, but other locations are also possible (e.g., at a base station).

[0060] Figure 3Exemplary components of provider underlying extensions within cloud provider networks and communication service provider networks, and their connectivity, are shown in more detail according to some embodiments. Provider underlying extension 300 provides the resources and services of the cloud provider network within CSP network 302, thereby extending the functionality of cloud provider network 100 to be closer to end-user devices 304 connected to the CSP network.

[0061] The provider underlying extension 300 similarly includes a logical separation between a control plane 306B and a data plane 308B, which extend the control plane 114A and data plane 116A of the cloud provider network 100, respectively. The provider underlying extension 300 can be pre-configured by the cloud provider network operator, for example, with an appropriate combination of hardware and software and / or firmware elements to support various types of compute-related resources, and in a manner that reflects the experience of using the cloud provider network. For example, one or more provider underlying extension location servers 310 can be provisioned by the cloud provider for deployment within the CSP network 302.

[0062] In some implementations, server 310 within provider underlying extension 300 may host certain local control plane components 314, such as components that enable provider underlying extension 300 to continue operating in the event of an interruption in the connection back to cloud provider network 100. Furthermore, certain controller functions may typically be implemented locally on data plane servers or even in the cloud provider data center, such as functions for collecting metrics for monitoring instance health and sending those metrics to a monitoring service, and functions for coordinating the transfer of instance status data during live migration. However, the control plane 306B functionality typically used by provider underlying extension 300 will remain within cloud provider network 100 to allow customers to utilize as much of the provider underlying extension's resource capacity as possible.

[0063] As shown in the diagram, the provider's underlying extension server 310 can host compute instances 312. Compute instances can be VMs, microVMs, or containers that package code and all its dependencies, allowing applications to run quickly and reliably across compute environments, including VMs. Therefore, a container is an abstraction of the application layer (meaning each container emulates a different software application process). Although each container runs isolated processes, multiple containers can share a common operating system, for example, by starting in the same virtual machine. In contrast, a virtual machine is an abstraction of the hardware layer (meaning each virtual machine emulates a physical machine that can run software). Virtual machine technology can use a single physical server to run the equivalent of multiple servers (each of which is called a virtual machine). While multiple virtual machines can run on a single physical machine, each virtual machine typically has its own copy of the operating system, as well as the application and its associated files, libraries, and dependencies. Virtual machines are often referred to as compute instances or simply "instances." Some containers can run on instances that run container agents, while others can run on bare-metal servers.

[0064] In some implementations, the execution of edge-optimized computing instances is supported by a lightweight virtual machine manager (VMM) running on server 310, which launches edge-optimized computing instances based on application configuration files. These VMMs enable the launch of lightweight microvirtual machines (microVMs) in fractions of a second. These VMMs also enable container runtimes and container orchestrators to manage containers as microVMs. Nevertheless, these microVMs also leverage the security and workload isolation provided by traditional VMs, such as running them as isolated processes via a VMM, as well as the resource efficiency offered by containers. As used herein, a microVM refers to a VM that is initialized using a finite device model and / or a minimal OS kernel supported by a lightweight VMM, and each microVM can have a low memory overhead of <5 MiB, allowing thousands of microVMs to be packaged onto a single host. For example, a microVM can have a streamlined version of the OS kernel (e.g., only with the required OS components and their dependencies) to minimize startup time and memory footprint. In one implementation, each process of the lightweight VMM encapsulates one and only one microVM. The process can run the following threads: API, VMM, and vCPU. The API thread is responsible for the API server and its associated control plane. The VMM thread exposes the machine model, the minimal legacy device model, the microVM metadata service (MMDS), and the VirtIO device emulation network and block device. Additionally, there are one or more vCPU threads (one vCPU thread per guest CPU core).

[0065] Additionally, server 310 can host one or more data volumes 324 if required by the client. Volumes can be provisioned in their own isolated virtual network within provider underlying extension 300. Compute instance 312 and any volume 324 together constitute provider network data plane 116A within provider underlying extension 308B.

[0066] A local gateway 316 can be implemented to provide network connectivity between the provider's underlying extension 300 and the CSP network 302. The cloud provider can configure the local gateway 316 with IP addresses on the CSP network 302 and exchange routing data with the CSP network component 320 (e.g., via Border Gateway Protocol (BGP)). The local gateway 316 may include one or more routing tables that control the routing of inbound traffic to the provider's underlying extension 300 and outbound traffic leaving the provider's underlying extension 300. The local gateway 316 can also support multiple VLANs where different parts of the CSP network 302 use separate VLANs (e.g., one VLAN label for the wireless network and another for the fixed network).

[0067] In some implementations of the provider-level extension 300, the extension includes one or more switches, sometimes referred to as top-of-rack (TOR) switches (e.g., in rack-based implementations). The TOR switches connect to CSP network routers (e.g., CSP network component 320), such as provider edge (PE) or software-defined wide area network (SD-WAN) routers. Each TOR switch may include an uplink link aggregation (LAG) interface to the CSP network router, with each LAG supporting multiple physical links (e.g., 1G / 10G / 40G / 100G). These links may run the Link Aggregation Control Protocol (LACP) and be configured as IEEE 802.1q trunks to enable multiple VLANs on the same interface. This LACP-LAG configuration allows the edge location management entity of the cloud provider network 100's control plane to add more peering links to the edge location without rerouting. Each of the TOR switches can establish an eBGP session with the carrier PE or SD-WAN router. The CSP can provide a Private Autonomous System Number (ASN) for the edge location and the CSP network 302 to facilitate the exchange of routing data.

[0068] Data plane traffic originating from provider underlying extension 300 can have multiple different destinations. For example, traffic addressed to a destination in data plane 116A of cloud provider network 100 can be routed via the data plane connection between provider underlying extension 300 and cloud provider network 100. Local network manager 318 can receive packets from compute instance 312 addressed to, for example, another compute instance in cloud provider network 100, and encapsulate the packets with a destination as the underlying IP address of a server hosting another compute instance, and then send them (e.g., via a direct connection or tunnel) to cloud provider network 100. For traffic from compute instance 312 addressed to another compute instance hosted in another provider underlying extension 322, local network manager 318 can encapsulate the packets with a destination as the IP address assigned to the other provider underlying extension 322, thereby allowing CSP network component 320 to handle packet routing. Alternatively, if CSP network component 320 does not support traffic between edge locations, local network manager 318 can address packets to a repeater in cloud provider network 100, which can then forward the packets to another provider underlying extension 322 via its data plane connection (not shown) to cloud provider network 100. Similarly, for traffic from compute instance 312 address to a location outside CSP network 302 or cloud provider network 100 (e.g., on the Internet), if CSP network component 320 allows routing to the Internet, local network manager 318 can encapsulate the packets with a source IP address corresponding to an IP address in the carrier address space assigned to compute instance 312. Otherwise, local network manager 318 can forward the packets to an Internet gateway in cloud provider network 100, which can provide Internet connectivity for compute instance 312. For traffic from compute instance 312 addressed to electronic device 304, local gateway 316 can use Network Address Translation (NAT) to change the source IP address of the packet from an address in the address space of the cloud provider network to an address in the address carrier network space.

[0069] The local gateway 316, local network manager 318, and other local control plane components 314 may run on the same server 310 as the managed computing instance 312, on a dedicated processor integrated with the edge location server 310 (e.g., on an offload card), or may be executed by a server separate from those servers that manage customer resources.

[0070] Figure 4An exemplary cloud provider network, including geographically dispersed provider underlying extensions (or “edge locations”), is illustrated according to some implementation schemes. As shown, the cloud provider network 400 can be formed as multiple regions 402, where a region is a separate geographical area in which the cloud provider has one or more data centers 404. Each region 402 may include two or more Availability Zones (AZs) interconnected via a dedicated high-speed network such as, for example, fiber optic communication connections. An Availability Zone refers to an isolated fault domain comprising one or more data center facilities that have separate power, separate networking, and separate cooling relative to other Availability Zones. The cloud provider may endeavor to position Availability Zones within a region, spaced sufficiently far apart that natural disasters, widespread power outages, or other unforeseen events do not simultaneously take down more than one Availability Zone. Customers can connect to resources within the Availability Zones of the cloud provider network via publicly accessible networks (e.g., the Internet, cellular networks, CSP networks). A switching center (TC) is the primary backbone location linking customers to the cloud provider network and may be co-located within other network provider facilities (e.g., Internet service providers, telecommunications providers). Each region may operate two or more TCs for redundancy.

[0071] The number of edge locations 406 can be significantly higher than the number of regional data centers or availability zones. This widespread deployment of edge locations 406 can provide low-latency connectivity to the cloud for a much larger group of end-user devices (compared to those groups of end-user devices that happen to be very close to a regional data center). In some implementations, each edge location 406 may peer to a portion of the cloud provider network 400 (e.g., a parent availability zone or regional data center). This peering allows various components operating within the cloud provider network 400 to manage the computing resources of the edge location. In some cases, multiple edge locations may be located or installed in the same facility (e.g., a separate rack for a computer system) and managed by different zones or data centers to provide additional redundancy. It should be noted that while edge locations are generally described herein as being within the CSP network, in some cases, such as when the cloud provider network facility is relatively close to the communication service provider facility, the edge location may remain physically located within the cloud provider network while being connected to the communication service provider network via fiber optic or other network links.

[0072] Edge location 406 can be structured in several ways. In some implementations, edge location 406 can be an extension of the cloud provider's network layer, including a limited amount of capacity provided outside of an availability zone (e.g., in a small data center of the cloud provider or in other settings located near customer workloads and potentially far from any availability zone). Such edge locations can be referred to as local regions (due to being more local or closer to a group of users than a traditional availability zone). Local regions can be connected to publicly accessible networks such as the Internet in various ways (e.g., directly, via another network, or via a dedicated connection to a region). While local regions typically have more limited capacity than a region, in some cases, local regions may have considerable capacity, such as thousands or more racks. Some local regions may use infrastructure similar to that of a typical cloud provider data center, rather than the edge location infrastructure described herein.

[0073] As indicated herein, a cloud provider network can be structured as multiple regions, each representing a geographical area in which the cloud provider clusters its data centers. Each region may also include multiple (e.g., two or more) Availability Zones (AZs) interconnected via dedicated high-speed networks, such as fiber optic connections. AZs can provide isolated fault domains comprising one or more data center facilities, which have separate power supplies, separate networking, and separate cooling relative to those data center facilities in another AZ. Preferably, AZs within a region are positioned far enough apart that the same natural disaster (or other fault-causing event) will not simultaneously affect or take down more than one AZ at a time. Customers can connect to the cloud provider network's AZs via publicly accessible networks (e.g., the Internet, cellular networks).

[0074] The parenting of an edge location as an Availability Zone (AZ) or region within a cloud provider's network can be based on numerous factors. One such parenting factor is data attribution. For example, to retain data originating from a CSP network in a particular country, an edge location deployed within that CSP network can serve as the parent of an AZ or region in that country. Another factor is service availability. For instance, some edge locations may have different hardware configurations, such as the presence of components like local non-volatile storage (e.g., solid-state drives), graphics accelerators, etc., for customer data. Some AZs or regions may lack services that utilize these additional resources, so an edge location can serve as the parent of an AZ or region that supports the use of those resources. Yet another factor is latency between the AZ or region and the edge location. While deploying edge locations within a CSP network offers latency benefits, those benefits can be offset by the edge location serving as the parent of a distant AZ or region, introducing significant latency to traffic from the edge location to the region. Therefore, edge locations are typically the parent of nearby (in terms of network latency) AZs or regions.

[0075] Figure 5 An exemplary environment is illustrated, in which computing instances are launched at edge locations within a cloud provider network, according to some implementation schemes. As shown, cloud provider network 500 includes hardware virtualization service 506 and database service 508. Cloud provider network 500 has multiple edge locations 510. In this example, the multiple edge locations 510 are deployed in each of one or more CSP networks 501. Edge locations 510-1 to 510-M are deployed in CSP network 501-1, while other edge locations (not shown) may be deployed in other CSP networks (e.g., 501-2 to 501-N). CSP network 501 may be different networks or network slices of the same CSP or networks of different CSPs.

[0076] Figure 5 The circles numbered "1" through "3" illustrate an exemplary process by which user 138 (e.g., a customer of a cloud provider's network) can initiate a computing instance at one of the edge locations 510. Figure 5 At circle "1", user 138 uses electronic device 134 to request identification of an available edge location. As indicated above, communication between electronic device 134 and provider network 100 (such as requesting identification of an edge location to launch an instance at the edge location) can be routed via interface 104 (such as via application programming interface (API) calls implemented as a website or application). In addition to serving as a front end for control plane services, interface 104 can also perform operations such as verifying the identity and permissions of the user initiating the request, evaluating the request, and routing it to the appropriate control plane service.

[0077] A request to identify edge locations may include zero or more parameters to filter, limit, or otherwise constrain the returned set of edge locations to fewer than all edge locations 510. For example, one such parameter could be an identifier of a specific CSP (e.g., when cloud provider network 500 has edge locations integrated with multiple CSPs). Another such parameter is an identifier of the specific network of the CSP (e.g., if the CSP has edge locations for 4G networks, 5G networks, etc.). Yet another such parameter might restrict the returned edge locations to those with certain hardware support (e.g., accelerators). Yet another such parameter might restrict the returned edge locations to those near or within a certain distance of a geographic indicator (e.g., city, state, zip code, geographic coordinates, etc.).

[0078] In the illustrated implementation, the request is processed by hardware virtualization service 506. Upon receiving a request, hardware virtualization service 506 retrieves the identity (if any) of the edge location that satisfies the request from edge location data 509. Exemplary edge location data 509 may be stored in a database provided by database service 508. For each edge location, edge location data 509 may include an identifier assigned to the edge location, an indication or identifier of the CSP network in which the edge location is deployed, and an indication or identifier of the geographic location of the edge location. As an example, a user might request the identification of edge locations within a 10-mile radius of New York City on the 5G network of CSP company X. Upon identifying edge locations that satisfy the user's request, hardware virtualization service 506 returns a list or set of edge locations to electronic device 134.

[0079] exist Figure 5 At circle "2", user 138 requests to start a compute instance at a specified edge location. Such a request can include various parameters, such as the type of instance to be started. Upon receiving the request, hardware virtualization service 506 can perform checks to ensure that the specified edge location has sufficient capacity to start the instance and perform other operations. It should be noted that in some implementations, hardware virtualization service 506 may avoid returning to an edge location at or near full resource capacity in response to the user request at circle "1" to avoid rejecting the request at circle "2".

[0080] exist Figure 5 At circle “3”, hardware virtualization service 506 issues control plane commands to the specified edge location to start the requested compute instance (e.g., via agent 130). For example, hardware virtualization service 407 may then issue commands to the VMM at the edge location or on the edge location server to start the client's compute instance.

[0081] Because a large number of edge locations may be deployed, customers may find it difficult to manually identify and select the appropriate edge locations for their applications. (Based on the reference...) Figure 6 and Figure 7 The described method can be executed by components of a cloud provider's network to select the edge location of a managed computing instance.

[0082] Figure 6 Another exemplary environment for launching virtualized computing resources (including VMs, microVMs, and / or containers) at the edge location of a cloud provider network, according to some implementation schemes, is shown. As shown, the cloud provider network 600 includes a hardware virtualization service 606, an edge location layout service 620, and a database service 622. The cloud provider network 600 has multiple edge locations 510 in various CSP networks 501 (e.g., such as those mentioned above). Figure 5 (As described).

[0083] Figure 6 The circles numbered "1" through "3" illustrate an exemplary process by which user 138 (e.g., a customer of a cloud provider's network) can initiate a computing instance at one of the edge locations 510. Figure 6 At circle "1", user 138 sends a request to hardware virtualization service 606 to start a compute instance. Here, the parameters of the request may include a geographic identifier and indications of latency constraints or requirements. Depending on the implementation, the geographic identifier can take various forms (e.g., geographic coordinates, postal codes, metropolitan areas, etc.). For example, a geographic identifier might be a postal code associated with region 698, coordinates within region 698, or a region corresponding to region 698 (e.g., a city area). The latency indicator can be specified based on the time (e.g., less than 10 milliseconds) between the device associated with the geographic identifier (e.g., in region 698) and the server ultimately selected as the host for the requested compute instance.

[0084] More complex launch requests from user 138 may include parameters specifying additional latency requirements. For example, the request may specify latency requirements for communication between the requested instance and the device associated with the geographic indicator (e.g., within a region) and between the requested instance and the cloud provider network region or availability zone whose parent is ultimately selected as an edge location for hosting. As another example, the request may request multiple instances distributed across multiple edge locations, specifying latency requirements for communication between the requested instance and the device associated with the geographic indicator, as well as between the edge locations.

[0085] Additional startup request parameters may include the number of compute instances to be started, the type of compute instances, and whether the compute instances (in the case of multiple compute instances) should be tightly packed together (e.g., on the same server or edge location) or spread out (e.g., on a server or edge location).

[0086] For reference Figure 5 In the described method, additional startup parameters can be provided to limit the edge location layout service 720's search for suitable edge locations (e.g., parameters that identify a specific CSP or a specific network of a CSP, parameters that identify the hardware requirements of the edge location, etc.).

[0087] In some implementations, the requested initiation is constrained to Figure 6 The parameters at circle "1" can be stored as part of an application profile. The application profile can include parameters related to executing the user workload at the provider's underlying extensions (e.g., the expected amount of compute resources to be dedicated to instances launched based on the profile, the expected latency and geographical constraints of the launched instances, instance layout and extension configuration, etc.). Cloud provider network customers may have previously created application profiles, which can then be referenced in requests such as those to launch instances at circle "1".

[0088] In some implementations, a parameter value that can be included in the application configuration file may be a value identifying a resource to be used as a template for launching a compute instance based on the application configuration file. For example, if a user creates a VM image, virtual device, container image, or any other type of resource that can be used to launch a compute instance (such as, for example, a VM, microVM, container, etc.), the user can provide an identifier for the resource (e.g., an identifier for a resource known to cloud provider network 100). In some implementations, the user can provide an identifier for the storage location of the resource that can be used to launch the compute instance (e.g., a URL or other identifier for a storage location within cloud provider network 100 or another location where the resource is stored).

[0089] In some implementations, another exemplary parameter that can be specified in the application profile includes parameters relating to compute resources to be dedicated to the profile-based instance. For example, a user can specify resource constraints based on CPU, memory, networking performance, or any other resource-related parameters (e.g., a user might specify that the application profile-based instance should be allocated two vCPUs, 8 GiB of memory, up to 10 Gbps of networking, or any other combination of resources) such that the requested resources are provided to the application profile-based instance (assuming the requested resources are available at any provider underlying extension location that satisfies other application profile constraints). In some implementations, a user can specify resource constraints based on a defined instance type (e.g., an instance type associated with a defined amount of CPU, memory, networking, etc., resources defined by cloud provider network 100). Other resource-related parameters may include block device mappings to be used by the launched instance, kernel version, etc.

[0090] In some implementations, other exemplary parameters include parameters relating to other aspects of deploying edge-optimized instances at provider underlays. For example, one communication service provider-related parameter that can be specified includes an identifier of a particular communication service provider (e.g., to indicate that a user wishes to launch an instance at a provider underlay associated with communication service provider A or B, but not at a provider underlay associated with communication service provider C). Yet another exemplary communication service provider-related parameter that can be specified includes one or more specific geographic locations where the edge-optimized instance is expected to launch (e.g., a provider underlay near downtown Austin, a provider underlay near the San Francisco Bay Area, a provider underlay in the Southwest or Northeast, etc.). Yet another exemplary parameter includes a latency profile for executing user workloads at the provider underlay, where the latency profile typically indicates the expected latency between the edge-optimized instance and the end user or other network points (e.g., at a PSE with 20 milliseconds or less latency to the end user, at a PSE near Los Angeles with 30 milliseconds or less latency to the end user, etc.).

[0091] In some implementations, other exemplary parameters that can be specified in the application configuration file include various networking configurations. For example, to enable communication between an intra-regional application running in a private network and an application running in a provider-level extension, the application configuration file can be configured to provide a private network endpoint to the intra-regional private network to invoke an edge-optimized instance. To enable bidirectional communication, customers can also provide a private network endpoint to their provider-level extension application, which can be used for communication from the provider-level extension to the region.

[0092] In some implementations, other exemplary parameters that can be specified in the application configuration file include scaling strategies used once one or more instances have been launched based on the application configuration file. For example, a user can specify inward and outward scaling strategies for their application in the application configuration file, where such strategies enable the adjustment of capacity in and on the provider's underlying scaling locations. In some implementations, when scaling inward, new capacity is launched in the same location under load by default and scaled to other locations, provided they meet client latency constraints (if any). For example, if no client latency limit is specified, new capacity can be added in the same location under load and scaled to other locations until the monitored metric falls below a scaling threshold.

[0093] exist Figure 6 At circle "2", hardware virtualization service 606 requests from edge location layout service 620 the identifier of candidate edge location 510 that satisfies the parameters of the user's startup request. Edge location layout service 622 can evaluate the parameters against latency data 609. Typically, latency data 609 provides an indication of latency between points within CSP network 501 (e.g., base stations providing connectivity within area 698 and edge location 510) and possibly between points within CSP network 501 and points in cloud provider network 600 (e.g., compute instances hosted by servers in cloud provider network data centers). Latency data 609 may also include geographic data about the location of various access points in CSP network 501 to allow edge location layout service 620 to associate user-specified geographic indicators with CSP network (e.g., coverage areas of base stations or other devices through which electronic devices access CSP network 501). Access points (sometimes referred to as entry points) include devices (e.g., base stations) through which CSP subscriber devices connect to the CSP network. The delay data 609 can be exported in several ways, some of which are described below. As shown, the delay data 609 is stored in a database hosted by the database service 622. In other embodiments, the delay data 609 can be obtained from the services of the CSP network (e.g., the edge location layout service 620 queries the services of the CSP network instead of querying the database service 622).

[0094] Upon receiving a request from hardware virtualization service 606 for suitable edge locations that meet customer requirements, edge location layout service 622 can access latency data 609 to identify which edge locations meet those requirements. The example is illustrative. Assume the user has provided a geographic indicator corresponding to region 698. Wireless CSP network 501 may include many base stations, some of which provide coverage for geographic region 698. The routes between those base stations and edge locations 510 may differ (e.g., some may have to traverse aggregation sites, such as aggregation site 206, some may have additional hops in the network path from the base stations to the edge locations, etc.). Latency data may include point-to-point latency between base stations and edge locations, and edge location layout service 620 can identify a set of candidate edge locations with communication latency that meets customer latency constraints based on those latency. For example, edge location layout service 620 may determine that latency 1 to edge location 510-1 meets customer constraints, while latency 2 to another edge location 510 does not. Therefore, the edge location layout service 620 returns edge location 510-1 as a candidate edge location to the hardware virtualization service 606.

[0095] In addition to identifying edge locations that meet customer latency requirements, the edge location layout service 622 can further narrow down suitable edge locations using other customer parameters (if specified) (e.g., edge locations for a specific CSP, specific networks of the CSP, etc.).

[0096] Based on the candidate edge locations (if any) returned by the edge location layout service 620, the hardware virtualization service 606 can return an error to the client if the request cannot be satisfied, or continue to start the compute instance. For example, if no edge location satisfies the client's latency requirements, or if the client has requested N compute instances distributed across N edge locations but fewer than N edge locations satisfy the client's latency requirements, the request may fail. Assuming the client's request can be satisfied, the hardware virtualization service 606 can issue control plane commands to the edge locations to start the requested instance, such as... Figure 6 The circle "3" indicates (for example, see the above for the...) Figure 5 (Description of the circle "3").

[0097] In some cases, the number of suitable edge locations returned by the edge location layout service 622 may exceed the number of compute instances requested by the client. In such cases, the hardware virtualization service 606 can proceed with additional selection criteria to choose which suitable edge location will be used to host the client-requested compute instances. The hardware virtualization service 606 can score each of the suitable edge locations using a cost function based on various criteria and select the “best” edge location based on its score relative to other edge locations. One such criterion is capacity cost; a PSE deployed in Manhattan may have a higher monetary cost than a PSE deployed in Newark, New Jersey (e.g., based on providing lower latency for users in Manhattan or increased demand for the site). Another such criterion is the available capacity on the suitable edge location. One way to measure available capacity is to track the number of compute instances previously launched at each edge location or each edge location server. The hardware virtualization service 606 can (e.g., in a database) track which edge locations have been previously used to launch compute instances and the resource consumption of those compute instances. Another way to measure available capacity is based on the resource utilization of the edge location or the edge location server. Agents or other processes running locally at the edge location or on the edge location server can monitor the utilization of the processors, memory, network adapters, and storage devices used to host computing instances and report the utilization data to the hardware virtualization service 606. The hardware virtualization service 606 can select an edge location with the highest capacity (or lowest utilization) from suitable edge locations returned by the edge location layout service 620.

[0098] Various methods are possible for obtaining latency data 609, including those described below. To facilitate a robust set of customer latency requirements, edge location layout service 622 may use one or more methods described herein, or other methods, to determine, for example, latency between end-user electronic devices and base stations, base stations and edge locations, base stations and cloud provider network areas or availability zone data centers, edge locations and edge locations, and edge locations and cloud provider network areas or availability zone data centers. Latency generally refers to the one-way time between a device sending a message to a receiver and the receiver receiving the message, or the round-trip time between a device issuing a request and subsequently receiving a response to said request. In some embodiments, latency data 609 provides or allows the derivation of latency between various points for layout determination by edge location layout service 622.

[0099] According to the first method, the CSP network may include a latency service. The latency service may periodically receive or otherwise monitor latency throughout the CSP network. The latency service may include an API through which the edge location layout service 622 can issue calls to obtain latency data 609. This approach may be referred to as a query-based approach. An exemplary API of the latency service receives one or more routes (e.g., specified via endpoints within the CSP network) and returns the latency of the routes. Providing identifiers of various endpoints in the CSP network (e.g., by IP address), the edge location layout service 622 can use the latency service of the CSP network to construct a view of point-to-point latency across the CSP network. For example, based on knowledge of various access points (e.g., base stations) of the CSP network, the coverage areas of the access points, and edge locations, the edge location layout service 622 can construct a latency dataset that correlates geographical areas with edge locations. Additionally, based on knowledge of various edge locations integrated with the CSP network, the edge location layout service 622 can also measure latency between the cloud provider network and each of the edge locations. For example, edge location layout service 622 can store or cache responses from latency services and other latency measurements in the database of database service 622.

[0100] According to the second method, the CSP can provide detailed information about its network topology, from which the edge location layout service 622 can derive information for making layout determinations based on distance and hop delay models between various points in the network. This method can be referred to as a model-based method. Network topology information can be provided in a graph or other suitable data structure representing, or converted to, the number of network hops and distances between network nodes (e.g., between base stations and edge locations, between edge locations, and between edge locations and cloud provider networks, which may be extended by the cloud provider using network topology information related to connectivity between the CSP network and the cloud provider network). Additionally, network topology information can include information related to the geographic location of the end-user device's access point to the network (e.g., base station coverage). Using a set of heuristics, the network topology information can be used to model various delays (e.g., point-to-point delays) through the CSP network to generate delay data 609. For example, heuristic methods could include estimated delays in signals between network nodes at a given distance (e.g., using the speed of light), modeled delays added by various hops through the network (e.g., due to processing delays at routers or other network devices), etc. Because network topology may change over time, CSPs can periodically provide updated network topology information.

[0101] According to the third approach, the CSP and / or cloud provider can establish a network of "publisher" nodes that collect latency data and report it to the edge location deployment service 622. Such publisher nodes can collect latency data in various ways, such as inspecting other devices, subscribing to events emitted by CSP network components, or periodically polling the CSP network API to collect QoS data. While similar to the query-based approach (which provides a more up-to-date view of network latency than the model-based approach), the third approach (referred to as the monitor-based approach) can be implemented with less reliance on the CSP (whether by gaining access to internal networking APIs such as latency services, requiring the CSP to deploy latency monitoring facilities that may not exist, or by relying on the CSP to obtain network topology data). For example, edge locations and / or end-user electronics may include applications that monitor latency to other devices. At the edge location, the application may be executed by a compute instance or as a control plane component. At the end-user electronics, the application may be a background process incorporated as part of a software development kit for deploying the application to the end-user device. In either case, the application may periodically obtain identifiers of other edge locations, base stations, or CSP network access points and / or electronic devices connected to the CSP network (e.g., via IP address) from the services of the cloud provider network or the CSP network, measure the latency of the identified devices (e.g., via authentication requests), and report the results to the edge location deployment service 622. In the case of end-user devices, the application may further report latency data between the end-user device and its access point to the CSP network (e.g., base station). The edge location deployment service 409 may aggregate and store the reported data as latency data 609.

[0102] Figure 7 Another exemplary environment is shown, in which a computing instance is launched at an edge location of a cloud provider network according to some implementation schemes. As shown, the cloud provider network 700 includes a hardware virtualization service 706, an edge location deployment service 720, and a database service 722. Although not shown, the cloud provider network 700 has multiple edge locations in the CSP network (e.g., such as those mentioned above). Figure 5 (As described).

[0103] Figure 7 The circles numbered "1" through "3" illustrate an exemplary process by which user 138 (e.g., a customer of a cloud provider's network) can initiate a computing instance at one of the edge locations 510. Figure 7At circle "1", user 138 sends a request to hardware virtualization service 606 to launch a compute instance. Here, the parameters of the request may include a device identifier and an indication of latency constraints or requirements. Depending on the implementation, the device identifier may take various forms (e.g., IMEI number, IP address, etc.). The latency indicator can be specified based on the time (e.g., less than 10 milliseconds) between the identified device and the server ultimately selected to host the requested compute instance. The launch request can be as described in the reference above. Figure 6 The various other parameters described.

[0104] exist Figure 7 At circle "2", hardware virtualization service 706 requests from edge location layout service 720 the identifier of candidate edge location 510 that meets the parameters of the user's startup request. Edge location layout service 720 continues to identify candidate edge locations 510, as described in the reference above. Figure 6 The circle "2" describes this. To this end, the edge location layout service 720 first uses the device identifier to obtain a geographic indicator associated with the device location. For example, the edge location layout service 720 can request a geographic indicator from the device location service 742 of the CSP network 701 (e.g., by providing an IP address or IMEI number). The device location service 742 can provide a geographic indicator for the identified device. As another example, the edge location layout service 720 can request a geographic indicator from the identified device 790. For example, the device identifier might be the IP address of the electronic device 790 performing the device location agent 744. The device location agent 744 can provide a geographic indicator for the electronic device 790. The edge location layout service 720 can use the geographic indicator along with user-specified delay constraints to identify candidate edge locations, as described above for... Figure 6 As described.

[0105] Based on candidate edge locations (if any) returned by the edge location layout service 720, the hardware virtualization service 706 can return an error to the client if the request cannot be satisfied, or continue to start the compute instance. For example, if no edge location satisfies the client's latency requirements, the request may fail. Assuming the client's request can be satisfied, the hardware virtualization service 706 can issue control plane commands to the edge location to start the requested instance, such as... Figure 7 The circle "3" indicates (for example, see the above for the...) Figure 5 and / or Figure 6 (Description of the circle "3").

[0106] It should be noted that in some implementations, the geographic indicator may be inferred based on the latency to the electronic device rather than from the device location service 742 or agent 744 to obtain a specific geographic indicator. For example, user 138 may provide a device identifier and latency requirements. In this case, the designated device may be used as an agent for determining the geographic indicator. For example, hardware virtualization service 706 or edge location layout service 720 may enable multiple other devices (not shown) to check the device's IP address from several known locations to infer the device's geographic location and thus the corresponding geographic indicator.

[0107] In addition to sending control plane commands to selected edge locations to initiate compute instances, hardware virtualization services 506, 606, and 706 can also send control plane commands to selected edge locations to associate IP addresses on the CSP network with the initiated compute instances. The IP addresses can be selected from a pool of IP addresses in the CSP network address space assigned to the PSE. For example, the initiated instance might be assigned IP address "A," which the PSE gateway advertises to the CSP network components, such that when a device connected via the CSP network sends a packet to address "A," the packet is routed to the PSE.

[0108] Figure 8 This is an exemplary environment, based on some implementation schemes, where compute instances are launched due to the mobility of electronic devices. In some cases, a cloud provider network can automatically launch compute instances at various edge locations deployed within a communications service provider network (CSP) to continue meeting customer-specified latency constraints, even if the movement of the electronic device changes its access point to the CSP network. As the mobile device changes its access point to the CSP network, the latency between those access points and specific edge locations deployed within the CSP network may change. This is provided as an example and for reference. Figure 2 The latency of electronic device 212 connecting to a compute instance hosted by edge location 216 via access point RAN 202 may be lower than the latency to a compute instance hosted by edge location 228 when connected via access point RAN 202, due to additional traffic routing via aggregation site 206. Conversely, another electronic device connected via access point RAN 204 may have higher latency to a compute instance hosted by edge location 216 compared to a compute instance hosted by edge location 228. If a cloud provider network customer provides latency constraints as part of launching a compute instance within an edge location of the CSP network, a change in the access point of a device connected via the CSP network may result in a violation of those constraints. This can occur, for example, when a cloud provider network customer launches a compute instance to provide low-latency connectivity to a specified device and that device subsequently changes its access point.

[0109] Return to Figure 8 The control plane components of the cloud provider network 800 manage edge locations 810-1 and 810-2 deployed within the CSP network 801. The compute instance 813 hosted by edge location 810-1 initially meets customer-specified latency constraints because device 890 is already connected to the CSP network 801 via access point 888. At some point, device 890 changes its access point to the CSP network 801 from access point 888 to access point 889, and Figure 8 The exemplary process of tracing circles numbered "1" to "8" initiates computational instance 815 to illustrate the movement of electronic device 890.

[0110] exist Figure 8 At circle "1", the mobility management component 862 of CSP network 801 manages the mobility of devices (including electronic device 890) connected to CSP network 801 as those devices move and may change access points to CSP network 801. Such changes in access points are referred to herein as "mobility events". Mobility management components are typically defined components in a wireless network, such as the Access and Mobility Management Function (AMF) of a 5G network or the Mobility Management Entity (MME) of a 4G or LTE network. For example, the detection of such mobility events in CSP network 801 may be based on a signal measured by electronic device 890 and reported to CSP network 801 periodically or when other conditions are met. These measurements may include, for example, received power or signal quality perceived by electronic device 890 from different geographical coverage areas (or "cells") provided by different access points (e.g., access points 888, 889). In some implementations, the mobility management component 862 and / or other components of the CSP network 801 may use these measurements to determine whether to transfer the electronic device 890 from one access point to another, and which access point is the optimal connection point.

[0111] In this example, electronic device 890 is moving, making its connection to CSP network 801 via access point 889 better or likely to be better than its connection to CSP network via access point 888. Figure 8At circle “2”, the mobility management component 862 of CSP network 801 provides an indication to edge location connectivity manager 811 of a mobility event involving electronic device 890. As indicated above, mobility management component 862 may make such a determination based on measurements received from electronic device 890 or on signal quality data otherwise obtained by said component. In some embodiments, the indication of a mobility event is an indication that the electronic device is actually moving from a first cell provided by first access point 888 to a second cell provided by second access point 889. In some embodiments, the indication of a mobility event includes one or more predictions that the electronic device 890 will move from a cell provided by first access point 888 to one or more other cells provided by other access points of CSP network 801. Such predictive mobility events may include the probability that the electronic device 890 will change its access point to one or more other cells and may include an indication of when the event will actually occur.

[0112] In some implementations, the mobility management component 862 sends mobility events to all or part of the total number of edge locations deployed to the CSP network 801. In other implementations, the edge location connection manager 811 at each edge location tracks electronic devices connected to computing instances hosted by the edge location and requests the mobility management component 862 to send updates related to those electronic devices.

[0113] exist Figure 8 At circle "3", the edge location connection manager 811 sends an indication of the mobility event and device-specific connection data to the edge location mobility service 830 of the cloud provider network 800. In some embodiments, the edge location connection manager 811 obtains some or all of the device-specific connection data from connection data 812 maintained by the edge location 810-1. Connection data 812 may include source and destination network addresses associated with the connection between the electronic device (e.g., electronic device 890) and the computing instance (e.g., computing instance 813), the time the connection was established, the status of the connection, the type of protocol used, etc. In some embodiments, the edge location connection manager 811 examines the connection data 812 to determine whether the electronic device associated with the received mobility event is connected to one or more of the computing instances hosted by the edge location 810-1, and then sends an indication of the mobility event and device-specific connection data of the electronic device to the edge location mobility service 830.

[0114] In some implementations, the edge location mobility service 830 determines that the communication latency between the electronic device 890 and the first computing instance 813 via the second access point 889 will not meet the latency constraint (and therefore a migration of the computing instance will occur to satisfy the latency constraint). The constraint may not be met because additional hops or distances are introduced by routing communication from the second access point 889 to the existing computing instance 813, or because the edge location 810-1 is not reachable from the second access point 889 (e.g., due to the network topology and configuration of the CSP network 801). The edge location mobility service 830 can retrieve data stored during the request to launch an instance (e.g., when a customer requests to launch an instance such as the one referenced above). Figure 5 , Figure 6 and Figure 7 The described instance (data stored in a database, such as part of an application configuration file) obtains latency constraints associated with the instance. Edge location mobility service 830 can obtain the latency between an existing computing instance and a new access point, which can be determined from latency data (e.g., latency data 609).

[0115] Although not shown, in some implementations, the edge location connection manager for a given edge location can send connection data to the edge location mobility service, and the mobility management component of the CSP network 801 can send mobility events to the edge location mobility service.

[0116] Assuming the new latency between compute instance 813 and access point 889 exceeds the latency constraint, edge location mobility service 830 sends an instance launch request to hardware virtualization service 806 of cloud provider network 800, such as... Figure 8 The circle "4" indicates this. As indicated above, the geographic indicator can take various forms depending on the implementation (e.g., geographic coordinates, postal codes, metropolitan areas, etc.). In some implementations, the geographic indicator is based on an indication of a mobility event provided to the edge location mobility service 830, for example, such that the geographic indicator corresponds to the location of the access point to which the electronic device 890 is moving or is expected to move. In some implementations, additional initiation parameters may include an identifier of a specific CSP or a specific network of the CSP, parameters identifying the hardware requirements of the edge location, etc., enabling the execution of a computational instance to migrate to an edge location with characteristics similar to edge location 810-1.

[0117] exist Figure 8At circle "5", hardware virtualization service 806 requests identifiers of candidate edge locations from edge location layout service 820, the candidate edge locations satisfying parameters of an initiation request received from edge location mobility service 830. Edge location layout service 820 can evaluate parameters against latency data available to edge location layout service 820. Typically, latency data provides an indication of latency between points within CSP network 801 (e.g., base stations providing connectivity within areas and edge locations) and between points possibly within CSP network 801 and points in cloud provider network 800 (e.g., compute instances hosted by servers in cloud provider network data centers). Latency data may also include geographic data about the location of individual access points in CSP network 801, allowing edge location layout service 820 to associate specified geographic indicators with CSP network (e.g., coverage areas of base stations or other equipment through which electronic devices access CSP network 801).

[0118] Upon receiving a request from hardware virtualization service 806 for a suitable edge location that meets various parameters (specified in the request), edge location layout service 820 can access latency data and other information to identify which edge locations meet those requirements. Based on candidate edge locations (if any) returned by edge location layout service 820, hardware virtualization service 806 can select one of the candidates, such as by evaluating the cost function of the candidates as described herein. It should be noted that, as indicated above, a mobility event may include the probability that electronic device 890 moves from access point 889 to 890 (even though a handover has not yet occurred). In this case, hardware virtualization service 806 can take the probability factor into account in the cost function to determine whether to initiate an instance. For example, if the probability of movement is low (e.g., <50%) and the resource utilization of the candidate edge location is high (e.g., there is sufficient unused resource capacity for 10 new instances out of a total capacity of 100 instances), hardware virtualization service 806 may choose to wait for the actual mobility event.

[0119] In some implementations, the identifier of the new access point 889 can be used as a proxy for a geographic indicator. Edge location mobility service 830 can receive the identifier from mobility management component 862 (possibly via edge location connection manager 811) and send the identifier to hardware virtualization service 806. Hardware virtualization service 806 can then send the identifier to edge location layout service 820, which can then use the access point identifier to estimate latency for candidate edge locations.

[0120] In this example, the edge location layout service 820 returns the identifier of edge location 810-2 as a candidate edge location, and if more than one candidate is returned, the hardware virtualization service 806 selects edge location 810-2. The hardware virtualization service 806 issues control plane commands to the local resource manager 814 at edge location 810-2 to start the requested instance, such as... Figure 8 The circle "6" indicates this. In some implementations, the computing instance 815 launched at edge location 810-2 in response to a mobility event associated with electronic device 890 may be based on the same resources (e.g., virtual machine image, container image, etc.) used by the computing instance 813 previously connected to the launching electronic device 890. Once launched, electronic device 890 can establish a connection with the computing instance 815 launched at edge location 810-2 and resume using any applications interacting with the device.

[0121] In some implementations, an IP address pool in the CSP network address space is reserved by the CSP network for one or more edge locations. IP addresses are assigned from the pool to compute instances launched at those edge locations. In this way, compute instances hosted at edge locations can be treated as another device on the CSP network, thereby facilitating traffic routing between electronic devices (e.g., electronic device 890) connected via the CSP network and compute instances hosted at edge locations. In some implementations, a control plane component, such as hardware virtualization service 806, assigns a new IP address from the pool to a new compute instance 815.

[0122] Hardware virtualization service 806 can return the identifier of the new compute instance 815 to edge location mobility service 830. In some implementations, edge location mobility service 830 can examine the connectivity data associated with the original compute instance 813 to determine whether to allow compute instance 813 to run or migrate it to compute instance 815. For example, if compute instance 813 is still communicating with other electronic devices, compute instance 813 can continue to support those other devices while electronic device 890 begins communicating with compute instance 815. In some implementations, edge location mobility service 830 can trigger a "migration" targeting compute instance 815 and source compute instance 813, such as... Figure 8The circle "7" indicates this. Migration typically refers to moving virtual machine instances (and / or other resources) between hosts. Different types of migration exist, including live migration and restart migration. During a restart migration, the customer experiences an interruption and effective power cycling of their virtual machine instances. For example, a control plane service can coordinate a restart migration workflow that involves tearing down the current compute instance on the origin host and subsequently creating a new compute instance on the new host. The instance is restarted by shutting down on the origin host and then restarting on the new host.

[0123] Live migration refers to the process of moving running virtual machines or applications between different physical machines without significantly disrupting the availability of the virtual machines (e.g., the end user will not notice the downtime). When the control plane performs a live migration workflow, it can create new "inactive" compute instances on the new host, while the original compute instance on the original host continues to run. Virtual machine state data (such as memory (including any state in memory of running applications), storage and / or network connectivity) is transferred from the original host with the active compute instance to the new host with the inactive compute instance. The control plane can transition the inactive compute instance to an active compute instance and degrade the original active compute instance to an inactive compute instance, after which the inactive compute instance can be discarded.

[0124] like Figure 8 The circle "8" indicates that state data migrating from compute instance 813 to compute instance 815 can be sent directly through CSP network 801. In other implementations, state data can traverse a portion of cloud provider network 800 (e.g., if one edge location cannot communicate with another edge location through CSP network 801).

[0125] It should be noted that edge location 810-1 includes a local resource manager 814, and edge location 810-2 includes an edge location connection manager 811 and connection data 812. Although Figure 8 The discussion envisions that electronic device 890 moves to a “closer” edge location 810-2, or vice versa, or that electronic device 890 may later move to another access point (not shown) that fails to meet the delay constraints in communication with edge location 810-2. Therefore, the description of the operation of edge location 810-1 can be applied to edge location 810-2, and vice versa.

[0126] Figure 9This is a flowchart illustrating operations of a method for launching compute instances at the edge of a cloud provider network, according to some embodiments. Some or all of the operations (or other processes described herein, or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions, and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that executes collectively on one or more processors, implemented by hardware, or implemented by a combination thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program that includes instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some embodiments, one or more (or all) of the operations are executed by one or more local control components (e.g., a local resource manager or other components that manage the launching, configuration, and termination of compute instances such as virtual machines or containers) deployed within a provider underlying extension of a communications service provider network in another diagram.

[0127] The operation includes: at block 902, receiving a message to launch a customer computing instance at a provider underlying extension of a cloud provider network embedded within a communication service provider network, wherein the message is received from a control plane service of the cloud provider network. The operation also includes: at block 904, launching a customer computing instance on a computer system of the provider underlying extension, the computer system having the capability to execute the customer computing instance, wherein the provider underlying extension communicates with the cloud provider network via the communication service provider network, and wherein the customer computing instance communicates with a subscriber's mobile device via the communication service provider network.

[0128] like Figure 2 As shown, a Cloud Provider Network Underlying Extension (PSE) can be deployed within a Communication Service Provider (CSP) network. These CSP networks typically provide subscriber devices with data connectivity to the CSP network and other networks, such as the Internet. The PSE may include computing resources (e.g., processors, memory, etc.) on which customers of the cloud provider network can launch compute instances, such as virtual machines or containers. Local management components of the PSE (such as a container engine or virtual machine manager) can manage the compute instances hosted using the PSE resources. Control plane components of the cloud provider network (such as hardware virtualization services) can issue commands to the local management components to launch instances. These commands can be routed through the CSP network via a secure tunnel between the cloud provider network and the PSE.

[0129] Deploying or integrating a PSE within a CSP network can reduce latency that might otherwise exist when hosting compute instances further away from the CSP network (e.g., in a regional data center within a cloud provider's network). For example, communication between a compute instance hosted by a PSE deployed within the CSP network and a mobile device can be routed entirely within the CSP network, without requiring traffic to leave the CSP network (e.g., via an internet switch).

[0130] Figure 10 This is a flowchart illustrating operations of another method for launching computing instances at the edge location of a cloud provider network, according to some embodiments. Some or all of the operations (or other processes described herein, or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions, and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that commonly executes on one or more processors, implemented by hardware, or implemented by a combination thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program that includes instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some embodiments, one or more (or all) of the operations are executed by one or more control plane services of the cloud provider network of another graph (e.g., hardware virtualization services 606, 706, edge location layout services 620, 720).

[0131] The operation includes, at box 1002, receiving a request from a customer to launch a compute instance at a service location within the cloud provider network, wherein the request includes latency requirements. As explained above, one of the advantages of deploying or embedding provider underlying extensions or edge locations within a communications service provider network is reduced latency between end-user devices and customer compute instances. To provide customers of the cloud provider network with the ability to utilize reduced latency, it is beneficial to allow customers to specify latency requirements or constraints that control where their compute instances are ultimately launched. Therefore, the cloud provider network may include interfaces such as APIs through which customers can launch instances based on requested latency requirements, as referenced above. Figure 6 and Figure 7 As described.

[0132] The operation includes, at block 1004, selecting a provider underlying extension from a plurality of provider underlying extensions of a cloud provider network to host a compute instance, wherein the selection is at least in part based on latency requirements, and wherein the selected provider underlying extension is connected to a communication service provider network and is at least in part controlled by services of the cloud provider network via a connection through at least a portion of the communication service provider network. (See reference...) Figure 6 and Figure 7As explained, edge location deployment services 620 and 720 can evaluate candidate edge locations to determine which edge locations meet the customer's latency requirements. To this end, the edge location deployment service obtains geographic indicators, which may be associated with a geographic area covered by one or more access points in the CSP network, and evaluates the latency from said one or more points to the edge locations deployed in the CSP network. Such geographic indicators may be provided by a request received in box 1002 (e.g., a geographic area specified by the customer, such as a city, postal code, etc.), or obtained, for example, by determining the location of the device identified by the request. Various techniques can be used to obtain latency values ​​or estimates between points in the CSP network (e.g., edge location to access point). The edge location deployment service can determine which edge locations, if any, meet the customer's latency requirements and return the candidate set to the hardware virtualization service. The set may include an indication of the latency margin between each of the edge locations in the set relative to the latency requirements. Using a cost function or other techniques to rank the candidate edge locations, the hardware virtualization service can select the edge location on which the requested compute instance is hosted. Factors that can be used in the selection process include available hardware capacity at candidate edge locations, overall capacity utilization, capacity cost, and latency margin relative to customer latency requirements.

[0133] The operation includes, at box 1006, sending a message to cause the selected provider underlying extension to launch a compute instance for the customer. Based on the selected provider underlying extension, the hardware virtualization service may (e.g., via a tunnel between the cloud provider network and the provider underlying extension deployed within the CSP network) issue one or more commands to the provider underlying extension to launch the requested instance.

[0134] Figure 11 This is a flowchart illustrating operations of a method for initiating a computing instance due to the mobility of an electronic device, according to some embodiments. Some or all of the operations (or other processes described herein, or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions, and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that commonly executes on one or more processors, implemented by hardware, or implemented by a combination thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program that includes instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some embodiments, one or more (or all) of the operations are executed by one or more control plane services (e.g., edge location mobility service 830, hardware virtualization service 806, edge location layout service 820) of another cloud provider network.

[0135] The operation includes, at block 1102, receiving a message including an indication of a mobility event associated with a mobile device on a communication service provider network, wherein the mobility event indicates a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point. (See reference...) Figure 8 As explained, when a device moves between different access points in a CSP network, the initial layout determination of the computing instance based on latency requirements may no longer meet those requirements. To continue meeting the latency requirements, the cloud provider network can respond to mobility events output by the mobility management components of the CSP network, such as access and mobility and mobility management functions (AMF) for 5G networks or mobility management entities (MMEs) for 4G or LTE networks. Such mobility events can be actual events (e.g., the mobile device has changed its connection point from a first access point to a second access point) or predictive events (e.g., the mobile device may connect to a second access point).

[0136] The operation includes: at block 1104, determining that at least a portion of the communication latency of the network path between the mobile device and the first computing instance via the second access point will not meet latency constraints, wherein the first computing instance is hosted by a first provider underlying extension of the cloud provider network. (See reference...) Figure 8 As described, not all mobility events will result in a breach of latency requirements. For example, a provider underlying extension hosting a compute instance may meet latency requirements for a set of access points on a CSP network, so a mobile device switching between those access points will not result in a breach of latency requirements. Edge location mobility service 830 may defer the launch of a new instance until a breach (or a predicted breach) of latency requirements occurs. For example, edge location mobility service 830 may evaluate latency data (e.g., latency data 609) between a first access point and a compute instance hosted by a first provider underlying extension, and between a second access point and a compute instance hosted by the first provider underlying extension.

[0137] The operation includes: at block 1106, identifying a second provider underlying extension of the cloud provider network that satisfies latency constraints for communication with the mobile device via the second access point. (See reference...) Figure 8 The described layout techniques (such as references) Figure 6 and Figure 7The techniques described can be used to address latency requirements met by a mobile device via a connectivity identifier of a second access point, or to request the initiation of a new instance, given latency requirements and an indication of a new (second) access point (e.g., based on a geographic identifier or an access point identifier identifying an access point within a CSP network). Hardware virtualization service 806 and edge location deployment service 820 can operate to identify candidate provider underlying extensions and select from those candidates to initiate computing instances.

[0138] The operation includes, at box 1108, sending a message to cause a second provider underlying extension to launch a second computing instance. Based on the selected provider underlying extension, the hardware virtualization service may (e.g., via a tunnel between the cloud provider network and the provider underlying extension deployed within the CSP network) issue one or more commands to the provider underlying extension to launch the requested instance.

[0139] Figure 12 An exemplary provider network (or “service provider system”) environment according to some implementations is illustrated. Provider network 1200 may provide resource virtualization to customers via one or more virtualization services 1210, which allow customers to purchase, lease, or otherwise obtain instances 1212 of virtualized resources (including, but not limited to, computing and storage resources) implemented on devices within one or more provider networks in one or more data centers. A local Internet Protocol (IP) address 1216 may be associated with resource instance 1212; the local IP address is the internal network address of resource instance 1212 on provider network 1200. In some implementations, provider network 1200 may also provide customers with public IP addresses 1214 and / or ranges of public IP addresses (e.g., Internet Protocol version 4 (IPv4) or Internet Protocol version 6 (IPv6) addresses) that can be obtained from provider 1200.

[0140] Typically, provider network 1200 can allow service provider customers (e.g., customers of one or more client networks 1250A to 1250C whose operations include one or more customer devices 1252) to dynamically associate at least some public IP addresses 1214 assigned or distributed to the customer with specific resource instances 1212 assigned to the customer via virtualization service 1210. Provider network 1200 can also allow customers to remap public IP addresses 1214 previously mapped to one virtualized computing resource instance 1212 assigned to the customer to another virtualized computing resource instance 1212 also assigned to the customer. For example, a service provider's (such as an operator of customer networks 1250A to 1250C) can use the virtualized computing resource instance 1212 and public IP addresses 1214 provided by the service provider to implement customer-specific applications and present those applications on an intermediate network 1240 such as the Internet. Then, other network entities 1220 on intermediate network 1240 can generate traffic to a target public IP address 1214 published by customer networks 1250A to 1250C; this traffic is routed to the service provider data center and, at the data center, routed via the network layer to the local IP address 1216 of the virtualized computing resource instance 1212, which is currently mapped to the target public IP address 1214. Similarly, response traffic from virtualized computing resource instance 1212 can be routed back to the intermediate network 1240 via the network layer to reach the source entity 1220.

[0141] As used herein, a local IP address refers to an internal or “private” network address of a resource instance, such as within a provider network. Local IP addresses may be within an address block reserved by the Internet Engineering Task Force (IETF) Comments Request for Comments (RFC) 1918 and / or have an address format specified by IETF RFC 4193, and may be variable within the provider network. Network traffic originating outside the provider network is not directly routed to a local IP address; instead, traffic uses a public IP address that maps to the local IP address of the resource instance. A provider network may include network devices or equipment that provide Network Address Translation (NAT) or similar functionality to perform mappings from public IP addresses to local IP addresses, and vice versa.

[0142] A public IP address is a variable network address on the Internet assigned to a resource instance by a service provider or by a customer. Traffic routed to a public IP address is translated via 1:1 NAT and forwarded to the corresponding local IP address of the resource instance.

[0143] Some public IP addresses may be assigned to specific resource instances by the provider's network infrastructure; these public IP addresses may be referred to as standard public IP addresses, or simply standard IP addresses. In some implementations, the mapping of standard IP addresses to the local IP addresses of resource instances is the default startup configuration for all resource instance types.

[0144] At least some public IP addresses can be assigned to or obtained by customers of Provider Network 1200; customers can then assign their assigned public IP addresses to specific resource instances assigned to them. These public IP addresses may be referred to as customer public IP addresses, or simply customer IP addresses. Instead of being assigned to resource instances by Provider Network 1200 as in the case of standard IP addresses, customer IP addresses can be assigned to resource instances by the customer, for example, via an API provided by the service provider. Unlike standard IP addresses, customer IP addresses are assigned to customer accounts and can be remapped to other resource instances by the corresponding customer as needed or desired. Customer IP addresses are associated with customer accounts, not with specific resource instances, and the customer controls the IP address until the customer chooses to release it. Unlike regular static IP addresses, customer IP addresses allow customers to mask resource instance or availability zone failures by remapping their public IP addresses to any resource instance associated with their customer account. For example, customer IP addresses enable customers to resolve customer resource instance or software issues by remapping their customer IP addresses to alternative resource instances.

[0145] Figure 13 This is a block diagram of an exemplary provider network that provides storage services and hardware virtualization services to customers according to some implementation schemes. Hardware virtualization service 1320 provides multiple computing resources 1324 (e.g., VMs) to customers. For example, computing resources 1324 may be leased or rented to customers of provider network 1300 (e.g., customers implementing customer network 1350). Each computing resource 1324 may be configured with one or more local IP addresses. Provider network 1300 may be configured to route packets from the local IP addresses of computing resources 1324 to public internet destinations, and from public internet sources to the local IP addresses of computing resources 1324.

[0146] Provider network 1300 can provide a client network 1350, for example, coupled to intermediate network 1340 via local network 1356, with the ability to implement virtual computing systems 1392 via hardware virtualization service 1320 coupled to intermediate network 1340 and provider network 1300. In some embodiments, hardware virtualization service 1320 may provide one or more APIs 1302 (e.g., web service interfaces) via which client network 1350 may access the functionality provided by hardware virtualization service 1320, for example, via console 1394 (e.g., web-based applications, standalone applications, mobile applications, etc.). In some embodiments, each virtual computing system 1392 at provider network 1300 and client network 1350 may correspond to computing resources 1324 that are leased, rented, or otherwise provided to client network 1350.

[0147] Clients can access the functionality of storage service 1310 from an instance of virtual computing system 1392 and / or another client device 1390 (e.g., via console 1394) via one or more APIs 1302, for example, to access and store data from storage resources 1318A to 1318N of virtual data storage areas 1316 (e.g., folders or "buckets", virtualized volumes, databases, etc.) provided by provider network 1300. In some embodiments, a virtualized data storage gateway (not shown) may be provided at client network 1350, which may locally cache at least some data (e.g., frequently accessed or critical data) and may communicate with storage service 1310 via one or more communication channels to upload new or modified data from the local cache, thereby maintaining the primary storage area (virtualized data storage area 1316) for data storage. In some implementations, users can install and access virtual data storage volumes 1316 via storage service 1310, which acts as a storage virtualization service, via virtual computing system 1392 and / or on another client device 1390, and these volumes may appear to the user as local (virtualized) storage devices 1398.

[0148] Although Figure 13 Although not shown, the virtualization service can also be accessed from resource instances within provider network 1300 via API 1302. For example, a customer, equipment service provider, or other entity can access the virtualization service from within a corresponding virtual network on provider network 1300 via API 1302 to request the allocation of one or more resource instances within the virtual network or another virtual network.

[0149] In some implementations, a system that implements some or all of the techniques described herein may include a general-purpose computer system (such as...) Figure 14 The computer system 1400 shown herein includes one or more computer-accessible media, or is configured to access said one or more computer-accessible media. In the illustrated embodiment, the computer system 1400 includes one or more processors 1410 coupled to system memory 1420 via an input / output (I / O) interface 1430. The computer system 1400 also includes a network interface 1440 coupled to the I / O interface 1430. Although Figure 14 Computer system 1400 is shown as a single computing device, but in various embodiments, computer system 1400 may include a single computing device or any number of computing devices configured to work together as a single computer system 1400.

[0150] In various embodiments, computer system 1400 may be a single-processor system including one processor 1410 or a multiprocessor system including several processors 1410 (e.g., two, four, eight, or another suitable number). Processor 1410 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1410 may be a general-purpose or embedded processor implementing any of a variety of instruction set architectures (ISAs), such as x86, ARM, PowerPC, SPARC, or MIPS ISA or any other suitable ISA. In a multiprocessor system, each processor 1410 may typically, but does not necessarily, implement the same ISA.

[0151] System memory 1420 may store instructions and data accessible by processor 1410. In various embodiments, system memory 1420 may be implemented using any suitable memory technology, such as random access memory (RAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash memory, or any other type of memory. In the illustrated embodiment, program instructions and data (such as the methods, techniques, and data described above) implementing one or more desired functions are shown stored within system memory 1420 as service code 1425 and data 1426. For example, service code 1425 may include code for implementing hardware virtualization services (e.g., 506, 606, 706, 806), edge location layout services (e.g., 620, 720, 820), edge location mobility services (e.g., 832), or other services or components shown in the figure. Data 1426 may include data such as latency data 609, application profiles, geographic data related to points within the CSP network, edge location data 509, etc.

[0152] In one embodiment, I / O interface 1430 may be configured to coordinate I / O traffic between processor 1410, system memory 1420, and any peripheral devices (including network interface 1440 or other peripheral interfaces) in a device. In some embodiments, I / O interface 1430 may perform any necessary protocols, timing, or other data transformations to convert data signals from one component (e.g., system memory 1420) into a format suitable for use by another component (e.g., processor 1410). In some embodiments, I / O interface 1430 may include devices that support attachment via various types of peripheral buses (e.g., variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard). In some embodiments, the functionality of I / O interface 1430 may be divided into two or more separate components, such as a northbridge and a southbridge. Furthermore, in some embodiments, some or all of the functionality of I / O interface 1430 (such as an interface to system memory 1420) may be directly integrated into processor 1410.

[0153] Network interface 1440 can be configured to allow data to pass between computer system 1400 and other devices 1460 attached to one or more networks 1450 (e.g., such as...). Figure 1 Exchange between other computer systems or devices shown. In various embodiments, network interface 1440 can support communication via any suitable wired or wireless general-purpose data network (e.g., various types of Ethernet). Additionally, network interface 1440 can support communication via telecommunications / telephone networks such as analog voice networks or digital fiber optic communication networks, via storage area networks (SANs) such as Fibre Channel SANs, or via I / O, any other suitable type of network, and / or protocol.

[0154] In some implementations, computer system 1400 includes one or more offload cards 1470 (including one or more processors 1475 and possibly one or more network interfaces 1440), which are connected using I / O interfaces 1430 (e.g., a version implementing the Peripheral Component Interconnect-Fast (PCI-E) standard or a bus for another interconnect such as Quick Path Interconnect (QPI) or Hyper Path Interconnect (UPI). For example, in some implementations, computer system 1400 may act as a host electronic device hosting computing instances (e.g., operating as part of a hardware virtualization service), and one or more offload cards 1470 may act as a virtualization manager capable of managing the computing instances running on the host electronic device. As an example, in some implementations, offload card 1470 may perform computing instance management operations such as pausing and / or unpausing computing instances, starting and / or terminating computing instances, performing memory transfer / copy operations, etc. In some implementations, these management operations may be performed by the offload card 1470 in cooperation with a hypervisor (e.g., upon request from the hypervisor) executed by other processors 1410A to 1410N of the computer system 1400. However, in some implementations, the virtualization manager implemented by the offload card 1470 may accommodate requests from other entities (e.g., from the computing instance itself) and may not cooperate with (or serve) any individual hypervisor.

[0155] In some embodiments, system memory 1420 may be an embodiment of a computer-accessible medium configured to store program instructions and data as described above. However, in other embodiments, program instructions and / or data may be received, transmitted, or stored on different types of computer-accessible media. Generally, computer-accessible media may include non-transitory storage media or memory media, such as magnetic or optical media, for example, a disk or DVD / CD coupled to computing device 1400 via I / O interface 1430. Non-transitory computer-accessible storage media may also include any volatile or non-volatile media, such as RAM (e.g., SDRAM, Double Data Rate (DDR) SDRAM, SRAM, etc.), read-only memory (ROM), etc., any of which may be included as system memory 1420 or another type of memory in some embodiments of computer system 1400. Furthermore, computer-accessible media may include transmission media or signals, such as electrical signals, electromagnetic signals, or digital signals, conveyed via communication media (such as networks and / or wireless links, such as those implemented via network interface 1440).

[0156] The various implementation schemes discussed or proposed herein can be implemented in a wide variety of operating environments. In some cases, the operating environment may include one or more user computers, computing devices, or processing devices that can be used to operate any of a number of applications. User or client devices may include any of a number of general-purpose personal computers, such as desktop or laptop computers running standard operating systems; and cellular, wireless, and handheld devices running mobile software and capable of supporting a number of networking and messaging protocols. Such systems may also include a number of workstations running a variety of commercially available operating systems and any of other known applications for purposes such as development and database management. These devices may also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and / or other devices capable of communicating via a network.

[0157] Most implementations utilize at least one network familiar to those skilled in the art to support communication using any of a variety of widely available protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), File Transfer Protocol (FTP), Universal Plug and Play (UPnP), Network File System (NFS), Public Internet File System (CIFS), Extensible Messaging and Presence Protocol (XMPP), AppleTalk, etc. The network may include, for example, a Local Area Network (LAN), a Wide Area Network (WAN), a Virtual Private Network (VPN), the Internet, an intranet, an extranet, a Public Switched Telephone Network (PSTN), an infrared network, a wireless network, and any combination thereof.

[0158] In implementations utilizing a web server, the web server can run any of a variety of server or middleware applications, including HTTP servers, File Transfer Protocol (FTP) servers, Common Gateway Interface (CGI) servers, data servers, Java servers, traffic application servers, etc. The server may also be able to respond to requests from user devices (such as those executed by the user device), and can be implemented in any programming language (such as...). The server may execute one or more scripts or programs (or one or more web applications) written in C, C#, or C++, or any scripting language (such as Perl, Python, PHP, or TCL), or combinations thereof. The server may also include a database server, including but not limited to commercially available database servers from Oracle(R), Microsoft(R), Sybase(R), IBM(R), etc. The database server may be relational or non-relational (e.g., "NoSQL"), distributed or non-distributed, etc.

[0159] The environment disclosed herein may include a variety of data storage areas and other memories and storage media as discussed above. These may reside in a variety of locations, such as on storage media local to (and / or residing in) one or more computers, or on any or all of the storage media located remotely from the computers across a network. In a particular set of embodiments, information may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing functions belonging to a computer, server, or other network device may be stored locally or remotely, as appropriate. Where the system includes computerized devices, each such device may include hardware elements that can be electrically coupled via a bus, including, for example, at least one central processing unit (“CPU”), at least one input device (e.g., mouse, keyboard, controller, touchscreen, or keypad) and / or at least one output device (e.g., display device, printer, or speaker). Such a system may also include one or more storage devices, such as hard disk drives, optical storage devices, and solid-state storage devices such as random access memory (RAM) or read-only memory (ROM), as well as removable media devices, memory cards, flash memory cards, etc.

[0160] Such devices may also include computer-readable storage medium readers, communication devices (e.g., modems, network interface cards (wireless or wired), infrared communication devices, etc.), and working memory, as described above. A computer-readable storage medium reader may be connected to or configured to receive computer-readable storage media, which represents remote, local, fixed, and / or removable storage devices, as well as storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information. Systems and various devices will also typically include multiple software applications, modules, services, or other elements residing within at least one working memory device, including operating systems and applications such as client applications or web browsers. It should be understood that alternative embodiments can have a large number of variations different from those described above. For example, custom hardware may also be used, and / or specific elements may be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connections to other computing devices, such as network input / output devices, may be employed.

[0161] Storage media and computer-readable media used to contain code or portions of code may include any suitable media known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information (such as computer-readable instructions, data structures, program modules or other data), including RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage devices, magnetic cartridges, magnetic tapes, disk storage devices or other magnetic storage devices, or any other media that can be used to store desired information and can be accessed by system devices. Based on the disclosure and teachings provided herein, those skilled in the art will understand other ways and / or methods for implementing various embodiments.

[0162] In the above description, various implementation schemes are described. Specific configurations and details are set forth for illustrative purposes to provide a thorough understanding of the implementation schemes. However, it will be apparent to those skilled in the art that the described implementation schemes can be practiced without specific details. Furthermore, well-known features may have been omitted or simplified to avoid obscuring the described implementation schemes.

[0163] In this document, bracketed text and boxes with dashed borders (e.g., large dashes, small dashes, dot dashes, and dots) are used to indicate optional operations that add additional features to some embodiments. However, this notation should not be construed as implying that these are the only options or optional operations, and / or that boxes with solid borders in some embodiments are not optional.

[0164] In various embodiments, reference numerals with suffix letters (e.g., 1318A to 1318N) can be used to indicate that the referenced entity may have one or more instances, and when multiple instances exist, each instance need not be identical, but may instead share some general characteristics or perform actions conventionally. Furthermore, unless explicitly stated to the contrary, the specific suffix used is not intended to imply the existence of a specific number of entities. Therefore, in various embodiments, two entities using the same or different suffix letters may or may not have the same number of instances.

[0165] References to "an implementation," "implementation," "exemplary implementation," etc., indicate that the implementation may include a specific feature, structure, or characteristic, but each implementation may not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Additionally, when a specific feature, structure, or characteristic is described in connection with an implementation, it should be understood that, whether explicitly described or not, implementing such a feature, structure, or characteristic in conjunction with other implementations is within the knowledge of those skilled in the art.

[0166] Furthermore, in the various embodiments described above, unless otherwise specifically indicated, disjunctive language such as “at least one of A, B, or C” is intended to mean A, B, or C or any combination thereof (e.g., A, B, and / or C). Thus, disjunctive language is generally not intended, nor should it be construed, as implying that a given embodiment requires at least one A, at least one B, or at least one C to each be present.

[0167] At least some embodiments of the disclosed technology can be described according to the following terms:

[0168] 1. A method comprising:

[0169] Receive a message from the mobility management component of a communication service provider network, the message including an indication of a mobility event associated with a mobile device of the communication service provider network, wherein the mobility event indicates a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point;

[0170] The communication latency of at least a portion of the network path between the mobile device and a first provider underlying extension of the cloud provider network via the second access point is determined to not meet the latency constraint, wherein the first provider underlying extension is deployed within the communication service provider network and hosts a first computing instance communicating with the mobile device.

[0171] A second provider underlying extension, identified as a cloud provider network deployed within the communication service provider network, that satisfies the latency constraints for communication with the mobile device via the second access point; and

[0172] Sending a message causes the underlying extension of the second provider to launch a second computing instance.

[0173] 2. The method according to Clause 1, wherein the first provider underlying extension and the second provider underlying extension are at least partially controlled by the control plane service of the cloud provider network.

[0174] 3. The method according to any one of Clauses 1 to 2, wherein the mobility event indicates the predicted probability of a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point.

[0175] 4. A computer-implemented method, comprising:

[0176] Receive a message including an indication of a mobility event associated with a mobile device on a communication service provider network, wherein the mobility event indicates a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point;

[0177] It is determined that the communication delay of at least a portion of the network path between the mobile device and the first computing instance via the second access point will not meet the latency constraint, wherein the first computing instance is hosted by a first provider underlying extension of the cloud provider network;

[0178] The second provider underlying extension that identifies the cloud provider network and satisfies the latency constraint for communication with the mobile device via the second access point; and

[0179] Sending a message causes the underlying extension of the second provider to launch a second computing instance.

[0180] 5. The computer implementation method according to Clause 4, wherein the first computing instance and the second computing instance are started from the same image.

[0181] 6. The computer implementation method according to any one of Clauses 4 to 5, wherein the first provider underlying extension and the second provider underlying extension are deployed within the communication service provider network and are at least partially controlled by the control plane service of the cloud provider network.

[0182] 7. The computer implementation method according to any one of Clauses 4 to 6, wherein the mobility event indicates a predicted probability of a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point.

[0183] 8. The computer implementation method according to Clause 7, wherein sending the message causes the second provider's underlying extension to launch a second computing instance based at least in part on the predicted probability being higher than a threshold.

[0184] 9. The computer implementation method according to any one of Clauses 4 to 8, further comprising sending a message to the first computing instance to initiate the transfer of state data from the first computing instance to the second computing instance.

[0185] 10. The computer implementation method according to Clause 9, wherein the status data is transmitted through at least a portion of the cloud provider's network.

[0186] 11. The computer implementation method according to any one of Clauses 4 to 10, further comprising sending another message to cause the first provider underlying extension to terminate the first computing instance.

[0187] 12. The computer implementation method according to any one of Clauses 4 to 11, wherein the delay constraint is specified by a customer of the cloud provider network.

[0188] 13. The computer implementation method according to any one of clauses 4 to 12, wherein the communication delay between the mobile device and the first computing instance fails to meet the delay constraint because the first computing instance is unreachable from the second access point.

[0189] 14. A system comprising:

[0190] A cloud provider network, comprising multiple provider underlying extensions deployed within a communication service provider network, wherein each of the multiple provider underlying extensions:

[0191] Connected to the cloud provider network via the communication service provider network;

[0192] This includes the capacity for hosting customer compute instances, and

[0193] Able to communicate with mobile devices of subscribers of the communication service provider network via the communication service provider network; and

[0194] One or more first electronic devices of the cloud provider network implementing one or more control plane services, the one or more control plane services including, when executed, instructions causing the one or more control plane services to perform the following operations:

[0195] Receive a message including an indication of a mobility event associated with a mobile device on a communication service provider network, wherein the mobility event indicates a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point;

[0196] It is determined that the communication delay of at least a portion of the network path between the mobile device and the first computing instance via the second access point will not meet the latency constraint, wherein the first computing instance is hosted by a first provider underlying extension among the plurality of provider underlying extensions;

[0197] Identify the second provider underlying extension among the plurality of provider underlying extensions that satisfies the latency constraint for communicating with the mobile device via the second access point; and

[0198] Sending a message causes the underlying extension of the second provider to launch a second computing instance.

[0199] 15. The system according to Clause 14, wherein the first computing instance and the second computing instance are launched from the same image.

[0200] 16. The system according to any one of Clauses 14 to 15, wherein the mobility event indicates a predicted probability of a change in the connection point between the mobile device and the communication service provider network from a first access point to a second access point.

[0201] 17. The system according to Clause 16, wherein sending the message causes the second provider's underlying extension to launch a second computing instance based at least in part on the predicted probability being higher than a threshold.

[0202] 18. The system according to any one of Clauses 14 to 17, wherein the one or more control plane services include, at execution, further instructions causing the one or more control plane services to send a message to the first computing instance to initiate the transfer of state data from the first computing instance to the second computing instance.

[0203] 19. The system according to Clause 18, wherein the status data is transmitted via at least a portion of the cloud provider's network.

[0204] 20. The system pursuant to any one of Clauses 14 to 19, wherein the latency constraint is specified by a customer of the cloud provider network.

[0205] Therefore, the specification and drawings should be considered illustrative rather than restrictive. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of this disclosure as set forth in the claims.

Claims

1. A computer-implemented method, comprising: A control plane component is executed in a first network to at least partially manage edge locations physically separated from the first network, wherein the edge locations are adapted to execute containers for one or more customers, and wherein a first edge location of the edge locations is embedded within a cellular communication service provider network; The control plane component receives a configuration file from the client, the configuration file including one or more constraints to be used when identifying the edge location where containers for the client's application are to be deployed. The control plane component determines that the customer container will be activated at the first edge location; The control plane component transmits a message to the first edge location to initiate the client container; The client container is launched on a computer system at the first edge location, wherein the computer system obtains a container image specified by the client and stored outside the first edge location, and launches the client container based on the container image, wherein the client container communicates with the mobile device of the subscriber of the cellular communication service provider via the cellular communication service provider network. The control plane component determines to initiate the second client container at the second edge location at least in part because the second edge location among the edge locations satisfies one or more constraints of the configuration file; as well as The control plane component transmits a message to the second edge location to start the second client container, thereby the second computing system at the second edge location also obtains the container image specified by the client and starts the second client container based on the container image.

2. The computer implementation method of claim 1, wherein determining to launch the client container at the first edge location occurs at least in part based on a latency requirement.

3. The computer implementation method of claim 1, wherein the client container communicates with the mobile device without crossing a network carrying traffic of users who are not subscribed to the cellular communication service provider's network.

4. The computer implementation method of claim 1, wherein the first edge location communicates with the first network via a tunnel through the cellular communication service provider network.

5. The computer implementation method of claim 1, wherein the first edge location is connected to at least one of: a network component of the radio access network of the cellular service provider network, a network component of the aggregation site of the cellular service provider network, or a network component of the core network of the cellular service provider network.

6. The computer-implemented method of claim 1, further comprising: After the client container is started, the one or more constraints of the configuration file are re-evaluated, thereby causing it to be determined that the first edge position no longer satisfies at least one of the one or more constraints, while the third edge position of the edge positions satisfies the one or more constraints. as well as The control plane component transmits a message to the third edge location, at which another client container for the application is launched.

7. The computer implementation method as described in claim 1, further comprising: Associating an Internet Protocol (IP) address with the client container, wherein the IP address comes from a pool of IP addresses in the address space of the cellular service provider network assigned to the first edge location.

8. The computer implementation method of claim 1, wherein the cellular communication service provider network includes a 5G cellular network.

9. A system comprising: One or more first electronic devices in a first network, the one or more first electronic devices performing a control plane assembly to at least partially manage edge locations physically separated from the first network, said edge locations being adapted to perform containers for one or more clients, and said said control plane assembly including instructions to cause the control plane assembly to perform the following operations when performed by said one or more first electronic devices: Receive a configuration file created by the customer, which includes one or more constraints used to identify where the application should be deployed; The client container is determined to be launched at the first edge location, which is embedded within a cellular communication service provider network, at least in part because the first edge location satisfies one or more constraints of the configuration file. Transmit a message to the first edge location to start the client container; The second client container is determined to be launched at the second edge location at least in part because the second edge location satisfies one or more constraints of the configuration file; as well as A message to start the second client container is transmitted to the second edge location, thereby the second computing system at the second edge location also obtains the container image specified by the client and starts the second client container based on the container image; as well as One or more second electronic devices at the first edge location, wherein the one or more second electronic devices are embedded within a cellular communication service provider network and include instructions that, when performed by the one or more second electronic devices, cause the one or more second electronic devices to perform the following operations: The customer container is activated, wherein one or more second electronic devices obtain a container image specified by the customer and stored outside the first edge location, and activate the customer container based on the container image, wherein the customer container communicates with the mobile device of the subscriber of the cellular communication service provider via the cellular communication service provider network.

10. The system of claim 9, wherein determining that the client container is launched at the first edge location occurs at least in part based on a delay requirement.

11. The system of claim 9, wherein the client container communicates with the mobile device without crossing a network carrying traffic of users who are not subscribed to the cellular communication service provider's network.

12. The system of claim 9, wherein the first edge location communicates with the first network via a tunnel through the cellular communication service provider network.

13. The system of claim 9, wherein the first edge location is connected to at least one of: a network component of the radio access network of the cellular service provider network, a network component of the aggregation site of the cellular service provider network, or a network component of the core network of the cellular service provider network.

14. The system of claim 9, wherein the control plane component further includes instructions that, when executed by the one or more first electronic devices, cause the control plane component to perform the following operations: After the client container is started, the one or more constraints of the configuration file are re-evaluated, resulting in the determination that the first edge position no longer satisfies at least one of the one or more constraints, while a third edge position among the edge positions satisfies the one or more constraints; and A message is transmitted to the third edge location to initiate another client container for the application at the third edge location.

15. The system of claim 9, wherein the one or more first electronic devices further includes instructions, when executed, to cause the one or more first electronic devices to associate an Internet Protocol IP address with the client container, wherein the IP address is derived from a pool of IP addresses in the address space of the cellular service provider network allocated to the first edge location.

16. The system of claim 9, wherein the cellular communication service provider network includes a 5G cellular network.

17. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of one or more computing devices, cause the one or more computing devices to perform operations including the following steps: A control plane component is executed in a first network to at least partially manage edge locations physically separated from the first network, wherein the edge locations are adapted to execute containers for one or more customers, and wherein a first edge location of the edge locations is embedded within a cellular communication service provider network; The control plane component receives a configuration file from the client, the configuration file including one or more constraints to be used when identifying the edge location where containers for the client's application are to be deployed. The control plane component determines to launch the client container at the first edge location at least in part because the first edge location among the edge locations satisfies one or more constraints of the configuration file; The control plane component transmits a message to the first edge location to initiate the client container; The client container is launched on a computer system at the first edge location, wherein the computer system obtains a container image specified by the client and stored outside the first edge location, and launches the client container based on the container image, wherein the client container communicates with the mobile device of the subscriber of the cellular communication service provider via the cellular communication service provider network. The control plane component determines to initiate the second client container at the second edge location at least in part because the second edge location among the edge locations satisfies one or more constraints of the configuration file; as well as The control plane component transmits a message to the second edge location to start the second client container, thereby the second computing system at the second edge location also obtains the container image specified by the client and starts the second client container based on the container image.

18. The non-transitory computer-readable storage medium of claim 17, wherein the cellular communication service provider network includes a 5G cellular network.

19. The non-transitory computer-readable storage medium of claim 17, wherein determining to launch the client container is based at least in part on a latency requirement.

20. The non-transitory computer-readable storage medium of claim 17, wherein the first edge location is connected to at least one of: a network component of the radio access network of the cellular service provider network, a network component of the aggregation site of the cellular service provider network, or a network component of the core network of the cellular service provider network.

Citation Information

Patent Citations

  • Traffic path change detection mechanism of mobile edge calculation

    CN108574728A

  • Edge server and method for the same

    JP2017017656A