Intelligent multi-operator network edge application deployment

SOADM service collects network quality metadata through an intelligent routing module, optimizes application deployment and traffic routing, solves the problem of mobile networks failing to provide consistent network quality, achieves stable and high-quality mobile edge computing connections, and improves user experience.

CN120937327APending Publication Date: 2025-11-11AMAZON TECH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202480021226.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-24
Filing Date
2024-02-23
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Mobile networks fail to provide application developers with consistent network quality guarantees, resulting in slow adoption and use of mobile edge computing and edge infrastructure, and limited latency and network performance for end users.

Method used

SOADM service collects network quality metadata through the intelligent routing module, provides intelligent routing capabilities, and dynamically adjusts application deployment and traffic routing to optimize network resource usage across different communication service providers and edge locations, simplifying the interaction between application developers and multiple CSPs.

Benefits of technology

It achieves stable, high-quality connections in different geographical locations and CSP networks, avoiding performance degradation due to network congestion and improving the service quality and user experience of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937327A_ABST
    Figure CN120937327A_ABST
Patent Text Reader

Abstract

Techniques for intelligent multi-operator network edge application deployment are described. Traffic addressed to applications implemented in a plurality of edge locations of a cloud provider network is initiated by a mobile user equipment device over a communication network using a first communication service provider (CSP). An edge location hosting the application may be selected from a plurality of such candidates as a destination for the traffic. The edge locations may be deployed in facilities of different CSPs. The traffic may be sent to the edge location using network addresses of the different CSPs to securely allow the traffic to enter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 174,083, filed February 24, 2023, entitled “Intelligent Multi-Carrier Network Edge Application Deployment,” the entire contents of which are incorporated herein by reference. Background Technology

[0003] Cloud computing platforms typically provide users with on-demand, managed computing resources. These resources (e.g., compute and storage capacity) are usually provided by large pools of capacity located in the data centers of a cloud provider. Users can request computing resources from the “cloud,” and the cloud can provide those resources to those users. Technologies such as virtual machines and containers are often used to allow users to securely share the capacity of a computer system. Attached Figure Description

[0004] Various examples according to this disclosure will now be described with reference to the accompanying drawings, in which:

[0005] Figure 1 The following is illustrated with some examples of a service-oriented application deployment management (“SOADM”) service that provides functionality for intelligent multi-carrier network edge application deployment.

[0006] Figure 2 An exemplary intra-regional cloud provider network deployment of an intelligent routing module for intelligent multi-carrier network edge application deployment is shown, based on some examples.

[0007] Figure 3 An exemplary edge network deployment of an intelligent routing module for intelligent multi-carrier network edge applications is shown, based on some examples.

[0008] Figure 4 An exemplary graphical user interface for operator network contract selection and customization for intelligent multi-operator network edge application deployment, provided by SOADM services, is shown, based on some examples.

[0009] Figure 5 The illustration shows an exemplary mobility-based application traffic rerouting performed by SOADM service for intelligent multi-carrier network edge application deployments, based on some examples.

[0010] Figure 6 This is a flowchart illustrating the operation of a method for deploying edge applications for intelligent multi-carrier networks, based on some examples.

[0011] Figure 7An SOADM service according to some implementation schemes is shown, which provides user-configurable multi-service application deployment and distribution of various types and locations across deployment zones.

[0012] Figure 8 An exemplary system according to some implementations is shown, the exemplary system including a cloud provider network and also including various edge locations of the cloud provider network.

[0013] Figure 9 An exemplary cloud provider network, including geographically dispersed edge locations, is shown according to some implementation schemes.

[0014] Figure 10 An exemplary system is shown, according to some implementation schemes, in which a cloud provider network underlying extension is deployed within a communications service provider network.

[0015] Figure 11 Exemplary components at edge locations within cloud provider networks and communications service provider networks, and the connectivity between them, are shown in more detail according to some implementation schemes.

[0016] Figure 12 This illustrates a dynamic service and service resource redistribution based on user configurations by SOADM services according to changes in end-user location and latency, according to some implementation schemes.

[0017] Figure 13 An exemplary multi-service application with a heterogeneous distribution strategy deployed and distributed by SOADM services is shown according to some implementation schemes.

[0018] Figure 14 The illustration shows a DNS-based service discovery and latency-based routing system provided by SOADM services according to some implementation schemes.

[0019] Figure 15 This paper illustrates a service discovery and latency-based routing system provided by SOADM services according to some implementation schemes.

[0020] Figure 16 The example provider network environment is shown based on some examples.

[0021] Figure 17 This is a block diagram of an example provider network offering storage services and hardware virtualization services to customers, based on some examples.

[0022] Figure 18 This is a block diagram illustrating an example computer system that can be used in some examples. Detailed Implementation

[0023] This disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media for deploying applications at the edge of intelligent multi-carrier networks. In some examples, a location-aware, service-oriented application deployment management (“SOADM”) service provides intelligent routing capabilities for end-user requests initiated using a typical third-party mobile network, thereby intelligently routing application user traffic to application resources deployed at edge locations within the same mobile network, at edge locations of different communication service providers (CSPs), or within different edge locations or cloud provider network areas. The SOADM service can collect network quality metadata from CSPs and edge locations to make improved routing decisions and / or allow customers to determine where they want to deploy applications, obtain data contracts from the required CSPs, and configure preferred radio resources with these CSPs in a simple manner (e.g., “network slicing”).

[0024] Currently, mobile networks fail to provide application developers with consistent network quality guarantees, forcing them to deal with large global communications (or "operator") service providers (CSPs) offering mobile data / voice services worldwide. Consequently, the adoption and use of mobile edge computing and edge infrastructure has been slow. In the examples disclosed herein, SOADM services can provide a form of intelligent network multiplexing that simplifies the process for application developers to receive improved network quality through their applications' Quality of Service (QoS) attributes, abstracting away the complexities of dealing with various CSPs.

[0025] In some examples, the SOADM service provides / utilizes interfaces to provide intelligent routing capabilities for end-user requests originating from mobile networks. For instance, the SOADM service can provide interfaces to request (or receive) wireless network QoS attributes (e.g., latency, bandwidth, throughput, etc.) from a CSP, which can be customized for local conditions specific to a particular geographic location (e.g., the edge location of a cloud provider network deployed within a CSP facility or its vicinity). Through this functionality, the SOADM service can obtain the QoS attributes of end-user connections regardless of which CSP is serving the end-user. Therefore, applications deployed to mobile network edge computing can now receive stable, high-quality connections and avoid the impact of certain performance degradation scenarios, such as mobile network congestion.

[0026] In some examples, a CSP may periodically and frequently publish latency data to an SOADM service on devices on its network. The SOADM service may also obtain data related to routing and end-user latency from other sources (e.g., from a content distribution network) to further optimize its routing decisions. In some examples, the SOADM service may include a bootstrapping module that can use this (and other) information to generate and return a ranked list of locations (hosted—or capable of hosting—application resources) best suited to serve end-user requests, regardless of who owns those locations.

[0027] Therefore, in the examples disclosed herein, the SOADM service can intelligently serve mobile terminal user requests across multiple potential CSPs and multiple potential edge computing locations by making intelligent routing decisions based on latency data or other relevant network or load factors.

[0028] For example, Figure 1 The document illustrates SOADM services that provide functionality for deploying intelligent multi-carrier network edge applications, based on several examples. Figure 1 In this context, SOADM service 102 is implemented within one or more regions 112A-112M of the cloud provider network 100, and in some examples, it is implemented as software executed by one or more computing devices at one or more geographical locations.

[0029] Provider Network 100 (or “cloud” provider network) provides users with the ability to use one or more of various types of computing-related resources, such as computing resources (e.g., executing virtual machine (VM) instances and / or containers, executing batch jobs, executing code without provisioning servers), data / storage resources (e.g., object storage, block-level storage, data archive storage, databases and database tables, etc.), network-related resources (e.g., configuring virtual networks (including compute resource groups), content distribution networks (CDNs), domain name services (DNS)), application resources (e.g., databases, application build / deployment services), access policies or roles, identity policies or roles, machine images, routers, and other data processing resources. These and other computing resources may be provided as services, such as hardware virtualization services that can execute computing instances, storage services that can store data objects, etc. Users of Provider Network 100 (or “customers”) may use one or more user accounts associated with a customer account, but these items may be used interchangeably to some extent depending on the use case. Users can interact with provider network 100 across one or more intermediate networks (e.g., the Internet) via one or more interfaces, such as through application programming interface (API) calls, via a console implemented as a website or application, etc. An API is an interface and / or communication protocol between a client and a server, such that if a client issues a request in a predefined format, the client should receive a response in a specific format or initiate a defined action. In a cloud provider network context, an API provides a gateway to enable customers to access cloud infrastructure by allowing customers to obtain data from or initiate actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted within the cloud provider network. APIs can also enable different services within the cloud provider network to exchange data with each other. Interfaces may be part of or act as the front end of the control plane of provider network 100, which includes "back-end" services that support and implement services that can be provided more directly to customers.

[0030] Therefore, a cloud provider network (or simply "the cloud") can refer to a large pool of accessible virtualized computing resources, such as computing, storage, and networking resources, applications, and services. The cloud provides convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and deployed in response to client commands. These resources can be dynamically provisioned and reconfigured to adapt to variable loads. Thus, cloud computing can be viewed as both applications delivered as a service over publicly accessible networks (e.g., the internet, cellular networks) and the hardware and software in the cloud provider's data centers that provide these services.

[0031] Cloud provider networks can be structured into multiple regions 112A-112M, where a region is a geographical area in which the cloud provider clusters its data centers. Each region comprises multiple (e.g., two or more) Availability Zones (AZs) interconnected via a private high-speed network (e.g., fiber optic communication connection). An AZ (also called a “zone”) provides an isolated fault domain comprising one or more data center facilities, which have separate power supplies, separate networking, and separate cooling relative to data center facilities in another AZ. A data center refers to the physical building or enclosure that houses the servers of the cloud provider network and provides power and cooling to the servers of the cloud provider network. Preferably, AZs within a region are positioned far enough apart that a natural disaster (or other event causing failure) will not simultaneously affect or take down more than one AZ.

[0032] Users can connect to the cloud provider's Availability Zone (AZ) via a publicly accessible network (e.g., the Internet, cellular networks), for example, through a Switching Center (TC). A TC is the primary backbone location linking users to the cloud provider's network and can be co-located with other network provider facilities (e.g., Internet Service Providers, telecommunications providers) and securely connected to the AZ (e.g., via encrypted or direct connections). Two or more TCs can operate in each region for redundancy. Regions connect to a global network comprising private networking infrastructure (e.g., fiber optic connections controlled by the cloud provider) connecting each region to at least one other region. The cloud provider's network delivers content from access points (or "POPs") located outside but networked with these regions, via edge locations and regional edge caching servers. This partitioning and geographical distribution of computing hardware enables the cloud provider's network to provide users with low-latency resource access globally with high fault tolerance and stability.

[0033] Generally, provider network services and operations can be broadly categorized into two types: control plane operations carried on the logical control plane and data plane operations carried on the logical data plane. The data plane represents the movement of user data through a distributed computing system, while the control plane represents the movement of control signals through the distributed computing system. The control plane typically includes one or more control plane components distributed across one or more control servers and implemented by one or more control servers. Control plane services typically include administrative operations such as system configuration and management (e.g., resource placement, hardware capacity management, diagnostic monitoring, system status information). The data plane includes user resources implemented on the provider network (e.g., compute instances, containers, block storage volumes, databases, file storage). Data plane services typically include non-administrative operations such as transferring user data to and from user resources. Control plane components are typically implemented on a separate set of servers from the data plane servers, and control plane and data plane services can be transmitted over separate / different networks.

[0034] To provide these and other computing resource services, provider network 100 typically relies on virtualization technology. For example, virtualization technology can provide users with the ability to control or use computing resources (e.g., “computing instances,” such as VMs using guest operating systems (O / S) that may or may not operate on top of the underlying host O / S; containers that may or may not operate within VMs; computing instances that can execute on “bare metal” hardware without an underlying hypervisor), where one or more computing resources can be implemented using a single electronic device. Thus, users can directly use computing resources hosted by the provider network (e.g., provided by a hardware virtualization service) to perform various computing tasks. Alternatively or additionally, users can indirectly use computing resources by submitting code to be executed by the provider network (e.g., via an on-demand code execution service), which in turn uses one or more computing resources to execute the code, typically without the user having any control or knowledge of the underlying computing instances involved.

[0035] As described herein, one type of service that a provider network can offer may be referred to as a "managed computing service," which executes code or provides computing resources for its users within a managed configuration. Examples of managed computing services include, for instance, on-demand code execution services, hardware virtualization services, container services, and so on.

[0036] On-Demand Code Execution Service (referred to in various examples as Function Compute Service, Function Service, Cloud Function Service, Function as a Service, or Serverless Compute Service) enables users of Provider Network 100 to execute their code on cloud resources without having to select or manage the underlying hardware resources used to execute the code. For example, a user can use On-Demand Code Execution Service by uploading their code and using one or more APIs to request the service to identify, provision, and manage any resources required to run the code. Therefore, in various examples, "serverless" functionality may include on-demand executable code provided by a user or other entity, such as the Provider Network itself. Serverless functionality is maintained within the Provider Network through On-Demand Code Execution Service and may be associated with a specific user or account or generally accessible to multiple users / accounts. Serverless functionality may be associated with a Uniform Resource Locator (URL), Uniform Resource Identifier (URI), or other reference that can be used to invoke the serverless functionality. Serverless functionality can be executed via computing resources such as virtual machines or containers when triggered or invoked. In some examples, serverless functionality may be invoked via Application Programming Interface (API) calls or specially formatted Hypertext Transfer Protocol (HTTP) request messages. Therefore, users can define the serverless functions that can be executed on demand without requiring them to maintain dedicated infrastructure to execute the serverless functions. Alternatively, resources maintained by the provider network 100 can be used to execute serverless functions on demand. In some examples, these resources may be maintained in a "ready" state (e.g., with a pre-initialized runtime environment configured to execute serverless functions), thereby allowing serverless functions to be executed almost in real time.

[0037] Hardware virtualization services (referred to in various implementations as elastic computing services, virtual machine services, compute cloud services, compute engines, or cloud computing services) enable users of Provider Network 100 to provision and manage computing resources, such as virtual machine instances. Virtual machine technology can use a physical server, for example, an equivalent form of running many servers (each referred to as a virtual machine) using a hypervisor, which may run at least on an off-grid card of the server (e.g., a card connected to the physical CPU via PCI or PCIe), and other components of the virtualization host may be available for some virtualization management components. Such an off-grid card of the host may include one or more CPUs that are not available for user instances but are dedicated to instance management tasks, such as virtual machine management (e.g., hypervisor), I / O virtualization of network-attached storage volumes, local migration management tasks, instance health monitoring, etc. Virtual machines are often referred to as compute instances or simply "instances." As used herein, provisioning a virtual computing instance typically includes: reserving resources (e.g., compute and storage resources) of the underlying physical computing instance for the client (e.g., from a pool of available physical computing instances and other resources); installing or launching the required software (e.g., an operating system); and making the virtual computing instance available to the client to perform client-specified tasks.

[0038] Another type of managed computing service can be a container service that allows users of a cloud provider's network to instantiate and manage containers, such as container orchestration and management services (referred to in various implementations as container services, cloud container services, container engines, or container cloud services). In some examples, container service 114 can be a Kubernetes-based container orchestration and management service (referred to in various implementations as Kubernetes container services, Azure Kubernetes service, IBM Cloud Kubernetes service, Kubernetes engine, or Kubernetes container engine). As mentioned in this article, containers package code and all its dependencies, so applications (also referred to in various container services as tasks, swarms, or clusters) can run quickly and reliably across various computing environments. A container image is a standalone executable package that includes everything needed to run an application process: code, runtime, system tools, system libraries, and settings. A container image becomes a container at runtime. Therefore, a container is an abstraction of the application layer (meaning each container emulates a different software application process). Although each container runs an isolated process, multiple containers can share a common operating system, for example, by starting within the same virtual machine. In contrast, a virtual machine is an abstraction of the hardware layer (meaning each virtual machine emulates a physical machine that can run software). While multiple virtual machines can run on a single physical machine, each virtual machine typically has its own copy of the operating system, along with the application and its associated files, libraries, and dependencies. Some containers can run on instances running container agents, and some can run on bare metal servers or on off-board cards of the server.

[0039] In some examples, SOADM service 102 abstracts the complexity of deploying distributed applications within / through a cloud provider network 100 that provides one or more possible deployment zones of one or more types. These deployment zones typically correspond to physical locations where the cloud provider network provides data centers or other computing power and can include various types of deployment zones, such as traditional cloud provider regions 112 and availability zones, as well as so-called “edge locations” 116 / 117 / 118 (e.g., cloud provider-operated edge locations, customer-operated edge locations, third-party-operated edge locations, and edge locations associated with CSPs). According to the examples described herein, SOADM service enables users to create service group configurations representing service-oriented applications (including their constituent services and dependent resources), and to specify distribution policies for deploying and / or redistributing application services and resources, among other configurations. Using such configurations, SOADM service 102 can automatically deploy and scale simple or complex, single or multi-service applications for users across any number of deployment zones and deployment zone types, based on user configuration.

[0040] Providing personalized and immersive experiences to end users presents a challenge for architectures using centralized processing in a single location, as latency to end users can impact the desired user experience. For example, an application deployed in one geographic location can provide responsive service to users nearby; however, end users in remote geographic locations may experience poor service due to significant communication latency caused by this distance. Furthermore, applications deployed at the edge, close to users, may be overwhelmed or suffer from outages or individual issues (e.g., associated with CSPs from which many end users derive network connectivity).

[0041] Furthermore, these issues can vary over time as geographical access patterns change throughout the day. For example, an application might be heavily used during the daytime in North America, while experiencing almost no concurrent usage in Asia during the same period. However, this usage could shift at different times when it is daytime in Asia and nighttime in North America. Therefore, application developers need ubiquitous computing proximity to their end users, even if the end users' locations may change over time. However, even in modern cloud networks, this distributed environment significantly increases the complexity of application development and operation.

[0042] As indicated herein, in some examples, SOADM service 102 allows users to configure applications, which typically include one or more services and dependent resources, where deployment and / or distribution policies instruct application developers how and where they wish to deploy the application. SOADM service 102 can also obtain various types of network metadata from various CSPs—such as associated jitter, bandwidth, latency, etc., related to a user's mobile service in a specific area—and / or data related to more "back-end" aspects (e.g., resource availability or performance within edge locations). Using this data, SOADM can manage the deployment, redistribution, and use of various components of an application on behalf of the user without requiring any (or a large number of) user intervention.

[0043] At a high level, SOADM service 102 may include a deployment engine 132 and a deployment monitor 134. The deployment engine 132 manages the deployment of application components (e.g., provisioning and deprovisioning application instances in different locations, possibly based on capacity information 156 provided by capacity service 136, indicating the locations where various available resources exist), while the deployment monitor 134 observes the operation of deployed applications and determines whether and when to modify the application deployment (e.g., based on user-configured configuration data, metrics obtained from monitoring service 138, or directly from edge locations). Further details will follow. Figure 7 We will introduce additional features in these areas.

[0044] However, SOADM service 102 may also include an intelligent routing module 111, which is deployed wholly or partially in region 112A (as intelligent routing module 111B) or one or more edge locations (e.g., edge location 116A). As described herein, intelligent routing module 111 can be used, for example, to intelligently route user traffic of an application to the location deemed most suitable based on the application's QoS requirements.

[0045] As shown in circle (1), in some examples, various CSPs may provide multiple sets of wireless network-related metrics 152 to cloud provider network 100. This transmission may occur periodically (e.g., according to a schedule such as every few minutes) and / or on demand, thereby allowing SOADM service 102 to request updates. The multiple sets of wireless network-related metrics 152 may include information relating to the wireless capabilities and / or current performance of a specific portion of its network, such as performance specific to the area surrounding the CSP's embedded edge location, for example, a set of metrics for edge location 116A, a corresponding set of metrics for edge location 116B (in a separate geographic location 114N), etc. It is noteworthy that this reporting may be made for a specific CSP for different geographic locations associated with the edge location associated with the CSP; similarly, other CSPs may also report these metrics, etc., for geographic locations associated with their CSP's embedded edge location. However, in some examples, CSPs may report metrics for other locations that are not directly associated with any edge location but may only correspond to specific administrative boundaries (e.g., at the granularity of community, city, zip code, state, or country).

[0046] The types of network metrics collected may vary in different examples based on the implementer's needs, but may include quality of service (QoS) type metrics, such as 5G bandwidth radio type attributes like bandwidth, jitter, latency, throughput, round-trip time, number of simultaneous connections, connection persistence, network availability, error rate (e.g., block error rate), etc. This data may be stored in a data storage system (e.g., a database, data lake, etc.), and in some examples, one or more specific types of metrics may be aggregated, summarized, and / or transformed, for example, to provide average values ​​within a time window, cumulative values ​​within a time window, maximum or minimum values ​​within a time window, etc. This storage and / or analysis may be performed by monitoring service 138 (which can directly receive metric 152), SOADM service 102 (which can directly receive or indirectly receive metric 152 through monitoring service 138), or another service of provider network 100.

[0047] In some examples, this data set or summary can be provided to users of the cloud provider's network through a console-type application supported by the console backend 160. This console allows users to explore real-time (or near-real-time and / or historical) information describing the various CSP networks that work with the cloud provider. For example, an application developer user role might want to deploy an application centered around a specific geographic location (such as a stadium, concert venue, office building, conference center, etc.). The application developer might then want to see which CSPs are providing services in that geographic area and determine which CSPs can provide the wireless services that meet the application's needs.

[0048] For example, application developers might have applications that support live sporting events and are highly latency-sensitive, such as augmented reality applications, which end users in the stadium might want to use frequently, while users elsewhere might not use them at all. Therefore, developers (e.g., user 104) could use their electronic devices 106, via a console partially provided by the console backend 160, to explore CSP network metric values ​​162 (and / or their transformations or summaries) to determine which CSPs they might want to obtain (e.g., purchase) service contracts from. For example, developers might want to identify CSPs that provide connectivity with at least a threshold amount of bandwidth (e.g., 500 megabits per second), latency less than a certain threshold amount (e.g., less than 10 milliseconds), availability greater than a certain threshold (e.g., 99.99%), jitter less than a certain threshold (e.g., less than 50 microseconds). Developers could browse summaries of these CSP metric values ​​162 in or near the location of interest (the sporting event) and use the console to obtain service contracts for that location. In some examples, developers can provide or configure their own set of QoS requirements for applications, which may or may not be independent of the contracts obtained, and these requirements can be used to filter out contracts that do not meet the requirements, and / or to determine the routing locations for end-user traffic (e.g., by selecting locations that should be able to meet the QoS requirements of these applications based on recently reported network metrics).

[0049] For example, developer user 104 might attempt to obtain a service contract from CSP 'A' for 100 gigabytes of traffic near Los Angeles and 100 gigabytes of traffic near San Francisco, and similarly, might also want to obtain a service contract from CSP 'B' for the same 100 gigabytes of traffic near Los Angeles and 100 gigabytes of traffic near San Francisco. Furthermore, the developer might attempt to obtain a service contract from CSP 'C' for 100 gigabytes of traffic near San Francisco. In this example, the contract could allow for traffic prioritization to and from a specific application—also known as application-aware network slicing, prioritized application workflows, etc. These contracts could also provide certain network quality guarantees, such as services with specific guaranteed bandwidth, latency, jitter, round-trip time, etc.

[0050] For the requested service contract (represented as a stored data structure – service contract 164), the cloud provider network 100 can send service plan configuration data 154 to each involved CSP at circle (3). This transmission may include a verification aspect, whereby the receiving CSP can confirm whether they are able to fulfill all aspects of the requested contract, and if not, can return an error or other indication that the contract cannot be fulfilled. However, if the CSP verifies that it can satisfy the requested contract (or, if verification is not performed or is not required), the CSP can make the contract effective. For example, the CSP can configure its resources to prioritize routing / processing of traffic associated with the application.

[0051] In some examples, during the exchange at circle (3), each involved CSP may provide back a set of one or more network addresses to SOADM service 102. These addresses can be used for inbound traffic to its associated CSP deployment edge location (e.g., edge location 116A, which is deployed within the CSP's facility and is "within" or accessible through the CSP's network). Therefore, this set of network addresses is within the private address space of the CSP network and may be provided individually, as part of an address range or network address block (e.g., a Classless Inter-Domain Routing (CIDR) block), and may be IPv4 addresses, IPv6 addresses, etc.

[0052] In some examples, an edge location deployed within a CSP's facility (or network) may include a carrier gateway that blocks traffic from entering the edge location unless it originates from an "allowed" address (e.g., an address within a defined address range of the CSP). This configuration helps protect the edge location from external attacks by third parties, among other things. Therefore, in a scenario where SOADM service 102 might expect to forward application traffic originating from one CSP network to a CSP-deployed edge location on a different CSP network, the traffic will be blocked because it originates from a device outside that CSP network.

[0053] In some examples, this can be addressed by simply disabling the carrier gateway or otherwise configuring it to be more lenient in accepting traffic into the edge location. However, this allows potentially more types of traffic to enter the edge location, introducing security risks. Alternatively, in some examples, when SOADM service 102 wants to send traffic originating elsewhere to the CSP's edge location, SOADM service 102 can use one of the network addresses as the traffic source in a Network Address Translation (NAT) type scheme. This arrangement allows SOADM service 102 to send traffic that will be allowed into the edge location, receive return traffic, and maintain and use the original source network address of the user equipment device involved to return a response to it.

[0054] For example, at circle (4), the user equipment device (electronic device 110) of application user 108 can send a request to the application service via the cellular wireless network of CSP 'A'. CSP 'A' may be able to identify traffic as belonging to an application and determine that it is associated with a priority application workflow (i.e., given priority due to being under a valid contract). For example, the traffic may include a custom header with one or more values ​​identifying its association with an application (e.g., an application identifier value), and / or include one or more values ​​identifying its association with a contract (e.g., a contract identifier), although many other configurations known or deducible by those skilled in the art may be used. Alternatively, the CSP's network may also perform various lookups to identify or confirm this information, such as by examining the destination IP address or other header or payload information in the traffic. Regardless of how the association of traffic with applications and contracts is discovered or confirmed, the CSP may prioritize / transmit the traffic and send it to the edge location. Since the source address is the address of electronic device 110 (or other CSP network element) configured by the CSP network, the operator gateway will allow the traffic to enter the edge location 116A.

[0055] The traffic can then be delivered to the intelligent routing module 111A. If an established route for the traffic flow (e.g., the connection between electronic device 110 and application instance 120) no longer exists, the intelligent routing module 111A can identify the optimal location (and thus the optimal application instance 120) to send the traffic. Alternatively or additionally, the intelligent routing module 111A can verify that the "default" or existing destination still meets the application's QoS requirements, and if not, attempt to find another candidate location as a destination that can meet the application's QoS requirements.

[0056] In various examples, this determination can be based on a variety of factors. For instance, the intelligent routing module 111A may be able to determine which locations host application resources (e.g., application instance 120), which may include various edge locations 116 / 117 / 118, locations in region 112A, etc., and / or determine which locations are allowed to host application information (and therefore, new application instances 120 can be launched). The intelligent routing module 111A may also identify the requested application QoS network requirements (provided by user 104, such as via the console as part of a service contract 164) and the CSP metric value 162 of the location corresponding to the hosted application resources. The intelligent routing module 111A may also identify the load on these application instances 120, such as execution latency, request count per unit time, etc. Using some or all of this information, the intelligent routing module 111A may be configured to make routing decisions about which location is likely to provide the best service for the application, for example, by identifying the location that should provide the lowest round-trip time / total latency, identifying the location that can provide the highest bandwidth, determining the location that currently has QoS metrics that meet the application QoS requirements provided by all customers, etc.

[0057] Therefore, the optimal destination for traffic could be edge location 116A itself, as shown in circle (6), which brings the "computation" portion of the application very close to user 108. Alternatively, it is possible that a different location might be more suitable, whether it is another edge location 117A (in circle (7A)) within the same geographic location 114A but hosted in (or next to) a different CSP 'B', another edge location 116B (in circle (7B)) in the same CSP 'A' but located in a different geographic location 114N, or another edge location 118B (in circle (7C)) in a different CSP located in a different geographic location 114N. In these cases, the intelligent routing module 111A can identify the source address from a set of candidate source addresses provided (or configured) by the associated CSP network and use this address (via NAT type translation) to send traffic to the corresponding CSP network with the edge location as the destination. In some examples, this transmission can be performed very quickly because all edge locations can be connected back to the cloud provider network 100 via efficient links (such as fiber optic cables or radios), and / or one of the edge locations 116 / 117 / 118 can be directly interconnected at the switching point.

[0058] In some cases, the smart routing module 111A can alternatively (mostly or even entirely) be implemented within area 112A of the cloud provider network 100, as shown in the smart routing module 210. Figure 2 An exemplary intra-regional cloud provider network deployment of an intelligent routing module for edge application deployment in a smart multi-carrier network, according to some examples, is shown. In this example, traffic received from electronic device 110 destined for an application provided by cloud provider network 100 is sent via a CSP network to edge location 116A. In this case, when the existing forwarding / routing path of the packets / flows is unknown, the traffic can be sent to intelligent routing module 210 within area 112A of cloud provider network 100.

[0059] At circle (1), the steering module 202 can determine the location to be used as the traffic destination. At circle (2), the location may include analysis of application location 212 (exit or allow, in which case an application instance can be deployed), multi-CSP multi-location network metadata 214 (e.g., QoS type information indicating latency, round-trip time information, congestion, etc.), and / or contract 216 data. As an example, the steering module 202 can identify which application location 212 is "closest" to the location of electronic device 110 (based on geographic coordinate type information included in the traffic, and / or based on a specific network address used as a source identifier for the traffic, which may be known to be associated with a specific geographic location) and meets contract / application requirements. As another example, the user may have provided an indication of which metric or set of metrics is most important, and selection can be made based on analysis of these metrics to identify the optimal location.

[0060] After selecting a location (here, edge location 118B), the redirection module 202 at circle (3) can notify the network bridging module 200 of the assigned network address, which comes from that specific CSP and can be used as the source address when forwarding traffic. This source address and the actual source address associated with the received traffic can be maintained in the address mapping 204 data structure, and the traffic can be sent to the application instance at edge location 118B (e.g., via a high-speed link, such as fiber optic cable, radio link, etc.). When the response is sent back to the intelligent routing module 210, it can (e.g., from address mapping 204) identify the original source network address and use this address to send the response back to the electronic device 110 via the original CSP network. In some examples, once the traffic is determined... destination Network elements (e.g., routers) within the CSP network or edge location 116A can provide a more direct indication of the final destination of the traffic, allowing the traffic to be forwarded more directly to edge location 118B, and / or an identifier of this destination can be provided along with the traffic, enabling the redirection module 202 (or network bridging module 200) to identify the destination and forward the traffic more quickly.

[0061] In some cases, the smart routing module can also be implemented elsewhere, such as within the edge location itself. Figure 3An exemplary edge network deployment of the intelligent routing module 310 for intelligent multi-carrier network edge application deployment is shown, based on some examples. As shown, the intelligent routing module 310 can be deployed in one or more edge locations 116 / 118 and perform many or all of the same operations as the deployment configuration in the aforementioned areas, although with some potential direct modifications. For example, changes made to application location 212 and / or contract 216, as well as updates to network metadata 214 reported by the CSP, can be provided as periodic (or event-driven, or routing module-requested) updates 302 sent from SOADM service 102 to the intelligent routing module 310. Traffic to be sent to other locations can then be passed via a link to area 112A at circle (a) and then to another edge location 118B, or traffic can be sent directly (e.g., via a direct connection, or even across one or more other networks) to another location, as shown in circle (B).

[0062] As mentioned earlier, network metrics data can be collected from multiple CSPs and their multiple locations and presented to users, which may be part of allowing users to obtain connectivity contracts. Figure 4 This illustrates an exemplary graphical user interface (GUI) for operator network contract selection and customization for intelligent multi-carrier network edge application deployment, provided by SOADM services, based on some examples. In this example, the GUI 400 (such as a web application) allows the user to browse a contract marketplace where connectivity contracts with different CSPs can be obtained for application use. In this example, a set of "active" (or purchased or acquired) contracts displays the name of each CSP, the geographical location of the contract, and the amount of traffic (500 gigabytes in this case). The user can select a "view" input element (a button in this case) to view further details and perform actions on the contract (e.g., cancel, modify, verify, etc.).

[0063] Users can also browse other possible contracts, such as by selecting a location of interest (here, via dropdown input element 410, "San Francisco Bay Area"), and the GUI can then display available contracts 415—here, the first contract comes from a CSP named "premier-mobile" and network metrics corresponding to the location and / or the QoS guarantees available for the contract (optionally, in a way that allows viewing, adding, or removing additional QoS metrics for the plan, such as by selecting the "More" button), and the second contract comes from a CSP named "yay-area-mobile" and network metrics corresponding to the location and / or the QoS guarantees available for the contract (again, optionally, in a way that allows viewing, adding, or removing additional QoS metrics for the plan, such as by selecting the "More" button). Each entry has an associated "Buy" input element, allowing the user to attempt to verify and / or acquire such a contract with the CSP. Of course, this GUI 400 is exemplary, and many other configurations are also possible, such as including "filters" that allow the user to provide the required requirements for the contract to filter out unmatched products.

[0064] In some examples, the smart routing module can also dynamically adjust the destination during a session. For example, Figure 5 This diagram illustrates an exemplary mobility-based application traffic rerouting performed by SOADM services for intelligent multi-carrier network edge application deployments, based on several examples. As shown, when a user is located in San Francisco, their traffic can initially be intelligently allocated to a first edge location 505 in the San Francisco Bay Area at time = T1. However, if the user moves, such as by car (which may itself be the user's device) or train, the intelligent routing module can dynamically redistribute the user's traffic to different locations to better suit the application.

[0065] Subsequently, as the user's device location changes, the intelligent routing module may reassess which location is "best." As shown in 510, at time = T2, the user device may be located midway between San Francisco and Los Angeles, and the intelligent routing module can determine, at least in part, that a different location might be more optimized based on the user device's location, possibly because the latency between the user's location and the different application locations is less than the latency provided by the original edge location 505. Therefore, the user's application traffic can instead be rerouted to the second edge location 515, shown here as a solid border and a white center, indicating that it is owned by a different CSP. Later, at time = T3, as the user continues to move, the intelligent routing module can again determine that the application has a new "best" destination. Therefore, at 520, the user's application traffic can be rerouted to the application instance in the third edge location 505 in the Los Angeles area.

[0066] As described in this article, these technologies can simply provide the “fastest path” for processing user radio traffic in a prioritized manner, while “back-end” traffic can be similarly processed optimally by application developers. Therefore, application developers’ QoS requirements can be ensured through management of the intelligent routing module, without requiring more proactive or continuous developer involvement.

[0067] Figure 6 This is a flowchart illustrating operations of a method for deploying intelligent multi-carrier network edge applications according to some examples. Some or all of operations 600 (or other processes, variations, and / or combinations thereof described herein) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. The code is stored, for example, on a computer-readable storage medium in the form of a computer program including instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operations 600 are executed by intelligent routing modules and / or SOADM services 102 of other diagrams.

[0068] Operation 600 includes: at block 602, receiving a request message addressed to an application that is at least partially implemented in multiple edge locations of a cloud provider network, wherein the request message is initiated by a mobile user equipment device using the communication network of a first communication service provider (CSP).

[0069] In some examples, the request message is received within: the second edge location of the first CSP; or within the area of ​​the cloud provider's network.

[0070] Operation 600 further includes, in block 604, selecting an edge location from a plurality of edge locations that meets the application's quality of service (QoS) requirements as the destination of the request message, wherein the edge location is deployed within the facility of the second CSP. In some examples, the edge location is selected as the destination based at least in part on the geographic location of the mobile user equipment device. In some examples, the QoS requirements define at least one of the following: maximum latency; maximum jitter; minimum bandwidth; minimum availability score; or maximum round-trip time.

[0071] Operation 600 further includes, in block 606, identifying a first network address from a set of one or more network addresses provided by a second CSP, wherein the set of network addresses is within the address space of the network of the second CSP.

[0072] Operation 600 also includes, in block 608, transmitting a request message to an edge location using a first network address as the source identifier. In some examples, transmitting the request message to the edge location occurs at least in part using fiber optic or radio transmission.

[0073] In some examples, operation 600 further includes: receiving a response message originating from within the edge location; and sending the response message back to the mobile user equipment device via the communication network of the first CSP.

[0074] In some examples, operation 600 further includes: receiving a second request message sent to an application, wherein the second request message is initiated by a second mobile user equipment device using the communication network of the first CSP; selecting a second edge location from a plurality of edge locations as the destination of the second request message, wherein the second edge location is deployed within the facility of the first CSP; and transmitting the second request message to a computing instance within the second edge location.

[0075] In some examples, operation 600 further includes: receiving from each of a plurality of CSPs, including a first CSP and a second CSP, a set of network quality of service characteristics corresponding to each of one or more edge locations deployed within the facility of the CSP; and providing a summary of at least some of the set of network quality of service characteristics to a user associated with the application via a user interface. In some examples, operation 600 further includes: receiving by the cloud provider network a request for a service plan from the second CSP, the service plan being associated with the second CSP, with the edge location of the second CSP, or with a geographic location including the edge location. In some examples, operation 600 further includes: transmitting a message to the second CSP to configure or reserve network resources on behalf of a client associated with the application.

[0076] In some examples, operation 600 further includes: determining that a second edge location from a plurality of edge locations is more suitable than an edge location for handling services associated with a mobile user equipment device, wherein the second edge location is deployed within another facility of the second CSP; and causing an additional request message initiated by the mobile user equipment device to be sent to the second edge location.

[0077] As indicated herein, in some implementations, the SOADM service allows users to configure applications, which typically include one or more services and dependent resources, where deployment and / or distribution policies indicate how and where application developers wish to deploy the application. Using this user configuration data, SOADM can manage the deployment and redistribution of various components of an application on behalf of the user without requiring any (or a large number of) user involvement. For example, Figure 1The illustration shows an SOADM service according to some implementation schemes, which provides user-configurable multi-service application deployments and distribution across various deployment zones and locations. Figure 1 In this context, SOADM service 102 is implemented within a cloud provider network 100 and, in some implementations, is implemented as software executed by one or more computing devices located at one or more geographic locations.

[0078] In some implementations, segments of the cloud provider network—referred herein to as edge locations (“ELs,” or alternatively, “Provider Underlying Extensions” or “Edge Zones”)—may be configured within a separate network or facility from the cloud provider network. For example, a cloud provider network typically comprises a physical network (e.g., metal enclosures, cabling, rack hardware) referred to as the underlying layer. An underlying layer can be viewed as a network structure containing the physical hardware that runs the services of the provider network. In some implementations, provider underlying “extensions” (or edge zones) form edge locations that act as extensions to the cloud provider network underlying layer, consisting of one or more servers located in a user or partner facility, a separate cloud provider-managed facility, a communications service provider facility, or a facility of another type that includes servers local to a nearby AZ or region of the cloud provider network via a network (e.g., a publicly accessible network, such as the Internet) or a direct network connection. Users can access edge locations through the cloud provider underlying layer or other networks and can use the same or similar application programming interfaces (APIs)—provided by the cloud provider network itself—to create and manage resources at the edge locations as they would use to create and manage resources in the regions of the cloud provider network.

[0079] As described above, one example type of edge location is formed by servers located within a user's or partner's facility. This type of edge location, located outside the cloud provider's network data center, can be referred to as an "outpost" of the cloud provider's network. Another example type of edge location is formed by servers located in a separate facility managed (or operated, owned, leased, etc.) by the cloud provider, but includes data plane capacity that is at least partially controlled by a separate control plane within the cloud provider's network (e.g., in an Availability Zone or a cloud provider's network region).

[0080] In some implementations, another example of an edge location is a location deployed within the network infrastructure of a communications service provider. Communications service providers typically include companies that have deployed networks through which end users obtain network connectivity. For example, communications service providers can include mobile or cellular network providers (e.g., operating 3G, 4G, 5G networks, etc.), wired internet service providers (e.g., cable, digital subscriber line (DSL), fiber optic, etc.), and WiFi providers (e.g., in locations such as hotels, cafes, airports, stadiums, arenas, cities, etc.). While the traditional deployment of computing resources in data centers offers various benefits due to centralization, physical constraints between end-user devices and those computing resources (such as network distance and the number of network hops) can prevent the achievement of very low latency. By installing or deploying capacity in the form of edge locations, cloud provider network operators can provide significantly reduced access latency to end-user devices for computing resources, in some cases single-digit millisecond latency. This low-latency access to computing resources can be a significant driver for improving the responsiveness of existing cloud-based applications and enabling next-generation applications such as game streaming, virtual reality, real-time rendering, industrial automation, autonomous vehicles, etc.

[0081] Therefore, as used herein, computing resources of a cloud provider network located outside of a cloud provider network region are referred to as cloud provider network edge locations, or simply edge locations, because they are closer to the “edge” from which end users connect to a network compared to other computing resources deployed in more centralized data centers (e.g., within the facilities of a cloud provider that implements part of the cloud network region). Such edge locations include one or more networked computing systems that provide computing resources to users of the cloud provider network in order to serve end users with lower latency than that achievable with computing instances hosted in more traditional data center sites. Edge locations deployed within a communications service provider network can also be referred to as “wavelength zones.”

[0082] keep going, Figure 2Exemplary systems according to some embodiments are illustrated, including a cloud provider network and various edge locations of the cloud provider network. As previously described, cloud provider network 100 (sometimes simply referred to as the "cloud") refers to a pool of network-accessible computing resources (such as computing, storage, and networking resources; applications, and services), which may be virtualized or bare metal. The cloud provides convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and released in response to user commands. These resources can be dynamically provisioned and reconfigured to be adapted to variable loads. Therefore, cloud computing can be viewed as both applications delivered as a service over publicly accessible networks (e.g., the Internet, cellular communication networks) and the hardware and software in the cloud provider's data centers providing these services.

[0083] Cloud provider network 100 can provide users with on-demand, scalable computing platforms via a network, for example, allowing users to have scalable “virtual computing devices” available for their use via computing servers (which provide computing instances via one or both of a central processing unit (CPU) and a graphics processing unit (GPU) optionally used with local storage devices) and block storage servers (which provide virtualized persistent block storage for specified computing instances). These virtual computing devices have the attributes of a personal computing device, including hardware (various types of processors, local memory, random access memory (RAM), hard disk and / or solid-state drive (SSD) storage devices), operating system options, networking capabilities, and pre-loaded application software. Each virtual computing device can also virtualize its console input and output (e.g., keyboard, monitor, and mouse). This virtualization allows users to connect to their virtual computing devices using computer applications (such as browsers, application programming interfaces (APIs), software development kits (SDKs), etc.) so that they can configure and use their virtual computing devices as they would do with a personal computing device. Unlike personal computing devices that have a fixed amount of hardware resources available to the user, the hardware associated with a virtual computing device can be scaled up or down depending on the resources the user needs.

[0084] As indicated above, a user (e.g., user 838) can connect to virtualized computing devices and other cloud provider network 100 resources and services via one or more networks (e.g., public networks) using various interfaces (e.g., APIs). An API refers to an interface and / or communication protocol between a client (e.g., electronic device 834) and a server, such that if the client issues a request in a predefined format, the client should receive a response in a specific format or cause a defined action to be initiated. In a cloud provider network context, an API provides a gateway to allow users to access cloud infrastructure by allowing them to obtain data from or initiate actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted within the cloud provider network. APIs also enable different services within the cloud provider network to exchange data with each other. Users may choose to deploy their virtual computing systems to provide network-based services for their own use and / or for use by their users or clients.

[0085] The cloud provider network 100 may include a physical network (e.g., metal enclosures, cables, rack hardware) referred to as the substrate. The substrate can be viewed as a network structure containing the physical hardware running the provider network's services. The substrate may be isolated from the rest of the cloud provider network 100; for example, it may not be possible to route from the substrate network address to addresses in the production network running the cloud provider services, or to user networks hosting user resources.

[0086] The cloud provider network 100 may also include an overlay network of virtualized computing resources running on the underlying layer. In at least some embodiments, a hypervisor or other device or process on the network layer uses encapsulation protocol technology to encapsulate and route network packets (e.g., client IP packets) between client resource instances on different hosts within the provider network via the network layer. The encapsulation protocol technology can be used on the network layer to route encapsulated packets (also referred to as network layer packets) between endpoints on the network layer via overlay network paths or routes. The encapsulation protocol technology can be viewed as providing a virtual network topology overlay on the network layer. Thus, network packets can be routed along the underlying network based on the construction in the overlay network (e.g., a virtual network that may be referred to as a Virtual Private Cloud (VPC), a port / protocol firewall configuration that may be referred to as a security group). A mapping service (not shown) can coordinate the routing of these network packets. The mapping service may be a regionally distributed lookup service that maps a combination of overlay Internet Protocol (IP) and network identifiers to the underlying IP, enabling distributed underlying computing devices to find out where to send packets.

[0087] For illustration, each physical host device (e.g., compute server 806, block storage server 808, object storage server 810, control server 812) may have an IP address in the underlying network. Hardware virtualization technology enables multiple operating systems to run simultaneously on the host computer, for example, as virtual machines (VMs) on compute server 806. A hypervisor or virtual machine monitor (VMM) on the host assigns host hardware resources among the various VMs on the host and monitors the execution of the VMs. Each VM may be configured with one or more IP addresses in the overlay network, and the VMM on the host may be aware of the IP addresses of the VMs on the host. The VMM (and / or other devices or processes at the network layer) may use encapsulation protocol technology to encapsulate network packets (e.g., client IP packets) and route said network packets between virtualized resources on different hosts within the cloud provider network 100 via the network layer. Encapsulation protocol technology may be used at the network layer to route encapsulated packets between endpoints at the network layer via overlay network paths or routes. Encapsulation protocol technology can be viewed as providing a virtual network topology overlayed on the network layer. In some implementations, the encapsulation protocol technology includes a mapping service that maintains a mapping directory that maps IP overlay addresses (e.g., IP addresses visible to the user) to underlying IP addresses (IP addresses not visible to the user), which can be accessed by various processes on the cloud provider's network for routing packets between endpoints.

[0088] As shown in the figure, in various implementations, the traffic and operations at the underlying layer of the cloud provider network can be broadly subdivided into two categories: control plane traffic carried on the logical control plane 814A and data plane operations carried on the logical data plane 816A. While the data plane 816A represents the movement of user data through the distributed computing system, the control plane 814A represents the movement of control signals through the distributed computing system. The control plane 814A typically includes one or more control plane components or services distributed across and implemented by one or more control servers 812. Control plane traffic typically includes administrative operations such as establishing isolated virtual networks for various users, monitoring resource utilization and health, identifying specific hosts or servers to start a requested compute instance, and provisioning additional hardware as needed. The data plane 816A includes user resources implemented on the cloud provider network (e.g., compute instances, containers, block storage volumes, databases, file storage). Data plane traffic typically includes non-administrative operations such as transferring data to and from user resources.

[0089] Control plane components are typically implemented on a separate set of servers from the data plane servers, and control plane traffic and data plane traffic can be sent over separate / different networks. In some implementations, control plane traffic and data plane traffic may be supported by different protocols. In some implementations, messages (e.g., packets) sent over the cloud provider network 100 include flags indicating whether the traffic is control plane traffic or data plane traffic. In some implementations, the payload of the traffic is examined to determine its type (e.g., control plane or data plane). Other techniques for distinguishing traffic types are possible.

[0090] As shown in the figure, data plane 816A may include one or more compute servers 806, which may be bare metal (e.g., a single tenant) or may be virtualized by a hypervisor to run multiple VMs (sometimes referred to as "instances") or microVMs for one or more users. These compute servers 106 may support virtualized computing services (or "hardware virtualization services") of a cloud provider network. In some embodiments, the virtualized computing service may be part of control plane 814A, allowing users to issue commands via interface 804 (e.g., API) to launch and manage compute instances (e.g., VMs, containers) of their applications. In some embodiments, the virtualized computing service provides virtual compute instances with different compute and / or memory resources. In one embodiment, each of the virtual compute instances corresponds to one of several instance types. Instance types may be characterized by their hardware type, compute resources (e.g., the number, type, and configuration of CPUs or CPU cores), memory resources (e.g., the capacity, type, and configuration of local memory), storage resources (e.g., the capacity, type, and configuration of locally accessible storage devices), network resources (e.g., the characteristics of their network interfaces and / or network capabilities), and / or other suitable descriptive characteristics. The instance type selection functionality can, for example (at least in part), select an instance type for a user based on input from the user. For instance, a user can select an instance type from a set of predefined instance types. As another example, a user can specify the desired resources for an instance type and / or the requirements of the workload the instance will run, and the instance type selection functionality can select the instance type based on such a specification.

[0091] Data plane 816A may also include one or more block storage servers 808, which may include persistent storage devices for storing customer data volumes and software for managing these volumes. These block storage servers 808 may support managed block storage services for cloud provider networks. In some embodiments, the managed block storage service is part of control plane 814A, allowing users to issue commands via interface 804 (e.g., API) to create and manage volumes of applications running on compute instances. Block storage servers 808 include one or more servers on which data is stored as blocks. A block is a sequence of bytes or bits, typically containing an integer number of records with a maximum length of block size. Block data is typically stored in a data buffer and read or written in whole blocks at a time. Generally, a volume can correspond to a logical collection of data, such as a set of data maintained on behalf of a user. A user volume can be viewed as an individual hard drive ranging in size from, for example, 1 GB to 1 terabyte (TB) or larger, consisting of one or more blocks stored on a block storage server. Although viewed as an individual hard drive, it should be understood that a volume can be stored as one or more virtualized devices implemented on one or more underlying physical host devices. A volume can be partitioned a small number of times (e.g., up to 16 times), with each partition hosted on a different host. Volume data can be replicated across multiple devices within a cloud provider's network to provide multiple copies of the volume (where such copies collectively represent the volume on the computing system). Volume replicas in a distributed computing system can beneficially provide automatic failover and recovery, for example, by allowing users access to a primary copy of the volume or secondary copies of the volume synchronized with the primary copy at the block level, so that failure of the primary or secondary copy does not prevent access to volume information. The primary copy's role can be to facilitate reads and writes on the volume (sometimes referred to as "input / output operations" or simply "I / O operations") and propagate any writes to the secondary copy (preferably synchronously along the I / O path, but asynchronous replication can also be used). The secondary copy can be updated synchronously with the primary copy and provide a seamless transition during failover operations, whereby the secondary copy assumes the role of the primary copy, and the former primary copy is designated as the secondary copy or a new replacement secondary copy is provisioned. While some examples in this document discuss primary and secondary copies, it should be understood that a logical volume can include multiple secondary copies. Compute instances can virtualize their I / O to the volume via clients. The client represents instructions that enable a compute instance to connect to a remote data volume (e.g., a data volume stored on a physically separate compute device accessible over a network) and perform I / O operations at the remote data volume. In some implementations, the client is implemented on an offload card of a server that includes the processing unit (e.g., CPU or GPU) of the compute instance.

[0092] Data plane 816A may also include one or more object storage servers 810, which represent another type of storage device within a cloud provider network. Object storage servers 810 include one or more servers on which data is stored as objects within resources called buckets and can be used to support managed object storage services within the cloud provider network. Each object typically includes the stored data, variable metadata enabling the object storage server to analyze various aspects of the stored object, and a globally unique identifier or key that can be used to retrieve the object. Each bucket is associated with a given user account. Users can store any number of objects in their buckets, can write, read, and delete objects in their buckets, and can control access to their buckets and the objects contained therein. Furthermore, in implementations with multiple different object storage service servers distributed across different regions described above, users can select the region (or multiple regions) of the storage bucket, for example, to optimize latency. Users can use buckets to store various types of objects, including machine images that can be used to boot VMs, and snapshots representing point-in-time views of volume data.

[0093] Edge location 802 provides the resources and services of cloud provider network 100 within a separate network, thereby extending the functionality of cloud provider network 100 to new locations (e.g., for reasons related to latency in communication with user devices, legal compliance, security, etc.). As indicated, such edge location 802 may include cloud provider network managed edge location 740 (e.g., formed by servers located in cloud provider managed facilities separate from those facilities associated with cloud provider network 100), communication service provider edge location 842 (e.g., formed by servers associated with communication service provider facilities), user managed edge location 844 (e.g., formed by servers located on-premises at user or partner facilities), and other possible types of underlying extensions.

[0094] As illustrated in example edge location 840, edge location 802 may similarly include a logical separation between a control plane 818B and a data plane 820B, which extend the control plane 814A and data plane 816A of the cloud provider network 100, respectively. In some embodiments, edge location 802 is pre-configured by the cloud provider network operator with an appropriate combination of hardware and software and / or firmware elements to support various types of compute-related resources and to do so in a manner that reflects the experience of using the cloud provider network. For example, one or more edge location servers may be provisioned by the cloud provider to be deployed within edge location 802. As described above, in some embodiments, cloud provider network 100 provides a predefined set of instance types, each instance type having different types and amounts of underlying hardware resources. Various sizes for each instance type may also be provided. To enable users in edge 802 to continue using the same instance types and sizes they use in the region, servers may be heterogeneous servers. Heterogeneous servers may simultaneously support multiple instance sizes of the same type and may also be reconfigured to host any instance type supported by their underlying hardware resources. The reconfiguration of heterogeneous servers allows for the immediate occurrence of available server capacity—that is, while other VMs are still running and consuming additional capacity on edge location servers. This improves the utilization of compute resources within edge locations by allowing for better packaging of running instances on servers and provides a seamless experience regarding instance usage across Cloud Provider Network 100 and Cloud Provider Network edge locations.

[0095] As shown in the figure, the edge server can host one or more compute instances 822. Compute instances 822 can be VMs, or containers that package code and all its dependencies, allowing applications to run quickly and reliably across compute environments, such as VMs. Additionally, the server can host one or more data volumes 824 if needed by the user. In the region of the cloud provider network 100, such volumes could be hosted on dedicated block storage servers. However, due to the possibility of significantly smaller capacity at the edge location 802 compared to the region, optimal utilization may not be provided if such dedicated block storage servers are included at the edge location. Therefore, block storage services can be virtualized in edge location 802, allowing one of the VMs to run block storage software and store the data in volume 824. Similar to the operation of block storage services in the region of the cloud provider network 100, volume 824 within edge location 802 can be replicated for persistence and availability. The volume can be provisioned in its own isolated virtual network within edge location 802. Computation instance 822 and any volume 824 together constitute the provider network data plane 816A, and the data plane extension 820B within edge location 802.

[0096] In some implementations, servers within edge location 802 may host certain local control plane components 826, such as components that enable edge location 802 to continue operating in the event of an interruption in the connection back to cloud provider network 100. Examples of such components include: a migration manager that can move compute instances 822 between edge location servers if availability needs to be maintained; and a key value data store that indicates the location of volume copies. However, the functionality of the control plane 818B for the edge location will generally remain within cloud provider network 100 to allow users to utilize as much of the resource capacity as possible from the edge location.

[0097] In some implementations, the migration manager has a centralized coordination component running in a region, and a local controller running on servers at edge locations (and servers in the cloud provider's data center). When a migration is triggered, the centralized coordination component can identify the target edge location and / or the target host, while the local controller can coordinate the data transfer between the source host and the target host. The described resource movement between hosts in different locations can take one of several migration forms. Migration refers to moving virtual machine instances (and / or other resources) between hosts in a cloud computing network or between hosts outside the cloud computing network and hosts within the cloud. Different types of migrations exist, including live migrations and restart migrations. During a restart migration, the user experiences an interruption and effective power restart of their virtual machine instance. For example, a control plane service can coordinate a restart migration workflow that involves tearing down the current domain on the source host and then creating a new domain for the virtual machine instance on the new host. The instance is restarted by shutting down on the source host and then restarting on the new host.

[0098] Live migration refers to the process of moving a running virtual machine or application between different physical machines without significantly disrupting the availability of the virtual machine (e.g., the end user will not notice the downtime). When the control plane performs a live migration workflow, it can create a new "inactive" domain associated with the instance, while the instance's original domain continues to run as the "active" domain. The virtual machine's memory (including any state in memory of the running application), storage, and network connectivity are transferred from the original host with the active domain to the destination host with the inactive domain. The virtual machine may be briefly paused to prevent state changes while transferring the memory contents to the destination host. The control plane can transition the inactive domain to become an active domain and degrade the original active domain to an inactive domain (sometimes called "flipping"), after which the inactive domain can be discarded.

[0099] Techniques for various types of migration involve managing critical phases—the time when virtual machine instances are unavailable to users—that should be kept as short as possible. This can be particularly challenging in currently disclosed migration techniques because resources are being moved between hosts in geographically separated locations connected by one or more intermediate networks. For live migration, the disclosed techniques can dynamically determine, for example, the amount of memory state data to be pre-copied (e.g., while the instance is still running on the source host) and post-copied (e.g., after the instance begins running on the target host) based on latency between locations, network bandwidth / usage patterns, and / or based on which memory pages the instance most frequently uses. Furthermore, the specific time for transferring memory state data can be dynamically determined based on network conditions between locations. This analysis can be performed by a migration management component in a region or by a migration management component running locally in a source edge location. If the instance has access to the virtualized storage device, both the source and target domains can be attached to the storage device simultaneously to enable uninterrupted access to its data during migration and in the event of a rollback to the source domain.

[0100] In some examples, server software running at edge location 802 may be designed by the cloud provider to run on the cloud provider's underlying network, and the software may be able to create a private copy ("shadow underlay") of the underlying network within the edge location by running unmodified at edge location 802 using a local network manager 828. The local network manager 828 may run on the server at edge location 802 and bridge the shadow underlay with the edge location 802 network, for example, by acting as a VPN endpoint or an endpoint between edge location 802 and proxies 830, 832 in the cloud provider network 100, and by implementing a mapping service (for traffic encapsulation and decapsulation) to associate data plane traffic (from data plane proxies) and control plane traffic (from control plane proxies) with appropriate servers. By implementing a local version of the provider network's underlay-overlay mapping service, the local network manager 828 allows resources in edge location 802 to communicate seamlessly with resources in the cloud provider network 100. In some implementations, a single local network manager may perform these actions for all servers hosting compute instances 822 at edge location 802. In other implementations, each of the servers hosting compute instance 822 has a dedicated local network manager. In a multi-rack edge location, inter-rack communication can be achieved through the local network manager, which maintains open tunnels between itself and others.

[0101] The provider's underlying extended location may utilize secure networking tunnels through the edge location 802 network to the cloud provider network 100, for example, to maintain the security of user data while traversing the edge location 802 network and any other intermediate networks (potentially including the public internet). Within the cloud provider network 100, these tunnels consist of virtual infrastructure components including isolated virtual networks (e.g., in an overlay network), a control plane agent 830, a data plane agent 832, and underlying network interfaces. In some implementations, such agents may be implemented as containers running on compute instances. In some implementations, each server in the edge location 802 hosting the compute instances may utilize at least two tunnels: one for control plane traffic (e.g., Constrained Application Protocol (CoAP) traffic) and one for encapsulated data plane traffic. A connectivity manager (not shown) within the cloud provider network manages their cloud provider network-side lifecycle, for example, by automatically provisioning these tunnels and their components as needed and maintaining them in a healthy operational state. In some implementations, a direct connection between the edge location 802 and the cloud provider network 100 may be used for control data plane and data plane communication. Compared to VPNs that go through other networks, direct connections offer constant bandwidth and more consistent network performance because their network path is relatively fixed and stable.

[0102] A control plane (CP) agent 830 can be configured within the cloud provider network 100 to represent a specific host at the edge location. The CP agent acts as an intermediary between the control plane 814A in the cloud provider network 100 and the control plane 818B at the edge location 802. Specifically, the CP agent 830 provides the infrastructure for tunneling management API traffic destined for the edge location server from the regional underlying layer to the edge location 802. For example, the virtualized compute service of the cloud provider network 100 may issue commands to the VMM of the server at the edge location 802 to start compute instance 822. The CP agent maintains a tunnel (e.g., a VPN tun) to the local network manager 828 at the edge location. Software implemented within the CP agent ensures that only well-formed API traffic leaves and returns to the underlying layer. The CP agent provides a mechanism to expose remote servers on the cloud provider's underlying layer while still protecting underlying security data (e.g., encryption keys, security tokens) from leaving the cloud provider network 100. The unidirectional control plane traffic tunnel imposed by the CP agent also prevents any (potentially compromised) devices from calling back to the underlying layer. CP proxies can be instantiated one-to-one with servers at edge location 802, or can manage control plane traffic for multiple servers at the same edge location.

[0103] Data plane (DP) proxy 832 can also be configured within cloud provider network 100 to represent a specific server in edge location 802. DP proxy 832 acts as a shadow or anchor point for the server and can be used by services within cloud provider network 100 to monitor the health of hosts (including their availability, used / idle compute and capacity, used / idle storage and capacity, and network bandwidth utilization / availability). DP proxy 832 also allows isolated virtual networks to cross edge location 802 and cloud provider network 100 by acting as proxies for servers in cloud provider network 100. Each DP proxy 832 can be implemented as a packet-forwarding compute instance or container. As shown, each DP proxy 832 can maintain a VPN tunnel with a local network manager 828 that manages traffic to the server represented by the DP proxy 832. The tunnel can be used to send data plane traffic between the edge location server and cloud provider network 100. Data plane traffic flowing between edge location 802 and cloud provider network 100 can be transmitted via the DP proxy 832 associated with the edge location. For data plane traffic flowing from edge location 802 to cloud provider network 100, DP proxy 832 can receive the encapsulated data plane traffic, verify its correctness, and allow it to enter cloud provider network 100. DP proxy 832 can then forward the encapsulated traffic directly from cloud provider network 100 to edge location 802.

[0104] A local network manager 828 provides secure network connectivity to agents 830, 832 established within the cloud provider network 100. Once a connection is established between the local network manager 828 and the agents, a user can issue commands via interface 804 to instantiate (and / or perform other operations using) computing instances using edge location resources in a manner similar to how such commands would be issued to computing instances hosted within the cloud provider network 100. From the user's perspective, the user can now seamlessly use local resources within the edge location (and resources located within the cloud provider network 100, if needed). Computing instances hosted on servers at edge location 802 can communicate with electronic devices located on the same network, and, as needed, with other resources hosted within the cloud provider network 100. A local gateway 846 can be implemented to provide network connectivity between edge location 802 and networks associated with the extension (e.g., the communication service provider network in the example of edge location 842).

[0105] There may be situations where data needs to be transferred between the object storage service and the edge location 802. For example, the object storage service may store machine images used to launch VMs, as well as snapshots representing point-in-time backups of volumes. The object gateway may be provided on an edge location server or dedicated storage device and offer users configurable bucket-by-bucket caching of the contents of object storage buckets in their edge locations to minimize the impact of edge location region latency on user workloads. The object gateway may also temporarily store snapshot data from snapshots of volumes in the edge location and then synchronize it with the object server in the region, where possible. The object gateway may also store machine images that users specify for use within the edge location or at the user's premises. In some implementations, data within the edge location may be encrypted with a unique key, and for security reasons, cloud providers may restrict key sharing from the region to the edge location. Therefore, data exchanged between the object storage server and the object gateway may utilize encryption, decryption, and / or re-encryption to maintain secure boundaries regarding encryption keys or other sensitive data. A transformation intermediary can perform these operations and can (on the object storage server) create edge location buckets using PSE encryption keys to store snapshot and machine image data.

[0106] In this manner, an edge location forms a "provider-level extension" because it provides the resources and services of the cloud provider's network outside of and closer to the user's equipment, within the traditional cloud provider's data center. Edge locations can be structured in various ways. In some implementations, an edge location can be an extension of the cloud provider's network layer, including a limited amount of capacity provided outside of an availability zone (e.g., in a small data center of the cloud provider or in other facilities located near the user's workload and potentially far from any availability zone). Such edge locations can be referred to as "far zones" (due to their distance from other availability zones) or "near zones" (due to their proximity to the user's workload). Near zones can be connected to publicly accessible networks such as the public internet in various ways (e.g., directly, via another network, or via a dedicated connection to the region). While near zones typically have more limited capacity than a region, in some cases, a near zone can have a considerable capacity, such as thousands or more racks.

[0107] In some implementations, an edge location is an extension of the underlying layer of a cloud provider network, consisting of one or more servers located locally at a user's or partner's facility. These servers communicate with nearby availability zones or areas of the cloud provider network via a network (e.g., a publicly accessible network such as the Internet). This type of underlying extension, located outside the cloud provider network's data center, can be referred to as an "outpost" of the cloud provider network. Some outposts may be integrated into a communications network, for example, as multi-access edge computing (MEC) sites, whose physical infrastructure is distributed across telecommunications data centers, telecommunications aggregation sites, and / or telecommunications base stations within the telecommunications network. In a local example, the limited capacity of an outpost may be available only to the user who owns the site (and any other accounts permitted by the user). In a telecommunications example, the limited capacity of an outpost may be shared among any number of applications (e.g., games, virtual reality applications, healthcare applications) sending data to users on the telecommunications network.

[0108] Edge locations may include data plane capacity that is at least partially controlled by the control plane of a nearby availability zone within the provider network. Thus, an availability zone group may include a “parent” availability zone and any “child” edge locations belonging to the parent availability zone (e.g., at least partially controlled by its control plane). Certain limited control plane functionality (e.g., features requiring low-latency communication with user resources, and / or features enabling the edge location to continue operating when disconnected from the parent availability zone) may also exist in some edge locations. Therefore, in the example above, an edge location refers to an extension of at least the data plane capacity located at the edge of the cloud provider network, close to user devices and / or workloads.

[0109] Figure 9 An exemplary cloud provider network comprising geographically dispersed edge locations is illustrated according to some embodiments. As shown, the cloud provider network 100 may be formed as a plurality of regions 112, wherein a region is a separate geographical area in which the cloud provider has one or more data centers 904. Each region 112 may include two or more AZs interconnected with each other via a private high-speed network such as, for example, fiber optic communication.

[0110] The number of edge locations 116 can be significantly higher than the number of regional data centers or Availability Zones (AZs). This widespread deployment of edge locations 116 provides low-latency connectivity to the cloud for a much larger group of end-user devices (compared to those groups of end-user devices that happen to be very close to a regional data center). In some implementations, each edge location 116 may peer to a portion of the cloud provider network 100 (e.g., a parent Availability Zone or regional data center). This peering allows various components operating within the cloud provider network 100 to manage the computing resources of the edge location. In some cases, multiple edge locations may be located or installed in the same facility (e.g., a separate rack for a computer system) and managed by different zones or data centers to provide additional redundancy. It should be noted that while edge locations are generally described herein as being within a CSP network, in some cases, such as when the cloud provider network facility is relatively close to the communication service provider facility, the edge location may remain physically located within the cloud provider network while being connected to the communication service provider network via fiber optic or other network links.

[0111] Edge location 116 can be structured in several ways. In some implementations, edge location 116 can be an extension of the underlying cloud provider network, including a limited amount of capacity provided outside of an Availability Zone (AZ), such as in a small data center of the cloud provider or in other facilities located close to customer workloads and potentially far from any AZ. Such edge locations may be referred to as “local areas” (because they are closer to or more conveniently located than traditional AZs, such as a large group of users, industries, or IT centers). Local areas can be connected to publicly accessible networks such as the Internet in various ways, such as directly, via another network, or via a private connection to a region. While local areas typically have more limited capacity than a region, in some cases, local areas can have considerable capacity, such as thousands or more racks. Some local areas may use infrastructure similar to that of a typical cloud provider data center, rather than the edge location infrastructure described herein.

[0112] The role of an edge location as a parent of an Availability Zone (AZ) or region within a cloud provider's network can be based on numerous factors. One such parenting factor is data sovereignty. For example, to keep data originating from a CSP network in a particular country within that country, an edge location deployed within that CSP network can serve as a parent of an AZ or region within that country. Another factor can be service availability. For instance, some edge locations may have different hardware configurations, such as the presence of components like local non-volatile storage devices (e.g., solid-state drives) for user data, graphics accelerators, etc. Some AZs or regions may lack services that utilize these additional resources, so an edge location can serve as a parent of an AZ or region that supports the use of those resources. Yet another factor can be latency between an AZ or region and the edge location. While deploying an edge location within a CSP network offers latency benefits, these benefits can be offset by having the edge location as a parent of a distant AZ or region, introducing significant latency to traffic from the edge location to the region. Therefore, edge locations are typically parents of nearby AZs or regions (in terms of network latency).

[0113] Figure 10 An exemplary system is illustrated, according to some implementations, in which a cloud provider network edge location is deployed within a communications service provider network. The CSP network 1000 typically includes downstream interfaces to end-user electronic devices and upstream interfaces to other networks (e.g., the Internet). In this example, the CSP network 1000 is a wireless “cellular” CSP network, which includes radio access networks (RAN) 1002, 1004, aggregation sites (AS) 1006, 1008, and a core network (CN) 1010. RAN 1002, 1004 include base stations (e.g., NodeB, eNodeB, gNodeB) providing wireless connectivity to electronic devices 1012. The core network 1010 typically includes functionalities related to the management of the CSP network (e.g., billing, mobility management, etc.) and relay traffic between the CSP network and other networks. Aggregation sites 1006, 1008 can be used to consolidate traffic from many different radio access networks into the core network and to route traffic originating from the core network to various radio access networks.

[0114] End-user electronic device 1012 Figure 10From left to right, a base station (or radio base station) 1014 is wirelessly connected to radio access network 1002. This type of electronic device 1012 is sometimes referred to as user equipment (UE) or customer premises equipment (CPE). Data traffic is typically routed to core network 1010 via a fiber optic transmission network consisting of multiple hops of Layer 3 routers (e.g., at aggregation sites). Core network 1010 is typically housed in one or more data centers. For data traffic destined for locations outside of communication network 1000, network components 1022-1026 typically include firewalls through which traffic can enter or leave CSP network 1000 to reach external networks, such as the Internet or cloud provider network 100. Note that in some embodiments, CSP network 1000 may include facilities that allow traffic to enter or leave from sites further downstream of core network 1010 (e.g., at aggregation sites or RANs).

[0115] Edge locations 1016-1020 (or “wavelength zones”) include computing resources that are managed as part of a cloud provider network but are installed or located at various points within the CSP network (e.g., locally in space owned or leased by the CSP). Computing resources typically provide a certain amount of compute and storage capacity that the cloud provider can allocate to its users. Computing resources may also include storage and accelerator capacity (e.g., solid-state drives, graphics accelerators, etc.). Here, edge locations 1016, 1018, and 1020 communicate with cloud provider network 100.

[0116] Typically, for example, in terms of network hops and / or distance, the farther the edge location is from the cloud provider network 100 (or closer to the device 1012), the lower the network latency between the computing resources within the edge location and the device 1012. However, physical site constraints typically limit the amount of computing capacity that can be installed at each point within the CSP, or determine whether computing capacity can be installed at all points. For example, edge locations within the core network 1010 can typically have a much larger footprint (in terms of physical space, power requirements, cooling requirements, etc.) compared to edge locations within RAN 1002, 1004.

[0117] The installation or location at the edge of a CSP network may vary depending on the specific network topology or architecture of the CSP network. For example... Figure 10As indicated, edge locations are typically connectable to any location on the CSP network where packet-based traffic (e.g., IP-based traffic) can be interrupted. Additionally, communication between a given edge location and the cloud provider network 100 is typically securely relayed to at least a portion of the CSP network 1000 (e.g., via secure tunnels, virtual private networks, direct connections, etc.). In the example shown, network component 1022 facilitates data traffic routing to and from edge location 1016 integrated with RAN 1002, network component 1024 facilitates data traffic routing to and from edge location 1018 integrated with AS 1006, and network component 1026 facilitates data traffic routing to and from edge location 1020 integrated with CN 1010. Network components 1022-1026 may include routers, gateways, or firewalls. To facilitate routing, the CSP may assign one or more IP addresses from the CSP network address space to each edge location.

[0118] In 5G wireless network development, edge locations can be considered a possible implementation of Multi-Access Edge Computing (MEC). Such edge locations can connect to various points within a CSP 5G network, which provide interruption for data traffic as part of the User Plane Function (UPF). Older wireless networks can also include edge locations. For example, in 3G wireless networks, edge locations can connect to the packet-switched network portion of the CSP network, such as connecting to the Serving General Packet Radio Service Support Node (SGSN) or the Gateway General Packet Radio Service Support Node (GGSN). In 4G wireless networks, edge locations can connect to the Serving Gateway (SGW) or Packet Data Network Gateway (PGW) as part of the core network or Evolved Packet Core (EPC).

[0119] In some implementations, traffic between edge location 1028 and cloud provider network 100 can bypass CSP network 1000 without being routed through core network 1010. For example, network component 1030 of RAN 1004 can be configured to route traffic between edge location 1016 of RAN 1004 and cloud provider network 100 without traversing aggregation site or core network 1010. As another example, network component 1031 of aggregation site 1008 can be configured to route traffic between edge location 1032 of aggregation site 1008 and cloud provider network 100 without traversing core network 1010. Network components 1030, 1031 may include gateways or routers with routing data to direct traffic from edge location to cloud provider network 100 to cloud provider network 100 (e.g., via direct connection or intermediate network 1034) and to edge location traffic from cloud provider network 100 to edge location.

[0120] In some implementations, an edge location may connect to more than one CSP network. For example, an edge location may connect to two CSP networks when two CSPs share or route traffic through a common point. For example, each CSP may allocate a portion of its network address space to the edge location, and the edge location may include a router or gateway that can distinguish traffic exchanged with each of the CSP networks. For example, traffic destined for an edge location from one CSP network may have different destination IP addresses, source IP addresses, and / or VLAN tags compared to traffic received from another CSP network. Traffic originating from the edge location destined for one of the CSP networks may similarly be encapsulated with appropriate VLAN tags, source IP addresses (e.g., from a pool allocated to the edge location from the destination CSP network's address space), and destination IP addresses.

[0121] It should be noted that, although Figure 10 An exemplary CSP network architecture includes a radio access network, aggregation sites, and a core network; however, the naming and structure of a CSP network architecture may differ between generations of wireless technologies, between different CSPs, and between wireless CSP networks and fixed-line CSP networks. Additionally, although Figure 10 Several locations where edge locations can be located within a CSP network are shown, but other locations are also possible (e.g., at a base station).

[0122] Figure 11Exemplary components of edge locations within a cloud provider network and a CSP network, and the connectivity between them, are shown in more detail according to some embodiments. Edge location 1100 provides the resources and services of the cloud provider network within CSP network 1102, thereby extending the functionality of cloud provider network 100 to be closer to end-user devices 1104 connected to the CSP network.

[0123] Edge location 1100 similarly includes a logical separation between a control plane 1106B and a data plane 1108B, which extend the control plane 814A and data plane 816A of the cloud provider network 100, respectively. Edge location 1100 may be pre-configured by the cloud provider network operator with appropriate combinations of hardware and software and / or firmware elements to support various types of compute-related resources and in a manner that reflects the experience of using the cloud provider network. For example, one or more edge location servers 1110 may be provisioned by the cloud provider for deployment within the CSP network 1102.

[0124] In some implementations, the server 1110 within edge location 1100 may host certain local control plane components 1114, such as components that enable edge location 1100 to continue operating in the event of an interruption in the connection back to cloud provider network 100. Furthermore, certain controller functions may typically be implemented locally on data plane servers or even in the cloud provider's data center, such as functions for collecting metrics for monitoring instance health and sending those metrics to a monitoring service, and functions for coordinating the transfer of instance status data during live migration. However, the control plane 1106B functionality for edge location 1100 will typically remain within cloud provider network 100 to allow users to utilize as much of the edge location's resource capacity as possible.

[0125] As shown in the figure, edge location server 1110 can host compute instance 1112. A compute instance can be a VM, a microVM, or a container that packages code and all its dependencies, allowing applications to run quickly and reliably across compute environments, including VMs. Therefore, a container is an abstraction of the application layer (meaning each container emulates a different software application process). Although each container runs an isolated process, multiple containers can share a common operating system, for example, by starting within the same virtual machine. In contrast, a virtual machine is an abstraction of the hardware layer (meaning each virtual machine emulates a physical machine capable of running software). Virtual machine technology can use a single physical server to run the equivalent of multiple servers (each referred to as a virtual machine). While multiple virtual machines can run on a single physical machine, each virtual machine typically has its own copy of the operating system, along with the application and its associated files, libraries, and dependencies. Virtual machines are often referred to as compute instances or simply "instances." Some containers can run on instances that run container agents, while others can run on bare-metal servers.

[0126] In some implementations, the execution of edge-optimized computing instances is supported by a lightweight virtual machine manager (VMM) running on server 1110, which launches edge-optimized computing instances based on application profiles. These VMMs enable the launch of lightweight microvirtual machines (microVMs) in fractions of a second. These VMMs also enable container runtimes and container orchestrators to manage containers as microVMs. Nevertheless, these microVMs also leverage the security and workload isolation provided by traditional VMs, such as running them as isolated processes via a VMM, as well as the resource efficiency offered by containers. As used herein, a microVM refers to a VM that is initialized using a finite device model and / or a minimal OS kernel supported by a lightweight VMM, and each microVM can have a low memory overhead of <5 MiB, allowing thousands of microVMs to be packaged onto a single host. For example, a microVM can have a streamlined version of the OS kernel (e.g., only with the required OS components and their dependencies) to minimize startup time and memory footprint. In one implementation, each process of the lightweight VMM encapsulates one and only one microVM. The process can run the following threads: API, VMM, and vCPU. The API thread is responsible for the API server and its associated control plane. The VMM thread exposes the machine model, the minimal legacy device model, the microVM metadata service (MMDS), and the VirtIO device emulation network and block device. Additionally, one or more vCPU threads can exist (one vCPU thread per guest CPU core).

[0127] Additionally, server 1110 may host one or more data volumes 1124 if required by the user. Volumes can be configured in their own isolated virtual network within edge location 1100. Compute instance 1112 and any volume 1124 together constitute the provider network data plane 816A, an extension of the data plane 1108B within edge location 1100.

[0128] A local gateway 1116 can be implemented to provide network connectivity between edge location 1100 and CSP network 1102. The cloud provider can configure the local gateway 1116 using IP addresses on CSP network 1102 and exchange routing data with CSP network component 1120 (e.g., via Border Gateway Protocol (BGP)). The local gateway 1116 may include one or more routing tables that control the routing of inbound traffic to edge location 1100 and outbound traffic leaving edge location 1100. The local gateway 1116 may also support multiple VLANs where different parts of CSP network 1102 use separate VLANs (e.g., one VLAN label for the wireless network and another VLAN label for the fixed network).

[0129] In some implementations of edge location 1100, the extension includes one or more switches, sometimes referred to as top-of-rack (ToR) switches (e.g., in rack-based implementations). The ToR switches connect to a CSP network router (e.g., CSP network component 1120), such as a provider edge (PE) or software-defined wide area network (SD-WAN) router. Each ToR switch may include an uplink link aggregation (LAG) interface to the CSP network router, with each LAG supporting multiple physical links (e.g., 1G / 10G / 40G / 100G). These links may run the Link Aggregation Control Protocol (LACP) and be configured as IEEE 802.1q trunks to enable multiple VLANs on the same interface. This LACP-LAG configuration allows the edge location management entity of the cloud provider network 100's control plane to add more peering links to the edge location without rerouting. Each of the ToR switches can establish an eBGP session with the carrier PE or SD-WAN router. The CSP may provide a Private Autonomous System Number (ASN) for the edge location and the CSP network 1102 to facilitate the exchange of routing data.

[0130] Data plane traffic originating from edge location 1100 may have multiple different destinations. For example, traffic addressed to a destination in data plane 816A of cloud provider network 100 may be routed via a data plane connection between edge location 1100 and cloud provider network 100. Local network manager 1118 may receive packets from compute instance 1112 addressed to, for example, another compute instance in cloud provider network 100, and encapsulate the packets with a destination as the underlying IP address of a server hosting another compute instance, which then sends them (e.g., via a direct connection or tunnel) to cloud provider network 100. For traffic from compute instance 1112 addressed to another compute instance hosted in another edge location 1122, local network manager 1118 may encapsulate the packets with a destination as the IP address assigned to the other edge location 1122, thereby allowing CSP network component 1120 to handle packet routing. Alternatively, if CSP network component 1120 does not support traffic between edge locations, local network manager 1118 may address packets to a repeater in cloud provider network 100, which may then forward the packets to another edge location 1122 via its data plane connection (not shown) to cloud provider network 100. Similarly, for traffic from compute instance 1112 address to CSP network 1102 or a location outside cloud provider network 100 (e.g., on the Internet), if CSP network component 1120 allows routing to the Internet, local network manager 1118 may encapsulate packets with a source IP address corresponding to an IP address in the carrier address space allocated to compute instance 1112. Otherwise, local network manager 1118 may forward packets to an Internet gateway in cloud provider network 100, which may provide Internet connectivity for compute instance 1112. For traffic originating from compute instance 1112 addressed to electronic device 1104, local gateway 1116 can use Network Address Translation (NAT) to change the source IP address of the packet from the address space of the cloud provider network to the address space of the address carrier network.

[0131] The local gateway 1116, the local network manager 1118, and other local control plane components 1114 may run on the same server 1110 as the managed computing instance 1112, on a dedicated processor integrated with the edge location server 1110 (e.g., on an offload card), or may be executed by a server separate from those servers that manage user resources.

[0132] Back Figure 1Application developers face numerous complexities when attempting to implement distributed applications using cloud provider networks. For example, it can be challenging to choose where to deploy service components to optimize latency and cost to end users, distribute applications and their data across multiple locations, optimize compute capacity across locations based on a global capacity budget, connect external (e.g., via mobile / internet) and internal (e.g., between microservices) client requests to the nearest possible location, and operate, monitor, and adapt the distributed application as different users use it over time. Therefore, in some implementations, SOADM service 102 provides an edge computing architecture that abstracts away these complexities, eliminating the need for developers to choose where to deploy their application components, manage location-specific deployment processes, or optimize capacity during traffic fluctuations. Instead, users can provide configuration data for the application, and SOADM service 102 can dynamically adapt the application to changing end-user locations and call volumes.

[0133] The SOADM service 102 implementation disclosed in this paper enables users to more easily build highly available and / or latency-sensitive applications that will run seamlessly across multiple deployment zones (and their various types) using, such as AZ, local zone (LZ), wavelength zone (WZ), etc., thereby abstracting away all the complexities that arise when managing applications across many locations.

[0134] To this end, in some implementations, SOADM service 102 provides users with a single management interface (e.g., via API, web-based console, etc.) to manage their highly distributed applications. This SOADM service 102 can then dynamically select locations to deploy user applications, orchestrate underlying compute and network resources, and streamline the collection of observable telemetry data to a central location. Thus, SOADM service 102 presents a new paradigm to users by using a deployment model that supports this distributed application model of infrastructure and application code.

[0135] In some implementations, SOADM Service 102 provides a global deployment experience, enabling users to specify certain aspects of application deployment behavior, such as deployment cadence, validation steps, automatic rollback configurations, etc., and then manage deployment locations and the underlying sequencing of low-level deployment activities. SOADM Service 102 can also deploy supporting infrastructure (such as VPNs, load balancers, server endpoints, routers, monitoring service alarm functions, etc.) to useful locations. Therefore, the SOADM Service 102 deployment model allows users to leverage deployment best practices and manage the complexity of deploying code and infrastructure updates across a dynamic set of deployment locations.

[0136] SOADM service 102 also provides latency-based scaling for user application components, allowing applications to scale to new locations to accommodate changes in localization needs, site availability, capacity availability, etc. For example, if a user's application is configured to be distributed across a distribution group (e.g., a logical set of one or more regions of a provider network), SOADM service 102 can scale application components to edge locations (e.g., local regions, wavelength zones) closer to where client connections flow into the application, such as locations near or within Los Angeles. SOADM service 102 can also be configured by the user to balance horizontal scaling decisions with user-specified constraints, such as the maximum number of locations the service can exist in, the deployment zone type the application can be placed in, and the maximum or minimum number of compute instances (or other compute resource units) for a specific component or the entire application. Therefore, SOADM service 102 can monitor and analyze global traffic data for user applications (as well as information about capacity utilization and deployment location "health") to identify better locations for deploying additional application capacity and / or removing application capacity from certain locations.

[0137] For example, at circle (1), user 104 (e.g., a software application developer) interacts with SOADM service 102 using their electronic device 106 (e.g., a computing device such as a laptop computer, personal computer, tablet computer, smartphone, etc.) to provide configuration data 720 for their application. In some implementations, such as via a web application or a standalone application, SOADM service 102 (via electronic device 106) provides a graphical user interface (GUI) to user 104, which the user can utilize to configure their application. In some implementations, the user uses another means (e.g., a text editor or other application) to provide configuration data, and electronic device 106 can make API calls to SOADM service 102.

[0138] To define the application's configuration data 720, one or more API calls or commands are transmitted at circle (1). Examples using multiple API calls are used in this specification (and are described in further detail herein with reference to the accompanying drawings); however, it should be understood that in various implementations, more, fewer, and / or different commands may be implemented based on the implementer's expectations, and therefore these calls are considered illustrative rather than restrictive.

[0139] In some implementations, configuration data 720 defines a service group 722 for the application, which may include one or more service configurations 724 (corresponding to each type of service / microservice / component of the application), zero or more service resource configurations 726 (associated with one of the services and corresponding to the resources the service depends on, such as virtual block storage volumes, data storage systems, databases, or other components), one or more distribution policies 728 (each associated with a service and indicating where the service can be deployed), one or more deployment configurations 730 (associated with the service group and indicating how the service is deployed), and potentially other types of configuration data.

[0140] For example, a user can create a service group 722 configuration associated with one or more service configurations 724, each service configuration representing a set of computing infrastructure (e.g., containers, virtual machines, executables, code, etc.) and the underlying infrastructure of application services.

[0141] A service group is a logical grouping of related services and the service resources they depend on. Services within a service group can communicate with each other over a dedicated network, and services can be restricted to accessing resources defined within the same service group. A service group effectively acts as a logical partition of services and resources.

[0142] The services used in this paper are the core compute constructs in the deployment model provided by SOADM service 102. A service represents a set of software dependencies (e.g., containers) encapsulated together and deployed to various locations along with supporting infrastructure (e.g., load balancers, block storage, etc.). Each instance of a service can be generated from a template called a service configuration. Service configurations enable users to define the characteristics of the underlying infrastructure of a service instance, including where they want to deploy it (via a distribution strategy) and how they want to deploy it (via a deployment configuration). When a service configuration is updated, in some implementations, SOADM service 102 generates a version identifier, which can then be referenced in the deployment. In some implementations, a service template is also encapsulated within the service configuration, defining a location-independent resource configuration that can be used to define a set of compute resources that constitute the application.

[0143] Users can also create one or more service resource configurations corresponding to service resources, which define the resources (e.g., resources from a cloud provider's network) that support the service and upon which the service depends for proper operation. In some cases, service resources are managed by SOADM service 102, but in others, they are manually managed by the user. Fully managed resources can be specified using a "startup template" (such as a CloudFormation template provided by the AWS Cloud Formation™ service) or initialized based on existing resources in the user's account. Either way, the end result is resources distributed by SOADM service 102 to best meet the needs of services that depend on them. In some implementations, self-managed resources may be defined by the user with reference to existing resources in the user's account, and in some implementations, they cannot be copied or modified by SOADM service 102, but in other implementations, the user indicates that resources can be copied, modified, etc.

[0144] Users can also create distribution policies 728 for a given service, which serve as a construct influencing the location where SOADM service 102 deploys the given service. Users can optionally provide a global minimum and / or maximum compute capacity (e.g., the minimum or maximum number of containers or VMs used for the service), and / or a list of distribution groups to target along with their weights, to inform SOADM service 102 how to allocate global capacity among these distribution groups.

[0145] A distribution group is a logical group of locations that share a user application capacity pool. A distribution group can consist of a single region within a cloud provider's network, including any local region and / or wavelength zone to which it belongs (e.g., those that at least partially depend on control plane components within that region), and / or a larger geographical unit consisting of a defined or derived group of regions—for example, a "United States" distribution group could include all regions within the United States. A key characteristic of distribution groups is that SOADM service 102 can consider all regions within a distribution group to be equivalent and therefore can distribute compute and other resources as needed to any location within the distribution group to maintain the capacity pool at the desired level. One benefit provided by logical distribution groups (such as "United States" or "Eastern United States") is that they can automatically incorporate new regions (logically belonging to these groups) upon creation, so user applications do not need to be reconfigured to use new regions or other deployment zone types or locations associated with the distribution group.

[0146] Users can also create a deployment configuration 730, which defines how they want updates deployed to their applications. As part of this configuration, users can specify attributes applicable to the global deployment and optionally provide configuration data for deployment activities at a specific level (e.g., network boundary group level). Therefore, deployment configuration 730 can specify the deployment type that SOADM service 102 will use, the rate at which deployment changes will be made, how deployments will be validated, etc. In some implementations, the deployment configuration has two main components: a set of configuration data items that manage the overall global deployment from start to finish, and a deployment unit configuration that specifies how the underlying deployment system will update and validate applications in each network boundary group.

[0147] With complete group configuration data 720 for the application, the user can instruct SOADM service 102 to run (or deploy) the application by sending a message carrying (or otherwise indicating) a command to run the application. In response, at circle (2), deployment engine 132 obtains configuration data 720 and optional capacity information (identifying different deployment zones, the available capacity therein (e.g., available container or VM “slots” for use in each deployment zone), performance and / or availability and / or network information about these deployment zones, etc.) from capacity service 136, and determines a set of initial locations to deploy the services associated with service group 722 to locations consistent with the information provided in configuration data 720, namely the number of service instances of the services in the service group and any required service resources deployed to specific permitted deployment zones 718 in a manner consistent with this data.

[0148] For example, SOADM service 102 can obtain configuration data 720 and determine which type of deployment zone (e.g., using only AZ, or possibly using AZ and local area, as well as outpost and wavelength areas) can be used for the service identifier, and which deployment groups or regions can be used (e.g., using only the "Western United States" region, or "any region within the U.S. deployment group"). Additionally, other user configuration preferences can be obtained, such as optional user-defined weights associated with specific deployment locations (e.g., deployment groups). Through a set of candidate locations, SOADM service 102 can identify a complete set of specific candidate deployment locations (e.g., AZ#1, AZ#2, AZ#3, local area #50, local area #55, wavelength area #1, wavelength area #2, wavelength area #3). Furthermore, SOADM service 102 can obtain capacity information indicating the available resources in each of these specific candidate deployment locations, such as the number and / or type of "time slots" available for service instances at these locations, and location / latency information indicating where these locations are situated from a geographical (e.g., within Los Angeles) or network (e.g., located in or connected to a specific cellular provider's network).

[0149] Based at least in part on this information, SOADM service 102 can determine where the resources needed to initially deploy the service group can be determined. This placement can be localized at the outset (e.g., deployed only to locations within a first region or associated with a first region), then scaled on demand based on client traffic / latency, distributed at the outset (e.g., by placing resources in a large number of locations, such as all locations or randomly sampled locations), and again scaled up or down based on traffic, or selected based on historical usage information (specific to the application, other applications, or the entire provider network) that indicates where the expected usage is highest (e.g., on a particular day and / or at a certain time of day).

[0150] Subsequently, deployment engine 132 can send a set of commands to deploy applications to some or all of these locations. This may include invoking other services of the cloud provider network 100, such as deployment services, compute services, etc. (not shown), to deploy compute resources (e.g., containers, VMs, serverless functions, code, etc.) to the necessary locations. Deployment may also include, for example, placing service resources in locations that are the same as or "near" (from a network latency perspective) as dependent services, configuring routing and network information, configuring security information, etc.

[0151] exist Figure 1In this example, the deployment causes a group of service instances 752A-752N of the example "first" service of the service group (represented by black squares) to reach various deployment zone locations 718—here, two instances reach CSP edge location 742A as shown in circle (3A), one instance reaches cloud provider network-managed edge location 740A as shown in circle (3B), and one instance reaches each of the first AZ 714A and the second AZ 714B of the first region 112A as shown in circle (3C). Given this deployment, the service is configured with distribution policy 728, indicating that these instances of the service are eligible to be deployed in multiple different "types" of deployment zones—namely AZ 714, cloud provider network-managed edge location 740 ("local zone"), and CSP edge location 742. We also specify that the service resource (represented by the black triangle) is configured to be required for this first service, and it is deployed within the first AZ 714A and the second AZ 714B of the first region 112A (e.g., the service resource may or may not be placed at the edge location 116), as service resource instance 750. Furthermore, in this example, the "second" service of the service group (represented by the black circle) is configured to be eligible for deployment in a single type of deployment zone—here only AZ 714—therefore, one instance of this service is deployed to the first AZ 714A and the second AZ 714B. For example, this might be a "backend" service (such as a matchmaking function for a game or a log database), which is relatively less sensitive to end-user latency and may therefore be limited to placement in an AZ, which typically has higher availability and a higher quantity and type of resources. Other underlying architecture configurations (e.g., routing, security, etc.) are also performed at this time until the application is ready for use.

[0152] At this point, the client (e.g., electronic device 110, which may or may not be operated by user 108) can use known endpoint lookup techniques to access the application, such as by device 110 calling an API to find "nearby" endpoints for the application and obtaining the Internet Protocol (IP) network address used by at least one deployed service instance. Subsequently, device 110 can use these network addresses to communicate with the application, and the application begins to run.

[0153] As time progresses and application components (e.g., service instances, service resource instances, etc.) are used, cloud provider network 100 generates and / or obtains metrics and / or logs, as shown in circle (5), detailing and / or summarizing this usage. Metrics and / or logs may be collected (or stored in) monitoring service 138 and then made available to SOADM service 102, provided directly to SOADM service 102, stored in a storage location (e.g., object storage location, database, etc.), and then accessed by SOADM service 102, etc. Deployment monitor 134 of SOADM service 102 can use this information to determine whether and when to modify the application deployment based on user-configured configuration data 720 (e.g., distribution policy 728 and / or deployment configuration 730).

[0154] For example, client information associated with the application's clients (e.g., source IP network address, geographic coordinates, source network identifier, number of requests, etc.) can be obtained and analyzed to identify the location of the application's users and the extent to which the geographic distribution of usage has changed. For example, at a first point in time, based on recent metrics / logs, deployment monitor 134 can determine that a threshold number of clients (e.g., greater than 90%) exist within a defined geographic region (e.g., the eastern half of the United States), and therefore the application should be deployed at its maximum (or full) scale within that geographic region. At a second point in time, deployment monitor 134 can obtain updated metrics and / or log information indicating that usage has become more dispersed over a recent period—e.g., 40% from the western United States, 40% from the eastern United States, and 20% from Europe. Using this information, along with user-configured configuration data 720, deployment monitor 134 can determine that the current deployment (e.g., only in the eastern United States) is insufficient, and a more optimized deployment would include fewer resources in the eastern United States, more resources in the western United States, and more resources in Europe. For example, similar to the initial placement process described earlier in this document, SOADM service 102 can generate an “optimal” placement for the services of a service group and determine whether this placement is different from the current placement (or substantially different from the current placement based on certain thresholds).

[0155] When the optimal placement differs (or the difference is large enough that redistribution is worthwhile because the benefits outweigh the cost of doing so), at circle (7), deployment monitor 134 can cause deployment engine 132 to redistribute the application accordingly, for example, by adding additional service instances (and possibly service resource instances) to one or more new deployment zones, and possibly terminating existing service instances (and possibly service resource instances) in existing deployment zones. In this example, deployment engine 132 could deploy additional service instances of the application's first service to the cloud provider's network-managed edge location 740N, and deploy service resource instances (which the first service depends on) to AZ 714M of the associated region 112N (the "parent" of the edge location 740N), as shown in circle (8). As mentioned above, such redistribution may also include shutting down (or terminating, deleting, etc.) the application's resources, such as when new resources (e.g., new service instances) are to be added to a new deployment zone, but doing so would cause the total number of resources (e.g., instances) to exceed the user-configured maximum value—therefore, when deploying additional resources, some corresponding existing resources can be removed to prevent exceeding the maximum value.

[0156] For further details, Figure 12 This illustrates a dynamic redistribution of services and service resources by an SOADM service based on changing end-user locations and latency, according to some implementation schemes. As shown in 1200, at a first time (time = T1), a large number of clients (e.g., via metrics / logs described herein) are detected to be located in various locations across the eastern half of the United States. In this scenario, multiple instances of the application's "square" service can be deployed to multiple different deployment zones grouped "near" these users. For example, a service instance 1802 in the northeast is "near" (in terms of latency) a large number of users, another service instance 1206 in the southeast is near another large number of users, and service instance 1804 in between is near other users. Furthermore, "triangle" service resource instances upon which the square service depends can be deployed geographically close to these instances (also in terms of latency / routing).

[0157] At a later point in time (time = T2) as shown in 1210, more users / clients may begin to appear in other locations further away from the deployment resources. At this time, these clients can connect to existing resources (e.g., instances 1802 / 1804 / 1206), but their latency may be much higher than that observed by clients in the east.

[0158] At certain times, such as when a threshold for determining the number of clients (e.g., based on metrics / logs) exists in different geographic regions, and / or when the latency of some users exceeds the threshold, deployment monitor 134 can determine to redistribute the application. For example, at time = T3 (as shown in 1220), SOADM service 102 may have redistributed the application by placing a new service instance 1222 (and another service resource instance) in the northwest and a new service instance 1224 in the southwest. In this case, service instance 1206 has been removed, possibly based on the maximum limit of service instances (e.g., only four “square” instances may exist as specified by the user) or based on SOADM service 102 determining that the number of users in the southeast is less than the threshold, which would require a more “local” deployment.

[0159] To further understand, Figure 13 A more specific example of a multi-service application is shown, illustrating an exemplary multi-service application with a heterogeneous distribution strategy deployed and distributed by SOADM services according to some implementation schemes.

[0160] As a more concrete example, consider a high-level architecture for a sample distributed gaming application for online multiplayer games. The core use cases of this application are (1) authenticating players, (2) matchmaking players participating in the same game session, (3) selecting a game server that provides equal latency to all players in the game session, (4) running the game on the selected game server (which involves latency-sensitive operations such as simulating the interaction between players and the game environment and distributing these updates to the physics engine of each player in the game session), (5) enabling players in the game session to talk to each other via voice, (6) tracking statistics specific to various game sessions, and (7) recording player statistics for long-term storage and analysis.

[0161] For this example, these use cases can be categorized into latency-sensitive and latency-insensitive use cases. We assume that use cases (4)-(6) are latency-sensitive, while (1)-(3) and (7) are latency-insensitive. From a system design perspective, this typically means that the former set of use cases needs to run closer to the end client, while the latter set can run in a more "centralized" location. Therefore, these functional categories can be viewed as "edge services" and "centralized services," respectively. The centralized service can be further decomposed into a front-end service that accepts requests from the end client and a back-end service responsible for coordinating with the edge service.

[0162] Therefore, this user might want to use deployment zone types, including AZ, local zone, and wavelength zone, to target the United States and Europe. In this example, the user may have already set up “central” zones (e.g., “us-east-1” in the US and “eu-west-1” in Europe) for their centralized service and coordinate their local edge services from there. The system can also use “global tables” (e.g., synchronized distributed database tables) as a means of sharing data across zones. In this example, the terminal client (e.g., located in Florida, USA) connects to the user’s service via an accelerator endpoint that routes to the nearest front-end service (e.g., centralized front-end service 1302A) and then connects to an edge service endpoint at an edge location in Miami, Florida (e.g., a wavelength zone or local zone).

[0163] More generally, in this diagram, a user may have deployed three services, one service resource, and an application across multiple deployment zones located in various regions and multiple edge locations within the cloud provider network 100. For example, an application (e.g., a chat application, a multiplayer video game, etc.) may have a primary service that handles the closest possible real-time interaction with the client—this service (referred to herein as “edge services” 1306A-1306N) may benefit from deployment in two edge locations 116A-116N (of the same or multiple types) and in various regions 112A-112S. This application may also require other services, such as a logically centralized front-end service 1302 (deployed as services 1302A-1302B in two regions 112A and 112P) and a logically centralized back-end service (also deployed as services 1304A-1304B in two regions 112A and 112P). In this example, both the centralized front-end service 1302 and the back-end service 1304 rely on service resources—here, a global data table 1308 (also deployed as global table instances 1308A-1308B in two regions 112A and 112P) that provides data storage and retrieval capabilities. With this setup, a client (e.g., a software application executed by electronic device 110 and optionally used by user 108) can seek to interact with the application by making a call at circle (1A) to a global front-end endpoint 1310 (e.g., a routing accelerator entity, such as AWS GlobalAccelerator™), which can route the client at circle (1B) to an instance of the centralized front-end service 1302, which can return to the client the network address associated with a specific edge service instance “near” the client (e.g., in terms of network latency, network hop count, geographical distance, etc.). Here, electronic device 110 will thus connect to a nearby edge service instance 1306G, as indicated by circle (2).

[0164] For software applications and systems developed using a service-oriented application architecture, several challenges exist in achieving efficient and user-friendly service discovery and application-layer communication routing between application services. In this context, service discovery and application-layer routing broadly refer to the ability of various services within a service-oriented application to locate each other (and related service resources) on a network and establish communication as needed. For example, an application might include a first front-end service A that communicates with a second back-end service B during operation, and service B might further depend on and communicate with services C and D, and so on, where each service can be deployed in any number of different deployment zones. As described herein, each of these services can also have one or more service resource dependencies, where these resources, likewise, can be created at any given time and exist in any number of different deployment zones, depending on a user-specified deployment configuration.

[0165] Complicating matters further, in many cases, the various services of an application may be associated with separate development teams, which might fragment services across different accounts on a cloud provider's network of 100, or create additional isolation boundaries around services during development. Therefore, creating a network encompassing a range of application services so that these services can easily discover and communicate with each other is a challenging task. One approach to facilitate service discovery and communication for such applications is to provide each service with a public IP address, which services can use to connect to each other. However, for security and other implementation reasons, application developers often prefer that at least some services not be accessible to client devices on the public internet. Furthermore, developers can easily implement their applications so that services can be referenced by name rather than IP address (e.g., using domain names that can be resolved by a Domain Name System (DNS) resolver, or other types of identifiers that can be resolved using a service registry or other mechanisms). Application developers also want to be able to establish this inter-service communication without needing to modify their code at different stages of the development pipeline (e.g., using the same service code whether on a developer's desktop, in gamma testing, or in production). In addition, it is generally desirable for each service to be able to discover dependent services or resources deployed in the deployment zone "closest" to the requested service (e.g., in terms of network latency or hop count).

[0166] Among other things, the challenges described above are addressed by service discovery and application-level networking features implemented within the context of SOADM service 102. In some implementations, these features provide latency-centric routing in part by using data reflecting latency estimates between deployment zones of a cloud provider's network. For example, if service A, operating in a deployment zone near Paris, France, is looking for the “closest” instance of a backend service B deployed in a first deployment zone near Dublin, Ireland, and a second deployment zone near Northern Virginia, the service can use deployment zone-to-deployment zone latency data to identify instances of the backend service in the first deployment zone associated with the lowest estimated latency, and thus can correspondingly route traffic from service A to service B. Furthermore, SOADM service 102 can provide such latency-based service discovery and application-level networking features based on DNS-based domain names assigned to application services, API-based service and resource discovery requests implemented using service registries or similar mechanisms, or any combination thereof, enabling users to easily integrate such features into service-oriented applications managed by SOADM service 102.

[0167] Figure 14 A DNS-based service discovery and latency-based routing system provided by SOADM service 102 according to some implementation schemes is illustrated. For example, numbered circles (1)-(7) illustrate an exemplary process in which a user provides SOADM service 102 with configuration data defining an application and relevant characteristics of the application's required deployment behavior. Once deployed by SOADM service 102, the illustrated process also includes a request sent by a first service of the application (e.g., front-end service 1400) to identify a second service of the application (e.g., a back-end service, where multiple instances of the back-end service, including back-end service instance 1402A and back-end service instance 1402B, have been deployed) as part of the service operation. In response to such a request, router 1404 identifies an instance of the second service associated with the lowest network latency estimate relative to the first service, wherein the second service instance may be instantiated at any number of different deployment areas associated with different latency estimates relative to the first service. Then, router 1404 will request routing from the first service to the identified instance of the second service 1402A, in which routing is requested without traversing the public Internet (e.g., using the network within cloud provider network 100).

[0168] As indicated herein, in some implementations, SOADM service 102 provides a global deployment experience, enabling users to specify certain aspects of application deployment behavior, such as deployment cadence, validation steps, automatic rollback configurations, etc., and then manage deployment locations and the underlying sequencing of low-level deployment activities. SOADM service 102 also manages the deployment and configuration of the infrastructure used to support application deployment, such as virtual private networks, load balancers, server endpoints, routers, etc., as well as other management operations.

[0169] In some implementation schemes, Figure 14 At circle (1) in the diagram, user 104 (e.g., a software application developer) interacts with SOADM service 102 using electronic device 106 to provide configuration data for an application that the user wishes to deploy within the cloud provider network 100 and optionally at associated edge locations (e.g., including edge locations 116A-116N). In some implementations, the configuration data defines a service group for the application, which includes one or more service configurations, zero or more service resource configurations, one or more distribution policies, one or more deployment configurations, and possibly other types of configuration data.

[0170] In some implementations, as part of configuring an application with SOADM service 102, a user can specify one or more of the application's services and resources as "private" services or resources (e.g., specified as a "visibility" flag associated with the service or resource). A private service or resource is one that is accessible only to other services and resources within the same application or service group (or more generally to other services or applications that can access the service or resource using a private identifier). To enable application services to easily reference and access other services designated as private, SOADM service 102 can associate some or all of the application's services and resources with user-friendly domain names or other identifiers that can be used as part of the service and resource discovery and routing processes described herein. For example, an application's backend service might be identified by the domain name "backend-service.my-servicegroup.soadm.example.com" or any other similar type of identifier. These service and resource identifiers can be automatically generated by SOADM service 102 or optionally configured by user 104, for example, as part of a service group configuration. Similarly, application resources can also be assigned user-friendly identifiers that are represented using domain name formats or other types of Uniform Resource Identifiers (URIs). For example, these user-friendly identifiers can be used consistently across application deployments within the application's code without requiring hard-coded addresses for various services and resources.

[0171] Once configured, users can instruct SOADM service 102 to run (or deploy) applications by sending requests or commands. In response, Figure 14 At circle (2) in the document, the deployment engine or other components of SOADM service 102 obtain configuration data and optionally capacity information from a capacity service that identifies different deployment zones, the available capacity therein (e.g., container or VM “slots” available for use in each deployment zone), performance and / or availability and / or network information about these deployment zones, and determines a set of initial locations to deploy services and resources associated with the service group, the set of initial locations following the information provided in the configuration data. SOADM service 102 then sends commands to deploy applications to the selected locations, as described elsewhere herein.

[0172] exist Figure 14 In this example, the deployment at circle (2) allows a group of service instances (represented by black squares) of the example "first" service of the service group to reach various deployment zones, edge locations managed by cloud providers' networks, etc. In the example shown, the backend service is deployed to at least each of zones 112A and 112N (where various other services and resources of the application may also be deployed to the same or different zones). Similarly, the frontend service 1400 is deployed to one or more edge locations 116A-116N. Other underlying architecture configurations can also be performed at this time (e.g., creation and configuration of virtual private network 1406, DNS resolver 1408, server endpoint 1410, router 1404, etc.). Once deployed, the application can be accessed by clients.

[0173] In some implementations, as part of the architecture configuration, SOADM service 102 also configures DNS records at circle (3) within the hosted virtual private network 1406 for use by DNS resolver 1408 to map domain names or other identifiers assigned to various services or resources to canonical domain names or other identifiers (e.g., globally unique identifiers for services or resources within cloud provider network 100) used by service endpoint 1410 and router 1404, thereby routing requests accordingly. The configuration at circle (3) may optionally also include configuration of zone-to-zone latency data 1412 and application deployment data 1414 (e.g., data indicating the deployment zones of various services and resources for which applications have been deployed) at router 1404. For example, this latency data 1412 may be generated by SOADM service 102 based on network latency measurements observed between clients in the various deployment zones or otherwise obtained. Latency data 1412 may be updated periodically, for example, based on a recurring schedule, in response to the addition or removal of deployment zones, or based on other conditions. Furthermore, in some implementations, when the application deployment changes over time, such as when services and resources are deployed to a new deployment zone, removed from an existing deployment zone, or a combination thereof, SOADM service 102 can update application deployment data 1414 at router 1404.

[0174] like Figure 14 As shown, VPC 1406 has one or more individual service endpoints 1410 for routing application service traffic, where each service endpoint can be used to route traffic for one or more other services of the application. As described above, service groups can be associated with DNS type identifiers, such as service-group-name.soadm.example.com, where the service group name is globally unique to the associated user. Furthermore, a corresponding subdomain of the service group domain can be assigned to each service of the application (e.g., service-name.service-group.soadm.example.com). These DNS names created for each service can then be resolved by DNS resolver 1408 to the service endpoints in the virtual private network, where the service endpoints forward traffic to router 1404. Router 1404 then uses zone-to-zone latency data 1412 and application deployment data 1414 to identify an instance of the requested service located in the deployment zone, the instance having the lowest estimated latency relative to the requested service.

[0175] For example, at circle (4), at some point in time, an instance of frontend service 1400 sends a request to identify a backend service, which could be a request to establish a data connection, an API request, etc. In some implementations, the request includes a domain name pointing to DNS resolver 1408 (e.g., “backend-service.my-servicegroup.soadm.example.com”). In this example, DNS resolver 1408 includes records that map an application-specific domain name to another domain name (e.g., a canonical domain name) to direct the request to a dedicated service endpoint (e.g., service endpoint 1410) and router 1404 at circle (5) for latency-based routing within the cloud provider network 100. In some implementations, the canonical domain name, the original service identifier domain name, or both are included in the header of the request to enable router 1404 to identify the specific service or resource requested.

[0176] In some implementations, at circle (6), router 1404 uses deployment area-to-deployment area latency data 1412 to identify an instance of a backend service associated with the lowest network latency estimate relative to frontend service 1400. In some implementations, router 1404 uses application deployment data 1414, indicating the deployment area where the backend service is currently deployed, and deployment area-to-deployment area latency data 1412 to identify a service instance with the lowest latency estimate. For example, router 1404 may identify the deployment area associated with the requested service based on data contained in the request, information about the service endpoint from which it forwards the request, information about the location of router 1404, or any combination thereof, and use pairwise latency estimates to identify the deployment area associated with the lowest latency estimate. Figure 14 In the example, backend service instance 1402A is identified as having the lowest latency estimate.

[0177] In some implementations, at circle (7), once identified, router 1404 routes the request from frontend service 1400 to an instance of backend service 1402A located in a deployment zone (e.g., area 112A), the deployment zone being associated with the lowest latency relative to the deployment zone where the requested service resides (e.g., one of edge locations 116A-116N). As described above, in some implementations, the request is routed through the network within the cloud provider network 100 without traversing the public internet, thereby improving communication latency and security.

[0178] Figure 15 This illustrates a service discovery and latency-based routing system based on an application programming interface (API) provided by a service-oriented application deployment management service, according to some implementation schemes. Figure 15Circles (1)-(6) in the diagram illustrate an exemplary process in which a first service 1500 deployed in a first deployment zone uses a discovery request 1504 sent to SOADM service 102 to discover instances of computing resources deployed in a second deployment zone. As described in more detail, SOADM service 102 maintains a service and resource registry 1506 and zone-to-zone latency data 1508 to identify instances of services or resources associated with the lowest latency relative to the requested service. Although Figure 15 The examples shown include instances of service discovery and access to resource dependencies, but a similar process involving discovery requests can be used to discover instances of another service.

[0179] and Figure 14 Similarly, in Figure 15 At circle (1), the user provides user-specified configuration data defining the application service group, where the configuration data includes service definitions corresponding to the services of the application. As described elsewhere herein, each of these service definitions may be associated with, for example, a distribution policy definition indicating a set of deployment zone types where the associated services can be deployed, as well as other configurations. At circle (2), SOADM service 102 deploys the application to multiple deployment zone locations within the cloud provider network 100 based on the configuration data. Figure 15 In the example, front-end service 1500 is deployed to one or more edge locations 116A-116N, and instances of compute resources (e.g., compute resources 1502A and compute resources 1502B) are deployed to each of regions 112A and 112N.

[0180] In some implementations, at circle (3), the front-end service 1500, intending to access the deployed compute resource instance, sends a discovery request 1504 to the SOADM service 102. For example, the discovery request 1504 may be specified as an API request supported by the SOADM service 102 (e.g., a “discoverResource” API or a “discoverService” API) and may include an identifier of the requested service or resource (e.g., using a domain name, resource identifier, or other identifier). In some implementations, the identifier of the service or resource may include an identifier of the service group to which the service or resource belongs. As indicated above, in some examples, these service and resource identifiers may be generated by the SOADM service 102, the provider network 100, or customized by the user. In some implementations, the request may also include information indicating the deployment zone where the requested service 1500 resides; for example, deployment zone identification information may be explicitly specified as part of the request or otherwise derived from information included in the request.

[0181] In some implementations, in response to discovery request 1504, at circle (4), SOADM service 102 uses service and resource registry 1506 and latency data 1508 to identify an instance of the requested service or resource associated with the lowest latency estimate relative to the requesting frontend service 1500. For example, the service registry is a database of application services and resources, as well as information about the deployment zone where the service and resource are deployed at any given time. For example, when a service or resource is started, SOADM service 102 may register the service and resource instance with service and resource registry 1506, and may unregister it when the service or resource is terminated. The application service can then query service and resource registry 1506 to find available instances of the service. In some implementations, service and resource registry 1506 invokes the service instance's health check API to verify its ability to handle requests.

[0182] As described above, SOADM service 102 identifies instances of services or resources in part by using latency data 1508 that indicates a latency estimate from deployment zone to deployment zone. Therefore, given the identifier of the deployment zone of the requested service, the identifier of the requested service or resource, and the deployment zone identifier of the currently deployed requested service and resource, service and resource registry 1506 can determine the deployment zone of the currently deployed requested service or resource relative to the requested service that has the lowest latency estimate. In some embodiments, SOADM service 102 returns the identifier of the identified service or resource, which can be used to route a request from frontend service 1500 to the appropriate instance of the requested service or asset. For example, the identifier may include a private identifier that is not discoverable on the Internet (e.g., a service or resource network address or other identifier generated by provider network 100) and is used to route traffic within provider network 100.

[0183] In some implementations, using the domain name or other service or resource identifier returned by SOADM service 102, at circle (5), frontend service 1108 sends a request to router 1510 to communicate with or otherwise access a service or resource. For example, the request could be a request to open a data connection, or an API request to access a service or resource, and includes the service or resource identifier returned by SOADM service 102. In some implementations, at circle "6", router 1502B then directs the request to the identified instance of the service or resource with the lowest latency relative to the requesting frontend service 1500.

[0184] As described in this article, users can define applications by providing configuration data to define single-service or multi-service service groups 722 of the application, thereby enabling SOADM service 102 to intelligently deploy and dynamically redistribute application components according to user preferences—potentially across many different deployment zones, which may have different deployment zone types.

[0185] Figure 16 An example provider network (or “service provider system”) environment is illustrated according to some examples. Provider network 1600 may provide resource virtualization to customers via one or more virtualization services 1610, which allow customers to purchase, lease, or otherwise obtain instances 1612 of virtualized resources (including, but not limited to, computing and storage resources) implemented on devices within one or more provider networks in one or more data centers. A local Internet Protocol (IP) address 1616 may be associated with resource instance 1612; the local IP address is the internal network address of resource instance 1612 on provider network 1600. In some examples, provider network 1600 may also provide public IP addresses 1614 and / or ranges of public IP addresses (e.g., Internet Protocol version 4 (IPv4) or Internet Protocol version 6 (IPv6) addresses) available to customers from provider 1600.

[0186] Typically, provider network 1600 may, via virtualization service 1610, allow service provider customers (e.g., customers operating one or more customer networks 1650A-1650C (or “client networks”) including one or more client devices 1652) to dynamically associate at least some public IP addresses 1614 assigned to the customer with specific resource instances 1612 assigned to the customer. Provider network 1600 may also allow customers to remap public IP addresses 1614 previously mapped to one virtualized computing resource instance 1612 assigned to the customer to another virtualized computing resource instance 1612 also assigned to the customer. Using the virtualized computing resource instances 1612 and public IP addresses 1614 provided by the service provider, service provider customers, such as operators of customer networks 1650A-1650C, may, for example, implement customer-specific applications and present these applications on an intermediate network 1640 (such as the Internet). Then, other network entities 1620 on the intermediate network 1640 can generate traffic destined for a public IP address 1614 published by customer networks 1650A-1650C; the traffic is routed to the service provider's data center, and at the data center, it is routed via the network substrate to the local IP address 1616 of the virtualized computing resource instance 1612 currently mapped to the destination public IP address 1614. Similarly, response traffic from the virtualized computing resource instance 1612 can be routed back to the intermediate network 1640 via the network substrate to reach the source entity 1620.

[0187] As used herein, a local IP address refers to an internal or “private” network address of a resource instance within, for example, a provider network. Local IP addresses may be within an address block reserved by the Internet Engineering Task Force (IETF) Request for Comments (RFC) 1918 and / or have an address format specified by IETF RFC 4193, and may vary within the provider network. Network traffic originating outside the provider network is not directly routed to a local IP address; instead, traffic uses a public IP address mapped to the local IP address of the resource instance. Provider networks may include networking devices or apparatuses that provide Network Address Translation (NAT) or similar functionality to perform mappings from public IP addresses to local IP addresses and from local IP addresses to public IP addresses.

[0188] A public IP address is a variable network address on the Internet assigned to a resource instance by a service provider or customer. For example, traffic routed to a public IP address via 1:1 NAT is used to translate traffic and then forward the traffic to the appropriate local IP address of the resource instance.

[0189] Some public IP addresses may be assigned to specific resource instances by the provider's network infrastructure; these public IP addresses may be referred to as standard public IP addresses, or simply standard IP addresses. In some examples, the mapping of standard IP addresses to the local IP addresses of resource instances is the default startup configuration for all resource instance types.

[0190] At least some public IP addresses can be assigned to or obtained by customers of Provider Network 1600; customers can then assign their assigned public IP addresses to specific resource instances assigned to them. These public IP addresses may be referred to as customer public IP addresses, or simply customer IP addresses. Instead of being assigned to resource instances by Provider Network 1600 as in the case of standard IP addresses, customer IP addresses can be assigned to resource instances by the customer, for example, via an API provided by the service provider. Unlike standard IP addresses, customer IP addresses are assigned to customer accounts and can be remapped to other resource instances by the corresponding customer as needed or desired. Customer IP addresses are associated with customer accounts, not specific resource instances, and the customer controls the IP address until the customer chooses to release it. Unlike regular static IP addresses, customer IP addresses allow customers to mask resource instance or availability zone failures by remapping their public IP addresses to any resource instance associated with their customer account. For example, customer IP addresses enable customers to resolve resource instance or software issues by remapping their customer IP addresses to alternative resource instances.

[0191] Figure 17 This is a block diagram of an example provider network environment that provides storage services and hardware virtualization services to customers, based on some examples. Hardware virtualization service 1720 provides customers with multiple computing resources 1724 (e.g., computing instances 1725 such as VMs). Computing resources 1724 may be provided as a service to customers of provider network 1700 (e.g., customers implementing customer network 1750). Each computing resource 1724 may be configured with one or more local IP addresses. Provider network 1700 may be configured to route packets from the local IP addresses of computing resources 1724 to public internet destinations, and to route packets from public internet sources to the local IP addresses of computing resources 1724.

[0192] Provider network 1700 may provide a client network 1750, for example, coupled to intermediate network 1740 via local network 1756, with the ability to implement virtual computing systems 1792 via hardware virtualization service 1720 coupled to intermediate network 1740 and provider network 1700. In some examples, hardware virtualization service 1720 may provide one or more APIs 1702, such as web service interfaces, via which client network 1750 may access the functionality provided by hardware virtualization service 1720, for example, via console 1794 of client device 1790 (e.g., web-based applications, standalone applications, mobile applications, etc.). In some examples, each virtual computing system 1792 at provider network 1700 and client network 1750 may correspond to computing resources 1724 that are leased, rented, or otherwise provided to client network 1750.

[0193] Clients can access the functionality of the storage service 1710 from an instance of the virtual computing system 1792 and / or another client device 1790 (e.g., via console 1794) via one or more APIs 1702, for example, to access and store data from storage resources 1718A-1718N of virtual data storage areas 1716 (e.g., folders or "buckets", virtualized volumes, databases, etc.) provided by the provider network 1700. In some examples, a virtualized data storage gateway (not shown) may be located at the client network 1750. This virtualized data storage gateway may locally cache at least some data (e.g., frequently accessed data or critical data) and may communicate with the storage service 1710 via one or more communication channels to upload new or modified data from the local cache, thereby maintaining the master data repository (virtualized data repository 1716). In some examples, a user can install and access volumes of a virtualized data repository 1716 via a storage service 1710 that acts as a storage virtualization service, through a virtual computing system 1792 and / or another client device 1790, and these volumes appear to the user as local (virtualized) storage 1798.

[0194] Although Figure 17 Not shown, but the virtualization service can also be accessed from resource instances within provider network 1700 via API 1702. For example, a customer, equipment service provider, or other entity can access the virtualization service from within a corresponding virtual network on provider network 1700 via API 1702 to request the allocation of one or more resource instances within said virtual network or another virtual network.

[0195] Explanatory System

[0196] In some examples, systems implementing some or all of the techniques described herein may include general-purpose computer systems (such as...) Figure 18 The illustrated computer system 1800 (also referred to as an electronic device or computing device) includes or is configured to access one or more computer-accessible media. In the illustrated example, computer system 1800 includes one or more processors 1810 coupled to system memory 1820 via input / output (I / O) interface 1830. Computer system 1800 also includes a network interface 1840 coupled to I / O interface 1830. Although Figure 18 Computer system 1800 is shown as a single computing device, but in various examples, computer system 1800 may include a single computing device or any number of computing devices configured to work together as a single computer system 1800.

[0197] In various examples, computer system 1800 may be a single-processor system including one processor 1810 or a multiprocessor system including several processors 1810 (e.g., two, four, eight, or another suitable number). Processor 1810 may be any suitable processor capable of executing instructions. For example, in various examples, processor 1810 may be a general-purpose or embedded processor implementing any of a variety of instruction set architectures (ISAs), such as x86, ARM, PowerPC, SPARC, or MIPS ISA or any other suitable ISA. In a multiprocessor system, each of the processors 1810 may often, but not necessarily, implement the same ISA.

[0198] System memory 1820 may store instructions and data accessible by processor 1810. In various examples, system memory 1820 may be implemented using any suitable memory technology, such as random access memory (RAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash memory, or any other type of memory. In the illustrated example, program instructions and data that implement one or more desired functions (such as the methods, techniques, and data described above) are shown stored in system memory 1820 as SOADM service code 1825 (e.g., executable to fully or partially implement SOADM service 102) and data 1826.

[0199] In some examples, I / O interface 1830 may be configured to coordinate I / O traffic between processor 1810, system memory 1820, and any peripheral devices within the device, including network interface 1840 and / or other peripheral interfaces (not shown). In some examples, I / O interface 1830 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., system memory 1820) into a format suitable for use by another component (e.g., processor 1810). In some examples, for instance, I / O interface 1830 may include devices that support attachment via various types of peripheral buses, such as the Peripheral Component Interconnect (PCI) bus standard or variants of the Universal Serial Bus (USB) standard. In some examples, for instance, the functionality of I / O interface 1830 may be split into two or more separate components, such as a northbridge and a southbridge. Additionally, in some examples, some or all of the functionality of I / O interface 1830 (such as an interface to system memory 1820) may be directly incorporated into processor 1810.

[0200] For example, network interface 1840 may be configured to allow computer system 1800 to communicate with other devices 1860 (such as, e.g., connected to one or more networks 1850) attached to it. Figure 1 The network interface 1840 can exchange data with other computer systems or devices shown. In various examples, for instance, the network interface 1840 may support communication via any suitable wired or wireless general-purpose data network, such as various types of Ethernet networks. Additionally, the network interface 1840 may support communication via telecommunications / telephone networks (such as analog voice networks or digital fiber optic communication networks), via storage area networks (SANs) (such as Fibre Channel SANs), and / or via any other suitable type of network and / or protocol.

[0201] In some examples, computer system 1800 includes one or more offload cards 1870A or 1870B (including one or more processors 1875 and possibly one or more network interfaces 1840), which are connected using I / O interfaces 1830 (e.g., a version of the Peripheral Component Interconnect Fast (PCI-E) standard or a bus for another interconnect such as Quick Path Interconnect (QPI) or Hyper Path Interconnect (UPI). For example, in some examples, computer system 1800 may act as a host electronic device hosting computing resources such as compute instances (e.g., operating as part of a hardware virtualization service), and one or more offload cards 1870A or 1870B act as a virtualization manager that manages the compute instances running on the host electronic device. As an example, in some examples, offload card 1870A or 1870B may perform compute instance management operations such as pausing and / or unpausing compute instances, starting and / or terminating compute instances, performing memory transfer / copy operations, etc. In some examples, these management operations may be performed by the offload card 1870A or 1870B in cooperation with a hypervisor (e.g., upon request from the hypervisor) executed by other processors 1810A-1810N of the computer system 1800. However, in some examples, the virtualization manager implemented by the offload card 1870A or 1870B may adapt to requests from other entities (e.g., from the computing instance itself) and may not cooperate with (or serve) any single hypervisor.

[0202] In some examples, system memory 1820 may be one example of a computer-accessible medium configured to store program instructions and data as described above. However, in other examples, program instructions and / or data may be received, transmitted, or stored on different types of computer-accessible media. Generally, computer-accessible media may include any non-transitory storage medium or memory medium, such as magnetic or optical media, for example, a disk or DVD / CD coupled to computer system 1800 via I / O interface 1830. Non-transitory computer-accessible storage media may also include any volatile or non-volatile medium that may be included as system memory 1820 or another type of memory in some examples of computer system 1800, such as RAM (e.g., SDRAM, Double Data Rate (DDR) SDRAM, SRAM, etc.), read-only memory (ROM), etc. Furthermore, computer-accessible media may include transmission media or signals transmitted via communication media (such as networks and / or wireless links), such as electrical signals, electromagnetic signals, or digital signals, such as those implemented via network interface 1840.

[0203] The various examples discussed or presented herein can be implemented in a wide variety of operating environments, in some cases of which may include one or more user computers, computing devices, or processing devices that can be used to operate any of a number of applications. User devices or client devices may include any of a number of general-purpose personal computers, such as desktop or laptop computers running standard operating systems, and cellular, wireless, and handheld devices running mobile software and capable of supporting a number of networking and messaging protocols. Such systems may also include a number of workstations running a variety of commercially available operating systems and any of other known applications for purposes such as development and database management. These devices may also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and / or other devices capable of communicating via a network.

[0204] Most examples use at least one network familiar to those skilled in the art to support communication using any of a wide range of widely available protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), File Transfer Protocol (FTP), Universal Plug and Play (UPnP), Network File System (NFS), Public Internet File System (CIFS), Extensible Messaging and Field Protocol (XMPP), AppleTalk, etc. The network may include, for example, a Local Area Network (LAN), a Wide Area Network (WAN), a Virtual Private Network (VPN), the Internet, an intranet, an extranet, the Public Switched Telephone Network (PSTN), an infrared network, a wireless network, and any combination thereof.

[0205] In examples using web servers, the web server can run any of a variety of server or middleware applications, including HTTP servers, File Transfer Protocol (FTP) servers, Common Gateway Interface (CGI) servers, data servers, Java servers, business application servers, etc. The server may also be able to execute programs or scripts in response to requests from user devices, such as by executing programs or scripts in any programming language (e.g., ...). One or more web applications consisting of one or more scripts or programs written in C, C#, or C++, or any scripting language (such as Perl, Python, PHP, or TCL), or combinations thereof. The server may also include a database server, including but not limited to commercially available database servers from Oracle(R), Microsoft(R), Sybase(R), IBM(R), etc. The database server may be relational or non-relational (e.g., "NoSQL"), distributed or non-distributed, etc.

[0206] The environment disclosed herein may include various data repositories and other storage and storage media as discussed above. These may reside in various locations, such as on (and / or in) storage media local to one or more computers, or on storage media of any or all computers remote from a network. In a particular set of examples, information may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing functions belonging to a computer, server, or other network device may be stored locally and / or remotely, depending on the circumstances. Where the system includes computerized devices, each such device may include hardware elements electrically coupled via a bus, including, for example, at least one central processing unit (CPU), at least one input device (e.g., mouse, keyboard, controller, touchscreen, or keypad), and / or at least one output device (e.g., display device, printer, or speaker). Such a system may also include one or more storage devices, such as hard disk drives, optical storage devices, and solid-state storage devices such as random access memory (RAM) or read-only memory (ROM), as well as removable media devices, memory cards, flash memory cards, etc.

[0207] Such devices may also include computer-readable storage medium readers, communication devices (e.g., modems, network interface cards (wireless or wired), infrared communication devices, etc.), and working memory as described above. A computer-readable storage medium reader may be connected to or configured to receive a computer-readable storage medium, which represents a remote, local, fixed, and / or removable storage device and storage medium for temporarily and / or more permanently accommodating, storing, transmitting, and retrieving computer-readable information. The systems and various devices will also typically include numerous software applications, modules, services, or other elements, including operating systems and applications such as client applications or web browsers, residing within at least one working memory device. It should be understood that alternative examples may have many variations different from those described above. For example, custom hardware may also be used, and / or specific elements may be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connectivity to other computing devices, such as network input / output devices, may be employed.

[0208] Storage media and computer-readable media used to contain code or code portions may include any suitable media known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile media, removable and non-removable media implemented in any way or technology to store and / or transmit information (such as computer-readable instructions, data structures, program modules or other data), including RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, optical disc-read-only memory (CD-ROM), digital universal disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by system devices. Based on this disclosure and the teachings provided herein, those skilled in the art will appreciate other ways and / or methods for implementing the various examples.

[0209] In the foregoing description, various examples have been described. Specific configurations and details have been elaborated for illustrative purposes to provide a thorough understanding of the examples. However, it will also be apparent to those skilled in the art that the examples can be practiced without these specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the described examples.

[0210] In this document, parenthesized text and boxes with dashed borders (e.g., large dashes, small dashes, dot dashes, and dots) are used to indicate optional aspects for adding additional features to some examples. However, this notation should not be interpreted as implying that these are the only options or optional operations, and / or that in some examples, boxes with solid borders are not optional.

[0211] The reference numerals with suffix letters (e.g., 1718A-1718N) can be used to indicate that one or more instances of the mentioned entity may exist in various examples, and when multiple instances exist, each instance need not be identical, but may share some general characteristics or function in a common form. Furthermore, unless explicitly indicated to the contrary, the suffixes used do not imply the existence of a specific number of entities. Thus, in various examples, two entities using the same or different suffix letters may have or not have the same number of instances.

[0212] The use of terms like "an example" or "example" indicates that the described example may include a specific feature, structure, or characteristic, but each example may not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same example. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example, it should be assumed that, whether explicitly stated or not, implementing such a feature, structure, or characteristic in conjunction with other examples is within the knowledge of those skilled in the art.

[0213] Furthermore, in the various examples described above, unless otherwise specifically indicated, the disjunctive linguistic intent of phrases such as “at least one of A, B, or C” is understood to mean A, B, or C, or any combination thereof (e.g., A, B, and / or C). Similarly, the linguistic intent of phrases such as “at least one or more of A, B, and C” (or “one or more of A, B, and C”) is understood to mean A, B, or C, or any combination thereof (e.g., A, B, and / or C). Therefore, disjunctive language is neither intended nor should be understood to imply that a given example requires the existence of at least one of A, at least one of B, and at least one of C.

[0214] As used herein, the term "based on" (or similar) is an open-ended term used to describe one or more factors that influence a determination or other action. It should be understood that this term does not exclude additional factors that may influence a determination or action. For example, a determination may be based solely on the listed factors or on said factors and one or more additional factors. Therefore, if action A is "based on" B, it should be understood that B is a factor influencing action A, but this does not preclude the action from also being based on one or more other factors, such as factor C. However, in some cases, action A may be entirely based on B.

[0215] Unless otherwise expressly stated, articles such as “a / an” should generally be interpreted as including one or more of the described items. Thus, phrases such as “a device configured to…” or “computing device” are intended to include one or more of the described devices. Such one or more described devices may be collectively configured to perform the described operations. For example, “a processor configured to perform operations A, B, and C” may include a first processor configured to perform operation A working in conjunction with a second processor configured to perform operations B and C.

[0216] Furthermore, the words “may” or “may” are used in a permissive sense (i.e., implying a possibility) rather than a mandatory sense (i.e., implying a requirement). The words “include,” “including,” and “includes” are used to indicate an open relationship and therefore imply, including but not limited to. Similarly, the words “have,” “having,” and “has” also indicate an open relationship and therefore imply, having but not limited to. Terms such as “first,” “second,” “third,” etc., as used herein, serve as labels for the nouns that follow them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise explicitly indicated. Similarly, the values ​​of such numerical labels are generally not used to indicate the required quantity of a particular noun in a claim set forth herein, and therefore, the term “fifth” generally does not imply the presence of four other elements unless those elements are explicitly included in the claim or their presence is otherwise sufficiently clear.

[0217] Implementations of this disclosure may be described in one or more of the following clauses:

[0218] Clause 1. A computer-implemented method comprising: deploying computing resources to an edge location of a cloud provider network for an application; receiving, by a routing module, a request message destined for an application at least partially implemented at the edge location, wherein the request message is initiated by a mobile user equipment device using a communication network of a first communication service provider (CSP) having a first network address space, wherein the routing module is implemented in one of the edge locations deployed in a facility of the first CSP, or in a region of the cloud provider network; selecting, by the routing module, a first edge location satisfying the quality of service (QoS) requirements of the application as the destination of the request message, wherein the first edge location is deployed within a facility of a second CSP, wherein the QoS requirements confirm one or more of the following: maximum latency, maximum jitter, minimum bandwidth, minimum availability score, or maximum round-trip time; identifying a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the second CSP's network, the address space being different from the first network address space of the first CSP's network; and transmitting the request message to the first edge location using the first network address as a source identifier.

[0219] Clause 2. The computer-implemented method as described in Clause 1, further comprising: receiving from each of a plurality of CSPs including the first CSP and the second CSP a set of network quality of service characteristics corresponding to each of one or more edge locations deployed within the facility of the CSPs; and providing a summary of at least some of the set of network quality of service characteristics to a user associated with the application via a user interface.

[0220] Clause 3. The computer-implemented method as described in Clauses 1 to 2, further comprising: receiving, by the cloud provider network, a request for a service plan from the second CSP, the service plan being associated with the second CSP, with the first edge location of the second CSP, or with a geographic location including the first edge location.

[0221] Clause 4. A computer-implemented method comprising: receiving a request message destined for an application at least partially implemented in a plurality of edge locations within a cloud provider network, wherein the request message is initiated by a mobile user equipment device using a communication network of a first communication service provider (CSP); selecting from the plurality of edge locations an edge location that meets the quality of service requirements of the application as the destination of the request message, wherein the edge location is deployed within the facilities of a second CSP; identifying a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the second CSP's network; and transmitting the request message to the edge location using the first network address as a source identifier.

[0222] Clause 5. The computer-implemented method as described in Clause 4, further comprising: receiving a response message originating from within the edge location; and sending the response message back to the mobile user equipment device via the communication network of the first CSP.

[0223] Clause 6. The computer-implemented method as described in Clauses 4 and 5, further comprising: receiving a second request message destined for the application, wherein the second request message is initiated by a second mobile user equipment device using the communication network of the first CSP; selecting a second edge location from the plurality of edge locations as the destination of the second request message, wherein the second edge location is deployed within the facility of the first CSP; and transmitting the second request message to a computing instance within the second edge location.

[0224] Clause 7. The computer-implemented method as described in Clauses 4 through 6, wherein the selection of the edge location as the destination is based at least in part on the geographic location of the mobile user equipment device.

[0225] Clause 8. A computer-implemented method as described in Clauses 4 through 7, wherein the quality of service requirement identifies at least one of the following: maximum latency, maximum jitter, minimum bandwidth, minimum availability score, or maximum round-trip time.

[0226] Clause 9. The computer-implemented method as described in Clauses 4 through 8, further comprising: receiving from each of a plurality of CSPs including the first CSP and the second CSP a set of network quality of service characteristics corresponding to each of one or more edge locations deployed within the facility of the CSPs; and providing a summary of at least some of the set of network quality of service characteristics to a user associated with the application via a user interface.

[0227] Clause 10. The computer-implemented method as described in Clause 9, further comprising: receiving, by the cloud provider network, a request for a service plan from the second CSP, the service plan being associated with the second CSP, with the edge location of the second CSP, or with a geographic location including the edge location.

[0228] Clause 11. The computer-implemented method as described in Clause 10, further comprising: transmitting a message to the second CSP to represent client configuration or reserve network resources associated with the application.

[0229] Clause 12. The computer-implemented method as described in Clauses 4 to 11, further comprising: determining that a second edge location from the plurality of edge locations is more suitable than the edge locations for processing services associated with the mobile user equipment device, wherein the second edge location is deployed within another facility of the second CSP; and causing an additional request message initiated by the mobile user equipment device to be sent to the second edge location.

[0230] Clause 13. The computer-implemented method as described in Clauses 4 through 12, wherein the transmission of the request message to the edge location occurs at least in part by using fiber optic or radio transmission.

[0231] Clause 14. The computer-implemented method as described in Clauses 4 through 13, wherein the reception of the request message occurs: within a second edge location of the first CSP; or within the area of ​​the cloud provider network.

[0232] Clause 15. A system comprising: a first or more electronic devices for implementing multiple edge locations of a multi-tenant provider network; and a second or more electronic devices for implementing a service-oriented application deployment management (SOADM) service in the multi-tenant provider network, the SOADM service including instructions that, when executed, cause the SOADM service to: receive a request message destined for an application at least partially implemented in the multiple edge locations, wherein the request message is initiated by a mobile user equipment device using a communication network of a first communication service provider (CSP); select an edge location from the multiple edge locations that meets the quality of service requirements of the application as the destination of the request message, wherein the edge location is deployed within the facilities of a second CSP; identify a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the network of the second CSP; and transmit the request message to the edge location using the first network address as a source identifier.

[0233] Clause 16. The system as described in Clause 15, wherein the SOADM service further includes instructions that, when executed, cause the SOADM service to: receive a response message originating from within the edge location; and send the response message back to the mobile user equipment device via the communication network of the first CSP.

[0234] Clause 17. The system as described in Clauses 15 and 16, wherein the selection of the edge location as the destination is based at least in part on the geographic location of the mobile user equipment device.

[0235] Clause 18. The system as described in Clause 17, wherein the selection of the edge location as the destination is also based at least in part on the resource availability associated with the edge location.

[0236] Clause 19. The system as described in Clauses 15 through 18, wherein the SOADM service further includes instructions that, when executed, cause the SOADM service to: receive from each of a plurality of CSPs including the first CSP and the second CSP a set of network quality of service characteristics corresponding to each of one or more edge locations deployed within the facility of the CSP; and provide a summary of at least some of the set of network quality of service characteristics to a user associated with the application via a user interface.

[0237] Clause 20. The system as described in Clause 19, wherein the SOADM service further includes instructions that, when executed, cause the SOADM service to: receive a request for a service plan from the second CSP, the service plan being associated with the second CSP, with the edge location of the second CSP, or with a geographic location including the edge location.

[0238] The specification and drawings should therefore be considered illustrative rather than restrictive. However, it will be apparent that various modifications and alterations may be made therein without departing from the broader scope of this disclosure as set forth in the claims.

Claims

1. A computer-implemented method, comprising: Receive request messages sent to applications that are at least partially implemented in multiple edge locations of a cloud provider's network, wherein the request messages are initiated by a mobile user equipment device using the communication network of a first communication service provider (CSP); The edge location that meets the service quality requirements of the application is selected from the plurality of edge locations as the destination of the request message, wherein the edge location is deployed within the facility of the second CSP; Identify a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the network of the second CSP; as well as The request message is transmitted to the edge location by using the first network address as the source identifier.

2. The computer-implemented method as described in claim 1, further comprising: Receive response messages originating from within the said edge location; as well as The response message is sent back to the mobile user equipment device via the communication network of the first CSP.

3. The computer-implemented method as described in claim 1, further comprising: Receive a second request message sent to the application, wherein the second request message is initiated by a second mobile user equipment device using the communication network of the first CSP; A second edge location is selected from the plurality of edge locations as the destination of the second request message, wherein the second edge location is deployed within the facility of the first CSP; The second request message is transmitted to the computing instance within the second edge location.

4. The computer-implemented method of claim 1, wherein at least one of the following: The selection of the edge location as the destination is based at least in part on the geographic location of the mobile user equipment device; or The quality of service requirements define at least one of the following: maximum latency, maximum jitter, minimum bandwidth, minimum availability score, or maximum round-trip time.

5. The computer-implemented method as described in claim 1, further comprising: Receives a set of network service quality characteristics corresponding to each of one or more edge locations deployed within the facility of the CSP, including the first CSP and the second CSP; as well as A summary of at least some of the set of network service quality characteristics is provided to users associated with the application through a user interface.

6. The computer-implemented method as described in claim 5, further comprising: The cloud provider network receives a request to obtain a service plan from the second CSP, the service plan being associated with the second CSP, with the edge location of the second CSP, or with a geographic location including the edge location; as well as The message is transmitted to the second CSP to represent client configuration or reserve network resources associated with the application.

7. The computer-implemented method as described in claim 1, further comprising: Determining a second edge location from the plurality of edge locations to be more suitable than the stated edge locations for handling services associated with the mobile user equipment unit, wherein the second edge location is deployed within another facility of the second CSP; and This causes an additional request message initiated by the mobile user equipment device to be sent to the second edge location.

8. The computer-implemented method of claim 1, wherein at least one of the following: The transmission of the request message to the edge location occurs at least in part through fiber optic or radio transmission; or The receipt of the request message occurs when: Within the second edge position of the first CSP; or Within the area of ​​the cloud provider's network.

9. A system comprising: The first or more electronic devices are used to implement multiple edge locations of a multi-tenant provider network; as well as A second or more electronic devices are configured to implement a service-oriented application deployment management (SOADM) service in the multi-tenant provider network, the SOADM service including instructions that, when executed, cause the SOADM service to: Receive a request message sent to an application that is at least partially implemented in the plurality of edge locations, wherein the request message is initiated by a mobile user equipment device using the communication network of a first communication service provider (CSP); The edge location that meets the service quality requirements of the application is selected from the plurality of edge locations as the destination of the request message, wherein the edge location is deployed within the facility of the second CSP; Identify a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the network of the second CSP; as well as The request message is transmitted to the edge location by using the first network address as the source identifier.

10. The system of claim 9, wherein the SOADM service further includes instructions that, when executed, cause the SOADM service to: Receive response messages originating from within the said edge location; and The response message is sent back to the mobile user equipment device via the communication network of the first CSP.

11. The system of claim 9, wherein the selection of the edge location as the destination is based at least in part on at least one of the following: the geographic location of the mobile user equipment device, or the resource availability associated with the edge location.

12. The system of claim 9, wherein the SOADM service further includes instructions that, when executed, cause the SOADM service to: Receives a set of network service quality characteristics corresponding to each of one or more edge locations deployed within the facility of the CSP, including the first CSP and the second CSP; The application provides users associated with it with a summary of at least some of the set of network service quality characteristics through a user interface; and A request is received to obtain a service plan from the second CSP, the service plan being associated with the second CSP, with the edge location of the second CSP, or with a geographic location including the edge location.

13. A computer-implemented method, comprising: For applications, computing resources are deployed to the edge of the cloud provider's network; The routing module receives a request message destined for an application at least partially implemented at the edge location, wherein the request message is initiated by a mobile user equipment device using a communication network of a first communication service provider (CSP) having a first network address space, wherein the routing module is implemented in one of the edge locations deployed in the facilities of the first CSP, or in a region of the cloud provider network. The routing module selects a first edge location from the edge locations that meets the service quality requirements of the application as the destination of the request message, wherein the first edge location is deployed within the facility of the second CSP, and wherein the service quality requirements confirm one or more of the following: maximum latency, maximum jitter, minimum bandwidth, minimum availability score, or maximum round-trip time. Identify a first network address from a set of one or more network addresses provided by the second CSP, wherein the set of network addresses is within the address space of the network of the second CSP, and the address space is different from the first network address space of the first CSP network; as well as The request message is transmitted to the first edge location by using the first network address as the source identifier.

14. The computer-implemented method of claim 1, further comprising: Receives a set of network service quality characteristics corresponding to each of one or more edge locations deployed within the facility of the CSP, including the first CSP and the second CSP; as well as A summary of at least some of the set of network service quality characteristics is provided to users associated with the application through a user interface.

15. The computer-implemented method of claim 1, further comprising: The cloud provider network receives a request to obtain a service plan from the second CSP, the service plan being associated with the second CSP, with the first edge location of the second CSP, or with a geographic location including the first edge location.

Citation Information

Patent Citations

  • Enhanced NEF function, MEC and 5G integration

    CN111684824A

  • Method and system for intent-driven deployment and management of communication services in wireless communication system

    CN114208136A

  • Mobility of cloud computing instances hosted within communication service provider network

    CN115004661A

  • Connectivity of cloud edge locations to communications service provider networks

    US11159344B1

  • User-configured multi-location service deployment and scaling

    US11425054B1