Code execution on distributed units

By executing code on distributed units of the radio access network, the high load problem caused by limited resources of mobile computing devices is solved, and low-latency and high-efficiency code execution is achieved, which extends the device's battery life and reduces the temperature.

CN119998791APending Publication Date: 2025-05-13AMAZON TECH INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380069072.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-08-29
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Due to resource limitations, mobile computing devices are difficult to effectively execute high-load codes, resulting in shortened battery life and excessive temperature problems. At the same time, they cannot access remote resources at low latency to unload processing tasks.

Method used

By implementing code execution services on distributed units of the radio access network, it utilizes its abundant computing resources to execute code difficult to process by mobile computing devices, thereby reducing the burden on mobile devices.

Benefits of technology

Implements code execution at low latency and high efficiency, extending the battery life of mobile devices, reducing device temperature and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998791A_ABST
    Figure CN119998791A_ABST
Patent Text Reader

Abstract

Systems and methods are described for implementing a distributed unit in a radio access network that executes code on behalf of a mobile device. The distributed unit may be implemented on an edge server physically in close proximity to the radio unit with few or no intermediary devices. Thus, the edge server may provide services to a mobile device, such as executing code in an execution environment on behalf of the mobile device with significantly lower latency than a further cloud-based server on the edge server. The edge server may preload a computing environment with code for which a mobile device is likely (e.g., since a particular application is executing on the mobile device) to execute, and may determine whether to execute code on the edge server or on a cloud provider network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In general, computing devices can be used to exchange information via a network. Computing devices can utilize wireless networks provided by service providers to facilitate information exchange according to one or more wireless communication protocols. For example, a service provider can maintain a wireless network that enables mobile computing devices to exchange information according to wireless telecommunication protocols such as 4G (LTE), 5G, and 6G protocols. A wireless network can be composed of separate network components, such as a radio unit that sends and receives radio signals within a specific geographic area, and a distributed unit that sends and receives data from the radio unit and connects to other communication networks. The radio unit can therefore receive data from the distributed unit and send the data to the mobile computing device, which can process the data using the computing resources of the mobile device. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Reference numerals may be repeated throughout the drawings to indicate correspondence between referenced elements.The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure.

[0003] Figure 1 is a block diagram depicting an example operating environment in which a mobile computing device can transmit and receive data via a radio access network and can request a distributed unit to execute code on behalf of the mobile computing device in accordance with aspects of the present disclosure.

[0004] Figure 2 is a flow diagram depicting an example interaction in which a mobile computing device may request a distributed unit to execute code on its behalf, and in which the distributed unit may provide results of executing such code, according to aspects of the present disclosure.

[0005] Figure 3 is a flow diagram depicting example interactions according to aspects of the present disclosure, in which a mobile computing device may notify a distributed unit that an application is running that may send a request to execute code, and in which the distributed unit prepares an environment for executing code that the application may request.

[0006] Figure 4 is a flow chart depicting an example routine for processing a request to execute code on a distributed unit in accordance with aspects of the present disclosure.

[0007] Figure 5 is a block diagram depicting the overall architecture of a computing device configured to implement a distributed unit for performing operations according to aspects of the present disclosure.

[0008] Figure 6 is a diagram illustrating an example functional split that may be selected for a 5G RAN according to various embodiments of the present disclosure. DETAILED DESCRIPTION

[0009] In general, aspects of the present disclosure relate to improving the performance of mobile computing devices. More specifically, aspects of the present disclosure relate to systems, methods, and computer-readable media associated with distributed units in a radio access network ("RAN") that accept requests from mobile computing devices to execute code on the distributed units and provide the results of executing such code back to the mobile device or another specified target. Thus, the distributed units can execute code that will overburden the limited resources of the mobile computing device (e.g., code representing a workload that, if executed on the mobile computing device, will consume a significant amount of the mobile computing device's battery life, cause the mobile computing device to exceed its recommended operating temperature, or otherwise cause the user experience of the user of the mobile computing device to be less than optimal). The distributed units of the RAN have significantly more computing resources than a typical mobile computing device, and as discussed in more detail below, the distributed units are typically separated from the mobile computing device only by radio units and antennas, thereby allowing the distributed units to execute code in scenarios where latency would prevent the mobile computing device from utilizing a less-close computing environment. Thus, the distributed units can implement code execution services with improved performance relative to cloud-based implementations of these services hosted on more distant servers because the more distant servers are separated from the mobile computing devices by a network infrastructure that places physical limits on minimum latency and that may vary in distance and latency.

[0010] The RAN includes antennas for over-the-air communications with mobile computing devices, as well as elements such as radio units ("RUs"), distributed units ("DUs"), and centralized units ("CUs"). In general, the RUs, DUs, and CUs convert analog radio signals received from the antennas into digital packets that can be routed through the network, and similarly convert digital packets into radio signals that can be sent through the antennas. This signal conversion is accomplished by a series of network functions, which can be distributed among the RUs, DUs, and CUs in various ways to achieve different balances of latency, throughput, and network performance. These network functions generally correspond to the lowest three network layers in the seven-layer OSI model of computer networking, and the distribution of these network functions among the RUs, DUs, and CUs can be referred to as the "functional split" of the RAN.

[0011] The RUs, DUs, and CUs may be geographically distributed, but typically at least the RUs will be located close to amplifiers, filters, towers, antennas, and other hardware used to send and receive radio signals. One or more CUs of a 5G network may be located away from antennas located in a more centralized location, as discussed in more detail below, and one or more DUs may be located near or co-located with one or more RUs. In some embodiments, the functional split of the RAN and the functions selected to run on one or more DUs may be factors in determining the location of one or more DUs relative to one or more CUs and one or more RUs.

[0012] RU, DU and CU may be provided in different ratios to each other. For example, multiple RUs may be connected to one DU, and multiple DUs may be connected to one CU. Each RU may provide coverage to different geographic areas, which may partially overlap with the area of ​​adjacent RUs to facilitate switching and provide seamless communication. The DU may typically receive data from the CU for delivery to the RU (and vice versa), and may process the data to encode or decode it, modulate or demodulate it, add or remove error correction or redundancy, or otherwise prepare the data received from the CU for air transmission by the RU, and vice versa.

[0013] The RU may typically implement "Layer 1" (which may be referred to herein as "L1," "physical layer," or "PHY") functions, which correspond to the first and lowest layer in the OSI model. In 5G and other wireless protocols, these functions are related to the transmission and reception of radio signals. Layer 1 functions may be further divided into "high PHY" functions (such as converting binary data to and from electrical pulses representing data) and "low PHY" or "RF" functions (such as converting electrical pulses to and from radio waves that can be wirelessly transmitted through an antenna, transmitting and receiving radio waves during a specified timing window and at a specified frequency, and so on). The DU may also implement some Layer 1 functions, such as encoding digital signals to be sent by the RU, adding redundancy and error correction codes to digital signals, mapping transmission channels to physical channels, and so on. In some embodiments, the DU may include dedicated hardware for performing Layer 1 functions, such as an accelerator card that performs signal processing functions. The DU may also implement "Layer 2" (which may be referred to herein as "L2" or "Data Link Layer") functions, which correspond to the second layer of the OSI model, such as mapping logical channels to transport channels, determining the timing of digital signal transmission of RUs within a timing window, and so on. The L2 function provides an interface between the higher transport layer and the physical layer. In 5G, the L2 layer has three sublayers: Medium Access Control (MAC), Radio Link Control (RLC), and Packet Data Convergence Protocol (PDCP). Each of these can be considered a network function. PDCP provides security for radio resource control (RRC) traffic and signaling data, sequence numbering and sequential delivery of RRC messages and IP packets, and IP packet header compression. The RLC protocol provides control of the radio link. The MAC protocol maps information between logical channels and transport channels.

[0014] The data link layer interfaces with the network layer ("layer 3" or "L3"). In 5G, the network layer is also called the radio resource control (RRC) layer and is responsible for functions such as packet forwarding, quality of service management, and the establishment, maintenance, and release of the RRC connection between the mobile computing device and the RAN. The CU may typically implement layer 3 functionality.

[0015] Figure 6 is a diagram illustrating various functional splits that may be selected for a 5G RAN. As shown, the functional split defines a different set of L1 and L2 functions that run on the RU relative to those that run on the CU and DU. L3 functions (not depicted) also run on the CU. For example, in a 5G RAN architecture that follows Split 7, the functionality of the baseband unit (BBU) used in previous generations of wireless networks is split into two functional units: a DU responsible for real-time L1 and L2 scheduling functions; and a CU responsible for non-real-time, higher L2 and L3 functions.

[0016] In some embodiments, certain components of the RAN may be implemented using a cloud provider network. A cloud provider network (sometimes referred to simply as a "cloud") refers to a network-accessible pool of computing resources (such as computing, storage, and networking resources, applications, and services) that may be virtualized or provided as "bare metal" hardware. The cloud can provide convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and released in response to customer commands. These resources can be dynamically provisioned and reconfigured to adjust to variable loads. Cloud computing can therefore be viewed as both applications delivered as services over a publicly accessible network (e.g., the Internet) and the hardware and software in the cloud provider data centers that provide those services. Other components of the RAN may require or benefit from the use of dedicated hardware (such as a radio transceiver or a dedicated signal processor).

[0017] A cloud provider network may provide users with an on-demand scalable computing platform over a network, for example, thereby allowing users to have scalable "virtual computing devices" (also referred to as virtual computing instances) available to them through their use of computing servers (which provide computing instances via one or both of a CPU and a GPU, optionally in conjunction with local storage) and block storage servers (which provide virtualized persistent block storage for a specified computing instance). These virtual computing devices have the attributes of a personal computing device, including hardware (processors of various types, local memory, random access memory ("RAM"), hard disk and / or solid-state drive ("SSD") storage devices), operating system selection, networking capabilities, and pre-loaded application software. Each virtual computing device may also virtualize its console input and output (e.g., keyboard, display, and mouse). This virtualization allows users to connect to their virtual computing devices using computer applications (such as browsers, application programming interfaces, software development kits, etc.) to configure and use their virtual computing devices as they would do with a personal computing device. Unlike personal computing devices that have a fixed amount of hardware resources available to the user, the hardware associated with the virtual computing device may be scaled up or down depending on the resources required by the user. An application programming interface (API) refers to an interface and / or communication protocol between a client and a server such that if a client makes a request in a predefined format, the client should receive a response in a specific format or initiate a defined action. In the context of a cloud provider network, an API provides a gateway to enable customers to access the cloud infrastructure by allowing them to obtain data from the cloud provider network or cause actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted in the cloud provider network. An API may also enable different services of a cloud provider network to exchange data with each other. Users may choose to deploy their virtual computing systems to provide network-based services for their own use and / or for use by their customers or clients.

[0018] The cloud provider network can be formed into multiple regions, where a region is a separate geographic area in which the cloud provider clusters data centers. Each region can include two or more availability zones connected to each other via a private high-speed network, such as a fiber-optic communication connection. An availability zone refers to an isolated fault domain that includes one or more data center facilities with power, networking, and cooling that are separate from those in another availability zone. Preferably, the availability zones within a region are located far enough away from each other so that the same natural disaster does not take more than one availability zone offline at the same time. Customers can connect to the availability zones of the cloud provider network via a publicly accessible network (e.g., the Internet, a cellular communication network). A switching center (TC) is the main backbone location that links customers to the cloud provider network and can be co-located at other network provider facilities (e.g., an Internet service provider, a telecommunications provider). To achieve redundancy, two TCs can be operated per region.

[0019] The cloud provider network can include a physical network (e.g., sheet metal boxes, cables, rack hardware) called the underlay. You can think of the underlay as the network fabric that contains the physical hardware that runs the provider network's services. The underlay can be isolated from the rest of the cloud provider network, for example, it may not be possible to route from an underlay network address to an address in the production network that runs the cloud provider's services, or to a customer network that hosts customer resources.

[0020] The cloud provider network may also include an overlay network of virtualized computing resources running on the underlay. Thus, network packets may be routed along the underlay network according to the constructs in the overlay network (e.g., VPC, security groups). A mapping service may coordinate the routing of these network packets. The mapping service may be a regionally distributed lookup service that maps a combination of an overlay IP and a network identifier to an underlay IP so that the distributed underlay computing devices can look up where to transmit the packet.

[0021] For illustration, each physical host (e.g., computing server, block storage server, object storage server, control server) may have an IP address in the underlying network. Hardware virtualization technology may enable multiple operating systems to run concurrently on a host computer, for example, as a virtual machine (VM) on a computing server. The hypervisor or virtual machine monitor (VMM) on the host allocates the hardware resources of the host among the various VMs on the host and monitors the execution of the VM. Each VM may be provided with one or more IP addresses in the overlay network, and the VMM on the host may know the IP address of the VM on the host. The VMM (and / or other devices or processes on the network bottom layer) may use encapsulation protocol technology to encapsulate network packets (e.g., client IP packets) and route the network packets between virtualized resources on different hosts within the cloud provider network through the network bottom layer. Encapsulation protocol technology may be used on the network bottom layer to route encapsulated packets between endpoints on the network bottom layer via an overlay network path or route. Encapsulation protocol technology may be viewed as providing a virtual network topology overlaid on the network bottom layer. The encapsulation protocol technology may include a mapping service that maintains a mapping directory that maps IP overlay addresses (e.g., public IP addresses) to underlying IP addresses (private IP addresses), which can be accessed by various processes on the cloud provider network for routing packets between endpoints.

[0022] In various embodiments, the traffic and operations underlying the cloud provider network can be generally divided into two categories: control plane traffic carried by a logical control plane and data plane operations carried by a logical data plane. The data plane represents the movement of user data through a distributed computing system, while the control plane represents the movement of control signals through a distributed computing system. The control plane typically includes one or more control plane components distributed across and implemented by one or more control servers. Control plane traffic typically includes management operations, such as establishing isolated virtual networks for various customers, monitoring resource usage and health, identifying specific hosts or servers to launch requested computing instances, and provisioning additional hardware as needed. The data plane includes customer resources (e.g., computing instances, containers, block storage volumes, databases, file storage) implemented on the provider network. Data plane traffic typically includes non-management operations, such as transmitting data to and from customer resources.

[0023] As shown, the data plane may include one or more computing servers, which may be bare metal (e.g., a single tenant) or may be virtualized by a hypervisor to run multiple VMs (sometimes referred to as "instances") for one or more customers. These computing servers may support virtualized computing services for the provider network. The provider may provide virtual computing instances with different computing and / or memory resources. In one embodiment, each of the virtual computing instances may correspond to one of several instance types. The instance type may be characterized by its hardware type, computing resources (e.g., the number, type, and configuration of central processing units [CPUs] or CPU cores), memory resources (e.g., the capacity, type, and configuration of local memory), storage resources (e.g., the capacity, type, and configuration of locally accessible storage devices), network resources (e.g., the characteristics of its network interface and / or network capabilities), and / or other suitable descriptive characteristics. Using the instance type selection functionality, an instance type may be selected for a customer, for example (at least in part) based on input from a customer. For example, a customer may select an instance type from a predefined set of instance types. As another example, a customer may specify the desired resources of an instance type and / or the requirements of the workload that the instance will run, and the instance type selection functionality may select an instance type based on such specifications.

[0024] The control plane components are typically implemented on a collection of servers separate from the data plane servers, and the control plane traffic and the data plane traffic can be transmitted over separate / different networks. In some embodiments, the control plane traffic and the data plane traffic can be supported by different protocols. In some embodiments, the message (e.g., packet) transmitted over the provider network includes a flag for indicating whether the traffic is control plane traffic or data plane traffic. In some embodiments, the payload of the traffic can be checked to determine its type (e.g., control plane or data plane). Other techniques for distinguishing traffic types are possible.

[0025] Some customers may desire to use resources and services of a cloud provider network, but prefer to provision these resources and services within their own network (e.g., on the customer's premises) for various reasons (e.g., latency of communication with customer devices, legal compliance, security, or other reasons). The techniques described herein enable a small portion of a cloud provider network (referred to herein as a "provider underlay extension" or PSE) to be provisioned within a customer's network. Customers can access their PSEs via the cloud provider underlay or their own network, and can create and manage resources in the PSE using the same APIs that they would use to create and manage resources in a region.

[0026] The PSE may be pre-configured with a suitable combination of hardware and software and / or firmware elements, for example, by a provider network operator to support various types of computing-related resources, and to do so in a manner that reflects the experience of using the provider network. For example, one or more PSE servers may be prepared in a customer network by a cloud provider. As described above, the provider network may provide a collection of predefined instance types, each of which has different types and quantities of underlying hardware resources. Each instance type of various sizes may also be provided. In order to enable customers to continue to use the same instance types and sizes used in their PSEs as they do in the region, the PSE server may be a heterogeneous server. Heterogeneous servers may concurrently support multiple instance sizes of the same type, and may also be reconfigured to host any instance type supported by its underlying hardware resources. The reconfiguration of heterogeneous servers may be performed instantly using the available capacity of the PSE server, which means that other VMs are still running and consuming other capacity of the PSE server. This may improve the utilization of resources within the PSE by allowing better packaging of instances running on physical hosts, and also provide a seamless experience of instance usage across regions and PSEs.

[0027] In one embodiment, the PSE server may host one or more VMs. Customers can use these VMs to host containers that package code and all its dependencies so that applications can run quickly and reliably from one computing environment to another. In addition, if the customer needs, the PSE server may host one or more data volumes. In the region, such volumes may be hosted on dedicated block storage servers. However, due to the possibility of having significantly smaller capacity at the PSE than in the region, if the PSE includes such dedicated block storage servers, it may not be possible to provide the best utilization experience. Therefore, the block storage service can be virtualized in the PSE so that one of the VMs runs the block storage software and stores the data of the volume. Similar to the operation of the block storage service in the region, the volumes within the PSE can be replicated to achieve persistence and availability. These volumes can be provisioned within its own VPC within the PSE. The VM and any volume together constitute an extension of the provider network data plane within the PSE.

[0028] In some implementations, the PSE server may host certain local control plane components, such as components that enable the PSE to continue to operate in the event of a disruption in connectivity back to the region. Examples of these components include a migration manager that can move VMs between PST servers if availability needs to be maintained; a key value data store that indicates where volume replicas are located; and a local VM placement component that can respond to requests for new VMs made via the customer network. However, in general, the control plane of the PSE will remain in the region to allow customers to use as much of the PSE's capacity as possible. In some embodiments, at least some of the VMs set up at the PSE and associated higher-level services that use such VMs as building blocks can continue to function even during periods of time when connectivity to the provider network data center is temporarily interrupted.

[0029] In the above manner, PSE forms an edge location because it provides resources and services of the cloud provider network outside the traditional cloud provider data center and closer to the customer device. The edge location as mentioned herein can be structured in several ways. In some implementations, the edge location can be an extension of the cloud provider network bottom layer, including a limited amount of capacity provided outside the availability zone (for example, in a small data center of a cloud provider or located near the customer workload and possibly away from other facilities in any availability zone). Such edge locations can be referred to as local areas (due to being more local or closer to a group of users than traditional availability zones). The local area can be connected to a publicly accessible network such as the Internet in various ways (for example, directly, via another network or via a private connection with a region). Although typically the local area will have a more limited capacity than a certain area, in some cases, the local area can have a considerable capacity, such as thousands or more racks.

[0030] Thus, in some embodiments, all or part of the RAN (including but not limited to edge servers implementing distributed units) may be implemented on edge location hardware, which may be physically closer to devices such as RUs and mobile computing devices, for example. In some embodiments, the edge location may be a provider bottom layer extension formed by one or more servers located internally at a customer or partner facility. One or more servers may communicate with a nearby availability zone or region of a cloud provider network via a network (e.g., a publicly accessible network such as the Internet). This type of provider bottom layer extension may be referred to as a "sentinel" of a cloud provider network. Some outposts may be integrated into a communication network. For example, an outpost may be integrated into a base station within a telecommunications network, or co-located with an RU of the network. The capacity of an internal outpost may be available only to customers who have premises (and any other accounts allowed by the customer). Similarly, the capacity of an outpost integrated into a telecommunications network may be shared between multiple applications (e.g., games, virtual reality applications, healthcare applications) that transmit data to users of the telecommunications network. An outpost integrated into a telecommunications network may also include data plane capacity, which may be at least partially controlled by a control plane implemented in a nearby availability zone of the cloud provider network. Thus, an availability zone group may include a "parent" availability zone and any "child" edge locations that are subordinate to the parent availability zone (e.g., at least partially controlled by its control plane). Certain limited control plane functionality (e.g., features required for low-latency communication with customer resources, and / or features that enable edge locations to continue to operate when disconnected from the parent availability zone) may also exist in some edge locations.

[0031] Therefore, RAN functionality can be implemented on hardware (e.g., edge location hardware) that is relatively close to the mobile computing device, providing significantly more computing resources than a typical mobile computing device, and is not subject to the power consumption restrictions that are typically present in the mobile computing device. However, the mobile computing device typically cannot access these computing resources. The mobile computing device that offloads work typically offloads the work to services provided on the cloud provider network, such as services that perform serverless computing functions or provide virtual machine instances. Relative to edge servers, these services can provide more computing resources and have fewer restrictions on power consumption. However, the servers that implement these services are typically separated from the mobile computing device by multiple network components (e.g., routers, gateways, load balancers, etc.). Each of these components will cause delays in the interaction between the mobile computing device and the server to which it will offload the work, making it impractical for the mobile computing device to offload work that requires low-latency processing. For example, it may be impractical for the mobile computing device to offload image or video processing, the generation of a user interface, or other generation or presentation of results that a user may interact with in real time or near real time. The performance and latency associated with these services may also vary significantly from one execution to the next due to rapid reconfiguration of resources that are not controlled by the user. For example, a serverless computing system may execute code on behalf of a user at different times in different environments, and the user may have limited or no control over these environments. For some cloud-based services, users may reserve dedicated resources to ensure consistent performance, but doing so represents an inefficient use of computing resources that are only used occasionally. As a result, mobile computing devices may be prevented from running applications that require a large amount of local processing power, or may be prevented from running such applications for extended periods of time due to issues such as battery consumption, device temperature under continuous heavy workloads, and may not have a viable option for offloading work to physically distant servers due to issues with latency and performance variability.

[0032] To address these issues, operators of radio access networks may implement distributed units that accept and satisfy requests to execute code on behalf of mobile computing devices. As discussed in more detail below, distributed units may implement services that would otherwise be provided as cloud services, such as serverless computing functions or images that execute virtual machine instances, but may be implemented with lower latency by implementing these services on distributed units, and may therefore make it more practical for mobile computing devices to utilize these services. In some embodiments, distributed units may implement software applications for which code execution is "split" between mobile computing devices and distributed units, such that some functions of the application are implemented only on the distributed unit, while other functions of the application are implemented only on the mobile computing device. In other embodiments, as described below, the distributed units described herein may allow applications on mobile computing devices to determine whether to execute code on the mobile computing device or on the distributed unit based on factors such as available processing power, bandwidth, battery, etc. The distributed unit can provide these services to many mobile computing devices simultaneously, because mobile computing devices typically have transient and temporary needs to offload work, and can provide these services at a consistent performance level (e.g., consistently low latency) relative to services implemented across multiple cloud-based servers and networks that vary in performance and latency contributions. Additionally, in some embodiments, the distributed unit can determine whether to execute code locally (e.g., in a computing environment on an edge server) or regionally (e.g., in a region of a cloud provider network) based on various criteria as described in more detail below.

[0033] It should be understood that the techniques described herein solve technical problems that arise particularly in the field of computer networks, and in particular solve problems that arise when a mobile computing device has limited local resources for executing code but cannot access remote resources with sufficiently low latency to allow its use. It should also be understood that the technical problems described herein are not analogous to any pre-Internet practice, and that the distributed units described herein improve the performance of the radio access network by enabling more efficient use of computing resources at both the mobile computing devices and the distributed units. Therefore, wireless network operators can use the techniques described herein to more efficiently utilize their radio access networks and provide wireless telecommunication services more efficiently.

[0034] Embodiments of the present disclosure will now be described with reference to the accompanying drawings, wherein like numerals refer to like elements throughout. The terms used in the descriptions presented herein are not intended to be interpreted in any limiting or restrictive manner, simply because the terms are utilized in conjunction with the detailed description of certain specific embodiments of the present invention. In addition, embodiments of the present invention may include several novel features, a single feature of which is not solely responsible for its desired properties or essential to practicing the invention described herein.

[0035] Figure 1 1 is a block diagram of an example operating environment in which edge server 130 may implement aspects of the present disclosure. In the depicted embodiment, edge server 130 includes a code data store 132, a distributed unit 134 of radio access network 120, and a plurality of execution environments 136 in which code may be executed on behalf of user device 102. Distributed unit 134 may operate based on communication with radio unit 122 and communication with network core 150 via backhaul network 140. Radio unit 122 may in turn communicate with user device 102 via air interface 110. Network core 150 may in turn communicate with devices on external networks, thereby enabling, for example, communication between user device 102 and servers or other computing devices on the Internet.

[0036] In general, user device 102 may be any device operable to communicate via air interface 110. Examples of user device 102 include mobile phones, tablet computing devices, laptop computing devices, wearable computing devices, desktop computing devices, personal digital assistants (PDAs), hybrid PDA / mobile phones, e-book readers, set-top boxes, voice command devices, cameras, digital media players, servers, etc. Illustratively, air interface 110 may be an air interface to any wireless network, including but not limited to a cellular telecommunications network, a Wi-Fi network, a mesh network, a personal area network, or any combination thereof. In some embodiments, air interface 110 may be an interface to a global system for mobile communications (GSM) network, a code division multiple access (CDMA) network, a long term evolution (LTE) network, or a combination thereof.

[0037] The radio unit 122 sends and receives data from the user devices 102 via the air interface 110, and acts as an endpoint for these user devices 102 to access the network core 150 via the radio access network 120 and the backhaul network 140. The radio unit 122 may correspond to a base station of a cellular telephone network, or in some embodiments may be deployed separately or independently from any existing cellular telephone network. In some embodiments, multiple radio units 122 may communicate with a single distributed unit 134.

[0038] The edge server 130 is referred to below as Figure 5 1 and 10. The distributed unit 134 is described and implemented in more detail and is generally responsible for at least some aspects of the physical layer and the data link layer of the radio access network 120. The distributed unit 134 receives data from the centralized unit (which in some embodiments is also fully or partially implemented by the edge server 130) and sends the data to the radio unit 122 for delivery to the user device 102. The edge server 130 also includes a code data storage 132 (which can illustratively be any non-transitory computer-readable storage medium) and a plurality of execution environments 136. In various embodiments, the execution environment 136 can be a virtual machine instance, a software "container" that provides an isolated runtime environment without providing hardware virtualization, or other environment in which code can be executed. In some embodiments, the execution environment 136 can be managed by a worker manager ( Figure 1 ), which may implement some of the functionality described herein, such as selecting a specific execution environment 136 in which to satisfy a specific request to execute code.

[0039] In some embodiments, edge server 130 may implement all or part of a serverless computing environment. As used herein, the term "serverless computing environment" is intended to refer to an environment in which the generation, configuration, and state of the underlying execution environment are abstracted from the user so that the user does not need to, for example, create an execution environment, install an operating system within the execution environment, or manage the state of the environment in order to execute the desired code in the environment. Similarly, the term "server-based computing environment" is intended to refer to an environment in which, in addition to executing the desired code in the environment, the user is at least partially responsible for managing the generation, configuration, or state of the underlying execution environment. Therefore, it will be understood by those skilled in the art that "serverless" and "server-based" may indicate the degree of user control over the execution environment in which the code is executed, rather than the actual absence or presence of a server.

[0040] The serverless computing environment may provide services that enable the user device 102 to submit or specify computer executable code to be executed in the execution environment 136. Each set of code on the serverless computing system may define a "task" and may implement specific functionality corresponding to the task when executed in the execution environment 136 of the serverless computing system. Individual implementations of tasks on the serverless computing execution system may be referred to as "executions" of the tasks (or "task executions"). The serverless computing system may further enable users to trigger the execution of tasks based on various potential events, such as detecting new data at a network-based storage system, transmitting an application programming interface ("API") call to the serverless computing system, or transmitting a specially formatted hypertext transfer protocol ("HTTP") packet to the serverless computing system. Therefore, users can utilize the serverless computing system to execute any specified executable code "on demand" without the need to configure or maintain the underlying hardware or infrastructure on which the code is executed. In addition, the serverless computing system can be configured to execute tasks in a rapid manner (e.g., in less than 100 milliseconds), thereby enabling "real-time" (e.g., with little perceptible delay to the end user) execution of tasks. The rapid execution of tasks in the serverless computing system, together with the proximity of the edge server 130 to the user device 102, can similarly facilitate the execution of tasks with little perceptible delay.

[0041] References to user code used herein may refer to any program code (e.g., program, routine, subroutine, thread, etc.) written in a specific programming language. In the present disclosure, the terms "code", "user code" and "program code" may be used interchangeably. For example, such user code may be executed in conjunction with a specific mobile application on a user device 102 to implement a specific function. As described above, a separate set of user codes (e.g., to implement a specific function) is referred to herein as a "task", and a specific execution of the code (including, for example, compiling code, interpreting code, or otherwise enabling code to be executed) is referred to as "task execution" or simply "execution". Tasks may be written in JavaScript (e.g., node.js), Java, Python and / or Ruby (and / or another programming language) by way of non-limiting examples. Tasks may be "triggered" in various ways to be executed on a serverless computing system. In one embodiment, a user or user device 102 may send a request to execute a task, which may generally be referred to as a "call" to execute a task. Such a call may include the user code to be executed (or its location) and one or more independent variables to be used to execute the user code. For example, the call may provide the user code of the task and a request to perform the task. In another example, the call may identify the previously uploaded task by the name or identifier of the previously uploaded task. In yet another example, the code corresponding to the task may be included in the call to the task, and uploaded to a separate location (e.g., code data storage area 132 or a storage location outside edge server 130) before the request is received by the serverless computing system. As described above, the code of the task can refer to additional code objects maintained at the serverless computing system by using the identifiers of those code objects, so that before executing the task, the code object is combined with the code of the task in the execution environment. The serverless computing system can change its execution strategy for the task based on where the task code is available when processing the call to the task. The request interface of the front end can receive a call from the user to execute the task as a hypertext transfer protocol secure (HTTPS) request. Moreover, when executing the task, any information (e.g., headers and parameters) included in the HTTPS request can also be processed and utilized. As discussed above, any other protocol including, for example, HTTP, MQTT, and CoAP can be used to transmit a message containing a task call to the request interface.

[0042] Although Figure 1 Not shown, edge server 130 may implement a placement service, a worker manager, and other components to support execution of serverless computing functions.

[0043] Edge servers 130 may communicate with network core 150 via backhaul network 140. Illustratively, backhaul network 140 may be any wired or wireless network or a combination thereof. Additionally, backhaul network 140 may include, but is not limited to, the Internet, a public or private intranet, a cellular telecommunications network, a Wi-Fi network, a wired network, a satellite network, a mesh network, a personal area network, a local area network (LAN), a wide area network (WAN), or one or more other public or private communication networks, or any combination thereof. In some embodiments, backhaul network 140 may be implemented entirely or partially within a cloud provider network, as described in more detail below.

[0044] The network core 150 is implemented in various embodiments in a nearby availability zone of a cloud provider network, in an internal edge location, at another edge location, or on other network-accessible computing resources. For example, the network core 150 may be implemented using spare capacity of an outpost that implements RAN functionality, or on a separate outpost. In some implementations, the network core may be implemented in a local area close to an edge location where a RAN is running. The network core 150 provides control plane functionality (as described above) and performs management and control functions of the telecommunications network, such as authenticating subscribers, applying usage policies, managing UE sessions, and other control plane functions. In some embodiments, the network core 150 may support multiple RANs, and may also support RANs from multiple tenants or customers. In other embodiments, the network core 150 may be implemented on dedicated computing resources of a specific wireless network operator, which may be provided within the premises or within the cloud provider network.

[0045] It should be understood that the example operating environment may include Figure 1 There may be more (or fewer) elements than those shown in the drawings. However, it is not necessary to show all of these elements in order to provide a feasible disclosure.

[0046] Figure 2 is a flow chart depicting an example interaction in which a user device 102 requests a distributed unit 134 to execute code on its behalf. Illustratively, Figure 2 The interaction depicted in allows user device 102 to offload code execution to distributed unit 134 (or, in some embodiments, to a serverless computing environment in a cloud provider network) and receive the results of executing the code in execution environment 136, which has superior computing resources than user device 102 but is located close enough to user device 102 to enable it to return results quickly. For example, execution environment 136 can be on physical hardware (e.g., hardware of edge server 130, as described below with reference to Figure 5The user device 102 may be implemented on a wireless radio 122 (described in more detail above), the physical hardware being physically close to the radio 122 (e.g., in the same rack) and / or communicating with the radio 122 via a single local area network within a single physical location. The interaction begins at (1), where the user device 102 sends a request for remote code execution to the radio 122. In some embodiments, the user device 102 may send the code to be executed. In other embodiments, the user device 102 may send an identifier, location, or other information that allows the code to be obtained (e.g., from a cloud-based data store). In addition, in some embodiments, the user device 102 may send data to be processed by the code as part of the request or independent of the request. The user device 102 may typically communicate directly with the radio 122 wirelessly, with only physical hardware (e.g., an antenna) separating the two.

[0047] In some embodiments, the request sent by the user device 102 may be generated by an application executed on the user device 102. In various embodiments, the application may determine whether to execute the code on the user device 102, in the execution environment 136 hosted by the distributed unit 134, or in an execution environment hosted by another platform (e.g., an environment at a region). For example, a determination may be made based on factors such as maximum latency requirements (e.g., how long does the application need before it must display the result) and / or the availability of various resources on the user device 102 (such as battery power, processor, memory, storage, etc.). In some embodiments, the application may consider whether the operating temperature of the user device 102 will exceed acceptable limits in the case of executing a sustained workload, or whether the predicted amount of computing resources that the execution code will utilize exceeds acceptable limits. In other embodiments, the user device 102 may generate and send a request to execute the code, and may similarly determine whether to request code execution.

[0048] At (2), the radio unit 122 relays the request to the distributed unit 134. In embodiments where the user device 102 sends a code identifier or location instead of sending a code, the distributed unit 134 (or, in some embodiments, the edge server 130) requests the code from the code data store 132 at (3). In some embodiments, the code data store 132 may store software images that can be loaded into a virtual machine instance and executed. Such images may include, for example, operating systems, runtime libraries, utilities, or other content. The image may also include user-specific code, or, in some embodiments, the request may include code to be loaded and executed after the image has been loaded onto the virtual machine instance. In other embodiments, the code data store 132 may store code such as serverless computing functions, which, as described above, may be executed in an execution environment 136 whose configuration is not specified or controlled by the user. In embodiments where the code is requested from the code data store 132, the code data store 132 returns the requested code at (4).

[0049] At (5), the distributed unit 134 (or, in some embodiments, another component of the edge server 130, such as a worker manager) selects an execution environment 136 from a plurality of execution environments 136 on the edge server 130. As described above, the execution environment 136 may be a virtual machine instance, a container, or other environment in which a request to execute code may be satisfied. In some embodiments, the execution environment 136 may be selected based on the code, the request, previous executions of the code, or other information. For example, an execution environment 136 having a particular set of computing resources allocated to it may be selected based on the amount of computing resources consumed during one or more previous executions of the code. As a further example, the request may specify a maximum acceptable delay or latency before a result should be returned, and the execution environment 136 may be selected based on a performance criterion indicating that it will return a result within that time frame. In some embodiments, the distributed unit 134 may determine whether to execute the code in an execution environment 136 hosted by the edge server 130 or in an execution environment hosted at a region. At (6), the distributed unit 134 sends the code (and, in some embodiments, data to be processed, configuration information, etc.) to the selected execution environment 136.

[0050] At (7), the execution environment 136 satisfies the request by executing the code and generating a result. Illustratively, the code may generate information or a user interface element presented in a user interface by the user device 102, and the result may be the generated information or element. For example, the code may process image data or video data captured by the user device 102, perform various conversions (e.g., color correction, image stabilization, inserting content for augmented reality display, etc.), and provide the result in real time or near real time for display to the user. As a further example, the code may process audio input and perform speech to text or translation services. In general, it should be understood that the code may perform any task or function that may utilize computing resources beyond (or in addition to) those available on the user device 102, or may replace computing resources available on the user device 102 to maintain battery life, processor availability, or other resources. It should also be understood that offloading code execution to the execution environment 136 hosted by the edge server 130 may allow the code to be executed and the response to be provided with significantly lower latency than executing the code in an environment more remote from the user device 102.

[0051] At (8), execution environment 136 sends the result of executing the code to distributed unit 134. In some embodiments, execution environment 136 may send the result in a format accepted and expected by user device 102. In other embodiments, distributed unit 134 (or another component of edge server 130) may convert the result provided by execution environment 136 into a format understood by user device 102. Distributed unit 134 may further format the result according to 5G or other standards for sending information over the air. At (9), distributed unit 134 may send the (formatted) result to radio unit 122, which may send the result to user device 102 over the air interface at (10).

[0052] It should be understood that Figure 2 is provided for purposes of example, and many variations on the depicted interactions are within the scope of the present disclosure. For example, distributed unit 134 may perform the interaction at (5) and select execution environment 136 before obtaining code from the data store. As a further example, as discussed above, another component of edge server 130 or edge server 130 itself may perform any or all of the interactions at (3), (4), (5), (6), and (8). Thus, Figure 2 It is to be construed as illustrative rather than restrictive.

[0053] Figure 31 is a flow chart depicting an example interaction in which a user device 102 may notify a distributed unit 134 that it may request code execution on the distributed unit 134 in the near future, and in which the distributed unit 134 may prepare an execution environment 136 in which the code that the user device may request is to be executed. For example, when the user device 102 is running a particular application that requests one or more code executions on the distributed unit 134, the distributed unit 134 may implement Figure 3 . For example, the interaction may be implemented when the user device 102 is running an autonomous reality (“AR”) application that frequently requests code execution to assist in placing objects in an AR stream. As a further example, the user device 102 may detect that it is running out of memory or storage, has a workload that exceeds the extent that it can handle using available computing resources, or that it should offload work to the distributed unit 134 to preserve its battery life. The interaction begins at (1), where the user device 102 sends a notification to the radio unit 122 that the user device 102 may request code execution on the distributed unit 134 in the near future. In some embodiments, the notification may specify the code that the user device 102 may request. For example, an application may be associated with one or more tasks that may be executed on the distributed unit 134 and stored in the code data storage area 132. Thus, the notification may identify these tasks (or, in some embodiments, may identify the application and thereby allow the distributed unit 134 to identify the task). In some embodiments, the user device 102 may have specific tasks that it offloads to the distributed unit 134 when certain conditions are met. For example, when the user has entered a sufficiently distant destination, the user device 102 may request that the distributed unit 134 execute code that generates a map display that includes the real-time location of the user device 102, such that continually redrawing the map on the user device 102 may drain available battery before the user reaches the destination. At (2), the radio unit 122 sends a notification to the distributed unit 134, which, as discussed above, may be implemented on hardware that is physically proximate to the radio unit 122.

[0054] In some embodiments, distributed unit 134 may determine whether to act on the notification received at (2). For example, distributed unit 134 may determine whether it has an idle execution environment 136 that may be configured in advance to execute code that user device 102 may request. As a further example, distributed unit 134 may evaluate previous executions of a particular application to determine the likelihood that user device 102 will actually make a request, and may configure execution environment 136 for user device 102 only if the likelihood exceeds a threshold. In further embodiments, distributed unit 134 may determine the latency of a regional execution environment, and then determine whether the request from user device 102 can be satisfied in the regional execution environment instead of the execution environment 136 of distributed unit 134. If distributed unit 134 determines not to act on the notification, interactions at (3), (4), (6), and (7) may be omitted.

[0055] At (3), the distributed unit 134 requests the code data store 132 to provide one or more codes for which the user device 102 may request execution. In some embodiments, as described above, the notification may identify multiple tasks that the user device may request to be executed on the distributed unit 134, and the distributed unit 134 may retrieve the code associated with some or all of these tasks. Illustratively, the distributed unit 134 may obtain the code associated with a subset of the tasks based on factors such as the time required to load the task, the maximum acceptable latency for returning the results of executing the task, the frequency or likelihood of the user device 102 requesting execution of the task (e.g., based on previous execution of the application requesting execution of the task, previous requests from the user device 102, or other historical data), or the number of available idle execution environments 136. At (4), the code data store 132 returns the one or more requested codes.

[0056] At (5), the distributed unit 134 (or, in some embodiments, another component of the edge server 130, such as a load balancer or a worker manager) selects one or more execution environments 136 to prepare for execution of the one or more codes obtained at (4). In some embodiments, the distributed unit 134 may cause the code to be loaded into the execution environment 136 so that the environment 136 is ready to execute the code immediately after receiving the request to execute the code. In other embodiments, the distributed unit 134 may provision the execution environment 136. For example, the notification may indicate that a request to execute a particular software image on a virtual machine instance is about to occur. Accordingly, the distributed unit 134 may provision and configure the virtual machine instance with the resources required to execute the software image, and in some embodiments may further load the software image in preparation for executing it. In some embodiments, the distributed unit 134 may determine whether to provision and / or configure the execution environment 136 based on factors such as workload (e.g., whether there are other pending requests to utilize the execution environment 136 on the edge server 130), the predicted amount of time before the user device 102 makes a request, or other factors. In some implementations, implementation of any or all of the interactions at (3), (4), (5), (6), and (7) may be postponed based on the expected arrival time of the request from the user device 102 .

[0057] At (6), the distributed unit 134 sends one or more codes that the user device 102 may request to one or more execution environments 136 (or, in some embodiments, to another component such as a worker manager). In some embodiments, the distributed unit 134 may also send configuration information, such as computing resources to be allocated to the virtual machine instance, the hardware that the virtual machine instance should emulate, etc. At (7), the one or more execution environments 136 are configured to execute the code that the user device 102 may request. In some embodiments, the code may be loaded and partially executed before a request to execute the code is received from the user device 102 (e.g., if there is an initialization step that can be performed without input).

[0058] In some embodiments, the user device 102 may "reserve" the execution environment 136 that it intends to use during the execution of a time-sensitive task on the user device 102, and thus may omit or pre-execute Figure 2 , such as the selection of an execution environment 136 at (5) and the interactions at (3), (4), and (6) for obtaining and loading code into the selected execution environment 136. In further embodiments, as discussed above, the distributed unit 134 may proactively reserve an execution environment 136 in response to the notification, depending on factors such as whether the idle environment 136 is available and is not expected to be used until the user device 102 needs it.

[0059] It should be understood that Figure 3 is provided for illustrative purposes, and many variations on the interactions depicted are within the scope of the present disclosure. For example, the distributed unit 134 may provision the execution environment 136 to perform a number of possible tasks that the user device 102 may request, and may therefore defer loading code for a particular task into the execution environment 136 until a request is received. As a further example, the user device 102 may send the requested code as part of the notification at (1), and may omit the interaction at (3) or (4). Still further, another component of the edge server 130 or the edge server 130 itself may perform any or all of the interactions at (3), (4), (5), and (6). Thus, Figure 3 It is to be construed as illustrative rather than restrictive.

[0060] Figure 4 4 is a flowchart depicting an example routine 400 for processing a request to execute code on a distributed unit according to aspects of the present disclosure. All or part of the example routine 400 may be performed, for example, by Figure 1 400. The example routine 400 begins at block 402, where a request to execute code on a distributed unit may be received from a user device via a radio unit. In some embodiments, as discussed above, the request may be to execute the code and return the result within a specified amount of time, and the routine 400 may determine whether code execution on a distributed unit is required to meet latency requirements. In addition, as discussed above, the request may include the code to be executed or an identifier or other reference that allows the code to be obtained.

[0061] At decision box 404, a determination may be made as to whether resources are available to execute the requested code. In some embodiments, the determination may be as to whether an execution environment is available. In other embodiments, the determination may be as to whether an execution environment can be provided without reducing the performance of the distributed unit at other tasks (such as sending and receiving data from the radio unit). For example, an edge server may provision a pool of execution environments, and then the distributed unit functionality may be implemented in a subset of these execution environments. The amount and / or size of the execution environment required to implement the distributed unit functionality may vary, for example, according to the amount of data to be sent to the radio unit within a given time frame, the amount of data to be received from the radio unit within a given time frame, and the like. In some embodiments, the edge server may reserve an execution environment or computing resources to ensure that the distributed unit functionality is implemented, and then make the remaining environment or resources available for executing the code upon request. If the determination at decision box 404 is that sufficient resources are available, then at box 406, an environment in which the code is executed may be selected. As described above, an environment may be selected based on the code to be executed, which in some embodiments may be a software image to be loaded onto a virtual machine instance having a specific configuration. In other embodiments, at block 406, the execution environment may be provisioned and / or configured.

[0062] At block 408, the code may be executed in the selected execution environment, which may produce an execution result. In some embodiments, as discussed above, the result may be in a format understood by the user device and may be sent to the user device without modification. In other embodiments, the result may require conversion or translation in order to be usable by the user device, and such conversion or translation may be performed. At block 410, the (processed) result is sent to the user device via the radio unit, after which routine 400 ends.

[0063] If the determination at decision box 404 is instead that sufficient resources are not available to execute the requested code, then at decision box 412, a determination may be made as to whether the code can be executed in the region of the cloud provider network. Illustratively, the determination may be as to whether the delay associated with executing the code in the region is below a maximum acceptable delay, which in some embodiments may be included in the request. If the determination is that the code can be executed in the region, then at box 414, the code may be sent to a regional data center, which may execute the code and generate a result that may be received at box 416. Routine 400 then branches to box 410, where the result may be sent to the user device via a radio unit. Alternatively, if the determination at decision box 412 is that the code cannot be executed in the region, then at box 418, the distributed unit may report to the user device that it cannot satisfy the request to execute the code in the requested time frame. In some embodiments, a determination may be made as to when sufficient resources will be available at the distributed unit and whether the result can then be obtained more quickly by waiting or by executing the code in the region. In other embodiments, blocks 412, 414, and 416 may be omitted, and the routine 400 may simply notify the user device 102 that code execution is unavailable. In some embodiments, the routine 400 may instead provide the user device with an estimated time as to when the result will be available, thereby allowing the user device to determine whether to cancel the request (and, in some embodiments, execute the code locally) or wait.

[0064] It should be understood that references to a "distributed unit" in example routine 400 may also refer to other components of the edge server or the edge server itself. That is, in some embodiments, the edge server may implement a distributed unit that receives requests to execute code and a worker manager that implements these requests, and the worker manager may therefore implement portions of example routine 400, such as blocks 404, 406, 408, 412, 414, and 416. It should also be understood that Figure 4 is provided for example purposes, and many variations on the example routine 400 are within the scope of the present disclosure. For example, in some embodiments, the routine 400 may include obtaining code from a data store or other source, as described in more detail above. As a further example, decision block 412 may be implemented before decision block 404, and the computing environment at the distributed unit may only be used when necessary. Therefore, Figure 4 It is to be construed as illustrative rather than restrictive.

[0065] Figure 5 An overall architecture of a computing system (referred to as edge server 130 ) implementing aspects of the present disclosure, such as a distributed unit that executes code on behalf of a user device, is depicted. Figure 5The overall architecture of the edge server 130 depicted in FIG. 1 includes an arrangement of computer hardware and software modules that can be used to implement various aspects of the present disclosure. The hardware modules can be implemented using physical electronic devices, as discussed in more detail below. The edge server 130 may include Figure 5 Those elements shown in more (or fewer) elements. However, it is not necessary to show all of these generally conventional elements in order to provide a feasible disclosure. In addition, Figure 5 The overall architecture shown in can be used to implement Figure 1 One or more of the other components shown in .

[0066] As shown, the edge server 130 includes a processor 502, an input / output device interface 504, a network interface 506, a data storage area 508, a layer 1 processing module 510, and a memory 520, all of which can communicate with each other via a communication bus 512. The network interface 506 can provide access to, for example, Figure 1 502 may be connected to one or more networks or computing systems of other components of the operating environment 100 depicted in FIG. The processor 502 may therefore receive information and instructions from other computing systems or services. The processor 502 may also provide output information to an optional display (not shown) via an input / output device interface 504. The input / output device interface 504 may also receive input from an optional input device (not shown). In some embodiments, the edge server 130 may also include a layer 1 processing module 510. The layer 1 processing module 510 may implement "layer 1" or "physical layer" functions related to the transmission and reception of radio signals, such as encoding digital signals to be sent by the RU, adding redundancy and error correction codes to digital signals, mapping transmission channels to physical channels, and the like. In some embodiments, the layer 1 processing module 510 may be implemented using dedicated hardware (not shown) for signal processing and other layer 1 functions (such as an accelerator card or other hardware).

[0067] The memory 520 may contain computer program instructions (which are grouped into modules in some embodiments) that the processor 502 executes to implement one or more aspects of the present disclosure. The memory 520 typically includes random access memory (RAM), read-only memory (ROM), and / or other persistent, auxiliary, or non-transitory computer-readable media. The memory 520 may store an operating system 522 that provides computer program instructions for use by the processor 502 in the general management and operation of the edge server 130. The memory 520 may also include computer program instructions and other information for implementing various aspects of the present disclosure. For example, in one embodiment, the memory 520 includes an interface module 524 that generates an interface for interacting with other computing devices (e.g., the radio unit 122) via an API, a CLI, and / or a Web interface.

[0068] In addition to and / or in conjunction with the interface module 524, the memory 520 may include the distributed unit 134 and the execution environment 136, which may be executed by the processor 502 to implement various aspects of the present disclosure, such as communicating with the radio unit 122 and executing code on behalf of the user device 102. The memory 520 may also include code 526, which may generally refer to any computer executable instructions that may be executed in the execution environment 136 to produce a result 528. The result 528 is executed in Figure 5 526 is depicted with dashed lines to distinguish that it is not necessarily executable by processor 502. In some embodiments, code 526 is also not executable by processor 502, but is only executable by a virtual processor emulated in one of execution environments 136. Additionally, in some embodiments, memory 520 may also include, for example, latency information, information about previous executions of code 526, notifications from user device 102, or other data.

[0069] Various example embodiments of the present disclosure may be described by the following terms:

[0070] Clause 1. A system comprising:

[0071] a radio unit in communication with an edge server located at an edge location of a cloud provider network, wherein the radio unit controls communications over an air interface with a plurality of mobile computing devices on a radio access network; and

[0072] The edge server located at the edge location, the edge server comprising a processor and a first computer executable instruction, wherein the processor executes the first computer executable instruction to implement a distributed unit, and the distributed unit performs operations including the following:

[0073] receiving, via the radio unit, a code execution request from a first mobile computing device of the plurality of mobile computing devices, the code execution request comprising a request for the edge server to execute second computer-executable instructions;

[0074] selecting a first virtual machine instance from a plurality of virtual machine instances hosted by the edge server to execute the second computer-executable instructions;

[0075] executing the second computer executable instructions in the first virtual machine instance to produce a result; and

[0076] The results are sent via the radio unit to a first mobile computing device.

[0077] Clause 2. The system of Clause 1, wherein the code execution request comprises the second computer-executable instruction.

[0078] Clause 3. The system of Clause 1, wherein the code execution request is generated by an application executing on the first mobile computing device.

[0079] Clause 4. The system of Clause 1, wherein the first mobile computing device generates the code execution request.

[0080] Clause 5. The system of Clause 1, wherein the operations further comprise: determining that the edge server has sufficient available computing resources to execute the second computer-executable instructions in the first virtual machine instance.

[0081] Clause 6. A computer-implemented method comprising:

[0082] receiving, by a distributed unit implemented on an edge server, a request from a mobile computing device to execute code in an execution environment hosted by the edge server, wherein the request is received via a radio unit;

[0083] selecting, by the edge server, a first execution environment from a plurality of execution environments hosted by the edge server to execute the code;

[0084] executing, by the edge server, the code in the first execution environment to generate a result; and

[0085] The results are provided by the distributed unit via the radio unit to a mobile computing device.

[0086] Clause 7. The computer-implemented method of Clause 6, wherein the first execution environment comprises a virtual machine instance.

[0087] Clause 8. The computer-implemented method of Clause 6, wherein the request comprises data to be processed by the code executed in the first execution environment.

[0088] Clause 9. The computer-implemented method of Clause 6, wherein the code comprises one or more of a software image or a serverless computing function.

[0089] Clause 10. The computer-implemented method of Clause 6, further comprising: obtaining the code from a data store.

[0090] Clause 11. A computer-implemented method as described in Clause 6, wherein the request to execute the code in the execution environment hosted by the edge server is generated at least in part based on one or more of: a maximum latency requirement, a battery charge of the mobile computing device, a temperature of the mobile computing device, a workload of the mobile computing device, available memory of the mobile computing device, available processors of the mobile computing device, available storage of the mobile computing device, or a predicted amount of computing resources that will be utilized when executing the code.

[0091] Clause 12. The computer-implemented method of Clause 6, wherein selecting the first execution environment from the plurality of execution environments is based at least in part on a predicted amount of computing resources that will be utilized when executing the code.

[0092] Clause 13. The computer-implemented method of Clause 12, wherein the predicted amount of computing resources is based at least in part on previous executions of the code.

[0093] Clause 14. The computer-implemented method of Clause 6, further comprising: receiving, by the distributed unit, an indication from the mobile computing device that the mobile computing device is executing an application, the application being operable to send the request to execute the code.

[0094] Clause 15. The computer-implemented method of Clause 14, further comprising: in response to receiving the indication that the mobile computing device is executing the application, configuring the first execution environment to execute the code prior to receiving the request.

[0095] Clause 16. The computer-implemented method of Clause 14, further comprising configuring a plurality of execution environments in response to receiving the indication that the mobile computing device is executing the application.

[0096] Clause 17. One or more non-transitory computer-readable media storing computer-executable instructions that, when executed by an edge processor comprising a processor, configure the edge processor to perform operations comprising:

[0097] receiving, via the radio unit, a request to execute code from a mobile computing device;

[0098] selecting a first execution environment from a plurality of execution environments hosted at a location of the edge server;

[0099] executing the code in the first execution environment to generate a result; and

[0100] The results are provided to the mobile computing device via the radio unit.

[0101] Clause 18. The one or more non-transitory computer-readable media of Clause 17, wherein the plurality of execution environments comprises a first plurality of execution environments hosted by the edge server and a second plurality of execution environments in a cloud provider network.

[0102] Clause 19. One or more non-transitory computer-readable media as described in Clause 17, wherein the one or more non-transitory computer-readable media store further computer-executable instructions, which when executed by the edge server configure the edge server to perform further operations, the operations comprising: determining, based at least in part on the request, to execute the code in an execution environment hosted by the edge server.

[0103] Clause 20. The one or more non-transitory computer-readable media of Clause 17, wherein the edge server is physically located near the radio unit.

[0104] Clause 21. One or more non-transitory computer-readable media as described in Clause 17, wherein the one or more non-transitory computer-readable media store further computer-executable instructions, which when executed by the edge server configure the edge server as a distributed unit.

[0105] It should be understood that not all objects or advantages may be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that certain embodiments may be configured to operate in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other objects or advantages that may be taught or suggested herein.

[0106] All of the processes described herein may be embodied in and entirely automatically via a software code module comprising one or more specific computer executable instructions executed by a computing system. The computing system may include one or more computers or processors. The code module may be stored in any type of non-transitory computer readable medium or other computer storage device. Some or all of the methods may be embodied in dedicated computer hardware.

[0107] Many other variations in addition to those described herein will be apparent in light of the present disclosure. For example, depending on the embodiment, certain actions, events, or functions of any of the algorithms described herein may be performed in a different order, added, merged, or completely omitted (e.g., not all of the described actions or events are necessary for the practice of the algorithm). In addition, in certain embodiments, actions or events may be performed concurrently (e.g., by multithreading, interrupt handling, or multiple processors or processor cores or on other parallel architectures), rather than sequentially. In addition, different tasks or processes may be performed by different machines and / or computing systems that may run together.

[0108] The various illustrative logic blocks and modules described in conjunction with the embodiments disclosed herein may be implemented or executed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. The processor may be a microprocessor, but in an alternative, the processor may be a controller, a microcontroller or a state machine, a combination thereof, or the like. The processor may include a circuit configured to process computer executable instructions. In another embodiment, the processor includes an FPGA or other programmable device that performs logic operations without processing computer executable instructions. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. Although this article is primarily described with respect to digital technology, the processor may also primarily include analog components. The computing environment may include any type of computer system, including but not limited to computer systems based on microprocessors, mainframe computers, digital signal processors, portable computing devices, device controllers, or computing engines within appliances (to name a few).

[0109] Unless expressly provided otherwise, conditional language, such as, inter alia, "can," "may," "might," or "could," should additionally be understood within the context as being generally used to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not. Thus, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or steps in any way, or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether such features, elements, and / or steps are included or will be performed in any particular embodiment.

[0110] Unless expressly provided otherwise, connective language such as the phrase "at least one of X, Y, or Z" should otherwise be understood within the context as being generally used to present that an item, term, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that certain embodiments require that at least one of X, at least one of Y, or at least one of Z each be present.

[0111] Any process description, element or box in the flowchart described herein and / or depicted in the accompanying drawings should be understood to potentially represent a module, segment or portion of code including one or more executable instructions for implementing a specific logical function or element in the process. It should be understood by those skilled in the art that alternative implementations are included within the scope of the embodiments described herein, wherein elements or functions may be deleted, executed in an order different from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved.

[0112] Unless expressly provided otherwise, articles such as "a" or "an" should generally be understood to include one or more of the described items. Thus, phrases such as "a device configured to..." are intended to include one or more of the recited devices. Such one or more recited devices may also be collectively configured to perform the stated statements. For example, "a processor configured to perform statements A, B, and C" may include a first processor configured to perform statement A working in conjunction with a second processor configured to perform statements B and C.

Claims

1. A system comprising: a radio unit in communication with an edge server located at an edge location of a cloud provider network, wherein the radio unit controls communications over an air interface with a plurality of mobile computing devices on a radio access network; as well as The edge server located at the edge location, the edge server comprising a processor and a first computer executable instruction, wherein the processor executes the first computer executable instruction to implement a distributed unit, and the distributed unit performs operations including the following: receiving, via the radio unit, a code execution request from a first mobile computing device of the plurality of mobile computing devices, the code execution request comprising a request for the edge server to execute second computer-executable instructions; selecting a first virtual machine instance from a plurality of virtual machine instances hosted by the edge server to execute the second computer-executable instructions; executing the second computer executable instructions in the first virtual machine instance to produce a result; as well as The results are sent via the radio unit to a first mobile computing device.

2. The system of claim 1, wherein the code execution request comprises the second computer-executable instruction.

3. The system of claim 1, wherein the code execution request is generated by an application executing on the first mobile computing device.

4. The system of claim 1, wherein the first mobile computing device generates the code execution request.

5. The system of claim 1, wherein the operations further comprise: A determination is made that the edge server has sufficient available computing resources to execute the second computer-executable instructions in the first virtual machine instance.

6. A computer-implemented method comprising: receiving, by a distributed unit implemented on an edge server, a request from a mobile computing device to execute code in an execution environment hosted by the edge server, wherein the request is received via a radio unit; selecting, by the edge server, a first execution environment from a plurality of execution environments hosted by the edge server to execute the code; executing, by the edge server, the code in the first execution environment to generate a result; as well as The results are provided by the distributed unit via the radio unit to a mobile computing device.

7. The computer-implemented method of claim 6, wherein the first execution environment comprises a virtual machine instance.

8. The computer-implemented method of claim 6, wherein the request includes data to be processed by the code executed in the first execution environment.

9. The computer-implemented method of claim 6, wherein the code comprises one or more of a software image or a serverless computing function.

10. The computer-implemented method of claim 6, further comprising: The code is obtained from the data storage area.

11. The computer-implemented method of claim 6, wherein the request to execute the code in the execution environment hosted by the edge server is generated based at least in part on one or more of: a maximum latency requirement, a battery charge of the mobile computing device, a temperature of the mobile computing device, a workload of the mobile computing device, available memory of the mobile computing device, available processors of the mobile computing device, available storage of the mobile computing device, or a predicted amount of computing resources that will be utilized when executing the code.

12. The computer-implemented method of claim 6, wherein selecting the first execution environment from the plurality of execution environments is based at least in part on a predicted amount of computing resources that will be utilized when executing the code.

13. The computer-implemented method of claim 12, wherein the predicted amount of computing resources is based at least in part on previous executions of the code.

14. The computer-implemented method of claim 6, further comprising: An indication is received by the distributed unit from the mobile computing device that the mobile computing device is executing an application, the application being operable to send the request to execute the code.

15. The computer-implemented method of claim 14, further comprising: In response to receiving the indication that the mobile computing device is executing the application, the first execution environment is configured to execute the code prior to receiving the request.

Citation Information

Cited By

  • Code execution on a distributed unit

    US12541405B2