Application-Centric Design for 5G and Edge Computing Applications

JP7686791B2Active Publication Date: 2025-06-02NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023569917
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-27
Filing Date
2022-04-28
Publication Date
2025-06-02
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

Existing 5G and edge computing technologies lack a coherent approach to manage both compute and network requirements of applications, leading to inefficiencies and suboptimal performance in dynamic environments.

Method used

An application-centric specification and runtime system that integrates compute and network resource management for applications, using an app slice specification and runtime components to ensure seamless operation across multi-layered 5G infrastructure.

Benefits of technology

Ensures that applications receive the necessary compute and network resources dynamically, optimizing performance and adaptability in complex and dynamic 5G environments, supporting real-time and high-bandwidth applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000023_0000
    Figure 00000023_0000
  • Figure 00000024_0000
    Figure 00000024_0000
Patent Text Reader

Abstract

A method is presented for specifying and running an application including multiple microservices in 5G slices in a multi-tiered 5G infrastructure, the method includes determining end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, etc., specifying a slice specification for a function including a network slice specification for the function and a compute slice specification for the function, and managing the compute and network requirements of the application simultaneously using runtime components including a resource manager, an application slice controller, and an application slice monitor, the resource manager maintaining a database and managing starting, stopping, updating, and deleting instances of the application.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 188,639, filed May 14, 2021, U.S. Provisional Patent Application No. 63 / 309,030, filed February 11, 2022, and U.S. Provisional Patent Application No. 17 / 730,499, filed April 27, 2022, the disclosures of which are incorporated herein in their entireties.

[0002] The present invention relates to 5G and edge computing applications, and more particularly to a unified application-centric specification called an app slice that takes into account both the compute and network requirements of an application. [Background technology]

[0003] The advent of 5G and edge computing allows applications to run closer to the source of data, enabling high-bandwidth, low-latency communication between "things" in the Internet-of-Things (IoT) and the edge computing infrastructure on which the applications run. However, 5G and edge computing have evolved independently, and edge computing infrastructure and frameworks, including 5G infrastructure and network capabilities and associated tools, are quite different. There is no coherent approach that considers the compute and network requirements of emerging 5G applications within a single environment. Summary of the Invention

[0004] A method is presented for specifying and running an application including multiple microservices in 5G slices in a multi-tiered 5G infrastructure, the method determining end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connection component, specifying a function slice specification including a function network slice specification and a function compute slice specification, using runtime components including a resource manager, an application slice controller, and an application slice monitor, and managing the compute and network requirements of the application simultaneously by the resource manager maintaining a database and managing starting, stopping, updating, and deleting instances of the application.

[0005] A non-transitory computer readable storage medium is presented that includes a computer readable program for specifying and executing an application including a plurality of microservices of a 5G slice in a multi-tiered 5G infrastructure, the computer readable program, when executed on a computer, causes the computer to determine end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connectivity component, specify a function slice specification including a function network slice specification and a function compute slice specification, and manage concurrently the compute and network requirements of the application using runtime components including a resource manager, an application slice controller, and an application slice monitor, and maintaining a database and managing starting, stopping, updating, and deleting instances of the application via the resource manager.

[0006] A system for specifying and executing an application including multiple microservices of 5G slices in a multi-tiered 5G infrastructure is presented, the system having a memory and one or more processors in communication with the memory configured to determine end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connectivity component, specify a function slice specification including a function network slice specification and a function compute slice specification, manage the compute and network requirements of the application simultaneously using runtime components including a resource manager, an application slice controller, and an application slice monitor, and maintain a database and manage starting, stopping, updating, and deleting instances of the application through the resource manager.

[0007] These and other features and advantages will become apparent from the following detailed description of exemplary embodiments, which is to be read in conjunction with the accompanying drawings.

[0008] In the present disclosure, preferred embodiments are described in detail with reference to the following drawings, as described below. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block / flow diagram of an exemplary app slice specification, in accordance with an embodiment of the present invention.

[0010] [Diagram 2] FIG. 2 is a block / flow diagram of exemplary components of an application specification in accordance with an embodiment of the present invention.

[0011] [Diagram 3] FIG. 3 is a block / flow diagram of an exemplary app slice runtime, in accordance with an embodiment of the present invention.

[0012] [Figure 4] FIG. 4 is a block / flow diagram illustrating a flow chart of a resource manager according to an embodiment of the present invention.

[0013] [Diagram 5] FIG. 5 is a block / flow diagram of an exemplary app slice controller, in accordance with an embodiment of the present invention.

[0014] [Figure 6] FIG. 6 illustrates an exemplary actual application for specifying and executing an application including multiple microservices in a 5G slice in a multi-tiered 5G infrastructure, in accordance with an embodiment of the present invention.

[0015] [Figure 7] FIG. 7 illustrates an exemplary processing system for specifying and executing an application, including multiple microservices, in a 5G slice in a multi-tiered 5G infrastructure, in accordance with an embodiment of the present invention.

[0016] [Figure 8] FIG. 8 is a block / flow diagram of an exemplary method for specifying and executing an application, including multiple microservices, in a 5G slice in a multi-tiered 5G infrastructure, in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Edge Computing is a term that refers to the placement of necessary compute, storage, switching and control functions relatively close to end users and Internet-of-Things (IoT) endpoints. Edge Computing significantly improves performance and the associated quality of experience for users, improving both efficiency and economics. Localizing applications to edge compute closer to the end user improves network transit latency. Latency and reliability are key drivers of performance improvement. Edge Computing enables data localization and efficient data processing. Additionally, industry and government regulations often require data localization for security and privacy reasons.

[0018] For performance reasons, it is often necessary to perform local processing of information to reduce the amount of traffic on transport resources. A generation ago, cloud computing enabled high-value enterprise services that could reach and scale globally, but with latency of minutes or seconds. Today, on-demand, time-shifted HD or 4K video is streamed from the cloud with latency of hundreds of milliseconds. In the future, new applications like haptic internet and virtual reality will require real-time response times of tens of milliseconds or even sub-milliseconds, and will reduce latency not by using cloud resources, but by using computing resources that are close to where the content is created and consumed, in the edge cloud.

[0019] Primary cloud compute environments will continue to operate and will be augmented with edge computing resources. Edge computing provides capabilities that enable the next generation of devices. The most critical data can be stored at the edge and the remaining data can be moved to a centralized facility. This allows edge technology to provide a real-time, high-speed experience to customers and provide flexibility to meet industrial requirements with centralized data storage.

[0020] Edge computing makes any device look and feel like a highly responsive device. Critical data can be processed at the edge of the network, directly on the device. Secondary systems and less urgent data are sent to the cloud for processing there. With Software Defined Networking (SDN), organizations have more flexibility to define rules about where and how data is processed to optimize application performance and user experience.

[0021] Edge computing, combined with 5G promising faster speeds and lower latency, offers a future of near real-time connectivity. Applications that interact with humans in real time require reduced latency between measurement and action. For example, with a response time of around 10 ms, humans can interact with distant objects with no discernible difference compared to interacting with local objects. Where humans expect speed, such as when they want to remotely control a visual scene or issue commands that expect a quick response, faster 1 ms response times are needed. Machine-to-machine communications such as Industry 4.0 require even faster sub-millisecond response times for closed-loop real-time control systems to automate processes such as quality control.

[0022] Moving data processing closer to the edge of the network also has security implications. SDN allows for the development of a layered approach to security that considers communication layer, hardware layer, and cloud security simultaneously. More specifically, in the network edge cloud, network functions virtualization (NFV) enables cloud-level dynamics and flexibility in network implementation. This is a key element to enable dynamic network slicing that will be beneficial for 5G services. Edge clouds are expected to be deployed at various levels of distribution and may be deployed in stages over time. Core data centers currently present in the network will continue to host centralized network functions.

[0023] 5G networks will enable integrated communication technology for a networked world. 5G will target a wide range of applications across various sectors, including industrial manufacturing, automotive, transportation, agriculture, healthcare, etc. 5G will natively support machine-to-machine communication and IoT connectivity, which has great potential to transform society. For example, the emergence of Industry 4.0 has given rise to several applications with extensive requirements to connect people, objects, processes, and systems in real time. Industry 4.0 requires networks across a wide range of industrial sectors, including manufacturing, oil and gas, power generation / distribution, mining, and chemical processing. Such networks are significantly different from traditional enterprise / consumer networks in terms of service requirements.

[0024] Although key connectivity requirements in terms of latency and throughput vary widely, 5G capabilities will enable a wide range of industrial applications such as remote operation, remote maintenance, augmented reality, and mobile workforce, as well as enterprise applications such as payment, haptics, V2X, and real-time monitoring. Often, these applications have latency requirements of less than 0.5-10 ms, very high data rate capacity requirements on the order of 10-1000 Mbps, and high density requirements on the scale of thousands of nodes. V2X applications have high reliability and very low latency requirements as vehicles need to make live or dead decisions while moving at high speeds.

[0025] Network slicing unlocks the potential of 5G for various verticals. Before the 5G era, cellular networks had an approach for a one-size-fits-all solution. The key principle behind network slicing is to instantiate multiple logical networks on a common physical fabric, with each logical network tailored to the individual requirements of the application. A network slice is a collection of network functions and specific radio access technology configurations customized for a specific use case. Slices are realized on a common infrastructure that shares compute, network and spectrum licenses. This allows for efficient utilization of infrastructure and resources, resulting in cost- and energy-efficient implementation. Network slicing achieves independence from a business, technical, functional and operational perspective. Network slicing can be seen as a means to create dedicated networks with predefined quality of service within a network to deliver new generation services. In simple terms, network slicing can be seen as a dedicated and independent private 5G network within a public 5G network. Slicing provides the ability to isolate traffic end-to-end, enabling full performance guarantees in multi-tenant and multi-service conditions. Network slicing also provides independence in terms of computing, storage and network resources.

[0026] Slice-based abstraction of emerging applications is key to achieve operational requirements in terms of real-time, reliability and responsiveness. Exemplary embodiments employ a real-time surveillance video analytics application with high throughput, low latency and reliability constraints for effective performance. Application requirements in terms of latency, bandwidth and reliability often change dynamically, affecting both network and compute requirements. Dynamic fine-tuning of network and compute parameters allows the underlying platform of the service to be constantly customized to meet changing needs. In 5G, two mechanisms for network slicing are specified. The first is soft network slicing and the second is hard network slicing. Soft network slicing is based on quality of service (QoS) techniques that dynamically allocate available network resources to different classes of traffic. In the case of long-term evolution (LTE), this is mainly achieved by user equipment assigning a QoS class index (QCI) to each traffic class, and in the case of 5G, it is achieved by using 5G QOS Identifier (5QI). Hard network slicing achieves slicing by using virtualization and functional separation of components.

[0027] Applications require compute and network resources to execute various application functions. Currently, network and compute resources are treated and managed independently. There is no coherent approach to consider them together for the benefit of the entire application. Networking vendors provide network resource guarantees without considering the compute requirements of the application and orchestration frameworks such as Kubernetes guarantee the provision of compute resources without considering the network requirements of the application. Furthermore, network resource guarantees are application agnostic while compute resource guarantees are made within a specific tier in a layered computing tier architecture. This siloed approach to compute and network resources does not work well for applications that require compute and network resources to be optimized together to enable the overall health and smooth running of applications within and across computing tiers.

[0028] Data needs to travel across the network at the speed and reliability required by applications, and at the same time, sufficient compute resources need to be available to process this data in real time to realize various application functions. If compute and network resources are treated independently, i.e., network resources are sufficient and data can flow through the network, but compute resources are insufficient to process the data, or compute resources are abundant but network resources are insufficient to move the network data, in both cases the application will be affected and unable to provide its function. The compute and network resource requirements of an application need to be statically identified, such as by profiling the application, and allowed for by the application even before it starts executing.

[0029] In addition to the static allocation of compute and network resources, it is also necessary to continuously monitor the operation of the application at run time to determine whether the statically allocated resources are sufficient to provide the application's functions. If they are not sufficient, such as due to changes in operating conditions, the static allocation of resources must be readjusted so that the application dynamically receives sufficient compute and network resources to continue operating smoothly in response to the new operating conditions. This dynamic adjustment of resources is important for the application and must consider both compute and network resources.

[0030] Therefore, a top-level abstraction is needed to have a unified view and simultaneously manage the compute and network resource requirements of an application. The top-level abstraction is called an app slice and takes into account the compute and network resource requirements of an application. Exemplary embodiments enable coherent specification and runtime of app slices that take into account and combine the compute requirements in the compute slices and the network requirements in the network slices.

[0031] Applications can be developed using monolithic or microservices-based architectures. In a monolithic architecture, the entire application is developed and deployed as a single entity, while in a microservices-based architecture, the application is broken down into smaller entities, or tasks or microservices, that are developed and deployed independently, and then interconnected to provide the functionality of the whole application. The App Slices specification is designed to cover both these types of architectures.

[0032] 1 shows an app slice specification 100 that includes a slice specification 105 for a top-level application, and if the application is split into smaller functions (microservices as functions), it includes compute and network slice specifications for each function (function slice specification 110). In a monolithic architecture there is only a single function, but in a microservices architecture there can be many functions.

[0033] With respect to the slice specification 105 of an application, this portion of the app slice specification 105 can be used to specify desired end-to-end application characteristics.

[0034] Spec 105 has four parameters:

[0035] With regard to latency parameters, each application 101 has a specific end-to-end latency requirement; that is, output should be returned within a specific amount of time. "Latency" in this case includes not only processing time, but also time spent in the network. This means that the total time it takes for data to be generated, sent over the network for processing, for the actual processing (computation) to occur, and for the output to be returned (again over the network) for one unit of work determines the end-to-end application latency. This desired "latency", specified in milliseconds, is the maximum tolerable end-to-end latency for the application. If the latency is higher than the specified value, it is of no use for the application's output.

[0036] Regarding the bandwidth parameter, based on the network characteristics of the application 101, it may require a certain amount of bandwidth. The bandwidth required by the application 101 is specified by this parameter and is expressed in kilobits per second (kbps).

[0037] Regarding the deviceCount parameter, the connection density of the application 101 is specified using this parameter. Connection density includes the total number of other devices to which the application 101 connects.

[0038] Regarding the confidence parameter, the confidence level regarding the application resource requirements is specified using this parameter. The value is between 0 and 1, where 0 is low confidence and 1 is fully confident.

[0039] These application-level slice specifications are translated into various types of slicing in 5G, such as "eMBB", "uRLLC" or "mMTC". The "eMBB" (enhanced Mobile Broadband) slice type is for applications that require high communication bandwidth. The "uRLLC" (ultra Reliable Low Latency Communications) slice type is for applications that require low latency and high reliability. The "mMTC" (Massive Machine Type Communications) slice type is for applications with high connection density.

[0040] With respect to a function's network slice specification 110, each function 111 requires certain network characteristics in order to continue to operate properly without degrading the quality of the output it produces. In particular, this applies to the data it receives. If the input data is received according to the function's needs, processing will be performed as required and output will be generated appropriately. These network characteristics required by the input side function 111 are specified as part of the function's network slice specification 112.

[0041] There are a total of four network parameters that form part of the network slice specification 112 of a function.

[0042] Regarding the latency parameter, this parameter specifies the maximum acceptable latency in milliseconds. This is the time the function expects to receive a packet, and if it fails to do so, the correctness of the output generated by the function is not guaranteed. It is OK if the actual latency is lower than this desired latency, but it cannot exceed it. In fact, the lower the latency is than the desired value, the better for the function 111.

[0043] With regard to the throughputGBR parameter, the function requires that the input data stream arrives at a certain rate, and this is the desired throughput (specified in kbps) that needs to be guaranteed in order for the function to perform properly. (GBR stands for Guaranteed Bit Rate.) This desired throughput is particularly useful in the case of streaming input data, where there is a continuous data stream that the function receives and needs to process at a certain rate in order to keep up with the incoming input stream and produce the correct output.

[0044] Regarding the throughputMBR parameter, this parameter specifies the maximum throughput (MBR stands for maximum bitrate) that the function can consume. Anything higher will not be used by the function.

[0045] Regarding the packetErrorRate parameter, one important aspect of network performance is how reliably the network can forward packets. The "packetErrorRate" parameter is the ratio of the number of packets received in error to the total number of packets received. Some functions 111 can tolerate packet errors at a certain rate, while other functions can tolerate packet errors at a different rate. The rate that a function can tolerate is specified using this parameter.

[0046] With respect to the function's compute slice specification 114, along with the network characteristics, the function 111 must also have certain compute characteristics that must be met in order for the function to execute well. If there are not enough resources available for the computation, the function will not execute properly even if the network characteristics are met. Therefore, both the network and compute requirements of the function 111 must be taken into account to achieve an overall smooth operation. This part of the slice specification is for the compute slices required by the function 111.

[0047] There are a total of five compute parameters that form part of a function's compute slice specification:

[0048] For the minCPUCores parameter, CPU resources are specified in terms of full CPU units. 1 represents either 1 vCPU / core on the cloud, or 1 hyperthread on a bare metal Intel processor. 1 CPU unit is divided into 1000 "millicpu", and the finest granularity that can be specified is "1m" (1millicpu). The "minCPUCores" parameter specifies the minimum number of CPU cores required by function 111. It is guaranteed as a function similar to "throughput GBR", the guaranteed bit rate of the network. "minCPUCores" can be specified as a decimal between 0 and 1, or as a number of millicpu or millicore. Specifying 100m is the same as specifying 0.1 for this parameter.

[0049] Regarding the Maximum CPU Cores (maxCPUCores) parameter, this parameter specifies the maximum number of CPU cores that can be used by the function 111. Any CPU resources exceeding this limit cannot be used by the function 111. This is similar to the "Throughput MBR", which is the maximum bit rate that the function 111 can consume. The unit of specification for "Maximum CPU Cores" is the same as that for "Minimum CPU Cores". In other words, it can be specified as a decimal between 0 and 1, and can also be specified in millipath units. Specifying 0.5 is the same as specifying 500m.

[0050] With regard to the minMemory parameter, memory resources are specified as bytes (a simple number) or as a fixed-point number with one of these subscripts E, P, T, G, M, K, or as the equivalent powers of two Ei, Pi, Ti, Gi, Mi, Ki. The parameter "minMemory" specifies the minimum amount of memory that the function 111 requires. If there is less memory available, the function 111 will not execute properly and may even crash. Therefore, to avoid this scenario, the function 111 can specify with this parameter the minimum amount of memory it requires to operate properly. Specifying 500M is roughly equivalent to specifying 500000000 (bytes) or 476.8MiB (mebibytes).

[0051] Regarding the maxMemory parameter, the maximum amount of memory that can be used by the function is specified by this parameter. The units are similar to "minMemory". Specifying 800M is roughly equivalent to specifying 800000000 (bytes) or 762.9 (mebibytes).

[0052] Regarding the tier parameter, this is an optional parameter that can be specified if the function must be executed at a specific tier in the computing fabric. It can have one of three values: "device", "edge" or "cloud". Its default value is "auto", indicating that the function 111 can be executed anywhere in the computing fabric. However, if this is not the case, this parameter can be used to specify the exact location of the tier in which the function 111 should be executed.

[0053] It should be noted that the tiers parameter of the Compute Slice specification 114 provides the ability to automatically map and execute functions 111 across multiple tiers. This type of functionality is not available out of the box in orchestration frameworks like Kubernetes, so additional consideration is required when mapping and executing functions across tiers in a computing stack.

[0054] It is a common programming paradigm to split individual functions of an application into microservices and then combine and interconnect the microservices to achieve the overall functionality of the application. Each microservice is called a function 111, and an application 101 can contain several interconnected functions 111.

[0055] The various components 200 of an application specification 100 are shown in Figure 2. First, an identifier for the application, called the application ID, is specified. This ID maps to a specific application and is used internally by the runtime system to obtain details about the application. Next, the name of the application is specified. Other metadata related to the application is then specified. This metadata may include the version number of the application, descriptions related to the application, URLs where details about the application can be found, the operating system and architecture on which the application runs, who maintains the application, etc. During instantiation, an instance of the application is created, which contains instances of each of its functions. Dependencies of these functions, function instances and instance connections are then specified.

[0056] The function dependency specification includes the various functions 202 that make up the application. For each function, a function ID, which is an identifier of the function, and a version number of the function are specified. The function instance specification includes the various function instances 204 that need to be spawned as part of the application. For each instance, the instance name, the function ID corresponding to the instance, and the spawn type of the instance need to be specified. The spawn type of the instance can be one of five spawn types such as new, reuse, dynamic, uniqueNodewide, and uniqueSitewide.

[0057] Each of these spawn types is explained below.

[0058] Regarding the "new" spawn type, the runtime system always creates a new instance of the function when this spawn type is specified.

[0059] For a "reuse" spawn type, the runtime system first checks whether there is another instance of the function that has already been executed with the same configuration. If so, the runtime system reuses that instance while the application is running. If no instance matching the configuration is found, a new instance is created by the runtime system.

[0060] For a "dynamic" spawn type, the runtime system does not create this instance at the start of the application, but rather the instance is created dynamically after the application has already begun execution.

[0061] For a "unique node-wide" spawn type, the runtime system first checks whether there are other instances of the function already running on the specified node / machine with the same configuration. If there are no other instances already running on the specified node / machine that match the instance's configuration, the runtime system creates a new instance. If there is an instance already running on the node / machine that matches the instance's configuration, the runtime system uses that instance during the execution of the application. For instances with this spawn type, only one instance of the function is created to run on the particular node.

[0062] For a "unique site-wide" spawn type, the runtime system first checks whether there is another instance of the function already running. If there is, the runtime system uses that instance during the execution of the application. If there is no instance already running, a new instance is spawned and started. For instances with this spawn type, only one instance of the function is spawned and runs across the site-wide distribution. The instance connection specification contains the connections between the various function instances. For each connection, the source instance, destination instance and binding information are specified, i.e., whether a binding of source or destination instances is specified. For each source instance and destination instance, the instance name and the connection endpoint name are specified.

[0063] After the app specification and app slice specification are written, the actual realization and execution is handled by the app slice runtime. The runtime 300 shown in FIG. 3 is located in the underlying compute and network infrastructure and is integrated with the application itself. The inputs to the runtime are the application specification and application slice specification 302, and the application slice configuration 304 used in the application instance and associated slice. Using these as inputs, the runtime system 300 with knowledge of the underlying infrastructure manages the creation or generation of the application instance with the provided configuration, the creation or generation of the appropriate slice with the requested configuration, allocates the requested compute and network resources to the individual function instances, schedules the instances at the appropriate tier with the appropriate slice, and monitors and ensures the overall smooth operation of the individual functions and the entire application. The runtime has three components: a resource manager 310, an app slice controller 312, and an app slice monitor 314.

[0064] The Resource Manager (RM) 310 is the heart of the runtime system 300 which manages the actual realization and execution in cooperation with slice controllers 312 and slice monitors 314. Application and slice specifications are received by the RM 310 and all requests to start, stop or update an instance of an application are also received by the RM 310. The RM 310 maintains a database 305 where all application and slice specifications, the configuration of the various instances, their status, details of the underlying compute and network infrastructure, etc. are stored.

[0065] 4 shows a flow chart 400 showing the procedure for RM 310 of any input. When an input 402 arrives, RM 310 first checks whether the input is for an application or slice specification or configuration (404). If it is a specification, that particular specification is stored in the database (406). There is no further action on the input and the procedure ends. If the input is for a configuration, the corresponding action is taken (408).

[0066] When starting or updating an application, the RM 310 checks whether the necessary compute and network resources requested in the configuration (410) are available in the underlying infrastructure. If they are available, the corresponding resources are allocated to the various function instances and the instances are scheduled to run (412). To run an instance of the application, the RM 310 retrieves the application specification from the database, creates or spawns all function instances based on the spawn type, creates all specified connections between the various instances, and finally allocates resources to these instances and schedules them to run in the underlying infrastructure. This is then updated in the database (416) terminating the procedure. If the action is to stop or delete, the corresponding function instances are stopped or deleted (414) and their status is updated in the database terminating the procedure. [Table 1]

[0067] When checking the availability of resources, RM310 first checks the application level slice specification, then for each individual function, RM310 follows the algorithm shown in Algorithm 1 above. Each function forming the application is checked one by one for the availability of resources in any of the tiers. These tiers are sorted such that the cheaper tiers are checked first, followed by the more expensive tiers. Thus, for each function, the compute resources (denoted by c_r) and network resources (denoted by n_r) requested are checked with the corresponding compute resources (denoted by tc_r) and network resources (denoted by tn_r) in the tier. All the parameters indicated in the compute slice specification (min CPU cores, max CPU cores, min memory, max memory and tier) are considered with the compute resource requirements, and all the parameters indicated in the network slice specification (latency, throughput GBR, throughput MBR and packet error rate) are considered with the network resource requirements. If the requested resources are less than the available resources, the resources (compute and network) of that tier are assigned to the function. For functions where a tier is explicitly specified and not automatic, resource availability is checked for that particular tier only, all other tiers are ignored. This is repeated for each function in all tiers, and the cheapest tier that meets the resource requirements of the function is assigned to that function. If the resource requirements for an application and all associated functions cannot be met, the RM 310 reports this, causes the application and associated functions to take appropriate action, and updates the DB accordingly. [Table 2]

[0068] As various functions are executed, RM310 periodically monitors the conditions of these functions and adjusts resources as necessary. To do this, RM310 checks all running functions at configurable intervals of seconds according to Algorithm 2 above. Specifically, it checks whether the resource requirements of the functions are met by the compute and network resources of the assigned layer. If for any reason, for example, due to changes in operating conditions / input content, network interruptions or hardware failures, the network or compute resources are found to be insufficient, RM310 attempts to find additional resources.

[0069] Again, as before, cheaper tiers are checked before more expensive tiers, and the cheapest tier that can satisfy the function's resource requirements is assigned to the function and the function is scheduled to run on the resources of this newly found tier. If there are no resources available in any tier, RM 310 reports this as an error for the particular function and leaves it up to the function to take appropriate action. RM 310 not only checks whether additional resources are required, but also checks whether too many resources have been allocated due to previously changed conditions and reduces resources if the conditions change again and less resources are now required. In such a case, RM 310 reduces the overall compute and network resource usage. Thus, RM 310 dynamically monitors and adjusts the compute and network resources of the function to ensure smooth operation. As a result, RM 310, in cooperation with application slice controller 312 and application slice monitor 314, performs static resource management at first and then dynamic resource management across tiers.

[0070] Note that at any point in time, a function is always provided with the compute and network resource requirements specified in the original specification. Additional resources are only granted and dynamically scaled as needed. The RM 310 communicates with the application slice controller 312 to configure the compute and network slices and execute the function on an underlying orchestration platform such as Kubernetes.

[0071] With respect to the application slice controller (ASC) 312 shown in FIG. 5, the ASC 312 manages slicing including compute and network slicing of the function according to instructions from the RM 310. When the RM 310 signals the ASC 312 to create a network slice, the ASC 312 communicates with the network slice interface 502 to create the network slice in the underlying network infrastructure 512. Since existing network vendors such as Celona do not provide admission control to allow the creation of a network slice, the exemplary method built a custom layer that operates on the Celona API and provides guarantees and admission control before allowing the creation of the network slice. This can lead to underutilization of the network if the actual usage is less than the requested usage, but the exemplary method requires this to guarantee the network. The exemplary method exposes this custom layer as a network slice interface of the ASC 312. Thus, by going through this customer's network slice interface layer, the ASC 312 creates a network slice that meets the requirements of the function including latency, throughput GBR, throughput MBR, and packet error rate. Based on these network requirements, appropriate QCI levels and priorities are selected to create network slices that meet the network requirements of the function.

[0072] Once ASC312 receives a signal to create a compute slice, it uses the underlying orchestration platform's functions via the compute slice interface 504 to associate the function's compute requirements with the underlying compute infrastructure 514. In particular, the minimum CPU cores, maximum CPU cores, minimum memory, and maximum memory are used to set the compute "request" and "limit" of the corresponding function container running on an orchestration platform such as Kubernetes, which provides admission control before granting the requested resources. In addition to creating these compute slices and network slices, ASC312 also manages the update and deletion of these slices. In case of a request to update or delete a network slice, ASC312 communicates with the underlying network slice interface to update or delete the particular network slice. In case the request is to update or delete a compute slice, ASC312 communicates with the orchestration platform's compute slice interface to update or delete the particular compute slice.

[0073] With respect to the App Slice Monitor (ASM), the ASM 314 monitors and maintains collection of various metrics for the compute slices and network slices created by the ASC 312. These metrics are made available to the RM 310 periodically, at specific configurable intervals, and on-demand, and are used by the RM 310 to make resource allocation and scheduling decisions. To obtain network slice metrics, the ASM 314 communicates with the network slice interface to collect metric data for individual network slices running in the system. To obtain compute slice metrics, the ASM 314 communicates with the orchestration platform's compute slice interface to collect metric data for individual compute slices running in the system. This network and compute slice data includes requested resources, currently used resources, overall usage history, and abnormal usage behavior. Such data helps the RM 310 to dynamically allocate resources and make scheduling decisions for already running functions, as needed.

[0074] In conclusion, the exemplary embodiments of the present invention present a unified application-centric specification called App Slice that considers both the compute and network requirements of an application. To realize this App Slice specification, the exemplary method proposes a novel App Slice runtime that ensures that an application always receives the required compute and network resources. The exemplary invention, together with the App Slice specification and runtime, helps to leverage new 5G applications in a multi-tiered, complex and dynamic 5G infrastructure.

[0075] Exemplary embodiments of the present invention further provide:

[0076] A system and method for specifying and executing an application that includes multiple microservices / functions in a 5G slice within a complex, dynamic, multi-tiered 5G infrastructure.

[0077] A system and method for specifying application level requirements and individual function level requirements that take into account network slice requirements and compute slice requirements.

[0078] A system and method for specifying an application structure including various functions to be utilized for execution in a 5G slice along with compute and network slice requirements, how they need to be executed, and what their interconnections are.

[0079] A system and method for actually implementing and executing specifications within a complex and dynamic 5G infrastructure using a runtime component, where application structure, application level and function level requirements, and application configuration are provided as inputs to the runtime system.

[0080] A system and method for processing various inputs using a resource manager that maintains a database and manages starting, stopping, updating and deleting instances of an application.

[0081] A system and method for checking application and function level requirements, including network slice and compute slice requirements, across various tiers in a multi-tier computing and networking fabric and, if available, allocating resources incrementally from the cheapest tier to the most expensive tier while ensuring that the requirements set forth in the specification are met.

[0082] A system and method for reporting back to an application when requirements cannot be met by the underlying compute and network infrastructure, thereby allowing the application to take appropriate action.

[0083] A system and method for periodically monitoring an application and dynamically adjusting the allocation of compute and network resources in case they are found to be insufficient for any reason (changing operating conditions / input content, network interruption, hardware failure, etc.) to ensure smooth end-to-end operation of the entire application.

[0084] A system and method for allocating cheaper tiers before more expensive tiers (while ensuring requirements are met) during dynamic adjustments of compute and network resources.

[0085] A system and method that exposes a unified layer (app slice controller) that interfaces with compute and networking infrastructure to manage network and compute slices within 5G infrastructure.

[0086] A system and method for exposing a unified layer to resource managers to simplify the processing of compute and network slice requests.

[0087] A system and method for monitoring a network, computing slices, and producing various metrics (such as requested resources, currently used resources, overall usage history, and anomalous usage behavior) that can be utilized by a resource manager to make dynamic resource allocation and scheduling decisions.

[0088] FIG. 6 is a block / flow diagram 800 of an actual application for specifying and executing an application including multiple microservices in a 5G slice in a multi-tiered 5G infrastructure in accordance with an embodiment of the present invention.

[0089] In one practical example, a facial recognition based video analytics application called real-time monitoring or watchlist is shown along with its app slice and application specification. Real-time monitoring applications can improve safety, security, and operational efficiency by governments and organizations leveraging face matching capabilities. The application can quickly and with high confidence identify known and unknown individuals under real-world challenges such as lighting, angle, facial hair, posture, glasses and other occlusions, motion, crowds, facial expressions, etc. The various components / functionality of this application along with the pipeline are shown in Figure 6.

[0090] In one practical application, a video feed from a camera 802 is decoded by a "video sensor" 804 and frames are made available to a "face detection" component 806, which detects faces 808 and makes them available to a "feature extraction" component 810. Unique face templates such as features are then extracted and made available to a "face matching" component 812, which compares and matches these features with a gallery of facial features 814 obtained by a "biometrics manager" component 816. All matching results are sent to an "alert manager" component 818 for storage 820 and also made available to any third party applications.

[0091] FIG. 7 is an exemplary processing system for specifying and executing an application, including multiple microservices of a 5G slice in a multi-tiered 5G infrastructure, in accordance with an embodiment of the present invention.

[0092] The processing system includes at least one processor (CPU) 904 operatively connected to other components via a system bus 902. Also operatively connected to the system bus 902 are a GPU 905, a cache 906, a read only memory (ROM) 908, a random access memory (RAM) 910, an input / output (I / O) adapter 920, a network adapter 930, a user interface adapter 940, and / or a display adapter 950. Additionally, the app slice 950 includes an application slice specification 952 and a function slice specification 954.

[0093] The storage devices 922 are operably connected to the system bus 902 by the I / O adapter 920. The storage devices 922 may be any type of disk storage device (e.g., magnetic disk storage device or optical disk storage device), solid state magnetic device, or the like.

[0094] The transceiver 932 is operably connected to the system bus 902 by the network adapter 930 .

[0095] User input device(s) 942 are operatively connected to system bus 902 by user interface adapter 940. User input device 942 may be any of a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, or a device incorporating the functionality of at least two of these devices. Of course, other types of input devices may be used while maintaining the spirit of the principles of the present invention. Of course, other types of input devices may be used while maintaining the spirit of the present invention. User input device(s) 942 may be the same type of user input device or may be different types of user input devices. User input device 942 is used to input information to and output information from the processing system.

[0096] A display device 952 is operatively connected to the system bus 902 by a display adapter 950 .

[0097] Of course, the processing system may include other elements (not shown) or omit certain elements, as would be readily apparent to one of ordinary skill in the art. For example, the processing system may include various other types of input and / or output devices, depending on the particular implementation, as would be readily apparent to one of ordinary skill in the art. For example, various wireless and / or wired input and / or output devices may be used. Furthermore, various configurations of additional processors, controllers, memory, etc. may be used, as would be readily apparent to one of ordinary skill in the art. These and other variations of the processing system will be readily apparent to one of ordinary skill in the art from the teachings of the present principles provided herein.

[0098] FIG. 8 is a block / flow diagram of an exemplary method for specifying and executing an application, including multiple microservices of a 5G slice in a multi-tiered 5G infrastructure, in accordance with an embodiment of the present invention.

[0099] The compute and network requirements of an application are managed simultaneously by:

[0100] In block 1001, end-to-end application characteristics are determined by using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connection component.

[0101] At block 1003, a slice specification for the function is specified, which includes a network slice specification for the function and a compute slice specification for the function.

[0102] In block 1005, runtime components including a resource manager, an application slice controller, and an application slice monitor are used, with the resource manager maintaining a database and managing starting, stopping, updating, and deleting instances of the application.

[0103] As used herein, the terms "data," "content," "information," and similar terms may be used interchangeably to refer to data that may be obtained, transmitted, received, displayed, and / or stored by various exemplary embodiments. Thus, the use of these terms should not be construed as limiting the spirit and scope of the disclosure. Additionally, where a computing device is described herein for receiving data from another computing device, the data may be received directly from the other computing device or indirectly via one or more intermediate computing devices, such as one or more servers, relays, routers, network access points, base stations, and the like. Similarly, where a computing device is described herein for transmitting data to another computing device, the data may be transmitted directly to the other computing device or indirectly via one or more intermediate computing devices, such as one or more servers, relays, routers, network access points, base stations, and / or the like.

[0104] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may be generally referred to herein as a "circuit," "module," "computer," "apparatus," or "system." Additionally, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code thereon.

[0105] Any combination of one or more computer readable media may be used. The computer readable medium may be a computer readable signal medium or a computer readable recording medium. The computer readable recording medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of computer readable recording media, including but not limited to, include one or more wires, portable computer diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fibers, portable compact disc read-only memories (CD-ROMs), optical data storage devices, magnetic data storage devices, or any suitable combination of the foregoing. In the context of this document, a computer readable recording medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus or device.

[0106] A computer-readable signal medium may include a propagated data signal in which computer-readable program code is embodied, for example in baseband or as part of a carrier wave. Such a propagated signal may be in any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable recording medium, but may be any computer-readable medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, device, or apparatus.

[0107] The program code embodied in the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0108] Computer program code for carrying out processes related to aspects of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0109] Aspects of the present invention are described below with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to generate a machine such that the instructions, executed by a processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks or modules of the flowcharts and / or block diagrams.

[0110] These computer program instructions can be stored on a computer-readable medium that can instruct a computer, other programmable data processing device, or other device to function in a particular manner, such that the instructions stored on the computer-readable medium generate a product including instructions for implementing the functions / acts specified in one or more blocks or modules of the flowcharts and / or block diagrams.

[0111] The computer program instructions can also be loaded into a computer, other programmable data processing apparatus or other device to generate a computer-implemented process such that a series of operational steps are executed on the computer, other programmable apparatus or other device, and the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / operations specified in the blocks or modules of the flowcharts and / or block diagrams.

[0112] As used herein, the term "processor" is intended to include any processing device, such as one that includes a central processing unit (CPU) and / or other processing circuitry. It should also be understood that the term "processor" may refer to one or more processing devices, and that various elements associated with a processing device may be shared by other processing devices.

[0113] The term "memory" as used herein is intended to include memory associated with a processor or CPU, such as, for example, RAM, ROM, fixed memory devices (e.g., hard drives), removable memory devices (e.g., diskettes), flash memory, etc. Such memory may be considered a computer-readable recording medium.

[0114] Furthermore, the term "input / output device" or "I / O device" as used herein is intended to include, for example, one or more input devices (e.g., keyboard, mouse, scanner, etc.) for inputting data into a processing unit and / or one or more output devices (e.g., speakers, displays, printers, etc.) for presenting results associated with a processing unit.

[0115] The foregoing should be understood in all respects as illustrative and exemplary, and not restrictive, and the scope of the invention disclosed herein should be determined not from the detailed description, but from the claims which are to be interpreted in accordance with the broadest possible interpretation permitted by the Patent Law. It should be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and that various modifications may be made by those skilled in the art without departing from the scope and spirit of the invention. Various other feature combinations may be implemented by those skilled in the art without departing from the scope and spirit of the invention. Although aspects of the invention have been described above with the fine detail and particularity required by the Patent Law, the scope of the claims which are sought to be protected by Letters Patent are set forth in the appended claims.

Claims

1. 1. A method for specifying and executing an application, the application including a plurality of microservices of a 5G slice in a multi-tiered 5G infrastructure, the method comprising: determining 1001 end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connection component; Specifying a slice specification for the function, including a network slice specification for the function and a compute slice specification for the function (1003); A method for simultaneously managing the compute and network requirements of an application using runtime components including a resource manager, an application slice controller, and an application slice monitor, the resource manager maintaining a database and managing (1005) starting, stopping, updating, and deleting instances of the application.

2. The method of claim 1 , wherein the application slice specifications include a latency parameter, a bandwidth parameter, a device count parameter, and a reliability parameter.

3. The method of claim 1 , wherein the network slice specifications of the function include a latency parameter, a throughput GBR parameter, a throughput MBR parameter, and a packet error rate parameter.

4. The method of claim 1 , wherein the compute slice specifications for the function include a minimum CPU core parameter, a maximum CPU core parameter, a minimum memory parameter, a maximum memory parameter, and a tier parameter.

5. 5. The method of claim 4, wherein the layer parameters automatically map and execute functions across multiple layers, and the resource manager initially performs static resource management and then cooperates with the application slice controller and the application slice monitor to perform dynamic resource management across layers.

6. The method of claim 1 , wherein the application slice controller manages compute slicing and network slicing for functions by using a network slice interface layer that provides guarantees and admission control prior to the creation of network slices.

7. 7. The method of claim 6, wherein the application slice monitor monitors and collects metrics regarding the compute slicing and the network slicing generated by the application slice controller, the metrics being made available to the resource manager periodically at specific configurable intervals.

8. A non-transitory computer-readable recording medium including a computer-readable program for specifying and executing an application, the application including a plurality of microservices of a 5G slice in a multi-tiered 5G infrastructure, When the computer-readable program is executed on a computer, the computer determining 1001 end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connection component; Specifying a slice specification for the function, including a network slice specification for the function and a compute slice specification for the function (1003); A non-transitory computer-readable storage medium using runtime components including a resource manager, an application slice controller, and an application slice monitor, the resource manager maintaining a database and managing (1005) starting, stopping, updating, and deleting instances of the application, thereby simultaneously managing the compute and network requirements of the application.

9. The non-transitory computer-readable medium of claim 8 , wherein the application slice specifications include a latency parameter, a bandwidth parameter, a device count parameter, and a reliability parameter.

10. The non-transitory computer-readable medium of claim 8 , wherein the network slice specifications of the function include a latency parameter, a throughput GBR parameter, a throughput MBR parameter, and a packet error rate parameter.

11. The non-transitory computer-readable medium of claim 8 , wherein the compute slice specifications for the function include a minimum CPU core parameter, a maximum CPU core parameter, a minimum memory parameter, a maximum memory parameter, and a tier parameter.

12. 12. The non-transitory computer-readable storage medium of claim 11, wherein the layer parameters automatically map and execute functions across multiple layers, and the resource manager initially performs static resource management and then cooperates with the application slice controller and the application slice monitor to perform dynamic resource management across layers.

13. 9. The non-transitory computer-readable storage medium of claim 8, wherein the application slice controller manages compute slicing and network slicing for functions by using a network slice interface layer that provides guarantees and admission control before the creation of network slices.

14. 14. The non-transitory computer-readable storage medium of claim 13, wherein the application slice monitor monitors and collects metrics regarding the compute slicing and the network slicing generated by the application slice controller, the metrics being made available to the resource manager periodically at specific configurable intervals.

15. A system for specifying and executing an application, the application including a plurality of microservices of a 5G slice in a multi-tiered 5G infrastructure, comprising: Memory, determining 1001 end-to-end application characteristics using an application slice specification including an application ID component, an application name component, an application metadata component, a function dependency component, a function instance component, and an instance connection component; Specifying a slice specification for the function, including a network slice specification for the function and a compute slice specification for the function (1003); one or more processors in communication with the memory configured to simultaneously manage the compute and network requirements of the application using runtime components including a resource manager, an application slice controller, and an application slice monitor, the resource manager maintaining a database and managing (1005) starting, stopping, updating, and deleting instances of the application; A system having

16. The system of claim 15 , wherein the application slice specifications include a latency parameter, a bandwidth parameter, a device count parameter, and a reliability parameter.

17. 16. The system of claim 15, wherein the network slice specifications of the function include a latency parameter, a throughput GBR parameter, a throughput MBR parameter, and a packet error rate parameter.

18. 16. The system of claim 15, wherein the compute slice specifications for the function include a minimum CPU core parameter, a maximum CPU core parameter, a minimum memory parameter, a maximum memory parameter, and a tier parameter.

19. 20. The system of claim 18, wherein the layer parameters automatically map and execute functions across multiple layers, and the resource manager initially performs static resource management and then cooperates with the application slice controller and the application slice monitor to perform dynamic resource management across layers.

20. The application slice controller manages compute slicing and network slicing for the function by using a network slice interface layer that provides guarantees and admission control prior to the creation of the network slice; 16. The system of claim 15, wherein the application slice monitor monitors and collects metrics regarding the compute slicing and the network slicing generated by the application slice controller, the metrics being made available to the resource manager periodically at specific configurable intervals.