Method and network node for action impact estimation on a cloud system

The method and network node for action impact estimation in cloud systems address the challenge of relying on human experts by simulating action impacts on cloud systems, enabling autonomous decision-making and efficient resource management.

WO2025120362A1PCT designated stage expired Publication Date: 2025-06-12TELEFONAKTIEBOLAGET LM ERICSSON (PUBL) +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2023/062421
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current cloud systems rely heavily on human experts for decision-making during remediation, which is costly and less feasible for distributed edge clouds, due to the lack of methods for evaluating the impact of actions on cloud systems.

Method used

A computer-implemented method and network node for action impact estimation on a cloud system, which involves receiving a request for impact estimation, obtaining the cloud system's status, updating a simulation to match the status, and estimating resource and time consumption, as well as Key Performance Indicator (KPI) values, associated with executing an action.

Benefits of technology

The solution enables autonomous evaluation of action impacts on cloud systems, allowing for more informed decision-making without relying on human experts, and is adaptable to frequent changes in cloud systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2023062421_12062025_PF_FP_ABST
    Figure IB2023062421_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a method for action impact estimation on a cloud system. The method comprises receiving a request for estimating at least one impact of an action on the cloud system. The method comprises obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The method comprises obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The method comprises obtaining, using the simulation of the cloud system, estimations of impacted applications' key performance indicator (KPI) values after executing the action.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND NETWORK NODE FOR ACTION IMPACT ESTIMATION ON ACLOUD SYSTEMTECHNICAL FIELD

[0001] The present disclosure relates to action evaluation in cloud environments.BACKGROUND

[0002] Cloud and edge cloud are key enablers of next generation networks, such as fifth generation (5G), sixth generation (6G) networks and beyond. The distributed, heterogeneous, highly dynamic, and large scaled characteristics of such a system makes it difficult to manage. Thus, it is important for such system to be autonomously or “self ’-managed, i.e., the system automatically handles service deployment and operation, including handling new service requests, detecting / predicting potential system / service issues, and taking actions to solve the problems.

[0003] When there is a new service requirement or when a new system / service problem is detected, such autonomous system needs to propose solutions, evaluate them, and make decisions. The proposal and decision making may be based on existing knowledge, some decision-making logics, or some learning procedures. For example, the TM forum IG1253 (Intent in Autonomous Networks, vl.3.0, August 15, 2022) defines an Intent Management Function (IMF) that operates an autonomous system using intents. It assumes that IMFs operate based on knowledge, make decisions about actions to be taken and have the means to execute the chosen actions.

[0004] When operating a cloud system or an edge cloud system, actions may include, but are not limited to, scaling up / down a service or a system component, introducing replicas, migrating a virtual machine or a container to a different host and reconfiguring or restarting a service or a system component.

[0005] Current cloud systems (data centers) mainly rely on human experts to make decisions for remediation when a problem occurs, which is expensive and less feasible for distributed edge clouds.SUMMARY

[0006] Although the knowledge based intent management function is defined to operate an autonomous network, actual methods and procedures on how to evaluate the effect of an action on a cloud system are missing.

[0007] In order to select the most appropriate actions to solve a current problem (e.g., a new service intent or a detected fault), a management system should estimate the potential impacts of an action, including how much resources the action will use, whether the action can solve the current problem or not, and after taking the action, whether there are any impacts on an existing service deployed in a cloud system. Measurable metrics for the potential impacts of an action, can be the estimated key performance indicators (KPIs) of services and the computing, storage and network resources usages and changes.

[0008] Therefore, there is a need for an action evaluation method that could evaluate different impacts of an action while it can adapt to the frequent changes in the cloud system.

[0009] There is provided a computer implemented method for action impact estimation on a cloud system. The method comprises receiving a request for estimating at least one impact of an action on the cloud system. The method comprises obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The method comprises obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The method comprises obtaining, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

[0010] There is provided a network node for action impact estimation on a cloud system. The network node comprises processing circuits and a memory, the memory containing instructions executable by the processing circuits whereby the network node is operative to receive a request for estimating at least one impact of an action on the cloud system. The network node is operative to obtain a status of the cloud system, and update a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The network node is operative to obtain, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The network node is operative to obtain, using the simulation of the cloud system, estimations ofimpacted applications’ key performance indicator (KPI) values after executing the action.

[0011] There is provided a non-transitory computer readable media having stored thereon instructions for action impact estimation on a cloud system. The instructions comprise receiving a request for estimating at least one impact of an action on the cloud system. The instructions comprise obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The instructions comprise obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The instructions comprise obtaining, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

[0012] The method and network node provided herein present improvements to the way action impact estimation on a cloud system operate.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram of an example action evaluation system.

[0014] Figure 2 is a flowchart of an example evaluation procedure executed by an action evaluation agent.

[0015] Figure 3 is a block diagram of an example action effect estimator architecture.

[0016] Figure 4 is a flowchart of an example action cost estimation procedure executed by an action effect estimator.

[0017] Figure 5 is a block diagram showing an example deployment as part of the TM IG1253 intent handling function.

[0018] Figure 6 is a schematic illustration of an example of 5G core service graph.

[0019] Figure 7 is a schematic illustration of an example of action evaluation for host resource overutilization in a cloud system where a 5G core service is deployed.

[0020] Figure 8 is a flowchart of a method for action impact estimation on a cloud system.

[0021] Figure 9 is a schematic illustration of a hardware in which steps and / or method described herein can be executed.

[0022] Figure 10 is a schematic illustration of a virtualization environment in which the different steps and hardware components described herein can be deployed.DETAILED DESCRIPTION

[0023] Various features will now be described with reference to the drawings to fully convey the scope of the disclosure to those skilled in the art.

[0024] Sequences of actions or functions may be used within this disclosure. It should be recognized that some functions or actions, in some contexts, could be performed by specialized circuits, by program instructions being executed by one or more processors, or by a combination of both.

[0025] Further, computer readable carrier or carrier wave may contain an appropriate set of computer instructions that would cause a processor to carry out the techniques described herein.

[0026] The functions / actions described herein may occur out of the order noted in the sequence of actions or simultaneously. Furthermore, in some illustrations, some blocks, functions or actions may be optional and may or may not be executed; these are generally illustrated with dashed lines.

[0027] At least some aspects of the techniques described herein may be implemented using artificial intelligence, which comprises a variety of techniques as would be apparent to a person skilled in the art, including machine learning techniques.Machine learning techniques include Neural Network (NN), or Artificial Neural Network (ANN), and both terms may be used interchangeably herein. In some contexts, an Artificial Neural Network could include biological portions.

[0028] Further, looking forward, in a virtual world (e.g., the metaverse, digital twins, etc.), the techniques described herein could be applied in relevant virtual scenarios.

[0029] An action evaluation system and a method to automate the evaluation of the impact of proposed actions for a cloud system are presented herein. The action may be a request for a new service or a fault remediation request. The impact includes the action’s resource and time consumption, the result in resource level changes in the cloud system, and the Key Performance Indicator (KPI) value changes of the affected applications deployed in the cloud system. Example KPIs for applications deployed in a cloud system include, but are not limited to: availability, which is the percentage of time the application is available for use by customers during a given period, response time, which is the amount of time taken by the application to respond to a request, throughput, which is the amount of data that can be processed by the application in a given time, Mean Time to Failure (MTTF), which is a maintenance metric thatmeasures the average amount of time a non-repairable asset operates before it fails, Mean Time to Recovery (MTTR), which is the average time it takes to recover from a product or system failure, resource utilization, which is the percentage of resources utilized by the application in the cloud system, including central processing unit (CPU), memory, storage, and network, and downtime, which is the amount of time the application is unavailable due to maintenance, upgrades, or unexpected events.

[0030] Herein, a method is described to automate the evaluation of the impact of proposed actions for a cloud system which includes two steps. In a first step, an Action Effect Estimator estimates the resource and time consumption of an action, and the new resource level of the cloud system after executing the action. This is done via simulating the action effect on a simulated cloud system including simulated hosts, background traffic, resources, scheduling, workloads, which reflects the up-to- date status of the cloud system.

[0031] In a second step, the method finds application KPI estimation models, or application model(s), for the affected applications from an Application Model Repository, and uses the model(s) to estimate application KPIs given the proposed actions.

[0032] The proposed solution can solve the problem of action evaluation for both new service request and fault remediation. It does not assume there is an anomaly in the cloud system, instead, it simulates the key characteristics of the current cloud system no matter if the system is in an anomaly state or in a normal state. Key characteristics may include any of hosts, background traffic, resources, scheduling, workloads, which reflects the up-to-date status of the cloud system. To successfully simulate the cloud system, a tradeoff may be made between accuracy and complexity.

[0033] The solution is suitable for highly dynamic cloud systems since the simulated cloud system synchronizes with the real cloud system when making an action evaluation. The solution can be used for large scaled and heterogenous edge cloud environments as well since it instantiates an Action Effect Estimator per cloud which reflects the current cloud status.

[0034] The solution outputs the impacts of an action on a cloud system including 1) cost of the action, i.e., resource and time consumptions, 2) resource impact, i.e., the resource level changes on hosts, virtual machines, containers, pods, etc. and 3) application impact, the KPI changes, which provides thorough information for action selection decision making.

[0035] Referring to figure 1, an Action Evaluation System 105 is described. The Action Evaluation System 105 evaluates the impacts of a given action on a cloud system 100. The evaluation is threefold. First, the cost of an action is evaluated, including resource and time consumption caused by the action execution. Second, the resource impact is evaluated, i.e., the resource level changes (e.g., resource released, resource increased, or resource used) in the cloud environment. Third, the KPI of the applications (including the directly or indirectly impacted applications) deployed in the cloud system are evaluated.

[0036] The solution can be used for evaluating actions for both new service requests and fault remediation request. The actions may include but are not limited to the following: scaling up / down / migrating / restarting a microservice or an application;- installing / reconfiguring an application; scaling up / down / migrating / restarting a system component (e.g., a host (with computing, input / output, storage resources), a container, a pod / a microservice); adding / reconfiguring a system component (e.g., a host, a network link, a container, a pod / a microservice); and- killing / stopping / tearing down a process, an application or a system component.

[0037] System overview

[0038] Still referring to figure 1 the components of the proposed Action Evaluation System 105 are described.

[0039] Input - The system takes an action evaluation request as input. The evaluation request includes two parameters: the action under estimation and the cloud system (e.g., identified through a cloud system identifier that is unique and discoverable, for example a URL) on which the action is about to execute. An action further includes an action name, a component on which that action is executed, and optionally parameters related to the action. For example, an action can be something like scale_up(microservicel, 3 instances), migrate(vml), restart(hostl, soft restart). A cloud system is a running system that an operator is operating. It can be identified by some unique cloud system identifier (ID).

[0040] Output - The action evaluation system outputs 1) the estimated applications’ Key Performance Indicator (KPI) values and 2) the estimated resources including the action execution cost and cloud system resource changed after action execution. The KPIs are user defined metrics for a specific application, e.g., service request delay,application throughput, success requests per second, to name a few. When a resource is measured, it is always measured via a resource value and a time value. The time can be a period during which the resource is consumed, or a delay after which a resource is released or added. The resource may further include computing, networking, and storage resources (e.g., 1 core, 5G memory, 5kb / sec, and 10G storage).

[0041] In Figure 1, three system components are defined. The Action Evaluation Agent 130 is responsible for handling action evaluation requests and outputting the action evaluation results. Upon receiving an action evaluation request, it requests the Action Effect Estimator 110 to provide a list of affected applications and their estimated available resources after executing the action, and the estimated resource consumptions for the action execution. Based on the affected applications and their available resources, it further searches the Application Model Repository 120 to obtain applications’ KPI estimation models, which will estimate the KPI values based on given resources. After that, the Action Evaluation Agent 130 will combine the results and output the evaluation results. An example procedure of Action Evaluation Agent 130 can be found further below.

[0042] The Action Effect Estimator 110 is responsible for estimating the impacts of an action on a cloud system. It obtains the up-to-date cloud status, simulates an action effect on the cloud system, and outputs the affected applications, available resources after action execution and the action resource and time consumptions. One Action Effect Estimator 110 is instantiated per cloud system, the detailed design is provided further below. For the cloud system and action, the action effect estimator 110 obtains a) a list of affected applications, b) their estimated available resources after action execution and c) the estimated action execution resource / time consumption.

[0043] The Application Model Repository 120 stores application KPI estimation models that are responsible for estimating KPI values of an application. The models estimate the KPI values given new resource level of the cloud system made available to the application after execution of the action. The models can be trained, and they can be artificial intelligence (AI) / machine learning (ML) models or other functions / logics that take available resources as input and that output the estimated KPI values. Details concerning the Application KPI Estimation Model can be found further below.

[0044] Action Evaluation Agent; Evaluation Procedure

[0045] Figure 2 illustrates the action evaluation procedure executed by the Action Evaluation Agent 130. Upon receiving an action evaluation request, the Action Evaluation Agent 130 checks whether there is an Action Effect Estimator 110 for the cloud system 100, and if no, it instantiates one, step 201. Once there is an Action Effect Estimator 110, the Action Evaluation Agent 130 uses the Action Effect Estimator 110 to estimate the action execution resource and time consumption, the impacted applications and their new available resources after action execution, step 202.

[0046] For the impacted applications, the Action Evaluation Agent 130 looks for their KPI estimation models in the Application Model Repository 120, step 203, and uses each model to estimate the KPIs of the application with given available resources after action execution, step 204.

[0047] Finally, the Action Evaluation Agent 130 combines the estimation results from step 202 and step 204, and outputs the estimated application KPI, the estimated action execution resource / time cost and the estimated new resources for the impacted applications after action execution, step 205. The outputted result can be used for the action evaluation request sender or some other decision-making entities for action selection. An example deployment can be found in the section entitled an action evaluation example.

[0048] Action Evaluation System Component Design; Action Effect Estimator

[0049] Figure 3 shows an example architecture for the action effect estimator 110; it is composed of the following entities.

[0050] The action effect estimator 110 includes a cloud simulation subsystem 320 that simulates 1) hosts, including resource level (computing (e.g., CPU, graphic processing unit (GPU)), storage, input / output (I / O), etc.) of each host, 2) workloads that runs on hosts, including virtual machines (VMs), containers, pods etc., 3) background traffic, and 4) scheduler that simulates the placement of workloads on hosts. If a new workload comes, the scheduler knows where it will be placed and for a given placement of workloads on hosts, and the scheduler is able to indicate how much resource is available to each workload. The cloud system scheduler makes the decisions for workload placement. As a simulated scheduler, it simulates the placement decision making logic. The scheduler collects the available resource information from each host, thus it is able to indicate how much resource is available to a work load. It may place a workload on a host based on the score of the host. Thescore can be evaluated based on the workload resource requirement, placement policies (e.g., affinity / anti -affinity rules) and host resource status.

[0051] The action effect estimator 110 includes a synchronizer 325 that is responsible for synchronizing the simulation subsystem to the running cloud system. The synchronizer handles the synchronization requests. If it is a first-time request when the action effect estimator 110 is instantiated, and the synchronizer 325 1) reads host configuration information, 2) initializes a host in the simulator, including background load, 3) reads workload placement data, 4) initializes workload placement in the simulator, 5) reads scheduler configuration and 6) initializes scheduler functions. If it is a follow up request, the synchronizer reads the current state of the cloud system, including the changes in host configuration, workload placement and scheduler configuration, and updates the cloud simulation subsystem 320.

[0052] The action effect estimator 110 includes an action effect repository 330 that stores the rules simulating the execution effects of an action, and especially, it defines the changes made to a simulation system when an action is executed. Expert defined “Action Knowledge Base” described in Li Wu, Johan Tordsson, Alexander Acker, Odej Kao. MicroRAS: Automatic Recovery in the Absence of Historical Failure Data for Microservice Systems. UCC 2020 - 13th IEEE / ACM International Conference on Utility and Cloud Computing, Dec 2020, Leicester, United Kingdom, can be an example implementation of it.

[0053] Examples of the action effect knowledge are as follows, where the operation can be predefined knowledge and the time delays can be knowledge learned from a running cloud system:- Action: scale up workload (microservice 1, instances=2):Effect: scheduling 1 instance of microservice 1 in hosts with a delay tO;- Action: restart_host(hostl, mode=soft):- Effects: migrating all the workload from hostl to other hosts, the workload on hostl unavailable for a period of t2, and no resource available on hostl for a period of tl- Action: Migrate_workload(microservice 2):- Effect: moving microservice 2 from its current deployed host to other hosts in the cloud system, and microservice 2 unavailable for a period of t3;- Action: Scale_up_host(host2, resource=[2 cores, 100G memory]):- Effect: extend host2 with 2 cores and 100G of memory with a delay t4.

[0054] Alternatively, action effect can be estimated using some ML models, which learn the action effect via a supervised learning method. The input of such models can be the cloud state, such as resource level of hosts, workload locations, action, and action parameters, and the output can be the cloud state after the action is executed.

[0055] Depending on the cloud and workload, some regression models or neural networks can be used, e.g., one may use a graphic neural network to model hosts and workloads status. The training data can be generated by executing various actions on various real cloud environments, and the data is collected before and after the action execution. To avoid affecting the applications running in the real cloud system, the actions can be tested using some simulated workload and avoiding the rush hours.

[0056] The action effect estimator 110 includes an application manifest repository 335 that stores the information of which workload belongs to which application.

[0057] The action effect estimator 110 includes an Action Cost Estimator 340 that is responsible for calculating the cost of an action. The cost of an action can be determined by resource changes caused by the action, including 1) unavailable resource R1 in a period of tl, 2) released resource R2 in a delay t2 and 3) increased resource R3 with delay t3 in the case of cloud system expansion. With the resource changes, an operator may calculate the cost C using some cost function fc, e.g., C = fc(Rl, tl, R2, t2, R3, t3).

[0058] Figure 4 shows an example logic for the action cost estimator 340. At step 401, the action cost estimator 340 gets effect rules from the Action-effect Knowledge Base 330 for the Action. At step 402, based on the effect rules, the workload resource changes, and the snapshot of the resource status before action execution, the Action Cost Estimator 340 gets the total unavailable resource R1 and time tl, the total released resource R2 and delay t2, and, if the cloud system is extended, total increased resource R3 and delay t3. At step 403, the cost C is calculated using the cost function fc as shown above, or another equivalent cost function, and the Action Cost Estimator 340 outputs the cost C.

[0059] As mentioned previously, the cost is determined by the effect of an action. It is a function of resource and time and includes 3 parts:1) resources R1 unavailable for a period of tl caused by the action (get from effect rule);2) released resources R2 after action execution, total action execution delay is t2 (get from workload resource changes removing the extra resource R3 if item 3) exists;3) extra resources R3 added by the action with delay t3 in the case of cloud system expansion (scale up action) (get from effect rule).

[0060] An example way to compute the cost could be with C = (Rl*tl - R2*t2 + R3*t3) / Rtotal, where R can be a function of computing, networking and storage resources.

[0061] Alternatively, regression ML model (e.g., logistic regression model, XGB Regressor, Random Forest Regressor) can be trained to estimate the action cost, via collecting data samples with input features, such as: action type, action parameters, cloud resource status, (from step 401 and 402), input features, and output calculated cost (from step 403)

[0062] The model can be trained for a specific cloud system that serves a specific cloud, or the training data can be collected from multiple different cloud systems and a model can be trained to be used for the multiple cloud systems.

[0063] Referring to figure 3, the action effect estimator 110 includes an Action Handler 345 which is responsible for handling action effect estimation requests. It communicates with the other entities in the action effect estimator 110 and outputs the action effect estimation result.

[0064] The procedure of action effect estimation, executed by the action effect estimator 110 is as follow. When receiving an action effect estimation request 300, the action handler 345 sends a synchronization request to the synchronizer 325, which updates the cloud simulation subsystem 320 with the up-to-date cloud system state, steps 301a and 301b. The action handler 345 snapshots, step 302, the state of the cloud simulation subsystem 320, and reads, step 303, from the action-effect repository 330 to get the (estimated) effect rules of the requested action. The action handler 345 applies, step 304, the effect rules on the cloud simulation subsystem, and reads, step305, resources of all the hosts / workloads, comparing with the recent snapshot and identifies all hosts / workloads whose available resource changed due to the action. The action handler 345 uses the application manifest repository 335 to determine, step306, the set of affected applications along with their new resources. The action handler 345 uses the action cost estimator 340 to determine, step 307, the resource cost of the action, and combines, step 308 the results from steps 306 and 307 andoutputs action cost, impacted applications, and the new available resources for each application, step 309.

[0065] Application KPI Estimation Model

[0066] An Application KPI Estimation Model is responsible for estimating application’s KPIs given resources. It can be some functions defined by an operator or some trained ML models. There is one model per application. An application may be deployed with different configurations and scales in different use cases. When training an ML model for KPI estimation, the KPIs need to be considered and modeled. For example, the designed input features may include available resources, traffic load, application scale, application configuration, and the output features are the values of the KPIs (e.g., response time, throughput, latency).

[0067] Depending on the type of the application, an ML model can be some simple regression models (e.g., SVM, linear regression) or some deep learning models (e.g., LSTM, CNN), or some graphic neural network (GNN) that can better learn interrelationship between microservices when the application is composed of multiple microservices.

[0068] Example Use Case

[0069] A cloud system operator can use the proposed action evaluation system and method to support decision making for a new service request or a failure handling request. The system can be deployed as part of a cloud management system. An example can be found in Figure 5, where the action evaluation agent / system 130 is deployed as part of the intent handling function 500 defined in TM Forum IG1253, where the Action-effect can be part of the knowledge and the action estimation can be part of the decision-making logic.

[0070] An Action Evaluation Example

[0071] In figure 5, an action evaluation example is illustrated for a host overutilization problem in a cloud system where a 5G core application is deployed. The example service graph 600 of the application can be found in figure 6. The nodes in the graph represent the 5G core network functions and the edges represent the connectivity between network functions. The functions include virtual radio access network (vRAN) which are the virtualized part of the radio access network, which may involve multiple VNFs, Access and Mobility Management function (AMF), Session Management function (SMF), User plane function (UPF), Policy Control Function (PCF), Authentication Server Function (AUSF), Unified Data Management(UDM), Application Function (AF), Network Exposure function (NEF), not illustrated, NF Repository function (NRF), not illustrated, and Network Slice Selection Function (NSSF), among others. For KPI estimation of the example application, a GNN model may be built with node features related to resource level, number of instances, configurations, and with edge features related to latency and traffic. Data can be collected with various traffic mode, resource level, number of instances of network functions and their corresponding KPI values to train the model.

[0072] If the action evaluation system 130 is deployed in an intent handling function 500, as per Figure 5, for example, a resource over-utilization issue may be the problem to be solved and some solution may propose two potential actions as shown and described in relation with Figure 7. In such a scenario, the action evaluation system 130 should follow the steps to evaluate the two potential actions. The action effect knowledge, in this case, is obtained from the repository 330, the action is applied to the cloud system simulator 320, and the evaluation result is the output for a decision-making logic 510.

[0073] In the example illustrated in figure 7, the cloud system under management comprises a plurality of hosts, each executing different functions. Host i 605 suffers from resource over utilization, with a 5G core call drop rate of 1%. Two actions can be executed for solving this problem: 1- restart (host i, soft) and 2- migrate (AMF1). The restart action effect is that the workloads are migrated in a 1 second delay, host i is unavailable for 2 minutes, and no resource is available on host i for 2 minutes. The migrate action effect is that the AMF1 is moved to other hosts, and AMF1 is unavailable for 1 second.

[0074] When applying the restart action, the AMF1 is migrated to host 2 615, UDM2 is migrated to host j 620, NSSF is migrated to host 1 610, and the resources of hosts 1, 2, i and j have changed. This has for consequence at resource level, that the microservices total resource availability changes by -5%, the resources of host i are released at 80%, the total resources of hosts 1, 2 and j are occupied by +20%, and the action cost is 5. The application affected is 5G core and the estimated call drop rate is 0.5%.

[0075] When applying the migrate action, the AMF1 migrates to host 2, and the resources of host i and host 2 have changed. This has for consequence at resource level, that the microservices total resource availability changes by +1%, the resources of host i are released at 30%, the resources of host 2 are occupied by +25%, and theaction cost is 0.05. The application affected is 5G core and the estimated call drop rate is 0.4%.

[0076] In this example, it can be observed that the migration action is with less cost than the host restart action, and it results in more resource available for pods and less host resource released, and both actions result in a better estimated application KPI. All the information generated by the solution provided herein helps the decisionmaking logic determine what is the most appropriate action to take to solve the current problem.

[0077] In the example of figure 5, the solution-proposing and the decision-making may also be deployed as part of the decision entity 520 in intent handling function 500.

[0078] Turning to figure 8, there is provided a computer implemented method 800 for action impact estimation on a cloud system. The method comprises receiving, step 801, a request for estimating at least one impact of an action on the cloud system. The method comprises obtaining, step 802, a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The method comprises obtaining, step 803, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The method comprises obtaining, step 804, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

[0079] The request for estimating the at least one impact of the action may be a service request or a fault remediation request. The simulation of the cloud system may comprise a simulation of resources, hosts, schedulers, background traffic, and workloads deployed in the cloud system.

[0080] Obtaining the estimations of resource consumption associated with executing the action and resource changes after executing the action, may further comprise taking a snapshot, step 805, of the simulation of the cloud system; reading, step 806, from an action-effect repository, estimated effect rules of the action; applying, step 807, the effect rules on the simulation of the cloud system; and comparing, step 808, resulting resources of the hosts and workloads deployed in the cloud system with the snapshot and identifying the hosts and workloads having changed due to the action.

[0081] The method may further comprise querying, step 809, an application manifest repository to determine a set of applications affected by the hosts and workloads having changed due to the action, and determining new resources assigned to the set of applications.

[0082] Obtaining the estimations of time consumption associated with executing the action and resource changes after executing the action may further comprise taking a snapshot, step 805, of the simulation of the cloud system; reading, step 806, from an action-effect repository, estimated effect rules of the action; applying, step 807, the effect rules on the simulation of the cloud system; and computing, step 810, a time consumption as a sum of unavailable resources R1 for a period of time tl, of released resources R2 after action execution time t2, and of optional increased resource R3 after resource expansion delay t3.

[0083] The method may further comprise determining, step 811, a resource cost for the action, and providing the resource cost for the action, the set of applications affected by the hosts and workloads having changed due to the action, and the new resources assigned to the set of applications. The resource cost C may be computed as a function of Rl, tl, R2, t2, R3 and t3.

[0084] Obtaining estimations of impacted applications’ key performance indicator (KPI) values after executing the action may comprise obtaining, step 812, and using KPI estimation models to obtain the estimations of the impacted applications’ KPI.

[0085] The method may further comprise, for the set of applications affected by the hosts and workloads having changed due to the action, querying, step 813, an application model repository to get KPI estimation models for the set of applications affected by the hosts and workloads having changed due to the action, and using the models to estimate the KPI values of each application in the set of applications given the new resources assigned to the set of applications.

[0086] There is also provided a method for managing a cloud system. The method comprises receiving a service request or a fault remediation request. The method comprises executing the method described above to obtain estimations of resource and time consumption associated with executing the action, resource changes after executing the action and estimations of impacted applications’ key performance indicator (KPI) values after executing the action. The method comprises using the estimations for deciding to execute, delay or reject the service request or the fault remediation request.

[0087] It should be noted that methods and steps described herein are, generally, computer implemented methods and steps. The term computer may be interpreted as having different meanings, such as explained next, for example.

[0088] Referring to figure 9, there is provided a network node (HW) 901 or distributed system 1000, in which functions and steps described herein can be implemented.

[0089] The network node 901 may be a server, a radio base station, edge node, or any other computing device which may be part of a cloud computing system, edge computing system, or which may be a standalone device.

[0090] The network node described herein is able to estimate the impacts of an action on a cloud system. This estimation can be used to manage the cloud system, schedule tasks, etc. and to decide if requests can be granted, could be delayed or should be rejected.

[0091] The network node 901 comprises processing circuitry 903 and memory 905. The memory 905 can contain instructions executable by the processing circuitry 903 whereby functions and steps described herein may be executed to provide any of the relevant features and benefits disclosed herein.

[0092] The network node 901 may also include non-transitory, persistent, machine- readable storage media 907 having stored therein software and / or instruction 909 executable by the processing circuitry 903 to execute functions and steps described herein. The network node may also include network interface(s) and a power source.

[0093] The instructions 909 may include a computer program for configuring the processing circuitry 903. The computer program may be stored in a physical memory local to the device, which can be removable, or it could alternatively, or in part, be stored in the cloud. The computer program may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0094] Referring to figure 10, there is provided a virtualization environment 1000 in which functions and steps described herein can be implemented.

[0095] The virtualization environment 1000 (which may go beyond what is illustrated in figure 10), may comprise systems, networks, servers, nodes, devices, etc., that are in communication with each other either through wire or wirelessly, e.g., through a network interface component (NIC) comprising physical network interface(s). Some or all of the functions and steps described herein may be implemented as one or morevirtual components (e.g., via one or more applications, components, functions, virtual machines, containers, etc.) executing on one or more physical apparatus in one or more networks, systems, environment, etc.

[0096] A virtualization environment provides hardware 1001 comprising processing circuitry 1003 and memory 1005. The memory 1005 can contain instructions executable by the processing circuitry 1003 whereby functions and steps described herein may be executed to provide any of the relevant features and benefits disclosed herein.

[0097] The hardware 1001 may also include non-transitory, persistent, machine- readable storage media 1007 having stored therein software and / or instruction 1009 executable by the processing circuitry 1003 to execute functions and steps described herein.

[0098] The instructions 1009 may include a computer program for configuring the processing circuitry 1003. The computer program may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media. The computer program may be stored in a physical memory local to the hardware 1001, which can be removable, or it could alternatively, or in part, be stored in the cloud. The computer program may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0099] Referring again to figures 9 and 10, there is provided a network node 901, 1001 for action impact estimation on a cloud system. The network node 901, 1001 comprises processing circuitry 903, 1003 and a memory 905, 1005. The memory contains instructions executable by the processing circuitry whereby the network node is operative to receive a request for estimating at least one impact of an action on the cloud system. The network node is operative to obtain a status of the cloud system, and update a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The network node is operative to obtain, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The network node is operative to obtain, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

[0100] The request for estimating the at least one impact of the action may be a service request or a fault remediation request. The simulation of the cloud system may comprise a simulation of resources, hosts, schedulers, background traffic, and workloads deployed in the cloud system.

[0101] The network node is further operative to take a snapshot of the simulation of the cloud system. The network node is operative to read, from an actioneffect repository, estimated effect rules of the action. The network node is operative to apply the effect rules on the simulation of the cloud system. The network node is operative to compare resulting resources of the hosts and workloads deployed in the cloud system with the snapshot and identify the hosts and workloads having changed due to the action.

[0102] The network node is further operative to query an application manifest repository to determine a set of applications affected by the hosts and workloads having changed due to the action, and determine new resources assigned to the set of applications.

[0103] The network node is further operative to take a snapshot of the simulation of the cloud system. The network node is operative to read, from an actioneffect repository, estimated effect rules of the action. The network node is operative to apply the effect rules on the simulation of the cloud system. The network node is operative to compute a time consumption as a sum of unavailable resources R1 for a period of time tl, of released resources R2 after action execution time t2, and of optional increased resource R3 after resource expansion delay t3.

[0104] The network node is further operative to determine a resource cost for the action, and provide the resource cost for the action, the set of applications affected by the hosts and workloads having changed due to the action, and the new resources assigned to the set of applications. The resource cost C may be computed as a function of Rl, tl, R2, t2, R3 and t3.

[0105] The network node is further operative to obtain and use KPI estimation models to obtain the estimations of the impacted applications’ KPI.

[0106] The network node is further operative to, for the set of applications affected by the hosts and workloads having changed due to the action, query an application model repository to get KPI estimation models for the set of applications affected by the hosts and workloads having changed due to the action, and use themodels to estimate the KPI values of each application in the set of applications given the new resources assigned to the set of applications.

[0107] Still referring to figures 9 and 10, there is provided a network node 901, 1001 operative to manage a cloud system comprising processing circuitry 903, 1003 and a memory 905, 1005. The memory contains instructions executable by the processing circuits whereby the network node is operative to receive a service request or a fault remediation request. The network node is operative to execute the method described herein to obtain estimations of resource and time consumption associated with executing the action, resource changes after executing the action and estimations of impacted applications’ key performance indicator (KPI) values after executing the action. The network node is operative to use the estimations for deciding to execute, delay or reject the service request or the fault remediation request.

[0108] There is provided a non-transitory computer readable media 907, 1007 having stored thereon instructions 909, 1009 for action impact estimation on a cloud system. The instructions comprise receiving a request for estimating at least one impact of an action on the cloud system. The instructions comprise obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system. The instructions comprise obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action. The instructions comprise obtaining, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

[0109] Modifications will come to mind to one skilled in the art having the benefit of the teachings presented in the foregoing description and the associated drawings. Therefore, it is to be understood that modifications, such as specific forms other than those described above, are intended to be included within the scope of this disclosure. The previous description is merely illustrative and should not be considered restrictive in any way. The scope sought is given by the appended claims, rather than the preceding description, and all variations and equivalents that fall within the range of the claims are intended to be embraced therein. Although specific terms may be employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

CLAIMS1. A computer implemented method for action impact estimation on a cloud system, comprising:- receiving a request for estimating at least one impact of an action on the cloud system; obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system; obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action; and obtaining, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

2. The method of claim 1, wherein the request for estimating the at least one impact of the action is a service request or a fault remediation request.

3. The method of claim 1, wherein the simulation of the cloud system comprises a simulation of resources, hosts, schedulers, background traffic, and workloads deployed in the cloud system.

4. The method of any one of claims 1 to 3, wherein obtaining the estimations of resource consumption associated with executing the action and resource changes after executing the action comprises:- taking a snapshot of the simulation of the cloud system;- reading, from an action-effect repository, estimated effect rules of the action; applying the effect rules on the simulation of the cloud system; and comparing resulting resources of the hosts and workloads deployed in the cloud system with the snapshot and identifying the hosts and workloads having changed due to the action.

5. The method of any one of claims 1 to 4, further comprising querying an application manifest repository to determine a set of applications affected by the hosts and workloads having changed due to the action, and determining new resources assigned to the set of applications.

6. The method of any one of claims 1 to 3, wherein obtaining the estimations of time consumption associated with executing the action and resource changes after executing the action comprises:- taking a snapshot of the simulation of the cloud system;- reading, from an action-effect repository, estimated effect rules of the action; applying the effect rules on the simulation of the cloud system; and computing a time consumption as a sum of unavailable resources R1 for a period of time tl, of released resources R2 after action execution time t2, and of optional increased resource R3 after resource expansion delay t3.

7. The method of claim 5, further comprising determining a resource cost for the action, and providing the resource cost for the action, the set of applications affected by the hosts and workloads having changed due to the action, and the new resources assigned to the set of applications.

8. The method of claim 7, wherein the resource cost C is computed as a function of Rl, tl, R2, t2, R3 and t3.

9. The method of any one of claims 1 to 4, wherein obtaining estimations of impacted applications’ key performance indicator (KPI) values after executing the action comprises obtaining and using KPI estimation models to obtain the estimations of the impacted applications’ KPI.

10. The method of claim 5, 7 or 8, further comprising, for the set of applications affected by the hosts and workloads having changed due to the action, querying an application model repository to get KPI estimation models for the set of applications affected by the hosts and workloads having changed due to the action, and using the models to estimate the KPI values of each application in the set of applications given the new resources assigned to the set of applications.

11. A method for managing a cloud system comprising:- receiving a service request or a fault remediation request; executing the method of claim 1 to obtain estimations of resource and time consumption associated with executing the action, resource changes after executing the action and estimations of impacted applications’ key performance indicator (KPI) values after executing the action; and- using the estimations for deciding to execute, delay or reject the service request or the fault remediation request.

12. A network node for action impact estimation on a cloud system comprising processing circuits and a memory, the memory containing instructions executable by the processing circuits whereby the network node is operative to:- receive a request for estimating at least one impact of an action on the cloud system; obtain a status of the cloud system, and update a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system; obtain, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action; and obtain, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

13. The network node of claim 12, wherein the request for estimating the at least one impact of the action is a service request or a fault remediation request.

14. The network node of claim 12, wherein the simulation of the cloud system comprises a simulation of resources, hosts, schedulers, background traffic, and workloads deployed in the cloud system.

15. The network node of any one of claims 12 to 14, further operative to:- take a snapshot of the simulation of the cloud system;- read, from an action-effect repository, estimated effect rules of the action; apply the effect rules on the simulation of the cloud system; and compare resulting resources of the hosts and workloads deployed in the cloud system with the snapshot and identify the hosts and workloads having changed due to the action.

16. The network node of any one of claims 12 to 15, further operative to query an application manifest repository to determine a set of applications affected by the hosts and workloads having changed due to the action, and determine new resources assigned to the set of applications.

17. The network node of any one of claims 12 to 14, further operative to:- take a snapshot of the simulation of the cloud system;- read, from an action-effect repository, estimated effect rules of the action; apply the effect rules on the simulation of the cloud system; and compute a time consumption as a sum of unavailable resources R1 for a period of time tl, of released resources R2 after action execution time t2, and of optional increased resource R3 after resource expansion delay t3.

18. The network node of claim 16, further operative to determine a resource cost for the action, and provide the resource cost for the action, the set of applications affected by the hosts and workloads having changed due to the action, and the new resources assigned to the set of applications.

19. The network node of claim 17, wherein the resource cost C is computed as a function of Rl, tl, R2, t2, R3 and t3.

20. The network node of any one of claims 12 to 15, further operative to obtain and use KPI estimation models to obtain the estimations of the impacted applications’ KPI.

21. The network node of claim 16, 18 or 19, further operative to, for the set of applications affected by the hosts and workloads having changed due to the action, query an application model repository to get KPI estimation models for the set ofapplications affected by the hosts and workloads having changed due to the action, and use the models to estimate the KPI values of each application in the set of applications given the new resources assigned to the set of applications.

22. A network node operative to manage a cloud system comprising processing circuits and a memory, the memory containing instructions executable by the processing circuits whereby the network node is operative to:- receive a service request or a fault remediation request; execute the method of claim 1 to obtain estimations of resource and time consumption associated with executing the action, resource changes after executing the action and estimations of impacted applications’ key performance indicator (KPI) values after executing the action; and- use the estimations for deciding to execute, delay or reject the service request or the fault remediation request.

23. A non-transitory computer readable media having stored thereon instructions for action impact estimation on a cloud system, the instructions comprising:- receiving a request for estimating at least one impact of an action on the cloud system; obtaining a status of the cloud system, and updating a simulation of the cloud system to match the status, the simulation of the cloud system comprising a simulation of applications running in the cloud system; obtaining, using the simulation of the cloud system, estimations of resource and time consumption associated with executing the action and resource changes after executing the action; and obtaining, using the simulation of the cloud system, estimations of impacted applications’ key performance indicator (KPI) values after executing the action.

Citation Information

Patent Citations

  • Techniques and system for optimization driven by dynamic resilience

    US10275331B1

  • Performance interference model for managing consolidated workloads in QOS-aware clouds

    US8732291B2