First node, second node, computer system and methods performed thereby for handling configuration of first resources
Patent Information
- Application Number
- PCT/EP2025/058564
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure EP2025058564_01102026_PF_FP_ABST
Abstract
Description
[0001] FIRST NODE, SECOND NODE, COMPUTER SYSTEM AND METHODS PERFORMED THEREBY FOR HANDLING CONFIGURATION OF FIRST RESOURCES
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to a first node and methods performed thereby for handling configuration of first resources. The present disclosure also relates generally to a second node, and methods performed thereby for handling the configuration of first resources. The present disclosure also relates generally to a computer system, and methods performed thereby for handling the configuration of first resources.
[0004] BACKGROUND
[0005] Computer systems in a communications network or communications system may comprise one or more nodes. A node may comprise a processing circuitry which, together with computer program code may perform different functions and actions, a memory, a receiving port, and a sending port. A node may be, for example, a server. Nodes may perform their functions entirely on the cloud.
[0006] The communications system may cover a geographical area which may be divided into cell areas, each cell area being served by a type of node, a network node in the Radio Access Network (RAN), radio network node or Transmission Point (TP), for example, an access node such as a Base Station (BS), e.g., a Radio Base Station (RBS), which sometimes may be referred to as e.g., gNB, evolved Node B (“eNB”), “eNodeB”, “NodeB”, “B node”, or Base Transceiver Station (BTS), depending on the technology and terminology used. The base stations may be of different classes such as e.g., Wide Area Base Stations, Medium Range Base Stations, Local Area Base Stations, and Home Base Stations, based on transmission power and thereby also cell size. A cell may be understood to be the geographical area where radio coverage may be provided by the base station at a base station site. One base station, situated on the base station site, may serve one or several cells. Further, each base station may support one or several communication technologies. The telecommunications network may also comprise network nodes which may serve receiving nodes, such as user equipments, with serving beams.
[0007] The standardization organization Third Generation Partnership Project (3GPP) is currently in the process of specifying a New Radio Interface called Next Generation Radio or New Radio (NR) or 5G-Universal Terrestrial Radio Access (UTRA), as well as a Fifth Generation (5G) Packet Core Network, which may be referred to as 5G Core Network (5GC), abbreviated as 5GC.
[0008] Figure 1 is a schematic diagram depicting a particular example of a 5G referencearchitecture, as defined by 3GPP, which may be used as a reference for the present disclosure. An Application Function (AF) 1 may provide a service in the communications system and may interact with the 3GPP Core Network through a Network Exposure Function (NEF) 2. The AF 1 may allow external parties to use the Exposure Application Programming Interfaces (APIs) offered by the network operator. In case the AF 1 is trusted, e.g., internal to the network operator, the AF 1 may interact with the 3GPP Core Network directly, with no NEF 2 involved. The NEF 2 may support different functionality. Specifically, the NEF 2 may support different Exposure APIs, e.g., sponsored Data, Quality of Service (QoS), etc., which may allow a content provider to request policies from the Mobile Network Operator (MNO), e.g., on a per user session and / or application basis. The NEF 2 may be understood to act as the entry point into the network of the operator, so an external AF, e.g., a content provider, may interact with the 3GPP Core Network through the NEF 2. A Unified Data Repository (UDR) 3 may store data grouped into distinct collections of subscription-related information: subscription data, policy data, structured data for exposure, and application data. A Unified Data Management Function (UDM) 4 may generate 3GPP 5G AKA Authentication Vectors, handle user identification handling, support a Serving Network Function (NF) Registration Management of a UE 5, e.g., storing the serving Access and Mobility Function (AMF) 6 for the UE 5, storing the serving Session Management Function (SMF) 7 for the UE's Protocol Data Unit (PDU) Session, etc., support retrieval of the UE's individual subscription data for slice selection, and handle subscription data for network exposure capabilities applicable to an individual UE 5 or a group of UEs. A Policy Control Function (PCF) 8 may support a unified policy framework to govern the network behavior. Specifically, the PCF 8 may provide Policy and Charging Control (PCC) rules to a Policy and Charging Enforcement Function (PCEF), not depicted in Figure 1. The SMF 71 User Plane function (UPF) 9 may enforce policy and charging decisions according to provisioned Policy and Charging Control (PCC) rules. The SMF 7 may support different functionalities, e.g., session establishment, modify and release, and policy related functionalities such as termination of interfaces towards policy control functions, charging data collection, support of charging interfaces and control and coordination of charging data collection at the UPF 9. Specifically, the SMF 7 may receive PCC rules from the PCF 8 and configure the UPF 9, e.g., for event reporting, accordingly through an N410 reference point, Packet Flow Control Protocol (PFCP) protocol as follows. The SMF7 may control the packet processing in the UPF 9 by establishing, modifying or deleting PFCP sessions and by provisioning, e.g., adding, modifying or deleting, Packet Detection Rules (PDRs), Forwarding Action Rules (FARs), Quality Enforcement Rules (QERs) and / or Usage Reporting Rules (URRs) per PFCP session, whereby a PFCP session may correspond to an individual PDU session or a standalone PFCP session not tied to any PDU session. Each Packet Data Rule (PDR) may contain a Packet Identifier (PDI) specifying the traffic filters or signatures againstwhich incoming packets may be matched. The UPF 9 may support handling of user plane traffic based on the rules received from the SMF 7, specifically, e.g., packet inspection, through PDRs, and different enforcement actions such as, e.g., traffic steering, QoS, charging / reporting, for example through FARs, QERs, and / or URRs. The PCF 4 may provide policy rules to the UE 5 through the AMF 6. The AMF 6 may manage access of the UE 5, for example, when the UE 5 may be connected through different access networks, and mobility aspects of the UE 5. A Charging Function (CHF) 11 may support charging related functionality, specifically online and offline charging. The PCF 8 may provide policy rules to the UE 5 through the AMF 6. The AMF 6 may support different functionality, e.g. Termination of Non-Access Stratum (NAS) signalling, NAS ciphering and integrity protection, registration management, connection management, mobility management, access authentication and authorization, and security context management. Particularly relevant for this disclosure, the AMF 6 may be used to convey information from / to the UE 5 through NAS signaling, including the UE Policies provided by the PCF 4. A Network Data Analytics Function (NWDAF) 12 may be understood to represent an operator managed network analytics logical function. The NWDAF 12 may be part of the 5GC architecture and may use the mechanisms and interfaces specified for 5GC and Operations, Administration and Maintenance (OAM). Also depicted in Figure 1 is a Network Slice Selection Function (NSSF) 13, Network Repository Function (NRF) 14, an Authentication Server Function (AUSF) 15, a 5G Equipment Identity Register (5G-EIR) 16, a Gateway Mobile Location Centre (GMLC) 17, a Location Management Function (LMF) 18, a Radio Access Network (RAN) 19, a and a Data Network (DN) 20. Each of the NSSF 13, the NEF 2, the NRF 14, the UDR 3, the PCF 8, the UDM 4, the CHF 11, the AF 1 , the NWDAF 12, the AUSF 15, the AMF 6, the SMF 7, the 5G-EIR 16, the GMLC 17, and the LMF 18, may have an interface through which they may be accessed, which as depicted in the Figure, may be, respectively: Nnssf 21 , Nnef 22, Nnrf 23, Nudr 24, Npcf 25, Nudm 26, Nchf 27, Naf 28, Nnwdaf 29, Nausf 30, Namf 31 , Nsmf 32, N5g-eir 33, Ngmlc 34 and Nlmf 35. Figure 1 also depicts an interface N1 36 between the UE 5 and the AMF 6, an N237 reference point, an interface N338 between the RAN 19 and the UPF 9, an interface N639 between the UPF 9 with the DN 20, and an N940 reference point of the UPF 9.
[0009] In a Network Function Virtualization (NFV) ecosystem, a Network Service (NS) request may generally contain a set of Virtual Network Functions (VNFs) with some dependencies and Service Level Agreement (SLA) requirements. An incoming NS may be either described by a chain of VNFs, commonly referred to as a Service Function Chain (SFC), or a more general graph topology, which may be known as Virtual Network Function - Forwarding Graph (VNF-FG). In cloud computing, a similar concept may be found in microservices, whereby an application may be composed of a set of microservices that may interact with each other to implement the functionality of an application.A good example of this may be understood to be a the 5G core architecture as realized by a dual-mode 5G core (5GC) such as that depicted in Figure 2. Figure 2 illustrates how 5GC may be understood to be based on microservice architecture, where there may be several different components that may interact with each other to provide core functionality. For a given procedure, e.g., registration, in order to meet the specified SLA, e.g., registration may need to be completed within 100msec, the appropriate amount of resources may need to be allocated to different components involved in processing and fulfilling this request. Figure 2 particularly depicts the components of 5GC Network Functions (NFs), see Figure 1 , grouped together into a number of working units, with Evolved Packet Core (EPC) network functions being depicted with solid lines, 5GC network functions being depicted with long dashed lines, and Gi-LAN functions being depicted in short dashed lines. As depicted in Figure 2, the data layer 41 comprises a cloud core data-storage manager 42, which may comprise an Unstructured Data Storage Function (USDF) 43 and a UDR 44. A control plane 45 may comprise a cloud core resource controller 46, which may comprise an NSSF 47, an Network Slice Admission Control Function (NSACF) 48, an NWDAF 49 and an NRF 50. The control plane 45 may also comprise a cloud core subscription manager 51 , which may comprise a UDM 52, a Home Subscriber Server (HSS) 53, an AUSF 54, and a 5G-Equipment Identity Register (EIR) 55. The control plane 45 may further comprise a cloud core policy controller 56, which may comprise a PCF 57, an NWDAF 58, a Policy and Charging Rules Function (PCRF) 59. The control plane 45 may additionally comprise a cloud core exposure server 60, which may comprise a NEF 61 and a Service Capability Exposure Function (SCEF) 62. The cloud core components depicted in Figure 2 may connect, in a Serviced Based Architecture (SBA) 63 arrangement, to a Packet Core Controller 64 and a Signalling Controller 65. The Packet Core Controller 64 may comprise an AMF 66, an Mobility Management Entity (MME) 67, an SMF 68, a Serving Gateway Control Plane Function (SGW-C) 69, an NWDAF 70 and a Packet Gateway Control Plane Function (PGW-C) 71. The Signalling Controller 65 may comprise a Security Edge Protection Proxy (SEPP) 71 , a Diameter Routing Agent (DRA) 72, a Binding Support Function (BSF) 73, and a Service Communication Proxy (SCP) 74. The Packet Core Controller 64 may be connected to an NR Stand Alone (SA), LTE / NR Non-Stand Alone (NSA) or a Global System for Mobile communications (GSM)ZWideband Code-Division Multiple Access (WCDMA) network 75. The Packet Core Controller 64 may be also connected to a user plane 76, which may comprise a Packet Core Gateway 77. The Packet Core Gateway 77 may comprise a UPF 78, a Serving Gateway User Plane Function (SGW-U) 79, a Packet Gateway User Plane Function (PGW-U) 80, an NWDAF 81, a Network Address Translation (NAT) 82, an Optimization Function (Opt) 83 and a Firewall Function (FW) 84. Interworking with legacy networks, e.g., Home Location Register (HLR)ZAuthentication Center (AuC), Base Station Controller (BSC)ZRadio Network Controller (RNC), Serving GPRS Support Node(SGSN), Internet Protocol Multimedia System (IMS), Diameter Agent (DA) and Diameter Edge Agent (DEA) functions may also be supported.
[0010] The software architecture depicted as an example in Figure 2 may be based on cloudnative design principles. In this architecture, each NF may offer one or more services to other NFs through some APIs. Various NFs may communicate with each other; hence, the performance of a given NF may depend on the performance of the NF that may offer services to the given NF. These cloud-native NFs may be virtualized (VNFs) or physical NFs, e.g., Physical Network Functions (PNFs).
[0011] Dynamic resource dimensioning may be understood to be one of the problem areas in resource allocation for such applications. Considering Service Function Chains (SFCs), there may be clear dependencies between VNFs, since the performance of a VNF may be understood to not depend only on the given VNF, but also on other VNFs in the VNF-FG, e.g., neighbors / dependent / supporter VNFs. This may become even more complicated when considering general and complex topologies, which may be understood to make developing accurate solutions for resource configuration a challenging task. Dynamic resource dimensioning in NFV, may be understood to refer to dynamically identifying the amount of resources needed by a VNF in a Network Service (NS), such that the end-to-end requirements on the NS may be met, while ensuring efficient resource utilization. This may be performed by taking into consideration the inter-dependencies between VNFs in a VNF-FG and the expected end-to end performance Key Performance Indicators (KPIs). The same may be understood to be true for microservices, where the challenge may be understood to be to determine the resources required by individual microservices, such that the whole application may achieve its KPI.
[0012] Most of the existing resource dimensioning methods may be understood to rely on offline training with complex emulation setups. Unfortunately, such environments do not accurately represent the state of a system while in production. Other methods attempt online training; however, this may be understood to be generally impractical, time-consuming, and expensive for extensive use. In most cases, these online methods may be used to estimate the resource dimensions for stand-alone VNFs or microservices, without considering their dependences.
[0013] In particular, some methods may rely on profiling data. Particularly, a profiling-based VNF resources configuration. Some methods have proposed profiling-based resource allocation for stand-alone VNFs. Even for a single VNF, the configurations to test may be in the thousands, and profiling for a single VNF may take many several days.
[0014] Other methods may rely on Machine Learning (ML)-based VNF resources configuration. Particularly, methods may rely on reinforcement Learning (RL)-based resources configuration in the cloud. The inter-dependency between a VNF and its neighbors has been considered with an online method to predict a future scaling decision [1], Q-learning has been used tooptimize resource allocation in the cloud with the aim to improve fault tolerance, energy consumption, and load balancing [2], Independent tasks of an application may be considered without explicitly assuming a graph with dependencies and relationships between application components.
[0015] Some of the methods aim at speeding up RL, for example, by using distributed Reinforcement Learning. Many simulators may be used, multiple Graphics Processing Units (GPUs), and increased batch size learning rates. Multi-agent simulation may be used to run in parallel. Several simulators may be used with varying level of fidelity, since it may be understood to be difficult to trust samples from a simulator.
[0016] Machine Learning
[0017] Machine learning (ML) may be understood as the study of computer algorithms that may improve automatically through experience. It is seen as a part of Artificial Intelligence (Al). ML algorithms may build a model based on sample data, known as "training data", in order to make predictions or decisions without being explicitly programmed to do so. ML algorithms may be used in a wide variety of applications, such as email filtering and computer vision, where it may be difficult or unfeasible to develop conventional algorithms to perform the needed tasks.
[0018] There may be basically three types of ML Algorithms: Supervised Learning, Unsupervised Learning, and Reinforcement Learning (RL).
[0019] Supervised Learning algorithms may comprise a target / outcome variable, or dependent variable, which may have to be predicted from a given set of predictors, that is, independent variables. Using this set of variables, a function may be generated that may map inputs to desired outputs. The training process may continue until the model may achieve a desired level of accuracy on the training data. Once an ML model may have been trained, an inference process may begin, whereby new data may be run through the ML model to calculate an output. Examples of Supervised Learning may be Regression, Decision Tree, Random Forest, KNN, Logistic Regression etc.
[0020] In Unsupervised Learning algorithms, there may be no target or outcome variable to predict / estimate. It may be used for clustering a population into different groups, which may be widely used for segmenting customers in different groups for specific intervention. Examples of Unsupervised Learning may be K-means, mean-shift clustering, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), Expectation-Maximization (EM) Clustering using Gaussian Mixture Models (GMM), Agglomerative Hierarchical Clustering, etc....
[0021] Cluster analysis or clustering may be understood as an ML technique which may comprise grouping a set of objects in such a way that objects in the same group, which may be called a cluster, may be understood to be more similar, in some sense, to each other than to those in other groups, that is, other clusters. It may be understood as a main task ofexploratory data mining, and a common technique for statistical data analysis, used in many fields, including pattern recognition, image analysis, information retrieval, bioinformatics, data compression, computer graphics and ML.
[0022] Reinforcement learning (RL) may be understood to be a type of ML where an agent may learn to make decisions by taking actions in an environment to achieve some goal. The agent may receive feedback in the form of rewards, which it may use to learn the best strategy, or policy, to accumulate the most reward overtime.
[0023] Reinforcement learning may involve the following. An agent may be understood to be a learner or decision maker that may interact with the environment. The environment may be understood to refer to the world that the agent may interact with and learn from. A state (s) may be understood to refer to a representation of the current situation that the agent may be in. It may be understood to be the context within which the agent may make decisions. An action (a) may be understood to refer to a choice made by the agent that may affect the state. A reward (r) : may be understood to refer to feedback from the environment in response to the actions taken by the agent. It may be a scalar signal that may indicate how well the agent is doing at a given moment. A Policy may be understood to be a strategy used by the agent, which may map states to actions. The policy may be deterministic, that is, always the same action for a given state, or stochastic, that is, probabilistic actions for a given state. A value function may be understood to refer to a function that may estimate how good it may be for the agent to be in a given state, or how good it may be to perform a certain action in a given state. The "goodness" may be typically measured as the expected return, e.g., the cumulative reward, that may be achieved. A Q-function, an Action-Value Function, may be understood to refer to a function that may estimate the value of taking a certain action in a given state, and then following the current policy thereafter. A model of the environment may be understood to refer to the fact that some RL approaches may involve learning a model that may predict the next state (s’) and the reward for the current state and action. This may be understood to allow for planning and reasoning about the future without needing to actually take the action.
[0024] RL may involve making decisions in regard to exploration vs. exploitation. In reinforcement learning, the agent may need to balance exploration, that is, trying new things to discover better rewards, with exploitation, that is, using known information to maximize rewards. This may be understood to be a trade-off in RL.
[0025] In the reinforcement learning problem, the state may change every time the agent may apply a new action. The problem may be represented in the following way: The agent may receive the state of the environment at a certain time (s). Then the agent may select on action (a) and apply it in the environment. When this action is applied, the environment may provide a reward (r) and change to a new state (s’), the reward and state may be providedfinally by an interpreter. In reinforcement learning, the term "interpreter" may be used to describe a component of a reinforcement learning system that may interpret the state of the environment and the actions of an agent. It may be understood to be the part of the system that may bridge the agent with the environment it may be trying to learn from.
[0026] This cyclic procedure may be understood to bring a sequence of states, actions and rewards: s1, a1,r1;...;sT,aT,rT. The agent may use different learning algorithms to learn the most appropriate action to take on every different state of the NW, e.g., policy-learning based, such as actor-critic approaches, or value-based learning, such as deep-q networks.
[0027] Existing methods to handle resource dimensioning may be impractical, time-consuming, costly and slow and may have limited adaptability and accuracy.
[0028] SUMMARY
[0029] As part of the development of embodiments herein, one or more challenges with the existing technology will first be identified and discussed.
[0030] Existing Al-based approaches for resource dimensioning methods heavily depend on preexisting training data, which is often unavailable or insufficient, limiting the adaptability and accuracy an ML model.
[0031] Online training methods that collect experience iteratively through real-time interaction with the environment may be understood to be impractical, as data collection in these scenarios may be costly, very slow and risk Service Level Agreement (SLA) violations when used on production environments.
[0032] Online training is time-consuming. For instance, collecting 20,000 data points may take over a year if each point requires deploying applications in a cloud environment, e.g., Kubernetes, waiting for pods and measured Key Performance Indicators (KPIs) to stabilize, and collecting data, a process that may take upwards of 30 minutes per data point.
[0033] Offline training approaches that rely on profiling data, in addition to the same challenge of collecting a large amount of data cited above, lack feedback from an actual system. Without accurate, real-time data that may truly reflect the current state of a system, these models may be understood to be unable to produce accurate resource dimensions, limiting their capabilities in real-world applications.
[0034] According to the foregoing, it is an object of embodiments herein to improve the handling configuration of resources in a communications system.
[0035] According to a first aspect of embodiments herein, the object is achieved by a computer-implemented method, performed by a first node. The method is for handling configuration of first resources. The first node operates in a communications system. The first node obtains respective data from one or more deployments in the communications system. The first node then determines whether or not the obtained respective data comprises a combination of values of one or more variables used to train a machine learning model to determine aconfiguration of first resources in the communications system. The combination of values lacks a similarity with first information previously used to train the machine learning model according to a criterion of similarity. The first node initiates saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the obtained respective data comprises the combination of values.
[0036] According to a second aspect of embodiments herein, the object is achieved by a computer-implemented method, performed by a second node. The method is for handling the configuration of first resources. The second node operates in the communications system. The second node receives a first indication from the first node operating in the communications system. The first indication instructs the second node to save the combination of values of the one or more variables used to train the machine learning model to determine the configuration of the first resources in the communications system. The second node sends a third indication to a fourth node operating in the communications system. The third indication requests the fourth node to provide a reward corresponding to the combination of values. The reward is to be used to train the machine learning model to determine the configuration of the first resources in the communications system. The second node then receives the reward from the fourth node.
[0037] According to a third aspect of embodiments herein, the object is achieved by a computer-implemented method, performed by a computer system comprising the first node and a fifth node. The method is for handling the configuration of first resources. The fourth node operates in the communications system. The computer system obtains, by the first node, the respective data from the one or more deployments in the communications system. The computer system determines, by the first node, whether or not the obtained respective data comprises the combination of the values of the one or more variables used to train the ML model to determine the configuration of first resources in the communications system. The combination of values lacks the similarity with the first information previously used to train the ML model according to a criterion of similarity. The computer system initiates, by the first node, saving the combination of values to train the machine learning model, with the proviso the result of the determination is that the obtained respective data comprises the combination of values. The computer system trains, by the fifth node 115, the machine learning model with the saved combination of values. The computer system then outputs, by the fifth node, a third indication of the trained machine learning model.
[0038] According to a fourth aspect of embodiments herein, the object is achieved by the first node, for handling the configuration of first resources. The first node is configured to operate in the communications system. The first node is configured to obtain the respective data from the one or more deployments in the communications system. The first node is further configured to determine whether or not the respective data configured to be obtainedcomprises the combination of values of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of first resources in the communications system. The combination of values is configured to lack the similarity with the first information configured to have been previously used to train the machine learning model according to the criterion of similarity. The first node is further configured to initiate saving the combination of values to train the machine learning model, with the proviso the result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values.
[0039] According to a fifth aspect of embodiments herein, the object is achieved by the second node, for handling the configuration of first resources. The second node is configured to operate in the communications system. The second node is configured to receive the first indication from the first node configured to operate in the communications system. The first indication is configured to instruct the second node to save the combination of values of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of first resources in the communications system. The second node is also configured to send the third indication to the fourth node configured to operate in the communications system. The third indication is configured to request the fourth node to provide the reward configured to correspond to the combination of values. The reward is configured to be used to train the machine learning model configured to determine the configuration of first resources in the communications system. The second node is further configured to receive the reward from the fourth node.
[0040] According to a sixth aspect of embodiments herein, the object is achieved by the computer system, for handling the handling the configuration of first resources. The computer system is configured to comprise the first node and the fifth node. The computer system is configured to operate in the communications system. The computer system is configured to obtain, by the first node, the respective data from the one or more deployments in the communications system. The computer system is further configured to determine, by the first node, whether or not the respective data configured to be obtained is configured to comprise the combination of values. The combination of values is of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of the first resources in the communications system. The combination of values is configured to lack the similarity with the first information configured to have been previously used to train the machine learning model. The computer system is further configured to initiate, by the first node, saving the combination of values to train the machine learning model. This is performed with the proviso the result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values. The computer system is further configured to train, by the fifth node, the machine learning model with thecombination of values configured to be saved. The computer system is further configured to output, by the fifth node the third indication of the trained machine learning model.
[0041] By obtaining the respective data from one or more deployments in the communications system, determining whether or not the obtained respective data comprises the combination of values lacking the similarity with the first information previously used to train the machine learning model, and initiating saving the combination of values, the first node may then enable to train, e.g., by the fifth node, the machine learning model, with the combination of values thereby enabling to achieve fast training times, significantly accelerating the learning process. Another advantage of may be understood to be that the first node may enable efficient usage of resources of production deployments even during training times. An additional advantage may be understood to be that the first node may enable minimal deployment overhead. The training process may be understood to be designed to avoid deployment and redeployment of network service for additional data points required to train the model. A further advantage may be understood to be that by leveraging production environments, the first node may enable to remove the need for unrealistic emulation setups.
[0042] By receiving the first indication instructing the second node to save the combination of values, the second node may enable that the machine learning model may be trained with the saved combination of values, and enable the advantages just described.
[0043] By the second node sending the third indication to the fourth node requesting to provide the reward corresponding to the combination of values and receiving the reward from the fourth node, the reward may be used to train the machine learning model and enable the advantages just described.
[0044] BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Examples of embodiments herein are described in more detail with reference to the accompanying drawings, according to the following description.
[0046] Figure 1 is a schematic diagram illustrating an example of a 5G Network Architecture,
[0047] according to existing methods.
[0048] Figure 2 is a schematic diagram illustrating an example of a dual-mode 5G Core portfolio, according to existing methods.
[0049] Figure 3 is a schematic diagram illustrating a non-limiting example of a communications system, according to embodiments herein.
[0050] Figure 4 is a flowchart depicting embodiments of a method in a computer system, according to embodiments herein.
[0051] Figure 5 is a flowchart depicting embodiments of a method in a first node, according to embodiments herein.Figure 6 is a flowchart depicting embodiments of a method in a second node, according to embodiments herein.
[0052] Figure 7 is a schematic diagram depicting a non-limiting example of a computer system, according to embodiments herein.
[0053] Figure 8 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system, according to embodiments herein.
[0054] Figure 9 is a schematic diagram depicting a non-limiting example of aspects of a method in a communications system, according to embodiments herein.
[0055] Figure 10 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system, according to embodiments herein.
[0056] Figure 11 is a schematic diagram depicting a non-limiting example of aspects of a method in a communications system, according to embodiments herein.
[0057] Figure 12 is a schematic diagram depicting a non-limiting example of aspects of a method in a communications system, according to embodiments herein.
[0058] Figure 13 is a graphic illustration of aspects of a method in a communications system,
[0059] according to embodiments herein.
[0060] Figure 14 is a block diagram depicting details of another non-limiting example of aspects of a method in a communications system, according to embodiments herein.
[0061] Figure 15 is a block diagram depicting details of another non-limiting example of a communications system, according to embodiments herein.
[0062] Figure 16 is a schematic block diagram illustrating two non-limiting examples, a) and b), of a first node, according to embodiments herein.
[0063] Figure 17 is a schematic block diagram illustrating two non-limiting examples, a) and b), of a second node, according to embodiments herein.
[0064] Figure 18 is a schematic block diagram illustrating two non-limiting examples, a) and b), of a computer system, according to embodiments herein.
[0065] DETAILED DESCRIPTION
[0066] As part of the development of embodiments herein, the challenges with the existing technology mentioned in the Summary section will first be described in further detail.
[0067] The 5G core, as for example depicted in Figure 1 , may include many NFs that may communicate with each other through various interfaces for different procedures. Different procedures may trigger different NFs. Understanding these interactions and interdependencies among the NFs may be understood to be important to achieve efficient resource configuration. One example of a procedure involved in 5GC may be the registration procedure of the UE. The registration procedure may be initiated from the UE and communicated to the 5GC through a 5G base station (gNB). When the registration request is received by an AMF at the5GC, the AMF may initiate the default Packet Data Unit (PDU) session setup. At this stage, different 5GC NFs may be involved to handle the registration request, including SMF, PCF, and AUSF. The AMF may authenticate the UE after obtaining the keys from the AUSF and then, may obtain subscription data from the UDM. The AMF may then create a policyassociation with the PCF. The PCF may register with the AMF so that it may be notified on events such as location change and communication failure. The AMF may then update the SMF context and may send an initial context setup request to activate the default PDU session. The SMF may assign an IP address and the tunnel Identifier (ID) to be used for sending uplink data and may select the UPF to be used for the session. The AMF may notify the SMF when the session is ready for uplink and downlink data transfer. Once the session is ready, the gNB may enable security between the UE and itself and activate the default PDU session. At this point, downlink and uplink data streams may be ready to flow between the UE and the Internet. Similar procedures may be followed for other tasks, such as service requests, mobility management as well as session management.
[0068] When deploying a 5GC, there may often be requirements specified in terms of minimum rate of requests to be handled per unit time, such that the average latency of individual requests may be below a specific threshold. For a given set of requirements, it may be necessary to understand how much resource to allocate to individual NFs. If the resources of an individual NF introduce a bottleneck, this may degrade the entire performance of the 5GC. Hence, a key challenge may be understood to be to know how much resources to be allocated to individual NFs while considering the interdependencies between the NFs. The purpose of resource configuration may be understood to be to hence to ensure that the NFs within 5GC may be allocated the correct amount of resource in order to achieve the expected performance with respect to the various operation requirements.
[0069] In 5GC, achieving efficient resource allocation may be understood to be relevant to meet stringent performance requirements across various NFs. However, existing Al-based resource allocation approaches face some limitations. Many rely on pre-existing training data, which is often insufficient or unavailable for 5GC-specific operations, effecting accuracy of the model to changing network conditions and configurations. Online training, which may be understood to involve real-time data collection through iterative environment interactions may be understood to be costly, time-consuming, and may impose SLA violations. Even with distributed agents, this process may span weeks, making it impractical. Offline training methods may be understood to also be restricted by the lack of real-time profiling data that may reflect the current state of the 5GC, reducing the accuracy and applicability of these models.
[0070] In WO2024062273, a system and a method were proposed that were able to automatically generate resource configuration for workloads. The approach followed the RL-based approach. It obtained as input a VNF-FG, the target or the expected performance / KPIof VNFs in the VNF-FG, and the dependency relationship between the VNFs. It then identified the amount of resources that may have been needed to accommodate VNFs in a VNF-FG. This was performed by taking into consideration the inter-dependencies between the VNFs in a VNF-FG and the heterogeneity of the current infrastructures, e.g., accelerators. The system was able to achieve this by training it with different target performances and different applications.
[0071] In WO2025012677, a system and method were described for automated resource configuration. The described system may be understood to have mapped the performance specification in the SLA to an amount of resources in the infrastructure to ensure that the cloud infrastructure or the Network Functions Virtualization Infrastructure (NFVI) resources were utilized efficiently, and the target performance of the application represented by VNF-FG was met. The relevant steps in this method were estimating the performance of each VNF and / or microservice per unit resource, estimating the weights of dependencies between VNFs and / or microservices for performance evaluation, and estimating the target KPI of each VNF and / or microservice given a target KPI for the entire VNF-FG.
[0072] However, both these approaches may be understood to have relied on a simulator to avoid the overhead associated with trial and error. In such cases, the simulator may be understood to only have been able to estimate the characteristic of the application for a given application performance and exchanged workload, which may be understood to result in service dependency obtained from existing profiling data. Hence, the accuracy of the generated approach may be understood to degrade if the profiling data does not reflect the current situation of the system when generating resource configuration, e.g., rush hours to normal hours, daily change in traffic and workload, etc.. In addition, taking these approaches and applying them on real environments may introduce high overhead and may be time consuming, considering the long time required to deploy an application, waiting it to stabilize, and then collect data.
[0073] Addressing these limitations may require an approach that may enable rapid training of RL models, may minimize operational overhead by optimizing training processes, and may avoid complex emulation setups by leveraging the real production environment for training. By integrating real-world 5GC data, the accuracy of a model may be enhanced.
[0074] Certain aspects of the present disclosure and their embodiments address one or more of the challenges identified with the existing methods and provide solutions to the challenges discussed.
[0075] Embodiments herein may be understood to address the problems identified with the existing methods and may be understood to relate to a system and method for 5GC resource dimensioning. More particularly, some embodiments herein may relate to a system and a method that may choose the correct resource configuration for 5G Core Cloud-Native NetworkFunctions (CNFs) for given target Key Performance Indicator (KPI) requirements and load. In a particular non-limiting example, the described approach may achieve this through several modules: a Training Manager, an Action Manager, and several online Deployments. The Training Manager may be understood to be responsible fortraining an RL model and keeping track of the accuracy of that model. The Action Manager may be understood to be responsible for generating data safely to be used by the T raining Manager to train the model. Data may be generated by selecting actions to increase and / or decrease resources. Deployments may include a few tests and several production 5GC deployments where the production 5GC deployments may be used to provide connectivity services to several users under several Communications Service Provider (CSPs).
[0076] One of the two key challenges embodiments herein may address may be understood to be the need for a large amount of accurate runtime data needed to train the RL model, which may be solved by the use of operational, that is, live, deployments. The second, and consequential, challenge may be understood to be the need for SLA violation avoidance during the training which may be solved by careful RL action selection.
[0077] The embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which examples are shown. In this section, embodiments herein are illustrated by exemplary embodiments. It should be noted that these embodiments are not mutually exclusive. Components from one embodiment or example may be tacitly assumed to be present in another embodiment or example and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. All possible combinations are not described to simplify the description.
[0078] Figure 3 depicts two non-limiting examples, in panels “a” and “b”, respectively, of a communications system 100, in which embodiments herein may be implemented. In some example implementations, such as that depicted in the non-limiting example of Figure 3a, the communications system 100 may be a computer network. In other example implementations, such as that depicted in the non-limiting example of Figure 3b, the communications system 100 may be implemented in a telecommunications system, sometimes also referred to as a telecommunications network, cellular radio system, cellular network, or wireless communications system. In some examples, the telecommunications system may comprise network nodes which may serve receiving nodes, such as wireless devices. The communications system 100 may for example be a network such as a 5G system, or a newer system supporting similar functionality, such as for example, a Sixth Generation (6G) system. In some examples, the communications system 100 may support, additionally, a Long-Term Evolution (LTE) network and may support other technologies such as a for example, e.g., LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), LTE Half-Duplex Frequency Division Duplex (HD-FDD), and LTE operating in an unlicensed band. Thetelecommunications system may also support other technologies, such as Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System Terrestrial Radio Access (UTRA) TDD, Global System for Mobile communications (GSM) network, GSM / Enhanced Data Rate for GSM Evolution (EDGE) Radio Access Network (GERAN) network, Ultra-Mobile Broadband (UMB), EDGE network, network comprising any combination of Radio Access Technologies (RATs) such as e.g. Multi-Standard Radio (MSR) base stations, multi-RAT base stations etc., any 3rd Generation Partnership Project (3GPP) cellular network, Wireless Local Area Network / s (WLAN) or WiFi network / s, Worldwide Interoperability for Microwave Access (WiMax), IEEE 802.15.4-based low-power short-range networks such as IPv6 over Low-Power Wireless Personal Area Networks (6LowPAN), Zigbee, Z-Wave, Bluetooth Low Energy (BLE), or any cellular network or system. The telecommunications system may for example support a Low Power Wide Area Network (LPWAN). LPWAN technologies may comprise Long Range physical layer protocol (LoRa), Haystack, SigFox, LTE for Machines (LTE-M), and Narrow-Band loT (NB-loT).
[0079] The communications system 100 may comprise a computer system 101.
[0080] The communications system 100, e.g., as part of the communications system 101, may comprise a plurality of nodes, whereof a first node 111, a second node 112, a third node 113, a fourth node 114 and a fifth node 115 are depicted in Figure 3. It may be understood that the communications system 100 may comprise and / or operate in communication with more nodes than those represented on Figure 3.
[0081] Any of the first node 111, the second node 112, the third node 113, the fourth node 114 and the fifth node 115 may be understood, respectively, as a first computer system, a second computer system, a third computer system, a fourth computer system and a fifth computer system. In some examples, any of the first node 111, the second node 112, the third node 113, the fourth node 114 and the fifth node 115 may be implemented as a standalone server in e.g., a host computer in the cloud 120, as depicted in the non-limiting example depicted in panel b) of Figure 3.
[0082] Embodiments herein may be understood to be designed to work with cloud platforms and as a result may be well suited for a cloud implementation and usage.
[0083] Any of the first node 111, the second node 112, the third node 113, the fourth node 114 and the fifth node 115 may in some examples be a distributed node or distributed server, with, for example, some of their respective functions being implemented locally, e.g., by a client manager, and some of their functions implemented in the cloud 120, by e.g., a server manager. Yet in other examples, any of the first node 111 , the second node 112, the third node 113, the fourth node 114 and the fifth node 115 may also be implemented as processing resources in a server farm.The first node 111 may be understood to be a node having a capability to acquire data from other nodes, e.g., core NFs. In some examples, the first node 111 may further have the capability to analyze the acquire data, and expose them to the subscriber(s). In a particular non-limiting example, wherein the communications system 100 may be a 5G network, the first node 111 may be a Network Data Analytics Function, e.g., the NWDAF as defined in 3GPP TS 23.501 and TS 23.288, SA2. In further particular examples, the first node 111 may implement a Data Analytics module. In some examples, the first node 111 may also have a capability to select actions, to increase / decrease resources. This may be performed, e.g., via an Action Selector module.
[0084] The second node 112, in some examples may be a node having a capability to collect data exposed by the first node 111, e.g., via the Data Analytics module. The second node 112 may be, in some examples, a Training Data Generator.
[0085] The third node 113 may be understood to be a node having a capability to read data from a memory and keep track of the accuracy of an ML model for different ranges of parameters. The third node 113 may be, for example, a Training Director module.
[0086] The fourth node 114 may be understood to be a node having a capability to maintain an ML model, where it may be saved. In a particular non-limiting example, e.g., wherein the communications system 100 may be a 5G network, the fourth node 114 may be a Model Server.
[0087] The fifth node 115 may be a node having a capability to perform machine-implemented learning procedures, which may be also referred to as “machine learning” (ML). The fifth node 115 may have a capability tomanage an agents capable of training a machine learning model with a set of data. The fifth node 115 may have a capability to manage an artificial neural network. The artificial neural network may be understood as a machine learning framework, which may comprise a collection of connected sub-nodes, where in each sub-node or perceptron, there may be an elementary decision unit. Each such sub-node may have one or more inputs and an output. The input to a sub-node may be from the output of another subnode or from a data source. Each of the sub-nodes and connections may have certain weights or parameters associated with it. In order to solve a decision task, the weights may be learnt or optimized over a data set which may be representative of the decision task. The most commonly used sub-node may have each input separately weighted, and the sum may be passed through a non-linear function which may be known as an activation function. The nature of the connections and the sub-node may determine the type of the neural network, for example a feedforward network, recurrent neural network etc. That the fifth node 115 may have the capability to manage the artificial neural network may be understood herein as having the to train a machine learning model, and once the machine learning model may have been trained, to optionally use this model for prediction. The fifth node 115 may have acapability to use, e.g., execute, the machine learning model, once the machine learning model may have already been trained.
[0088] The fifth node 115 may support running python / Java with Tensorflow or Pytorch, Theano etc... The fifth node 115 may also have GPU capabilities.
[0089] Any of the first node 111 , the second node 112, the third node 113, the fourth node 114 and the fifth node 115 may co-localize, or be the same node.
[0090] The communications system 100 may also comprise a device 130. The device 130 may be also known as a e.g., user equipment (UE), wireless device, mobile terminal, wireless terminal and / or mobile station, mobile telephone, cellular telephone, or laptop with wireless capability, an Internet of Things (loT) device, or a Customer Premises Equipment (CPE), just to mention some further examples. The device 130 in the present context may be, for example, portable, pocket-storable, hand-held, computer-comprised, ora vehicle-mounted mobile device, enabled to communicate voice and / or data, via a RAN, with another entity, such as a server, a laptop, a Personal Digital Assistant (PDA), or a tablet, a Machine-to-Machine (M2M) device, an Internet of Things (loT) device, e.g., a sensor or a camera, a device equipped with a wireless interface, such as a printer or a file storage device, modem, Laptop Embedded Equipped (LEE), Laptop Mounted Equipment (LME), USB dongles, CPE or any other radio network unit capable of communicating over a radio link in the communications system 100. The device 130 may be wireless, i.e., it may be enabled to communicate wirelessly in the communications system 100 and, in some particular examples, may be able to support transmission using beamforming. The communication may be performed e.g., between two devices, between a device and a radio network node, and / or between a device and a server. The communication may be performed e.g., via a RAN and possibly one or more core networks, comprised, respectively, within the communications system 100.
[0091] The communications system 100 may comprise one or more radio network nodes, whereof a radio network node 140 is depicted in Figure 3b. The radio network node 140 may typically be a base station or Transmission Point (TP), or any other network unit capable to serve a wireless device or a machine type node in the communications system 100. The radio network node 140 may be e.g., a 5G gNB, a 4G eNB, or a radio network node in an alternative 5G radio access technology, e.g., fixed or WiFi. The radio network node 140 may be e.g., a Wide Area Base Station, Medium Range Base Station, Local Area Base Station, and Home Base Station, based on transmission power and thereby also coverage size. The radio network node 140 may be a stationary relay node or a mobile relay node. The radio network node 140 may support one or several communication technologies, and its name may depend on the technology and terminology used. The radio network node 140 may be directly connected to one or more networks and / or one or more core networks.The communications system 100 covers a geographical area which may be divided into cell areas, wherein each cell area may be served by a radio network node, although, one radio network node may serve one or several cells.
[0092] The first node 111 may communicate with the second node 112 over a first link 151, e.g., a radio link or a wired link. The first node 111 may communicate with the third node 113 over a second link 152, e.g., a radio link or a wired link. The first node 111 may communicate with any of the fifth node 115 over a third link 153, e.g., a radio link or a wired link. The second node 112 may communicate with the fourth node 114 over a fourth link 154, e.g., a radio link or a wired link. The first node 111 may communicate with the radio network node 140 over a fifth link 155, e.g., a radio link or a wired link. The second node 112 may communicate with the radio network node 140 over a sixth link 156, e.g., a radio link or a wired link. The fifth node 115 may communicate with the radio network node 140 over a seventh link 155, e.g., a radio link or a wired link. The radio network node 140 may communicate with the device 130 over an eighth link 158, e.g., a radio link or a wired link.
[0093] Any of the first link 151 , the second link 152, the third link 153, the fourth link 154, the fifth link 155, the sixth link 156, the seventh link 157 and / or the eighth link 158 may be a direct link or it may go via one or more computer systems or one or more core networks in the communications system 100, or it may go via an optional intermediate network. The intermediate network may be one of, or a combination of more than one of, a public, private, or hosted network; the intermediate network, if any, may be a backbone network or the Internet, which is not shown in Figure 3.
[0094] Although terminology from Long Term Evolution (LTE) / 5G has been used in this disclosure to exemplify the embodiments herein, this should not be seen as limiting the scope of the embodiments herein to only the aforementioned system. Other wireless systems supporting similar or equivalent functionality may also benefit from exploiting the ideas covered within this disclosure.
[0095] In future telecommunication networks, e.g., in the sixth generation (6G), the terms used herein may need to be reinterpreted in view of possible terminology changes in future technologies.
[0096] In general, the usage of “first”, “second”, “third”, “fourth”, “fifth”, “sixth”, “seventh” and / or “eighth” herein may be understood to be an arbitrary way to denote different elements or entities and may be understood to not confer a cumulative or chronological character to the nouns they modify.
[0097] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject matter disclosed herein, the disclosed subject matter should not be construed as limited to only the embodiments set forth herein; rather, these embodiments areprovided by way of example to convey the scope of the subject matter to those skilled in the art.
[0098] Embodiments of a computer-implemented method, performed by the computer system 101 comprising the first node 111 and the fifth node 115, will now be described with reference to the flowchart depicted in Figure 4. The method may be understood to be for handling configuration of first resources. The computer system 101 operates in the communications system 100.
[0099] In some embodiments, the communications system 100 may be a 3GPP network. In some embodiments, the communications system 100 may be a 5G network.
[0100] Several embodiments are comprised herein. In some embodiments, all the actions may be performed. In some embodiments, three or more actions may be performed. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description. Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. A non-limiting example of the method performed by the computer system 101 is depicted in Figure 4. The actions of the method performed by the computer system 101 may be performed in a different order than that depicted in Figure 4.
[0101] In Figure 4, optional actions are represented with dashed lines.
[0102] Action 401
[0103] It may be noted that the ultimate goal of embodiments herein may be understood to be to train an ML model, e.g., an RL model, to determine a configuration of first resources in the communications system 100. The first resources may be for dimensioning a network function virtualization system of the communications system 100. As a non-limiting example, the RL may be trained to choose the correct resource configuration for 5G Core Cloud-Native Network Functions (CNFs) for given target Key Performance Indicator (KPI) requirements and load.
[0104] Forthat purpose, in this Action 401, the computer system 101 obtains, by the first node 111, respective data from one or more deployments in the communications system 100.
[0105] A deployment may be understood as network infrastructure, e.g., antennas, base stations, software systems, etc. that may have been implemented. The deployment may comprise physical devices, installed software components, and network configuration.
[0106] The one or more deployments may be operational deployments. Operational deployments may include both production deployment and test deployments.The one or more deployments may comprise or be core deployments. These deployments may be understood to be mostly production environments where 5GC CNFs may be running.
[0107] In some examples, the one or more deployments may also comprise some test / laboratory environments.
[0108] The data obtained may be understood to be respective data, as the first node 111 may obtain data from each of the one or more deployments.
[0109] The data may be understood to comprise current KPI of the deployment, e.g., average control or signalling latency, data packet processing latency and reliability, current resource configuration of CNFs, and the load on the system. For example, the data may comprise current KPIs of deployments, current resources, and the load on the system.
[0110] Obtaining may be understood as receiving, collecting, fetching, retrieving or similar. Action 401 may be performed by a Data Analytics function or module. The first node 111, in this Action 401 may monitor and gather the data from the one or more deployments.
[0111] In some examples, the first node 111, e.g., via the Data Analytics module, may acquire data from core NFs of the communications system 100. For 5GC deployments, this module is implemented by NWDAF.
[0112] Action 402
[0113] The first node 111 may then read the runtime data as obtained in Action 401 , e.g., from the Data Analytics module. This data may include current KPI of the deployment, current resource utilization, and current load. Then the the first node 111, e.g., via an Action Selector module, may analyze the data. Particularly, the first node 111 may check whether there may be new information useful for training the ML model, such as changes in the load or changes in the KPIs.
[0114] The ML model may be understood to comprise one or more variables. Each of the one or more variables may have different values or ranges of values.
[0115] In this Action 402, the computer system 101 determines, by the first node 111, whether or not the obtained respective data comprises a combination of the values of the one or more variables used to train the ML model to determine the configuration of first resources in the communications system 100.
[0116] Determining may be understood as calculating, estimating, deriving or similar.
[0117] The combination of values lacks a similarity with first information previously used to train the ML model according to a criterion of similarity.
[0118] In some embodiments, the combination of values may be determined to lack the similarity based on one of the following options. According to one option, the combination of values may be determined to lack the similarity based on an Euclidean distance betweenrespective vectors representing the combination of values and the first information exceeding a first threshold. According to another option, the combination of values may be determined to lack the similarity based on an error output by the ML model when using the combination of values as input exceeding a second threshold.
[0119] As mentioned earlier, in some embodiments, one or more of the following may be applied. According to one option, the first resources may be for dimensioning the network function virtualization system of the communications system 100.
[0120] According to another option, the one or more deployments may be operational deployments.
[0121] In some examples, the one or more deployments may comprise some test / laboratory environments that may be used by the first node 111, e.g., via the Action Selector, for further exploration.
[0122] Action 403
[0123] In this Action 403, the computer system 101 initiates, by the first node 111, saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the obtained respective data comprises the combination of values. That is the combination of values that lacks the similarity with the first information previously used to train the ML model according to the criterion of similarity.
[0124] Initiating may be understood as triggering, enabling, facilitating, the enforcing by another node, or starting the managing itself.
[0125] The initiating saving in this Action 403 may comprise sending a first indication to the second node 112 operating in the communications system 100. The first indication may instruct the second node 112 to save the combination of values. This may be for future training. That is, in this Action 403, the first node 111 may expose the data, e.g., the combination of values to one or more subscriber(s). This data may be exposed to the second node 112.
[0126] This Action 403 may be performed by, e.g., the Action Selector. The second node 112 may be, in some examples, a Training Data Generator. The second node 112 may be understood to be responsible of collecting the data from core deployments exposed by the first node 111, e.g., via the Data Analytics module.
[0127] Action 404
[0128] In this Action 404, the computer system 101 may receive, by the second node 112, the first indication from the first node 111. The first indication may instruct the second node 112 to save the combination of values. The combination of values may be saved in a memory, which may be referred to herein as a Replay Memory. The memory wherein the combination ofvalues may be saved, e.g., the Replay Memory, may be understood to be a memory, which may store, e.g., RL agent experiences. An RL agent may be trained by selecting a batch from this memory buffer to form a learning batch.
[0129] The second node 112 may perform some pre-processing before saving the combination of values, to be used as training data for the ML model, in the Replay Memory to make the data usable by the ML model, e.g., the RL model.
[0130] Action 405
[0131] In this Action 405, the computer system 101 may send, by the second node 112, an indication referred to herein as a third indication, to a fourth node 114 comprised in the computer system 101. The third indication may request the fourth node 114 to provide a reward corresponding to the combination of values. The reward may be to be used to train the ML model to determine the configuration of first resources in the communications system 100.
[0132] The fourth node 114 may be a Model Server. The fourth node 114 may be understood to be responsible for maintaining the ML model, e.g., the RL model, where it may be saved and used for inferencing. The fourth node 114 may give access to the ML model. In one example, the ML model may then be pushed to the different deployments to be hosted and used locally by each deployment. In another example, the model may be hosted in a central place where it may be accessed through inferencing request. Depending on the flavor of the ML model, e.g., the RL, the fourth node 114 may be responsible for more than one model. For instance, in the case of Deep RL (DRL), the fourth node 114 may be responsible of a Q Network model and a Target Network model.
[0133] In some examples, the third indication may be an inferencing request. Particularly, the second node 112 may send the inferencing request to the fourth node 114, to get the reward for the multiple generated data points.
[0134] Based on the received third indication, the fourth node 114 may then calculate the reward based on the performance of the respective one or more deployments. The determination of the reward may be performed, for example in terms of the selected resource configuration and the given target KPI, which may be considered as the objective function. The target KPI may be latency, throughput, or packet loss, for instance. The action leading to a better objective function may be associated with a larger reward. In RL, the long-term cumulative reward may be the optimization target of RL. And the immediate rewards may be the performance feedback of the CNF on the resulted new resource configuration.
[0135] In one example, the reward may be calculated based on how close the system may be to meeting the target KPI for control latency, data latency, and reliability. For each of these target KPIs, an error may be calculated based on the deviation from the target KPI. Then thereward may be measured by the sum of these errors.
[0136] Action 406
[0137] In this Action 406, the computer system 101 may receive, by the second node 112, the reward from the fourth node 114. The ML model may be trained with the received reward.
[0138] In an example, the reward may be returned in this Action 406 to the Training Data Generator.
[0139] Action 407
[0140] In this Action 407, the computer system 101 may save, by the second node 112, the combination of values and the received reward in a memory of the communications system 100. The memory may be, for example, the Replay Memory.
[0141] In some examples, the second node 112 may save information indicating the state, next state, and reward. Some pre-processing may be required at this stage to make the data suitable for the ML model, e.g., the RL, model.
[0142] Action 408
[0143] In this Action 408, the computer system 101 may determine, by a third node 113 comprised in the computer system 101, another indication, referred to herein as a second indication, to be sent to the first node 111.
[0144] The third node 113 may be, for example, a T raining Director module. The third node 113, e.g., the Training Director may be understood to be responsible for reading data from another memory or second memory, that may be referred to herein as a Deployment Parameters Repository. The third node 113 may keep track of the accuracy of the model for different ranges of parameters. Based on the model accuracy, the third node 113 may decide for each deployment, of the one or more deployments, how to select actions. This information may then be instructed to the first node 111, e.g., the Action Selector.
[0145] The second indication may correspond to a three-phase approach to be performed by the first node 111, e.g., via the Action Selector.
[0146] It may be noted that the ultimate goal may be understood to be to train the ML model, e.g., an RL model that may be able to accurately calculate the rewards for all possible actions that may be taken for a given deployment e.g., of 5GC. For this to happen, the ML model trainer, that is, the fifth node 115, may need to be fed with data that may explore all possible combinations of inputs e.g., resource allocations, load, service KPI, etc.. A goal of the first node 111, e.g., via the Action Selector, may be understood to be to ensure that these combinations may be explored in a distributed manner. A challenge in this process, may be that since some combinations may result in violating the SLA of a particular deployment, themethod may need to ensure that such combinations may not be applied to a production environment.
[0147] As an overview, the third node 113, e.g., via the Training Director, may act differently towards production deployments and test deployments. For the production deployments, third node 113 may start with a first phase, referred to herein as the green phase. Once the third node 113 may be confident that the model is accurate in this area, the third node 113 may instruct the first node 111, e.g., via the Action Selector, to enter a second phase and explore an area referred to herein as the orange area. For the test deployments, the third node 113, e.g., via the Training Director, may instruct the first node 111, e.g., via the Action Selector, to enter a third phase and explore other areas referred to herein as the so-called red areas, where SLAs may be violated. What the second indication may indicate may depend to the different phases just explained.
[0148] The second indication may indicate one of the following.
[0149] Phase 1 - Green Phase:
[0150] According to one option, the second indication may indicate that the first node 111 is to select a respective action to be taken on a respective set of second resources to be used by the communications system 100 on the respective one or more deployments. Particularly, according to this one option the second indication may indicate that the first node 111 is to select the respective action within a respective first action space. The first action space may correspond, e.g., map, to respective one or more parameters that may be defined in a repository of parameters for the respective one or more deployments. The respective one or more deployments may be production deployments. The repository of parameters, or Deployment Parameters Repository, may include the original configuration settings for each deployment, such as, e.g., the minimum amount of resources that may be required for each 5GC CNF, and the maximum amount of resources that may be used for 5GC CNFs in that deployment. This may be given by an expert or defined based on a dimensioning tool and may be based on statistics. It may be noted that the minimum amount may be generally determined by dimensioning the resources for some expected peak load and some additional safety margin, making the system wasteful at most times and under dimensioned for when the system may experience an unexpected surge in load.
[0151] In the first or green phase, the focus may be understood to be on exploring non-critical, safe operational areas within the defined parameters in the Deployment Parameters repository for the different deployments. It may be noted that here, production deployments, e.g., 5GC deployments may be used, that may be already giving service to end users, to train the ML model, e.g., the RL model. An action taken may either result in an increase or decrease of the resource(s) allocated to the communications system 100, e.g., to 5GC CNF(s). In the green phase, the method may ensure that the resulting resource allocated to any CNF may be withina prescribed limit that may be retrieved from the Deployment Parameter repository. In the context of an s-greedy policy of generating actions, this may mean that if the resulting action may cause the resources allocated to the communications system 100, e.g., the CNFs, to fall outside of the allowable range, that action may be rejected and another action may be generated repeatedly until a suitable action is found.
[0152] A challenge in this phase may be the unexplored areas within the deployment space. For such unexplored areas, the ML model, e.g., the RL model, may struggle to make accurate estimations due to lack of data. For instance, certain configurations with too small resources may not be explored because such resource reduction may violate the SLA. Alternatively, unexplored areas may be found when the target KPI may be very stringent.
[0153] For each deployment, the third node 113, e.g., via the Training Director, may estimate the ML model accuracy for all resource values allowed in the Deployment Parameter repository. Once the ML model may be deemed accurate for these values, then third node 113 may instruct the first node 111, e.g., the Action Selector, to move to the orange phase.
[0154] At the end of this phase, an ML model may be obtained that may be accurate for resource values allowed in the repository.
[0155] Phase 2 - Orange Phase:
[0156] According to another option, the second indication may indicate that the first node 111 may have to select the respective action within the respective first action space corresponding to the one or more respective parameters defined in the repository of parameters for the one or more deployments, applying a safety margin. The respective one or more deployments may be production deployments.
[0157] The safety margin may allow reducing resources to more than the minimum required amount of resources without violating the SLA. For example, if the target KPI is latency of 10ms, then a safety margin of 10% may mean that it may be OK for the latency to be at most 9ms, while a safety margin of 80% may mean that the latency may be at most 2ms.
[0158] The third node 113 may also be responsible for deciding the safety margin to be used by the first node 111, e.g., the Action Selector of each deployment.
[0159] In the orange phase, the third node 113, e.g., via the Training Director, may allow the first node 111, e.g., the Action Selector, to explore parameter spaces outside of those that may be defined in the Deployment Parameter repository, subject to safety margins.
[0160] Specifically, the first node 111, e.g., via the Action Selector, may be allowed to reduce the resource(s) allocated CNFs to values below the minimum prescribed in the repository, such that the expected KPIs may not violate the safety margin. The third node 113, e.g., via the Training Director, may gradually decrease the safety margin to a configurable value, e.g., 10%, as it may get more confident about the accuracy of the ML model.To do this without violating the SLA for a given deployment a description is provided in Action 412.
[0161] When the third node 113, e.g., via the Training Director, may be confident that the ML model may be accurate enough for the current safety margin, third node 113 may gradually reduce the safety margin and continue the process. The reduction may be implementation dependent. In one example, the margin may be reduced in fixed steps, e.g., 5%. In another, the margin may be reduced in relation to the accuracy of the ML model.
[0162] The reduction of the margin may be performed repeatedly until a configured target safety margin may be reached. Alternatively, the margin may be increased if a current safety margin is violated because of model inaccuracy.
[0163] Once the configured minimum target safety margin may be reached, the system may continue to run the Action Selection in exploitation mode, e.g., e=0 for epsilon greedy algorithm, while monitoring the accuracy of the RL model. If it detects model inaccuracy, e.g., caused by a change in traffic mix, it may increase the safety margin to a safe value.
[0164] Phase 3 - Red Phase:
[0165] According to yet another option, the second indication may indicate that the first node 111 may have to select the respective action within a respective second action space that is to result in a violation of a Service Level Agreement (SLA) if applied in the respective one or more deployments. The second indication may further indicate that the first node 111 may have to apply the selected respective action from the respective second space in a test deployment.
[0166] The goal of this phase may be understood to be to explore parameter combinations that may likely result in SLA violations in production deployments. The third node 113, e.g., via the Training Director, may instruct the first node 111, e.g., via the Action Selector, of only test deployments to run in this phase. The exact way the actions may be taken in this phase may be implementation dependent. In one example, the ML model, e.g., the RL, may be run normally without any safety margins, for randomly selected SLA targets. In another example, the first node 111, e.g., via the Training Director, may select the SLA targets and reverse safety margins where the goal may be to take actions such that the performance KPI may be made to be worse than the target SLA. For example, if the target KPI was 10ms, a reverse margin of 20% may require that the actual latency may be above 12ms.
[0167] Action 409
[0168] In this Action 409, the computer system 101 may send, by the third node 113, the determined second indication, to the first node 111.Action 410
[0169] In this Action 410, the computer system 101 may receive, by the first node 111 , the second indication from the third node 113.
[0170] Action 411
[0171] In this Action 411 , the computer system 101 may select, by the first node 111 , a respective action to be taken on a respective set of second resources to be used by the communications system 100 on the respective one or more deployments, with the proviso that the initiating in Action 403 of the saving of the combination of values may have been or the result of the determination may be that the obtained respective data lacks the combination of values. In such embodiments, the selecting in Action 411 may be based on one or more respective parameters from the respective one or more deployments.
[0172] In some embodiments, the selecting in this Action 411 of the respective action may be based on the received second indication.
[0173] The first node 111, e.g., via the Action Selector, may be responsible for selecting actions to increase and / or decrease resources.
[0174] The details of how specific actions may be normally selected for specific deployments may be implementation dependent. In one example, action selection may be performed based on traditional epsilon greedy (s--greed) policy. Traditional s--greed policy may be defined as follows:
[0175] a. With probability s, the first node 111 may select a random action a, called exploration.
[0176] b. With probability 1-s, the first node 111 may select an action that may have a maximum reward, called exploitation.
[0177] The exploration probability may generally be gradually reduced from a high value, e.g., 1, to a small value, e.g., 0.05, during the lifetime of the training. The specific way may be implementation dependent.
[0178] The first node 111, e.g., via the Action Selector, may obtain its deployment parameters from the Deployment Parameters Repository, such as the target KPI, the minimum resource configurations of the given deployment’s CNFs, etc. Based on these parameters, the first node 111, e.g., via the Action Selector, may select an action. Selecting an action, that is, decreasing / increasing resources, may be performed in accordance with the methodology described next.
[0179] The actions may be selected using different strategies. One example may be following pre-defined scaling factors e.g., [0.5, 1.2], where 0.5 may be understood to mean decrease the resources amount by half, and 1.2 may be understood to mean increase the resources amountby a factor of 1.2. However, increasing and / or decreasing the resources may be performed in a way to not violate the SLA, as described in Action 411.
[0180] To enable that the third node 113, e.g., via the Training Director, may gradually decrease the safety margin to a configurable value without violating the SLA for a given deployment, the first node 111, e.g., via the Action Selector, may perform the following, illustrated later in Figure 5.
[0181] The third node 113, e.g., via the Training Director, may start by defining a large-enough safety margin. In one example, this may be manually specified. In another, this may be determined by using the minimum resources on all CNFs and estimating and / or measuring the resulting KPI and calculating the margin between the target KPI and the resulting KPI.
[0182] In another example, where several KPIs may be used in the SLA, the KPI closest to the SLA target may be used to define this margin. This safety margin may be assigned to the first node 111, e.g., the Action Selector.
[0183] The first node 111, e.g., via the Action Selector, may estimate the resulting KPI from an action using different methods. In one example, the first node 111 may use the estimated reward it may get from the ML model, e.g., the RL model, as an indication. This may be the reward calculated by the fourth node 114 as described in action 405 In another example, the first node 111 may use spline approximation to estimate the KPI based on current and historical values. In a third one, the first node 111 may use a combination of both approaches.
[0184] In one example, the problem of accurately calculating the rewards for all possible actions that may be taken for a given deployment may be modeled as a MDP, which may be understood to be a mathematical framework for modeling decision-making problems. In other examples, the problem of accurately calculating the rewards for all possible actions that may be taken for a given deployment may be modeled as Markov Games where each agent may have its own set of actions.
[0185] An MDP may be defined by a tuple (S, A, p, r), where S may be understood to be a finite set of states, A may be understood to be a finite set of actions, p may be understood to be a transition probability from state s to state s’ after action a may have been executed, and r may be understood to be the immediate reward obtained after action a may have been performed. A term IT may be denoted as a “policy” which may be a mapping from a state to action. The goal of an MDP may be understood to be to find an optimal policy by training agents to observe the state of the environment and take actions to ultimately maximize the reward function. The state space, action space, and reward function may be defined as follows.
[0186] A) State Space: considering there may be n NFs in a network service, e.g., CNFs in 5GC, in one example, the state space may be composed of the current resources configuration (res) of each NF, the current control / data latency and reliability of the networkservice, the target control / data latency and reliability of the given network service, and the load of the system. The state then may be represented by:
[0187] state = {res-, res2, resn, ctrllatency, datalatency, reliability, T_ctrliatency, T_dataiatency, T —reliability, load~^
[0188] For the 5GC example, NF1 may be considered to beAMF. Then, rest may be the number of CPU cores allocated to each AMF, control latency may be a measure of the delay in the signaling messages and control commands in the core and the data latency may refer to the delay experienced in the actual transmission of data packets between a UE and data endpoints through the core. The reliability may be understood to refer to the core ability to deliver uninterrupted service to UEs under the given load. The load may be understood to represent the number of UEs per second, which may refer to the rate at which the user devices may initiate network interactions.
[0189] B) Action Space: selecting an action in Action 411 may be understood to mean selecting or changing a resource configuration for NFs and moving the system to a new state space. In one example, an action may indicate increasing / decreasing a resource of a NF by x% with predefined minimum and maximum values, or not changing the resource value. Each action may cause the state to change to a new state. For the above example, a joint action may be considered, which may be composed of the actions to be applied to each NF in the network service.
[0190] A [alfa2, ... an]
[0191] In one example, an action, aron NF 1 (AMF) may indicate increase or decrease the resources of AMF by one of the following pre-defined scaling factors: [0.5, 1, 2], Here, 0.5 may be understood to mean decrease the resources by half, e.g., decrease from 3 CPU cores to 1.5, 1 may be understood to mean keep the current resource configuration, and 2 may be understood to be mean double the resources, e.g., increase from 3 to 6 CPU cores.
[0192] In another example, an action a on NF 1 (AMF) may indicate adding or subtracting specific values for resources. For example, instead of scaling by predefined factors, the resources of NF1 (AMF) may be increased or decreased by specific amount, such as adding 1 CPU core or subtracting 0.5 CPU cores.
[0193] Yet, in a third example, a hybrid approach may be used, where both predefined scaling factors and fixed value adjustments may be combined. For instance, resources may first be scaled by a factor, e.g., doubled, or halved, and then fine-tuned by adding or subtracting specific values, e.g., adding 0.5 CPU cores after doubling. This may allow flexibility and precision in resource dimensioning.
[0194] C) Reward Function: the reward function may be used to guide the RL agent(s) to find the optimal resource configuration. The reward may be calculated in terms of the selectedresource configuration and the given target performance, which may be seen as the objective function. The target performance part may be optimizing latency, throughput, or packet loss, for instance. The action leading to a better objective function may be associated with a larger reward. In RL, the long-term cumulative reward may be understood to be the optimization target of RL. And the immediate rewards may be the VNFs performance feedback on the resulted new resource configuration.
[0195] In one example, the reward may be calculated based on how close the system may be to meeting the target performance for control latency, data latency, and reliability. For each of these target performances, an error (rel_err) may be calculated based on the deviation from the target. Then the reward may be measured by the sum of these errors.
[0196] <
[0197]
[0198] Action 412
[0199] In this Action 412, with the proviso the one or more deployments are production deployments, the computer system 101 may refrain, by the first node 111, from selecting a first action, with the proviso the first action may be to violate the SLA.
[0200] As stated earlier, for each action proposed by the action generation method used, e.g., epsilon-greedy, the first node 111, e.g., via the Action Selector, may estimate, based on the reward calculated by the fourth node 114 as described in Action 405, whether the KPI that may result after taking the action may violate the safety margin or not. If it does, the first node 111 may reject that action and continue to generate and check until it may find a suitable one and apply it to the deployment.
[0201] Action 413
[0202] In this Action 413, the computer system 101 may apply, by the first node 111 , the selected action on the respective one or more deployments as part of a training of the machine learning model.
[0203] The method in Actions 401-413 may be iterated until no more data points may be needed.Action 414
[0204] In this Action 414, the computer system 101 trains, by the fifth node 115, the machine learning model with the saved combination of values.
[0205] The fifth node 115, e.g., via a Model Trainer, may be responsible for deciding when to train the ML model based on the amount of data availability in the Replay Memory and the defined batch size, training the ML model, and updating the ML model in the fourth node 114, e.g., the model server.
[0206] The fifth node 115, e.g., via a Model Trainer, may be monitor the Replay Memory. The fifth node 115 may then check whether there may be sufficient data to retrain the ML model. The sufficiency criterion may be included in the Model & Training Configuration Data repository. Once there may be sufficient data in the Replay Memory, the fifth node 115, e.g., via a Model Trainer, may retrieve the ML model(s) from the fourth node 114, e.g., the Model Server. The fifth node 115, e.g., via a Model Trainer, may then update this ML model(s).
[0207] Updating an ML model may mean retraining an ML model using the data in the Replay Memory or updating the weights of the target network, which may be the case when using DRL. Once the model may be trained, it may be updated in the Model Server to be used for inferencing.
[0208] The fifth node 115 may manage a Training Manager. The training manager may comprise, e.g., in a database, model & training configuration data. This database may comprise the data and metadata that may be required for managing and training the RL model. It may include the model architecture details, e.g., number of layers, neurons, etc., model metadata, e.g., framework used such as TensorFlow, required libraries, etc., hyperparameters, e.g., learning rate, batch size, etc.
[0209] Action 415
[0210] In this Action 415, the computer system 101 outputs, by the fifth node 115, a third indication of the trained machine learning model.
[0211] Embodiments of a computer-implemented method, performed by the first node 111, will now be described with reference to the flowchart depicted in Figure 5. The method may be understood to be for handling configuration of first resources. The first node 111 operates in the communications system 100.
[0212] In some embodiments, the communications system 100 may be a 3GPP network. In some embodiments, the communications system 100 may be a 5G network.
[0213] Several embodiments are comprised herein. In some embodiments, all the actions may be performed. In some embodiments, three or more actions may be performed. It should benoted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description. Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. A non-limiting example of the method performed by the first node 111 is depicted in Figure 5. The actions of the method performed by the first node 111 may be performed in a different order than that depicted in Figure 5.
[0214] In Figure 5, optional actions are represented with dashed lines.
[0215] The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the computer system 101 and will thus not be repeated here to simplify the description. For example, the ML model may be an RL model.
[0216] Action 501
[0217] In this Action 501 , the first node 111 obtains the respective data from the one or more deployments in the communications system 100.
[0218] Action 502
[0219] In some embodiments, in this Action 502, the first node 111 determines whether or not the obtained respective data comprises the combination of values of one or more variables used to train the machine learning model to determine the configuration of first resources in the communications system 100. The combination of values lacks a similarity with the first information previously used to train the machine learning model according to the criterion of similarity.
[0220] In some embodiments, one or more of the following may be applied. According to one option, the first resources may be for dimensioning the network function virtualization system of the communications system 100.
[0221] According to another option, the one or more deployments may be operational deployments.
[0222] In some embodiments, the combination of values may be determined to lack the similarity based on one of the following options. According to one option, the combination of values may be determined to lack the similarity based on the Euclidean distance between respective vectors representing the combination of values and the first information exceeding the first threshold. According to another option, the combination of values may be determined to lack the similarity based on the error output by the machine learning model when using the combination of values as input exceeding the second threshold.Action 503
[0223] In this Action 503, the first node 111 initiates saving the combination of values to train the machine learning model, with the proviso the result of the determination is that the obtained respective data comprises the combination of values.
[0224] Initiating may be understood as triggering, enabling, facilitating, the enforcing by another node, or starting the managing itself.
[0225] The initiating saving in this Action 503 may comprise sending the first indication to the second node 112 operating in the communications system 100. The first indication may instruct the second node 112 to save the combination of values.
[0226] Action 504
[0227] In this Action 504, the first node 111 may receive the second indication from the third node 113 operating in the communications system 100. The second indication may indicate one of: i) the first node 111 is to select the respective action within the respective first action space corresponding to the respective one or more parameters defined in the repository of parameters for the respective one or more deployments, wherein the respective one or more deployments may be production deployments, ii) the first node 111 is to select the respective action within the respective first action space corresponding to the one or more respective parameters defined in the repository of parameters for the one or more deployments, applying the safety margin, wherein the respective one or more deployments may be production deployments, and iii) the first node 111 is to select the respective action within the respective second action space that is to result in a violation of the SLA, if applied in the respective one or more deployments, and wherein the first node 111 may have to apply the selected respective action from the respective second space in the test deployment. The selecting in Action 505 of the respective action may be based on the received second indication.
[0228] Action 505
[0229] In this Action 505, the first node 111 may select the respective action to be taken on the respective set of second resources to be used by the communications system 100 on the respective one or more deployments, with the proviso that the initiating in Action 503 of the saving of the combination of values may have been or the result of the determination may be that the obtained respective data may lack the combination of values. The selecting in Action 505 may be based on one or more respective parameters from the respective one or more deployments.Action 506
[0230] In this Action 506, with the proviso the one or more deployments may be production deployments, the first node 111 may refrain from selecting the first action, with the proviso the first action may be to violate the SLA.
[0231] Action 507
[0232] In this Action 507, the first node 111 may apply the selected action on the respective one or more deployments as part of the training of the machine learning model.
[0233] In some embodiments, the combination of values may be determined to lack the similarity based on one of the following options. According to one option, the combination of values may be determined to lack the similarity based on the Euclidean distance between respective vectors representing the combination of values and the first information exceeding the first threshold. According to another option, the combination of values may be determined to lack the similarity based on the error output by the machine learning model when using the combination of values as input exceeding the second threshold.
[0234] The method in Actions 501-507 may be iterated until no more data points may be needed.
[0235] Embodiments of a computer-implemented method performed by the second node 112, will now be described with reference to the flowchart depicted in Figure 6. The method may be understood to be for handling the configuration of first resources. The second node 112 operates in the communications system 100.
[0236] The method may comprise the following actions. Several embodiments are comprised herein. In some embodiments, the method may comprise all the actions. In other embodiments, the method may comprise three or more actions. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description. It should be noted that the examples herein are not mutually exclusive.
[0237] Components from one example may be tacitly assumed to be present in another example and it will be obvious to a person skilled in the art how those components may be used in the other examples.
[0238] The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the computer system 101 and will thus not be repeated here to simplify the description. For example, the ML model may be an RL model.Action 601
[0239] In this Action 601, the second node 112 receives the first indication from the first node 111 operating in the communications system 100. The first indication instructs the second node 112 to save the combination of values of the one or more variables used to train the machine learning model to determine the configuration of the first resources in the communications system 100.
[0240] The first resources may be for dimensioning the network function virtualization system of the communications system 100.
[0241] Action 602
[0242] In this Action 602, the second node 112 sends the third indication to the fourth node 114 operating in the communications system 100. The third indication requests the fourth node 114 to provide the reward corresponding to the combination of values. The reward being to be used to train the machine learning model to determine the configuration of the first resources in the communications system 100.
[0243] Action 603
[0244] In this Action 603, the second node 112 receives the reward from the fourth node 114.
[0245] Action 604
[0246] In this Action 604, the second node 112 may save the combination of values and the received reward in a memory of the communications system 100.
[0247] Figure 7 is a schematic diagram depicting a non-limiting example of the computer system 101 operating in the communications system 100, according to embodiments herein. In this non-limiting example, the communications system 100 is a 5G network. The communications system 100 may comprise deployments 701. The deployments 701 may comprise the Deployment Parameters Repository 702, which may include the original configuration settings for each deployment, such as the minimum amount of resources required for each 5GC CNF and the maximum amount of resources that may be used for 5GC CNFs in that deployment. This may be given by an expert or defined based on a dimensioning tool. It may be noted that the minimum amount may be generally determined by dimensioning the resources for some expected peak load and some additional safety margin, making the system wasteful at most times and under dimensioned for when the system may experience unexpected surge in load. The deployments 701 may comprise a Core Deployment 703. These deployments may be mostly production environments where 5GC CNFs may be running. The deployments 701 may also comprise some test and / or laboratory environmentsthat may be used by the first node 111, e.g., as Action Selector, for further exploration. The Core Deployment 703 may comprise the Data Analytics module of the first node 111. The Data Analytics module of the first node 111 may acquire data from core NFs, analyze them, and expose them to the subscriber(s). This data may be exposed, e.g., via the first node 111, for example, via the Action Selector, to a Training Data Generator module of the second node 112 and saved in the Replay Memory 704. For 5GC deployments, the Data Analytics module of the first node 111 may be implemented by NWDAF. The communications system 100 may comprise a Training Manager 705. The Training Manager 705 may comprise the Model & Training Configuration Data database 706. The Model & Training Configuration Data database 706 may comprise the data and metadata that may be required for managing and training the RL model. It may include the model architecture details, e.g., number of layers, neurons, etc., model metadata, e.g., framework used such as TensorFlow, required libraries, etc., hyperparameters, e.g., learning rate, batch size, etc. The Training Manager 705 may comprise the Model Trainer module of the fifth node 115. The Model Trainer module of the fifth node 115 may be understood to be responsible of deciding when to train the model based on the amount of data availability in the Replay Memory 704 and the defined batch size, training the model, and updating the model in the model server module of the fourth node 114. The Training Manager 705 may also comprise the Replay Memory 704, which may be understood to be a memory, which may store the RL agent experiences. The RL agent may be trained by selecting a batch from this memory buffer to form a learning batch. The communications system 100 may further comprise an Action Manager 707. The Action Manager 707 may comprise the Training Director module of the third node 113. The Training Director module of the third node 113 may be understood to be responsible of reading data from the Deployment Parameters Repository 702. The Training Director module of the third node 113 may keep track of the model accuracy for different ranges of parameters. Based on the model accuracy, the Training Director module of the third node 113 may decide for each deployment how to select actions, and this information may be instructed to the Action Selector. It may also be responsible of deciding the safety margin to be used by the Action Selector of each deployment. The safety margin may allow reducing resources to more than the minimum required amount of resources without violating the SLA. The Action Manager 707 may further comprise the Model Server module of the fourth node 114, which may be understood to be responsible for maintaining the RL model, where it may be saved and used for inferencing. This module may provide access to the ML model. In one example, the model may be pushed to the different deployments to be hosted and used locally by each deployment. In another example, the model may be hosted in a central place where it can be accessed through inferencing request. Depending on the flavor of the RL, the module may be responsible for than one model. For instance, in the case of Deep RL (DRL), this module maybe responsible of a Q Network model and a Target Network model. The Action Manager 707 may further comprise the Training Data Generator 704, which may be responsible for collecting the data from core deployments exposed by Data Analytics module, such as current KPIs of deployments, current resources, and the load on the system. It may perform some preprocessing before saving the training data in the Replay Memory 704 to make the data usable by the RL model. The Action Manager 707 may further comprise the Action Selector module of the first node 111, which may be understood to be responsible for selecting actions to increase / decrease resources. The actions may be selected using different strategies, one example may be following pre-defined scaling factors e.g., [0.5, 1.2], where 0.5 may mean decrease the resources amount by half, and 1.2 may mean increase the resources amount by a factor of 1.2. However, increasing / decreasing the resources may need to be done in a way to not violate the SLA, as previously described.
[0248] Figure 8 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system 100 to trigger action selection, according to embodiments herein. These steps may be understood to happen per deployment. At 1 , the first node 111, via the Data Analytics module, may, according to Action 401 and Action 501, monitor and gathers data from the given deployment. At 2, the first node 111 , via the Action Selector module, may read runtime data from the Data Analytics. This data may include current KPI of the deployment, current resource utilization, and current load. Then, at 3, the first node 111 , via the Action Selector module, may, according to Action 402 and Action 502, check whether there may be new information useful for model training, such as changes in the load or changes in the KPIs. At 4, if there is new information, then, at 4.1, the first node 111, via the Action Selector, may, according to Action 403 and Action 503, send this information to the second node 112, via the Training Data Generator, and instruct the fourth node 114, via the Training Data Generator to save this information for future training. At 4.2, the second node 112, via the Training Data Generator, may, according to Action 405, send an inferencing request to the fourth node 114, via the Model Server, to get the reward for the multiple generated data points. At 4.3, the fourth node 114, via the Model Server, may calculate the reward based on the performance of the deployment and this reward may be returned, according to Action 406, to the second node 112, via the Training Data Generator. At 4.4, this information is saved, according to Action 407, to the Replay Memory . This information may indicate the state, next state, and reward. Some pre-processing may be required at this stage to make the data suitable for the RL model. At 5, the first node 111, via the Action Selector, may then select an action in this step. The first node 111, via the Action Selector, may get its deployment parameters from the Deployment Parameters Repository, such as the target KPI, the minimum resource configurations of the given deployment’s CNFs, etc. Based on theseparameters, the first node 111, via the Action Selector, may, according to Action 411 and Action 505, select an action. Selecting an action, e.g., decreasing / increasing resources, may be performed in accordance with the methodology described in Action 405. At 6., finally, the selected action may be applied on the given deployment. And the method may go back to step 1 until no more data points may be needed.
[0249] Figure 9 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system 100 to trigger action selection, according embodiments herein. At 1 , the first node 111 may, according to Action 401 , read runtime data from the Data Analytics. Then, at 2, the first node 111 , may, according to Action 402 and Action 502, check whether there may be new information useful for model training, such as changes in the load or changes in the KPIs. At 3, if there is new information, the fourth node 114 may calculate the reward based on the performance of the deployment and this reward may be saved at 4, according to Action 407. At 5, the first node 111 may according to Action 411 and Action 505, select an action. At 6., finally, the selected action may be applied on the given deployment. And the method may go back to step 1 until no more data points may be needed.
[0250] Figure 10 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system 100 to train the model, according embodiments herein. At 1, the fifth node 115, via the Model Trainer, may, monitor the Replay Memory. Then, at 2, the fifth node 115, may check whether there may be sufficient data to retrain the model. The sufficiency criterion may be included in the Model & Training Configuration Data repository. At 3, once there may be sufficient data in the Replay Memory, at 3.1 , the fifth node 115, via the Model Trainer may retrieve the model(s) from the fourth node 114, via the Model Server. At 3.2, the fifth node 115, via the Model Trainer may then, according to Action 415, update this model(s). Updating a model may be understood to mean retraining a model using the data in the Replay Memory or updating the weights of the target network, which may be the case when using DRL. At 3.3, once the model may be trained, it may be updated in the he fourth node 114, via the Model Server, to be used for inferencing. At the end, the method may go back to step 1 and the fifth node 115 may continue monitoring.
[0251] Figure 11 is a schematic diagram depicting a non-limiting example of signalling between nodes in the communications system 100 to train the model, according embodiments herein. At 1, the fifth node 115 may, monitor the Replay Memory. Then, at 2, the fifth node 115, may check whether there may be sufficient data to retrain the model. At 3, once there may be sufficient data in the Replay Memory, the fifth node 115 may retrieve the model(s) from the fourth node 114. At 4, the fifth node 115 may then, according to Action 415, update thismodel(s). Updating a model may be understood to mean retraining a model using the data in the Replay Memory or updating the weights of the target network, which may be the case when using DRL. At 5, once the model may be trained, it may be pushed to the fourth node 114 to be used for inferencing. At the end, the method may go back to step 1 and the fifth node 115 may continue monitoring.
[0252] Figure 12 is a flowchart depicting aspects of the orange phase described earlier, according to embodiments herein, as performed by the first node 111 , via the Action selector, so that the third node 113, via the Training Director, may, gradually decrease the safety margin to a configurable value, without violating the SLA for a given deployment. The third node 113, e.g., via the Training Director, may, according to Action 411, start by defining a large-enough safety margin (step 1). In one example, this may be manually specified. In another, this may be determined by using the minimum resources on all CNFs and estimating and / or measuring the resulting KPI and calculating the margin between the target KPI and the resulting KPI. In another example, where several KPIs may be used in the SLA, the KPI closest to the SLA target may be used to define this margin. This safety margin may be assigned to the first node 111, e.g., the Action Selector (Step 2). For each action proposed by the action generation method used, e.g., epsilon-greedy, the first node 111, e.g., via the Action Selector, may estimate whether the KPI that may result after taking the action may violate the safety margin or not (step 3). If it does, the first node 111, according to Action 412, may reject that action and continue to generate and check until it may find a suitable one and apply it to the deployment. The first node 111, e.g., via the Action Selector, may estimate the resulting KPI from an action using different methods. In one example, the first node 111 may use the estimated reward it may get from the ML model, e.g., the RL model, as an indication. In another example, the first node 111 may use spline approximation to estimate the KPI based on current and historical values. In a third one, the first node 111 may use a combination of both approaches. When the third node 113, e.g., via the Training Director, (step 4 and 5) may be confident that the ML model may be accurate enough for the current safety margin, it may, according to Action 408, gradually reduce (step 6) it and continue the process. The reduction may be implementation dependent. In one example, the margin may be reduced in fixed steps, e.g., 5%. In another, the margin may be reduced in relation to the accuracy of the ML model. The reduction of the margin may be performed repeatedly until a configured target safety margin may be reached. Alternatively, the margin may be increased (step 7) if a current safety margin is violated (step 3) because of model inaccuracy (step 5). Once the configured minimum target safety margin may be reached, the system may continue to run the Action Selection in exploitation mode, e.g., e=0 for epsilon greedy algorithm, while monitoring the accuracy of the RL model (step 4). If third node 113, e.g., via the Training Director, detectsmodel inaccuracy, e.g., caused by a change in traffic mix, it may increase the safety margin to a safe value (step 7) and go to step 2.
[0253] Figure 13 is a graphic illustration of the different areas of the three-phase method that may be performed by the first node 111, e.g., via the Action Selector, according to embodiments herein. In the green phase, the method may ensure that the resulting resource allocated to any CNF may be within a prescribed limit that may be retrieved from the Deployment Parameter repository. If the resulting action may cause the resources allocated to CNFs to fall outside of the allowable range, that action may be rejected and another action may be generated repeatedly until a suitable action may be found. At the end of this phase, a model may be obtained that may be accurate for resource values allowed in the repository. r_min may be understood to be the minimum amount of resources that may be required for each CNF to operate correctly. It may be understood to be deployment specific and may be maintained in the deployment repository. The values of r_min may be determined by an expert, or using predefined tools, or by statistical analysis, or by profiling tools that assess the resource demands of each CNF. Performance may be understood to refer to KPIs that may quantify the efficiency and effectiveness of 5GC system e.g., control latency, data latency, and reliability. In the orange phase, the third node 113 may allow the first node 111 to explore parameter spaces outside of those defined in the Deployment Parameter repository, subject to safety margins. For each action proposed by the action generation method used, the first node 111, e.g., via the Action Selector, may estimate whether the KPI that may result after taking the action may violate the safety margin or not. If it does, it may reject that action and continue to generate and check until it may find a suitable one and apply it to the deployment. The goal of the red phase may be understood to be to explore parameter combinations that may likely result in SLA violations in production deployments. In an example, the third node 113, e.g., via the Training Director, may selects the SLA targets and reverse safety margins where the goal may be to take actions such that the performance KPI may be made to be worse than the target SLA, e.g., if target KPI was 10ms, a reverse margin of 20% may require that the actual latency may be above 12ms.
[0254] Figure 14 is a schematic block diagram describing the steps that may be taken in each phase by the first node 111, e.g., via the Action Selector. All the steps are already described above, this Figure is used to provide more clarification. For the sake of the clarity, Figure 10 also includes the steps performed by the first node 111, e.g., via the Data Analytics, and the fourth node 114, via the Model Server following, up the steps performed by first node 111, e.g., via the Action Selector. The left flow shows the actions that may be taken during the green phase, the middle flow shows the actions that may be taken during the orange phase and theright flow shows the actions that may be taken during the red phase. Based on the second indication received from the third node 113, the first node 111, e.g., via the Action Selector may select actions according to Action 411 , which, during the green phase, may exceed a configured minimum, during the orange phase, may be within the specified safety margin, and during the red phase, may be regardless of the safety margin. The first node 111, e.g., via the Action Selector may then apply the action on core deployments, according to Action 413. During the green and orange phases, this may be on production and test deployments. During the red phase, this may be on test deployments only. The first node 111, e.g., via the Data Analytics, may then in each phase, get the deployment’s current KPI, resources and load. The fourth node 114, via the Model Server, may then in each phase, calculate the reward for a given target KPI, and calculate the reward for a range of target KPIs.
[0255] Figure 15 is a schematic block diagram of a computer system 101 operating in a communications system 100 deployment according to embodiments herein. Particularly, Figure 15 shows the deployment architecture for the computer system 101. The modules may be implemented by a Core Network Orchestrator 801 and an NWDAF. In addition, a MultiOperator Service 802 may manage the Core Network Orchestrators 801 of the different deployments. Their description is provided next. The NWDAF may be a Network Data Analytics Function, e.g., as defined in 3GPP TS 23.501, version 19.2.1 and TS 23.288, version 19.1.0 SA2. It may be the first node 111 and implement the Data Analytics module in Figure 7. It may acquire data from core NFs, analyze them, and expose them to the subscriber(s). This data may be exposed to the second node 112, via the Training Data Generator, and saved in Replay Memory. The Multi-Operator Service 802 may be understood to be a central entity responsible of managing several modules. Its responsibilities may include sending the action selection policy to the first node 111, e.g., via the Action Selector, of each deployment, such as operating in green phase, orange phase, or red phase. Also, getting training data from different deployments and saving them to the Replay Memory. This data may be used to train the RL model in the fifth node 115, via the Model Trainer. In one example, it may send the trained model to each deployment to be used by the first node 111, e.g., via the Action Selector. In another example, the model may be saved centrally and accessed by the different deployments through inferencing requests. It may also share the model parameters with the fourth node 114, via the Model Server such as number of layers for Deep Neural Network (DNN) and with Model Trainer such as when to start training and batch size. The Core Network Orchestrator 801 may be understood to be a local orchestrator responsible of its deployment, e.g., current “CN Ericsson Orchestrator”. It may be responsible for collecting parameters from the NWDAF, such as current KPIs, resources, load, and sending the resource configuration generated by the first node 111, e.g., via the Action Selector, to the coredeployment. It may also send the training data to the Multi-Operator Service 802 to be used for training the RL model.
[0256] As a summarized overview of the foregoing, embodiments herein may be understood to describe a system and a method that may enable to choose the correct resource configuration for 5G Core CNFs for a given target KPI requirements and load. This may be understood to be achieved by using existing 5GC deployments within production environments to train models using real world data. The building the RL model may be performed in a safe way, such that actions taken in the context of RL exploration / exploitation may not result in SLA violation. Embodiments herein may be understood to be able to interact with and retrieve data from a plurality of production environments accelerating the training and making the model generalizable. Embodiments herein may be understood to generate several training data points from a single data point retrieved from a deployment, further speeding up the training process without sacrificing accuracy.
[0257] Certain embodiments disclosed herein may provide one or more of the following technical advantage(s), which may be summarized as follows.
[0258] One advantage of embodiments herein may be understood to be that they may enable rapid training. Embodiments herein may be understood to enable AI / ML agent to achieve fast training times, significantly accelerating the learning process.
[0259] Another advantage of embodiments herein may be understood to be that they may enable an experience transfer across deployments. Embodiments herein may be understood to enable the transfer of experience from on deployment to another.
[0260] A further advantage of embodiments herein may be understood to be that they may enable generation of accurate training data. Embodiments herein may be understood to enables generating several training data points from a single data point collected from a production deployment.
[0261] Yet another advantage of embodiments herein may be understood to be that they may enable efficient resource usage. Embodiments herein may be understood to enable efficient usage of resources of production deployments even during training times.
[0262] An additional advantage of embodiments herein may be understood to be that the approach may be understood to create minimal deployment overhead. The training process may be understood to be designed to avoid deployment and redeployment of network service for additional data points required to train the model.
[0263] A further advantage of embodiments herein may be understood to be eliminating emulation requirements. By leveraging production environments, embodiments herein may be understood to remove the need for unrealistic emulation setups.Yet another advantage of embodiments herein may be understood to be that they may enable real world data integration. The model may be understood to be built and refined using real deployment data, enhancing the reliability and applicability of the training results.
[0264] Figure 16 depicts an example of the arrangement that the first node 111 may comprise to perform the method described in Figure 4, Figure 5, Figure 7, Figure 8, Figure 9, Figure 12, Figure 13, Figure 14 and / or Figure 15. The first node 111 may be understood to be for handling the configuration of first resources. The first node 111 is configured to operate in the communications system 100.
[0265] Several embodiments are comprised herein. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description.
[0266] Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments.
[0267] In Figure 16, an optional component is indicated with dashed lines.
[0268] The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the computer system 101 and will thus not be repeated here to simplify the description. For example, the ML model may be configured to be an RL model.
[0269] The first node 111 is configured to obtain the respective data from the one or more deployments in the communications system 100.
[0270] The first node 111 is further configured to determine whether or not the respective data configured to be obtained comprises the combination of values of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of first resources in the communications system 100. The combination of values is configured to lack the similarity with the first information configured to have been previously used to train the machine learning model according to the criterion of similarity.
[0271] The first node 111 is also configured to initiate saving the combination of values to train the machine learning model, with the proviso the result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values.
[0272] In some embodiments, the first node 111 may be further configured with the two following configurations.
[0273] In some embodiments, the first node 111 may be further configured to select the respective action configured to be taken on the respective set of second resources configured to be used by the communications system 100 on the respective one or more deployments. This may be performed with the proviso that the initiating of the saving of the combination ofvalues may be configured to have been or the result of the determination may be configured to be that the respective data configured to be obtained may be configured to lack the combination of values. The selecting may be configured to be based on one or more respective parameters from the respective one or more deployments.
[0274] In some embodiments, the first node 111 may be further configured to apply the action configured to be selected on the respective one or more deployments as part of the training of the machine learning model.
[0275] In some embodiments, the initiating saving may be configured to comprise sending the first indication to the second node 112 configured to operate in the communications system 100. The first indication may be configured to instruct the second node 112 to save the combination of values.
[0276] In some embodiments, the first node 111 may be further configured with the two following configurations.
[0277] In some embodiments, the first node 111 may be further configured to receive the second indication from the third node 113 configured to operate in the communications system 100. The second indication may be configured to indicate one of: i) the first node 111 is to select the respective action within the respective first action space configured to correspond to respective one or more parameters configured to be defined in the repository of parameters for the respective one or more deployments; the respective one or more deployments may be configured to be production deployments in such embodiments, ii) the first node 111 may be configured to select the respective action within the respective first action space configured to correspond to the one or more respective parameters configured to be defined in the repository of parameters for the one or more deployments, applying the safety margin; the respective one or more deployments may be configured to be production deployments in such embodiments, and iii) the first node 111 may be to select the respective action within the respective second action space that may be configured to result in a violation of an SLA, if applied in the respective one or more deployments; in such embodiments, the first node 111 may be configured to apply the respective action configured to be selected from the respective second space in a test deployment. The selecting of the respective action may be configured to be based on the second indication configured to be received.
[0278] In some embodiments, one or more of the following may apply: a) the first resources may be configured to be for dimensioning the network function virtualization system of the communications system 100, and b) the one or more deployments may be configured to be operational deployments.
[0279] In some embodiments, with the proviso the one or more deployments may be configured to be production deployments, the first node 111 may be further configured with to refrain from selecting the first action, with the proviso the first action may be configured to violate the SLA.In some embodiments, the combination of values may be configured to be determined to lack the similarity based on one of the following: a) the Euclidean distance between respective vectors configured to represent the combination of values and the first information exceeding the first threshold, and b) the error configured to be output by the machine learning model when using the combination of values as input exceeding the second threshold.
[0280] The embodiments herein in the first node 111 may be implemented through one or more processors, such as a processing circuitry 1601 in the first node 111 depicted in Figure 16, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardware component. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the first node 111. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the first node 111.
[0281] The first node 111 may further comprise a memory 1602 comprising one or more memory units. The memory 1602 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the first node 111.
[0282] In some embodiments, the first node 111 may receive information from, e.g., the second node 112, the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a receiving port 1603. In some embodiments, the receiving port 1603 may be, for example, connected to one or more antennas in first node 111. In other embodiments, the first node 111 may receive information from another structure in the communications system 100 through the receiving port 1603. Since the receiving port 1603 may be in communication with the processing circuitry 1601 , the receiving port 1603 may then send the received information to the processing circuitry 1601. The receiving port 1603 may also be configured to receive other information.
[0283] The processing circuitry 1601 in the first node 111 may be further configured to transmit or send information to e.g., the second node 112, the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a sending port 1604, which may be in communication with the processing circuitry 1601, and the memory 1602.
[0284] Those skilled in the art will also appreciate that the units comprised within the first node 111 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured withsoftware and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1601 , perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0285] The first node 111 may be configured to perform any of the Actions described in relation to Figure 4, Figure 5, Figure 7, Figure 8, Figure 9, Figure 12, Figure 13, Figure 14 and / or Figure 15, e.g., by means of the processing circuitry 1601 within the first node 111, configured to perform any of such actions.
[0286] Also, in some embodiments, different units comprised within the first node 111 may be configured to perform the different actions described above, implemented as one or more applications running on one or more processors such as the processing circuitry 1601.
[0287] Thus, the methods according to the embodiments described herein for the first node 111 may be respectively implemented by means of a computer program 1605 product, comprising instructions, i.e., software code portions, which, when executed on at least one processing circuitry 1601, cause the at least one processing circuitry 1601 to carry out the actions described herein, as performed by the first node 111. The computer program 1605 product may be stored on a computer-readable storage medium 1606. The computer-readable storage medium 1606, having stored thereon the computer program 1605, may comprise instructions which, when executed on at least one processing circuitry 1601, cause the at least one processing circuitry 1601 to carry out the actions described herein, as performed by the first node 111. In some embodiments, the computer-readable storage medium 1606 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1605 product may be stored on a carrier containing the computer program 1605 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readable storage medium 1606, as described above.
[0288] The first node 111 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the first node 111 and other nodes or devices, e.g., the second node 112, the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0289] In other embodiments, the first node 111 may comprise a radio circuitry 1607, which may comprise e.g., the receiving port 1603 and the sending port 1604.The radio circuitry 1607 may be configured to set up and maintain at least a wireless connection with the second node 112, the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. Circuitry may be understood herein as a hardware component.
[0290] Hence, embodiments herein also relate to the first node 111, operative to operate in the communications system 100. The first node 111 may comprise the processing circuitry 1601 and the memory 1602, said memory 1602 containing instructions executable by said processing circuitry 1601 , whereby the first node 111 is further operative to perform the actions described herein in relation to the first node 111, e.g., in Figure 4, Figure 5, Figure 7, Figure 8, Figure 9, Figure 12, Figure 13, Figure 14 and / or Figure 15.
[0291] Figure 17 depicts an example of the arrangement that the second node 112 may comprise to perform the method described in Figure 4, Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 13, and / or Figure 15. The second node 112 may be understood to be for handling the configuration of first resources. The second node 112 is configured to operate in the communications system 100.
[0292] Several embodiments are comprised herein. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description.
[0293] Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the computer system 101 and will thus not be repeated here to simplify the description. For example, the ML model may be configured to be an RL model.
[0294] In Figure 17, an optional component is indicated with dashed lines.
[0295] The second node 112 is configured to receive the first indication from the first node 111 configured to operate in the communications system 100. The first indication is configured to instruct the second node 112 to save the combination of values of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of first resources in the communications system 100.
[0296] The second node 112 is also configured to send the third indication to the fourth node 114 configured to operate in the communications system 100. The third indication is configured to request the fourth node 114 to provide the reward configured to correspond to the combination of values. The reward is configured to be used to train the machine learningmodel configured to determine the configuration of first resources in the communications system 100.
[0297] The second node 112 is further configured to receive the reward from the fourth node 114.
[0298] In some embodiments, the second node 112 may be further configured with the following configuration.
[0299] In some embodiments, the second node 112 may be further configured to save the combination of values and the reward configured to be received in the memory of the communications system 100.
[0300] In some embodiments, the first resources may be configured to be for dimensioning the network function virtualization system of the communications system 100.
[0301] In some embodiments, with the proviso the one or more deployments may be configured to be production deployments, the first node 111 may be further configured with to refrain from selecting the first action, with the proviso the first action may be configured to violate the SLA.
[0302] In some embodiments, the combination of values may be configured to be determined to lack the similarity based on one of the following: a) the Euclidean distance between the respective vectors configured to represent the combination of values and the first information exceeding the first threshold, and b) the error configured to be output by the machine learning model when using the combination of values as input exceeding the second threshold.
[0303] The embodiments herein in the second node 112 may be implemented through one or more processors, such as a processing circuitry 1701 in the second node 112 depicted in Figure 17, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardware component. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the second node 112. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the second node 112.
[0304] The second node 112 may further comprise a memory 1702 comprising one or more memory units. The memory 1702 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the second node 112.
[0305] In some embodiments, the second node 112 may receive information from, e.g., the first node 111 , the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a receiving port 1703. In some embodiments, thereceiving port 1703 may be, for example, connected to one or more antennas in second node 112. In other embodiments, the second node 112 may receive information from another structure in the communications system 100 through the receiving port 1703. Since the receiving port 1703 may be in communication with the processing circuitry 1701, the receiving port 1703 may then send the received information to the processing circuitry 1701. The receiving port 1703 may also be configured to receive other information.
[0306] The processing circuitry 1701 in the second node 112 may be further configured to transmit or send information to e.g., the first node 111, the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a sending port 1704, which may be in communication with the processing circuitry 1701, and the memory 1702.
[0307] Those skilled in the art will also appreciate that the units comprised within the second node 112 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1701 , perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0308] The second node 112 may be configured to perform any of the Actions described in relation to Figure 4, Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 13, and / or Figure 15, e.g., by means of the processing circuitry 1701 within the second node 112, configured to perform any of such actions.
[0309] Also, in some embodiments, different units comprised within the second node 112 may be configured to perform the different actions described above, implemented as one or more applications running on one or more processors such as the processing circuitry 1701.
[0310] Thus, the methods according to the embodiments described herein for the second node 112 may be respectively implemented by means of a computer program 1705 product, comprising instructions, i.e., software code portions, which, when executed on at least one processing circuitry 1701, cause the at least one processing circuitry 1701 to carry out the actions described herein, as performed by the second node 112. The computer program 1705 product may be stored on a computer-readable storage medium 1706. The computer-readable storage medium 1706, having stored thereon the computer program 1705, may comprise instructions which, when executed on at least one processing circuitry 1701, cause the at least one processing circuitry 1701 to carry out the actions described herein, asperformed by the second node 112. In some embodiments, the computer-readable storage medium 1706 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1705 product may be stored on a carrier containing the computer program 1705 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readable storage medium 1706, as described above.
[0311] The second node 112 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the second node 112 and other nodes or devices, e.g., the first node 111 , the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0312] In other embodiments, the second node 112 may comprise a radio circuitry 1707, which may comprise e.g., the receiving port 1703 and the sending port 1704.
[0313] The radio circuitry 1707 may be configured to set up and maintain at least a wireless connection with the first node 111 , the third node 113, the fourth node 114, the fifth node 115, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. Circuitry may be understood herein as a hardware component.
[0314] Hence, embodiments herein also relate to the second node 112, operative to operate in the communications system 100. The second node 112 may comprise the processing circuitry 1701 and the memory 1702, said memory 1702 containing instructions executable by said processing circuitry 1701, whereby the second node 112 is further operative to perform the actions described herein in relation to the second node 112, e.g., in Figure 4, Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 13, and / or Figure 15.
[0315] Figure 18 depicts an example of the arrangement that the computer system 101 may comprise to perform the method described in in Figure 4, Figure 7, Figure 10, Figure 11 , and / or Figure 15. The computer system 101 may be configured to comprise the first node 111 and the fifth node 115. The computer system 101 may be understood to be for handling the configuration of first resources. The computer system 101 is configured to operate in the communications system 100.
[0316] Several embodiments are comprised herein. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description.
[0317] Components from one embodiment may be tacitly assumed to be present in anotherembodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the computer system 101 and will thus not be repeated here to simplify the description. For example, the ML model may be configured to be an RL model.
[0318] In Figure 18, an optional component is indicated with dashed lines.
[0319] The computer system 101 is configured to obtain, by the first node 111 , the respective data from the one or more deployments in the communications system 100.
[0320] The computer system 101 is further configured to determine, by the first node 111, whether or not the respective data configured to be obtained is configured to comprise the combination of values of the one or more variables configured to be used to train the machine learning model configured to determine the configuration of the first resources in the communications system 100. The combination of values is configured to lack the similarity with the first information configured to have been previously used to train the machine learning model.
[0321] In some embodiments, the computer system 101 is further configured to initiate, by the first node 111, saving the combination of values to train the machine learning model. This is performed with the proviso the result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values.
[0322] In some embodiments, the computer system 101 is further configured to train, by the fifth node 115, the machine learning model with the combination of values configured to be saved.
[0323] In some embodiments, the computer system 101 is further configured to output, by the fifth node 115 the third indication of the trained machine learning model.
[0324] In some embodiments, the computer system 101 may be further configured with the following two configurations.
[0325] In some embodiments, the computer system 101 may be further configured to select, by the first node 111 , the respective action to be taken on the respective set of second resources configured to be used by the communications system 100 on the respective one or more deployments. This may be performed with the proviso that the initiating of the saving of the combination of values may be configured to have been or the result of the determination may be that the respective data configured to have been obtained lacks the combination of values. The selecting may be configured to be based on the one or more respective parameters from the respective one or more deployments.
[0326] In some embodiments, the computer system 101 may be further configured to apply, by the first node 111 , the action configured to be selected on the respective one or more deployments as part of the training of the machine learning model.In some embodiments, wherein the initiating saving may be configured to comprise sending the first indication to the second node 112 configured to operate in the communications system 100, the first indication being configured to instruct the second node 112 to save the combination of values, the computer system 101 may be further configured with the following three configurations.
[0327] In some embodiments, the computer system 101 may be further configured to receive, by the second node 112, the first indication from the first node 111.
[0328] In some embodiments, the computer system 101 may be further configured to send, by the second node 112, the third indication to the fourth node 114 configured to be comprised in the computer system 101. The third indication may be configured to request the fourth node 114 to provide the reward configured to correspond to the combination of values. The reward may be configured to be to be used to train the machine learning model to determine the configuration of first resources in the communications system 100.
[0329] In some embodiments, the computer system 101 may be further configured to receive, by the second node 112, the reward from the fourth node 114. The machine learning model may be configured to be trained with the reward configured to be received.
[0330] In some embodiments, the computer system 101 may be further configured with the following configuration.
[0331] In some embodiments, the computer system 101 may be further configured to save, by the second node 112, the combination of values and the reward configured to be received in the memory of the communications system 100.
[0332] In some embodiments, the computer system 101 may be further configured with the following three configurations.
[0333] In some embodiments, the computer system 101 may be further configured to determine, by the third node 113 configured to be comprised in the computer system 101 , the second indication to be sent to the first node 111. The second indication may be configured to indicate one of: i. the first node 111 is to select the respective action within the respective first action space configured to correspond to the respective one or more parameters configured to be defined in the repository of parameters for the respective one or more deployments; in such embodiments, the respective one or more deployments may be configured to be production deployments, ii) the first node 111 is to select the respective action within the respective first action space configured to correspond to one or more respective parameters defined in the repository of parameters for the one or more deployments, applying the safety margin; in such embodiments, the respective one or more deployments may be configured to be production deployments, and iii) the first node 111 is to select the respective action within the respective second action space that may be configured to result in the violation of the Service Level Agreement if applied in the respective one or more deployments; in such embodiments, thefirst node 111 may be configured to apply the selected respective action from the respective second space in a test deployment.
[0334] In some embodiments, the computer system 101 may be further configured to send, by the third node 113, the second indication configured to be determined, to the first node 111.
[0335] In some embodiments, the computer system 101 may be further configured to receive, by the first node 111 , the second indication from the third node 113. The selecting of the respective action may be configured to be based on the second indication configured to be received.
[0336] In some embodiments, one or more of the following may apply: a) the first resources may be configured to be for dimensioning the network function virtualization system of the communications system 100, and b) the one or more deployments may be configured to be operational deployments.
[0337] In some embodiments, with the proviso the one or more deployments may be configured to be production deployments, the computer system 101 may be further configured to refrain, by the first node 111, from selecting the first action, with the proviso the first action may be configured to violate the SLA.
[0338] In some embodiments, the combination of values may be configured to be determined to lack the similarity based on one of the following: a) the Euclidean distance between respective vectors configured to represent the combination of values and the first information exceeding the first threshold, and b) the error configured to be output by the machine learning model when using the combination of values as input exceeding the second threshold.
[0339] The embodiments herein in the first node 111 may be implemented through an arrangement as described in relation to Figure 10.
[0340] The embodiments herein in the fifth node 115 may be implemented through one or more processors, such as a processing circuitry 1801 in the fifth node 115 depicted in Figure 18, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardware component. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the fifth node 115. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the fifth node 115.
[0341] The fifth node 115 may further comprise a memory 1802 comprising one or more memory units. The memory 1802 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the fifth node 115.In some embodiments, the fifth node 115 may receive information from, e.g., the first node 111 , the second node 112, the third node 113, the fourth node 114, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a receiving port 1803. In some embodiments, the receiving port 1803 may be, for example, connected to one or more antennas in fifth node 115. In other embodiments, the fifth node 115 may receive information from another structure in the communications system 100 through the receiving port 1803. Since the receiving port 1803 may be in communication with the processing circuitry 1801, the receiving port 1803 may then send the received information to the processing circuitry 1801. The receiving port 1803 may also be configured to receive other information.
[0342] The processing circuitry 1801 in the fifth node 115 may be further configured to transmit or send information to e.g., the first node 111, the second node 112, the third node 113, the fourth node 114, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100, through a sending port 1804, which may be in communication with the processing circuitry 1801, and the memory 1802.
[0343] Those skilled in the art will also appreciate that the units comprised within the fifth node 115 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1801 , perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0344] The fifth node 115 may be configured to perform any of the Actions described in relation in Figure 4, Figure 7, Figure 10, Figure 11, and / or Figure 15, e.g., by means of the processing circuitry 1801 within the fifth node 115, configured to perform any of such actions.
[0345] Also, in some embodiments, different units comprised within the fifth node 115 may be configured to perform the different actions described above, implemented as one or more applications running on one or more processors such as the processing circuitry 1801.
[0346] Thus, the methods according to the embodiments described herein for the fifth node 115 may be respectively implemented by means of a computer program 1805 product, comprising instructions, i.e., software code portions, which, when executed on at least one processing circuitry 1801, cause the at least one processing circuitry 1801 to carry out the actions described herein, as performed by the fifth node 115. The computer program 1805 product may be stored on a computer-readable storage medium 1806. The computer-readable storage medium 1806, having stored thereon the computer program 1805, maycomprise instructions which, when executed on at least one processing circuitry 1801, cause the at least one processing circuitry 1801 to carry out the actions described herein, as performed by the fifth node 115. In some embodiments, the computer-readable storage medium 1806 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1805 product may be stored on a carrier containing the computer program 1805 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readable storage medium 1806, as described above.
[0347] The fifth node 115 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the fifth node 115 and other nodes or devices, e.g., the first node 111, the second node 112, the third node 113, the fourth node 114, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0348] In other embodiments, the fifth node 115 may comprise a radio circuitry 1807, which may comprise e.g., the receiving port 1803 and the sending port 1804.
[0349] The radio circuitry 1807 may be configured to set up and maintain at least a wireless connection with the first node 111 , the second node 112, the third node 113, the fourth node 114, the radio network node 140, the device 130, another node or user equipment, and / or another structure in the communications system 100. Circuitry may be understood herein as a hardware component.
[0350] Hence, embodiments herein also relate to the fifth node 115, operative to operate in the communications system 100. The fifth node 115 may comprise the processing circuitry 1801 and the memory 1802, said memory 1802 containing instructions executable by said processing circuitry 1801, whereby the fifth node 115 is further operative to perform the actions described herein in relation to the fifth node 115, e.g., in Figure 4, Figure 7, Figure 10, Figure 11 , and / or Figure 15.
[0351] In some examples, the computer system 101 may be further configured to comprise any of the second node 112, e.g., as described in relation to Figure 17, the third node 113, and the fourth node 114. Any of the third node 113 and the fourth node 114 may respectively comprise arrangements similar to those described for any of the first node 111, second node 112 and fifth node 115, and may be respectively configured as described in relation to Figure 18 and / or Figure 4.
[0352] When using the word "comprise" or “comprising”, it shall be interpreted as non- limiting, i.e., meaning "consist at least of'.The embodiments herein are not limited to the above-described preferred embodiments. Various alternatives, modifications and equivalents may be used. Therefore, the above embodiments should not be taken as limiting the scope of the invention.
[0353] Generally, all terms used herein are to be interpreted according to their ordinary meaning in the relevant technical field, unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any methods disclosed herein do not have to be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or where it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, wherever appropriate. Likewise, any advantage of any of the embodiments may apply to any other embodiments, and vice versa. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following description.
[0354] As used herein, the expression “at least one of:” followed by a list of alternatives separated by commas, and wherein the last alternative is preceded by the “and” term, may be understood to mean that only one of the list of alternatives may apply, more than one of the list of alternatives may apply or all of the list of alternatives may apply. This expression may be understood to be equivalent to the expression “at least one of:” followed by a list of alternatives separated by commas, and wherein the last alternative is preceded by the “or” term.
[0355] Any of the terms processor and circuitry may be understood herein as a hardware component.
[0356] As used herein, the expression “in some embodiments” has been used to indicate that the features of the embodiment described may be combined with any other embodiment or example disclosed herein.
[0357] As used herein, the expression “in some examples” has been used to indicate that the features of the example described may be combined with any other embodiment or example disclosed herein.
[0358] REFERENCES
[0359] 1. Jalodia, Nikita, Shagufta Henna, and Alan Davy. "Deep reinforcement learning for topology-aware VNF resource prediction in NFV environments." 2019 IEEE Conference on Network Function Virtualization and Software Defined Networks (NFV-SDN). IEEE, 2019.Belgacem, Ali, Said Mahmoudi, and Maria Kihl. "Intelligent multi-agent reinforcement learning model for resources allocation in cloud computing." Journal of King Saud University-Computer and Information Sciences (2022).
Claims
59CLAIMS:
1. A computer-implemented method, performed by a first node (111), the method being for handling configuration of first resources, the first node (111) operating in a communications system (100), the method comprising:- obtaining (501) respective data from one or more deployments in the communications system (100),- determining (502) whether or not the obtained respective data comprises a combination of values of one or more variables used to train a machine learning model to determine a configuration of first resources in the communications system (100), wherein the combination of values lacks a similarity with first information previously used to train the machine learning model according to a criterion of similarity, and- initiating (503) saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the obtained respective data comprises the combination of values.
2. The method according to claim 1 , further comprising:- selecting (505) a respective action to be taken on a respective set of second resources to be used by the communications system (100) on the respective one or more deployments, with the proviso that the initiating (503) of the saving of the combination of values has been or the result of the determination is that the obtained respective data lacks the combination of values, wherein the selecting (505) is based on one or more respective parameters from the respective one or more deployments, and- applying (507) the selected action on the respective one or more deployments as part of a training of the machine learning model.
3. The method according to any of claims 1-2, wherein the initiating (503) saving comprises sending a first indication to a second node (112) operating in the communications system (100), the first indication instructing the second node (112) to save the combination of values.
4. The method according to any of claims 2 or 3, further comprising:- receiving (504) a second indication from a third node (113) operating in the communications system (100), the second indication indicating one of:60i. the first node (111) is to select the respective action within a respective first action space corresponding to respective one or more parameters defined in a repository of parameters for the respective one or more deployments, wherein the respective one or more deployments are production deployments,ii. the first node (111) is to select the respective action within the respective first action space corresponding to the one or more respective parameters defined in the repository of parameters for the one or more deployments, applying a safety margin, wherein the respective one or more deployments are production deployments, and Hi. the first node (111) is to select the respective action within a respective second action space that is to result in a violation of a Service Level Agreement, SLA, if applied in the respective one or more deployments, and wherein the first node (111) is to apply the selected respective action from the respective second space in a test deployment, and wherein the selecting (505) of the respective action is based on the received second indication.
5. The method according to any of claims 1-4, wherein one or more of:- the first resources are for dimensioning a network function virtualization system of the communications system (100), and- the one or more deployments are operational deployments.
6. The method according to claim 4, wherein, with the proviso the one or more deployments are production deployments, the method further comprises:- refraining (506) from selecting a first action, with the proviso the first action is to violate the SLA.
7. The method according to any of claims 1-6, wherein the combination of values is determined to lack the similarity based on one of:- an Euclidean distance between respective vectors representing the combination of values and the first information exceeding a first threshold, and - an error output by the machine learning model when using the combination of values as input exceeding a second threshold.
618. A computer-implemented method, performed by a second node (112), the method being for handling configuration of first resources, the second node (112) operating in a communications system (100), the method comprising:- receiving (601) a first indication from a first node (111) operating in the communications system (100), the first indication instructing the second node (112) to save a combination of values of one or more variables used to train a machine learning model to determine a configuration of first resources in the communications system (100),- sending (602) a third indication to a fourth node (114) operating in the communications system (100), the third indication requesting the fourth node (114) to provide a reward corresponding to the combination of values, the reward to be used to train the machine learning model to determine the configuration of first resources in the communications system (100), and - receiving (603) the reward from the fourth node (114).
9. The method according to claim 8, wherein the method further comprises:- saving (604) the combination of values and the received reward in a memory of the communications system (100).
10. The method according to any of claims 8-9, wherein the first resources are for dimensioning a network function virtualization system of the communications system (100).
11. The method according to any of claims 8-10, wherein the combination of values is determined to lack the similarity based on one of:- an Euclidean distance between respective vectors representing the combination of values and the first information exceeding a first threshold, and - an error output by the machine learning model when using the combination of values as input exceeding a second threshold.
12. A computer-implemented method, performed by computer system (101) comprising a first node (111) and a fifth node (115), the method being for handling configuration of first resources, the computer system (101) operating in a communications system (100), the method comprising:- obtaining (401), by the first node (111), respective data from one or more deployments in the communications system (100),62- determining (402), by the first node (111), whether or not the obtained respective data comprises a combination of values of one or more variables used to train a machine learning model to determine a configuration of first resources in the communications system (100), wherein the combination of values lacks a similarity with first information previously used to train the machine learning model,- initiating (403), by the first node (111), saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the obtained respective data comprises the combination of values, - training (414), by the fifth node (115), the machine learning model with the saved combination of values, and- outputting (415), by the fifth node (115) a third indication of the trained machine learning model.
13. The method according to claim 12, further comprising:- selecting (411), by the first node (111), a respective action to be taken on a respective set of second resources to be used by the communications system (100) on the respective one or more deployments, with the proviso that the initiating (403) of the saving of the combination of values has been or the result of the determination is that the obtained respective data lacks the combination of values, wherein the selecting (411) is based on one or more respective parameters from the respective one or more deployments, and- applying (413), by the first node (111), the selected action on the respective one or more deployments as part of a training of the machine learning model.
14. The method according to any of claims 12-13, wherein the initiating (403) saving comprises sending a first indication to a second node (112) operating in the communications system (100), the first indication instructing the second node (112) to save the combination of values, and wherein the method further comprises:- receiving (404), by the second node (112), the first indication from the first node (111),- sending (405), by the second node (112), a third indication to a fourth node (114) comprised in the computer system (101), the third indication requesting the fourth node (114) to provide a reward corresponding to the combination of values, the reward being to be used to train the machine learning model to determine the configuration of first resources in the communications system (100), and63receiving (406), by the second node (112), the reward from the fourth node (114), and wherein the machine learning model is trained with the received reward.
15. The method according to claims 13 and 14, wherein the method further comprises:- saving (407), by the second node (112), the combination of values and the received reward in a memory of the communications system (100).
16. The method according to any of claims 13 or 15, further comprising:- determining (408), by a third node (113) comprised in the computer system (101), a second indication to be sent to the first node (111), the second indication indicating one of:i. the first node (111) is to select the respective action within a respective first action space corresponding to respective one or more parameters defined in a repository of parameters for the respective one or more deployments, wherein the respective one or more deployments are production deployments,ii. the first node (111) is to select the respective action within the respective first action space corresponding to one or more respective parameters defined in the repository of parameters for the one or more deployments, applying a safety margin, wherein the respective one or more deployments are production deployments, and Hi. the first node (111) is to select the respective action within a respective second action space that is to result in a violation of a Service Level Agreement if applied in the respective one or more deployments, and wherein the first node (111) is to apply the selected respective action from the respective second space in a test deployment - sending (409), by the third node (113), the determined second indication, to the first node (111), and- receiving (410), by the first node (111), the second indication from the third node (113), and wherein the selecting (411) of the respective action is based on the received second indication.
17. The method according to any of claims 12-16, wherein one or more of:- the first resources are for dimensioning a network function virtualization system of the communications system (100), and- the one or more deployments are operational deployments.
18. The method according to claim 16, wherein, with the proviso the one or more deployments are production deployments, the method further comprises:- refraining (412), by the first node (111), from selecting a first action, with the proviso the first action is to violate the SLA.
19. The method according to any of claims 12-18, wherein the combination of values is determined to lack the similarity based on one of:- an Euclidean distance between respective vectors representing the combination of values and the first information exceeding a first threshold, and - an error output by the machine learning model when using the combination of values as input exceeding a second threshold.
20. A first node (111), for handling configuration of first resources, the first node (111) being configured to operate in a communications system (100), the first node (111) being further configured to:- obtain respective data from one or more deployments in the communications system (100),- determine whether or not the respective data configured to be obtained comprises a combination of values of one or more variables configured to be used to train a machine learning model configured to determine a configuration of first resources in the communications system (100), wherein the combination of values is configured to lack a similarity with first information configured to have been previously used to train the machine learning model according to a criterion of similarity, and- initiate saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values.
21. The first node (111) according to claim 20, being further configured to:- select a respective action configured to be taken on a respective set of second resources configured to be used by the communications system (100) on the respective one or more deployments, with the proviso that the initiating of the saving of the combination of values is configured to have been or the result of the determination is configured to be that the respective data configured to be obtained is configured to lack the combination of values, wherein the selectingis configured to be based on one or more respective parameters from the respective one or more deployments, and- apply the action configured to be selected on the respective one or more deployments as part of a training of the machine learning model.
22. The first node (111) according to any of claims 20-21 , wherein the initiating saving is configured to comprise sending a first indication to a second node (112) configured to operate in the communications system (100), the first indication being configured to instruct the second node (112) to save the combination of values.
23. The first node (111) according to any of claims 21 or 22, being further configured to:- receive a second indication from a third node (113) configured to operate in the communications system (100), the second indication being configured to indicate one of:i. the first node (111) is to select the respective action within a respective first action space configured to correspond to respective one or more parameters configured to be defined in a repository of parameters for the respective one or more deployments, wherein the respective one or more deployments are configured to be production deployments, ii. the first node (111) is configured to select the respective action within the respective first action space configured to correspond to the one or more respective parameters configured to be defined in the repository of parameters for the one or more deployments, applying a safety margin, wherein the respective one or more deployments are configured to be production deployments, andHi. the first node (111) is to select the respective action within a respective second action space that is configured to result in a violation of a Service Level Agreement, SLA, if applied in the respective one or more deployments, and wherein the first node (111) is configured to apply the respective action configured to be selected from the respective second space in a test deployment,and wherein the selecting of the respective action is configured to be based on the second indication configured to be received.
24. The first node (111) according to any of claims 21-23, wherein one or more of:- the first resources are configured to be for dimensioning a network function virtualization system of the communications system (100), and66the one or more deployments are configured to be operational deployments.
25. The first node (111) according to claim 24, wherein, with the proviso the one or more deployments are configured to be production deployments, the first node (111) is further configured to:- refrain from selecting a first action, with the proviso the first action is configured to violate the SLA.
26. The first node (111) according to any of claims 21-25, wherein the combination of values is configured to be determined to lack the similarity based on one of:- an Euclidean distance between respective vectors configured to represent the combination of values and the first information exceeding a first threshold, and - an error configured to be output by the machine learning model when using the combination of values as input exceeding a second threshold.
27. A second node (112), for handling configuration of first resources, the second node (112) being configured to operate in a communications system (100), the second node (112) being further configured to:- receive a first indication from a first node (111) configured to operate in the communications system (100), the first indication being configured to instruct the second node (112) to save a combination of values of one or more variables configured to be used to train a machine learning model configured to determine a configuration of first resources in the communications system (100),- send a third indication to a fourth node (114) configured to operate in the communications system (100), the third indication being configured to request the fourth node (114) to provide a reward configured to correspond to the combination of values, the reward being configured to be used to train the machine learning model to determine the configuration of first resources in the communications system (100), and- receive the reward from the fourth node (114).
28. The second node (112) according to claim 27, wherein the second node (112) is further configured to:- save the combination of values and the reward configured to be received in a memory of the communications system (100).6729. The second node (112) according to any of claims 27-28, wherein the first resources are configured to be for dimensioning a network function virtualization system of the communications system (100).
30. The second node (112) according to any of claims 27-29, wherein the combination of values is configured to be determined to lack the similarity based on one of:- an Euclidean distance between respective vectors configured to represent the combination of values and the first information exceeding a first threshold, and - an error configured to be output by the machine learning model when using the combination of values as input exceeding a second threshold.
31. A computer system (101), configured to comprise a first node (111) and a fifth node (115), the computer system (101) being configured to be for handling configuration of first resources, the computer system (101) being configured to operate in a communications system (100), the computer system (101) being further configured to:- obtain, by the first node (111), respective data from one or more deployments in the communications system (100),- determine, by the first node (111), whether or not the respective data configured to be obtained is configured to comprise a combination of values of one or more variables configured to be used to train a machine learning model configured to determine a configuration of first resources in the communications system (100), wherein the combination of values is configured to lack a similarity with first information configured to have been previously used to train the machine learning model,- initiate, by the first node (111), saving the combination of values to train the machine learning model, with the proviso a result of the determination is that the respective data configured to be obtained is configured to comprise the combination of values,- train, by the fifth node (115), the machine learning model with the combination of values configured to be saved, and- output, by the fifth node (115) a third indication of the trained machine learning model.
32. The computer system (101) according to claim 31, being further configured to:- select, by the first node (111), a respective action to be taken on a respective set of second resources configured to be used by the communications system (100) on the respective one or more deployments, with the proviso that the68initiating of the saving of the combination of values is configured to have been or the result of the determination is that the respective data configured to have been obtained lacks the combination of values, wherein the selecting is configured to be based on one or more respective parameters from the respective one or more deployments, and- apply, by the first node (111), the action configured to be selected on the respective one or more deployments as part of a training of the machine learning model.
33. The computer system (101) according to any of claims 31-32, wherein the initiating saving is configured to comprise sending a first indication to a second node (112) configured to operate in the communications system (100), the first indication being configured to instruct the second node (112) to save the combination of values, and wherein the computer system (101) is further configured to:- receive, by the second node (112), the first indication from the first node (111), - send, by the second node (112), a third indication to a fourth node (114) configured to be comprised in the computer system (101), the third indication being configured to request the fourth node (114) to provide a reward configured to correspond to the combination of values, the reward being configured to be to be used to train the machine learning model to determine the configuration of first resources in the communications system (100), and - receive, by the second node (112), the reward from the fourth node (114), and wherein the machine learning model is configured to be trained with the reward configured to be received.
34. The computer system (101) according to claims 32 and 33, wherein the computer system (101) is further configured to:- save, by the second node (112), the combination of values and the reward configured to be received in a memory of the communications system (100).
35. The computer system (101) according to any of claims 32 or 34, is further configured to:- determine, by a third node (113) configured to be comprised in the computer system (101), a second indication to be sent to the first node (111), the second indication being configured to indicate one of:i. the first node (111) is to select the respective action within a respective first action space configured to correspond to respective one or more69parameters configured to be defined in a repository of parameters for the respective one or more deployments, wherein the respective one or more deployments are configured to be production deployments, ii. the first node (111) is to select the respective action within the respective first action space configured to correspond to one or more respective parameters defined in the repository of parameters for the one or more deployments, applying a safety margin, wherein the respective one or more deployments are configured to be production deployments, andHi. the first node (111) is to select the respective action within a respective second action space that is configured to result in a violation of a Service Level Agreement if applied in the respective one or more deployments, and wherein the first node (111) is configured to apply the selected respective action from the respective second space in a test deployment,- send, by the third node (113), the second indication configured to be determined, to the first node (111), and- receive, by the first node (111), the second indication from the third node (113), and wherein the selecting of the respective action is configured to be based on the second indication configured to be received.
36. The computer system (101) according to any of claims 31-35, wherein one or more of:- the first resources are configured to be for dimensioning a network function virtualization system of the communications system (100), and- the one or more deployments are configured to be operational deployments.
37. The computer system (101) according to claim 35, wherein, with the proviso the one or more deployments are configured to be production deployments, the computer system (101) is further configured to:- refrain, by the first node (111) from selecting a first action, with the proviso the first action is configured to violate the SLA.
38. The computer system (101) according to any of claims 31-37, wherein the combination of values is configured to be determined to lack the similarity based on one of:- an Euclidean distance between respective vectors configured to represent the combination of values and the first information exceeding a first threshold, and70- an error configured to be output by the machine learning model when using the combination of values as input exceeding a second threshold.