Apparatus, network node, system, methods, for context-aware communication network optimization

The apparatus and method for context-aware communication network optimization address the inefficiencies of existing RL techniques by adapting control policies based on current network states and contextual information, enhancing the robustness and cost-effectiveness of RAN optimization.

WO2025136157A1PCT designated stage expired Publication Date: 2025-06-26TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2023/051275
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing reinforcement learning (RL) techniques for Radio Access Network (RAN) optimization do not effectively account for varying performance objectives such as service type, deployment region, network load, and other factors, leading to inefficiencies and high computational costs.

Method used

An apparatus and method for context-aware communication network optimization that defines control policies for network nodes, uses observed data to train these policies, and adapts actions based on current network states and contextual information, including action fusion parameters.

Benefits of technology

This approach enables dynamic adaptation of control policies, improving the robustness and safety of RL in RAN optimization, while allowing existing trained policies to be used for new network contexts or performance objectives, thus reducing operational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2023051275_26062025_PF_FP_ABST
    Figure SE2023051275_26062025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus (110), a network node (120), methods and computer programs are disclosed The apparatus is for context-aware communication network optimization. The apparatus being configured to: define one or more control policies; obtain observed data; obtain, for each control policy, a trained control policy for the network node, each trained control policy is based on the observed data; obtain a current state of the communication network; obtain, from each trained control policy, a predicted action, wherein the predicted actions are based on the current state; obtain current contextual information related to the communication network; obtain, for each trained control policy, an action fusion parameter, the action fusion parameters are based on the current contextual information; obtain a singular action based on the predicted actions, the current contextual information, and the action fusion parameters; and send, to a network node, the singular action to be performed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]APPARATUS, NETWORK NODE, SYSTEM, METHODS, FOR CONTEXT-AWARECOMMUNICATION NETWORK OPTIMIZATION TECHNICAL FIELD The invention relates to an apparatus for context-aware communication network optimization, a network node for context-aware communication network optimization, a system for context-aware communication network optimization, corresponding methods, corresponding computer programs, and a computer readable storage medium. BACKGROUND Radio Access Network (RAN) utilises multiple control algorithms so as to improve traffic performance. These control algorithms include modulation and coding scheme (MCS) selection, power control and dynamic beam selection. Each control algorithm embodies a certain desired policy that dynamically controls operation of the radio link. Reinforcement Learning has been used in optimization and control of real-worldsystem. In particular, Wiering, M.A. and Van Hasselt, H., 2008. Ensemble algorithmsin reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4), pp.930-936. discloses several ensemble methods that combine multiple different RL algorithm in a single agent. SUMMARYA disadvantage in prior art is that ensemble techniques are one way of addressing RLalgorithm and hyperparameter selection for a target application use-case. However,prior art ensemble technique doest not work well for RAN optimization. Indeed, in RANsystems, the performance objective vary depending on the service type, deployment region, network load, and other factors. Prior art RL techniques do not take these factors into account during the learning phase or the inference phase. A further disadvantage in prior art is that reward function is construed to take into account one or more RAN metrics, for example service type or network load. Thus, the RL model is specific to the RAN contexts and reward formulations employed during training. However, the RL model is likely to encounter new RAN contexts and the reward formulations can also evolve over time. This would require new RL models to be trained from scratch, which would lead to high computational and operational costs. A further disadvantage in prior art is that RAN context is utilized as part of the state formulation. In particular, a single model is trained to optimize for all possible RAN contexts, by training the RL agent across a large range of RAN environment configuration. However, this RL is difficult to train and generalize, since the state space scales with the size of the RAN context, which is high-dimensional.An object of the invention is to improve a communication network.According to a first aspect of the invention, an apparatus is provided. The apparatus is for context-aware communication network optimization. The apparatus is configured to define one or more control policies for a network node in a communication network.The apparatus is configured to obtain observed data. The apparatus is configured toobtain, for each of the one or more control policies, a trained control policy for the network node. Each of the trained control policies is based on the observed data. The apparatus is configured to obtain a current state of the communication network. The apparatus is configured to obtain, from each of the trained control policies, a predicted action. Each of the predicted actions is based on the current state of the communication network. The apparatus is configured to obtain current contextual information related to the communication network. The apparatus is configured to obtain, for each of the trained control policies, an action fusion parameter. Each of the action fusion parameters is based on the current contextual information. The apparatus is configured to obtain a singular action to be performed by the network node. The singular action is based on each of the predicted actions, the current contextual information, and each of the action fusion parameters. The apparatus isconfigured to send, to the network node, the singular action to be performed. Herebyis achieved, that a singular action to be performed within the communication isprovided to the network node. In particular, the singular action is based on predictedactions obtained by trained control policies and based on the current state of thecommunication network, the current contextual information of the communicationnetwork, and action fusion parameters based on the current contextual information.Thus, the singular action is adapted based on the contextual information and the status of the communication network. According to an embodiment of the first aspect, the apparatus is configured to obtain, from the network node, a response of the network node from performing the singular action. The apparatus is configured to obtain, for each of the one or more trainedcontrol policies, a second trained control policy for the network node. Each of thesecond trained control policies is based on the observed data and the response. The apparatus is configured to obtain a second current state of the communication network. The apparatus is configured to obtain, from each of the second trained controlpolicies, a second predicted action. Each of the second predicted actions is based onthe second current state of the communication network. The apparatus is configuredto obtain a second current contextual information related to the communicationnetwork. The apparatus is configured to obtain, for each of the second trained control policies, a second action fusion parameter. Each of the second action fusion parameters is based on the second contextual information. The apparatus is configured to obtain a second singular action to be performed by the network node. The second singular information is based on each of the second predicted actions, the second current contextual information, and each of the second action fusion parameters. The apparatus is configured to send, to the network node, the second singular action to be performed.According to an embodiment of the first aspect, obtaining the observed data comprisesreceiving the observed data from the network node.According to an embodiment of the first aspect, obtaining the trained control policycomprises training each of the one or more control policies with the observed data.According to an embodiment of the first aspect, obtaining the current state of thecommunication network comprises receiving the current state of the communication network from the network node.According to an embodiment of the first aspect, obtaining the current contextualinformation comprises receiving the current contextual information from the networknode.According to an embodiment of the first aspect, obtaining the predicted actioncomprises running each of the trained control policies. The current state is an input to the one or more trained control policies, and each of the predicted actions is an output of the respective one or more trained control policies.According to an embodiment of the first aspect, the current contextual informationrelates to key performance indicators of the communication network.According to an embodiment of the first aspect, obtaining the action fusion parameterfor each of the one or more trained control policies comprises: performing a gridsearch; performing a gradient descent; performing Bayesian optimisation; and / orperforming reinforcement learning. According to a second aspect of the invention, a network node is provided. Thenetwork node is in a communication network. The network node is for context-awarecommunication network optimization. The network node is configured to send observed data. The network node is configured to send a current state of the communication network. The network node is configured to send current contextual information. The network node is configured to receive, from an apparatus, a singular action. The singular action is based on one or more predicted actions based on the current state of the communication network, the current contextual information, and one or more action fusion parameters based on the current contextual information. The network node is configured to perform the singular action. According to an embodiment of the second aspect, the network node is configured to send, to the apparatus, a response of the communication network from performing the singular action. The network node is configured to send a second current state of the communication network. The network node is configured to send a second current contextual information. The network node is configured to receive, from the apparatus, a second singular action, the second singular action is based on one or more second predicted actions based on the second current state of the communication network, the second current contextual information, and one or more second action fusion parameters based on the second current contextual information and obtained by the apparatus. The network node is configured to perform the second singular action.According to an embodiment of the second aspect, the current contextual informationrelates to key performance indicators of the communication network. According to a third aspect of the invention, a system is provided. The system is forcontext-aware communication network optimization. The system comprises theapparatus according to the first aspect of the invention. The system comprises the network node according to the second aspect of the invention. According to a fourth aspect of the invention, a method is provided. The method isperformed by an apparatus. The method is for context-aware communication networkoptimization. The method comprises defining one or more control policies for anetwork node in a communication network. The method comprises obtaining observeddata. The method comprises obtaining, for each of the one or more control policies, atrained control policy for the network node. Each of the trained control policies is basedon the observed data. The method comprises obtaining a current state of thecommunication network. The method comprises obtaining, from each of the trainedcontrol policies, a predicted action. Each of the predicted actions is based on thecurrent state of the communication network. The method comprises obtaining current contextual information related to the communication network. The method comprises obtaining, for each of the trained control policies, an action fusion parameter. Each of the action fusion parameters is based on the current contextual information. The method comprises obtaining a singular action to be performed by the network node.The singular action is based on each of the predicted actions, the current contextualinformation, and each of the action fusion parameters. The method comprises sending, to the network node, the singular action to be performed. According to an embodiment of the fourth aspect, the method comprises obtaining, from the network node, a response of the network node from performing the singular action. The method comprises obtaining, for each of the one or more trained control policies, a second trained control policy for the network node. Each of the second trained control policies is based on the observed data and the response. The methodcomprises obtaining a second current state of the communication network. Themethod comprises obtaining, from each of the second trained control policies, a second predicted action. Each of the second predicted actions is based on the secondcurrent state of the communication network. The method comprises obtaining asecond current contextual information related to the communication network. The method comprises obtaining, for each of the second trained control policies, a second action fusion parameter. Each of the second action fusion parameters is based on thesecond contextual information. The method comprises obtaining a second singularaction to be performed by the network node. The second singular information is basedon each of the second predicted actions, the second current contextual information, and each of the second action fusion parameters. The method comprises sending, to the network node, the second singular action to be performed.According to an embodiment of the fourth aspect, obtaining the observed datacomprises receiving the observed data from the network node.According to an embodiment of the fourth aspect, obtaining the trained control policycomprises training each of the one or more control policies with the observed data.According to an embodiment of the fourth aspect, obtaining the current state of thecommunication network comprises receiving the current state of the communication network from the network node.According to an embodiment of the fourth aspect, obtaining the current contextualinformation comprises receiving the current contextual information from the network node.According to an embodiment of the fourth aspect, obtaining the predicted actioncomprises running each of the trained control policies. The current state is an input tothe one or more trained control policies, and each of the predicted actions is an output of the respective one or more trained control policies.According to an embodiment of the fourth aspect, the current contextual informationrelates to key performance indicators of the communication network.According to an embodiment of the fourth aspect, obtaining the action fusionparameter for each of the one or more trained control policies comprises: performinga grid search; performing a gradient descent; performing Bayesian optimisation;and / or performing reinforcement learning.According to a fifth aspect of the invention, a method is provided. The method isperformed by a network node in a communication network. The method is for context-aware communication network optimization. The method comprises sending observeddata. The method comprises sending a current state of the communication network.The method comprises sending current contextual information. The method comprisesreceiving, from an apparatus, a singular action. The singular action is based on oneor more predicted actions based on the current state of the communication network, the current contextual information, and one or more action fusion parameters basedon the current contextual information. The method comprises performing the singularaction. According to an embodiment of the fifth aspect, the method comprises sending, to the apparatus, a response of the communication network from performing the singularaction. The method comprises sending a second current state of the communicationnetwork. The method comprises sending a second current contextual information. Themethod comprises receiving, from the apparatus, a second singular action. The second singular action is based on one or more second predicted actions based on the second current state of the communication network, the second current contextual information, and one or more second action fusion parameters based on the second current contextual information and obtained by the apparatus. The method comprisesperforming the second singular action.According to an embodiment of the fifth aspect, the current contextual informationrelates to key performance indicators of the communication network. According to a sixth aspect of the invention, a method is provided. The method is performed by a system. The method is for context-aware communication networkoptimization. The method comprises performing method steps according to the fourthaspect of the invention. The method comprises performing method steps according to the fifth aspect of the invention. According to a seventh aspect of the invention, a computer program is provided. The computer program comprises instructions, which when executed on an apparatus, causes the apparatus to perform the method according to the fourth aspect of the invention. According to an eighth aspect of the invention, a computer program is provided. The computer program comprises instructions, which when executed on a network node, causes the network node to perform the method according to the fifth aspect of the invention. According to a ninth aspect of the invention, a computer program is provided. The computer program comprises instructions, which when executed on a system, causes the system to perform the method according to the sixth aspect of the invention.According to a tenth aspect of the invention, a computer readable storage medium isprovided. The computer readable storage medium comprises a computer programaccording to the seventh aspect of the invention, the eighth aspect of the invention,and / or the ninth aspect of the invention. At least one or more embodiments advantageously enable to use the communicationnetwork context to dynamically adapt the control policies being used. Thus, the RLperformed is safe and robust. At least one or more embodiments advantageously enable to improve overall performance by exploiting the communication network context for aggressive learning. At least one or more embodiments advantageously enable to use the communicationnetwork context to fusion the outputs of the control policies at inference time. Thisallows to use existing trained control policies for new communication network contextor performance objectives.Further objectives of, features of, and advantages with, the invention will becomeapparent when studying the following detailed disclosure, the drawings, and theappended claims. Those skilled in the art realize that different features of the inventioncan be combined to create embodiments other than those described in the following. BRIEF DESCRIPTION OF THE DRAWINGS The above, as well as additional objects, features and advantages of the invention will be better understood through the following and non-limiting detailed description of embodiments of the invention, with reference to the appended drawings, in which: Figure 1 shows an embodiment of a communication network.Figure 2 shows a flowchart depicting an embodiment of a method performed by anapparatus. Figure 3 shows an embodiment of the method performed by the apparatus. Figure 4 shows a further embodiment of the method performed by the apparatus.Figure 5 shows an embodiment of the further method performed by the apparatus.Figure 6 shows a flowchart depicting an embodiment of a method performed by anetwork node.Figure 7 shows a flowchart of a method performed by a system.Figure 8 shows an illustrative example of messages exchanges between theapparatus and the network node.Figure 9 shows a block diagram of the apparatus.Figure 10 shows a block diagram of the network node.Figure 11 shows a block diagram of the system.Figure 12 shows a block diagram of the apparatus.Figure 13 shows a block diagram of the network node.Figure 14 shows a block diagram of the system.Figure 15 shows a block diagram illustrating a virtualization environment in whichmethod steps implemented by some embodiments may be virtualized. All figures are schematic, and generally show parts which are necessary in order to elucidate the invention, wherein other parts may be omitted or merely suggested. DETAILED DESCRIPTION The invention will now be described more fully herein with reference to the accompanying drawings, in which certain embodiments are shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Prior art ensemble techniques are one way of addressing RL algorithm and hyperparameter selection for a target application use-case. However, prior artensemble technique does not work well for RAN optimization. Indeed, in RANsystems, the performance objective varies depending on the service type, deploymentregion, network load, and other factors. Prior art RL techniques do not take these factors into account during the learning phase or the inference phase. Further, prior art’s reward function is construed to take into account one or more RAN metrics, for example service type or network load. Thus, the RL model is specific to the RAN contexts and reward formulations employed during training. However, the RLmodel is likely to encounter new RAN contexts and the reward formulations can alsoevolve over time. This would require new RL models to be trained from scratch, which would lead to high computational and operational costs.To overcome these problems, the solution to be discloses provides an apparatus forcontext-aware communication network optimization. The apparatus defines one ormore control policies for a network node in a communication network. The apparatusobtains observed data. The apparatus obtains, for each of the one or more controlpolicies, a trained control policy for the network node. Each of the trained control policies is based on the observed data. The apparatus obtains a current state of the communication network. The apparatus obtains, from each of the trained controlpolicies, a predicted action. Each of the predicted actions is based on the current stateof the communication network. The apparatus obtains current contextual information related to the communication network. The apparatus obtains, for each of the trainedcontrol policies, an action fusion parameter. Each of the action fusion parameters isbased on the current contextual information. The apparatus obtains a singular action to be performed by the network node. The singular action is based on each of the predicted actions, the current contextual information, and each of the action fusionparameters. The apparatus sends, to the network node, the singular action to beperformed. The present invention allows to use the communication network context to dynamicallyadapt the control policies being used. Thus, the RL performed is safe and robust. Thepresent invention further allows to enable to use the communication network context to fusion the outputs of the control policies at inference time. This allows to use existing trained control policies for new communication network context or performance objectives. In Figure 1, an embodiment of a communication network 100 is provided. Thecommunication network 100 may be a radio access network (RAN), as defined byThird Generation Partnership Project (3GPP). RAN is responsible for the GSM EdgeRAN (GERAN), Universal Terrestrial RAN (UTRAN), Evolved UTRAN (E-UTRAN),New Generation RAN (NG-RAN), and beyond, including their internal structures andfunctions, of systems for evolved GERAN, UTRAN, E-UTRAN, NG-RAN, and beyond.The communication network 100 comprises a network node 120. The network node120 may be an equipment capable, configured, arranged and / or operable to communicate directly or indirectly with a wireless device and / or with other network nodes or equipment in the wireless network to enable and / or provide wireless access to the wireless device and / or to perform other functions (e.g., administration) in the wireless network. Examples of network nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, Node Bs, evolved Node Bs (eNBs) and NR NodeBs (gNBs)). Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and may then also be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be arelay node or a relay donor node controlling a relay. A network node may also includeone or more (or all) parts of a distributed radio base station such as centralized digital units and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS). Yet further examples ofnetwork nodes include multi-standard radio (MSR) equipment such as MSR BSs,network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi-cell / multicast coordination entities (MCEs), core network nodes (e.g., MSCs, MMEs), O&M nodes, OSS nodes, SON nodes, positioning nodes (e.g., E-SMLCs), and / or MDTs. As another example, a network node may be a virtual network node as described in more detail below. More generally, however, network nodes may represent any suitable device (or group of devices) capable, configured, arranged, and / or operable to enable and / or provide a wireless device with access to the wireless network or to provide some service to a wireless device that has accessed the wireless network. The communication network 100 may also comprise an apparatus 110. The apparatus110 may be an equipment capable, configured, arranged and / or operable tocommunicate directly or indirectly with a wireless device and / or with other network nodes or equipment in the wireless network to enable and / or provide wireless access to the wireless device and / or to perform other functions (e.g., administration) in the wireless network. Examples of network nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, Node Bs, evolved Node Bs (eNBs) and NR NodeBs (gNBs)). Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and may then also be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be a relay node or a relay donor node controlling a relay. A network node may also include one or more (or all) parts of a distributed radio base station such as centralized digital units and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS). Yet further examples of network nodes include multi-standard radio (MSR) equipment such as MSR BSs, network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi-cell / multicast coordination entities (MCEs), core network nodes (e.g., MSCs, MMEs), O&M nodes, OSS nodes, SON nodes, positioning nodes(e.g., E-SMLCs), and / or MDTs. As another example, the apparatus 110 may be avirtual network node as described in more detail below. More generally, however, network nodes may represent any suitable device (or group of devices) capable, configured, arranged, and / or operable to enable and / or provide a wireless device with access to the wireless network or to provide some service to a wireless device that has accessed the wireless network. The network node 120 and the apparatus 110 may communicate over one or more wireless communication protocols, such as Institute of Electrical and ElectronicsEngineers (IEEE) 802.11, Code Division Multiple Access (CDMA), Wideband CDMA(WCDMA), Global System for Mobile Communications (GSM), Long Term Evolution (LTE), UMTS Terrestrial Radio Access Network (UTRAN), WiMax, or the like. In Figure 2, a flowchart depicting an embodiment of a method 200 is provided. The method 200 is performed by the apparatus 110. The method 200 is for context-aware communication network 100 optimization.The method 200 comprises defining 210 one or more control policies 310 for thenetwork node 120 in the communication network 100. An embodiment of the step 210 is illustrated in the Figure 3. Illustrated in Figure 3 are control policy_1 and control policy_ 2, however the skilled person would understand that it is possible to have only one control policy, and a plurality of control policy. The one or more control policies 310 may be table-based Q-learning models which learn values for a range of states and associated actions. The one or more control policies 310 may be Reinforcement Learning policy models that are parameterized. The Reinforcement Learning policy models may be using artificial neural networks. The artificial neural networks may be, but not limited to, DeepQ Network based Reinforcement Learning models, or proximal policy optimizationmodels. The one or more control policies 310 may be non-parametric models. The non- parameter models may be, but not limited to, employing kernel techniques, or a Gaussian process regression. The one or more control policies 310 may be heuristics-based policies. The heuristics- based policies may be, but not limited to, based on rule-based control, or empirical dynamics models.The one or more control policies 310 may be classical control policies. A classicalcontrol policy may be a Proportional-integral-derivative (PID) controller. The one ormore control policies 310 may be model predictive control. The one or more controlpolicies 310 may be signal processing and heuristics based implementation. In the event, the one or more control policies 310 are Reinforcement Learning policy models, each of the one or more control policies 310 may be of a different type of Reinforcement Learning policy model. For example, control policy_1 may be a Deep Q network based Reinforcement Learning model, and control policy_2 may be proximal policy optimization model. In particular, the reinforcement learning policy models may different in respect of artificial neural network architecture. The artificial neural network architecture may be fully connected, convolutional, and / or residual. Alternatively, or additionally, hyperparameter of the Reinforcement Learning policy models may vary. The hyperparameters may be numbers of layers, activation function, and / or learning rate. Further, the one or more control policies 310 are controlling different behaviour. For example, a choice of modulation and coding scheme for transmitted data packets may be controlled by the one or more control policies 310. The transmit power level ofwireless nodes may also be controlled by the one or more control policies 310. Hereby,the diversity in the control policies’ behaviour allows to more efficiently find predicted actions, and singular action.The method 200 comprises obtaining 220 observed data 320. An embodiment of thestep 220 of the method 200 is provided in Figure 3. The observed data 320 comprisesone or more observed states of the communication network. An embodiment of thestep 220 of the method 200 is provided in the Figure 3. The step 220 of method 200of obtaining 220 the observed data 320 may comprise receiving the observed data320 from the network node 120. The network node 120 may have collected observed data 120. Alternatively, or additionally, the apparatus 110 may be collected to the observed data 320.The observed data 320 may comprises one or more observed states of thecommunication network 100. Alternatively, or additionally, the observed data 320 maycomprise one or more observed actions sent to the network node 120 to be performed.Alternatively, or additionally, the observed data 320 comprises one or more observedresponses from the network node 120 performing the one or more observed actions. Alternatively, or additionally, the observed data 320 may comprise one or more next observed states of the communication network 100, where the one or more next observed states correspond to the state of the communication network 100 after performing the one or more observed actions.In an embodiment, the observed data 320 comprises tuples of [observed state,observed action, observed reward, observed next state]. The method 200 comprises obtaining 230, for each of the one or more control policies310, a trained control policy 330. Each of the trained control policies 330 is based onthe observed data 320. An embodiment of the step 230 of the method 200 is provided in Figure 3. Only two trained control policies (trained policy_1 and trained policy_2) are illustrated, however the skilled person would understand that it is possible to have only one trained policy, or a plurality of trained policies. In an embodiment, the step 230 of the method 200 of obtaining 230 comprises training each of the one or morecontrol policies 310 with the observed data 320. The one or more control policies 310are trained to optimize a problem. For example, the control policy_1 is trained with the observed data 320 to obtain trained control policy_1, and the control policy_2 is trainedwith the observed data 320 to obtain trained control policy_2. In an embodiment, thetraining of each of the one or more control policies 310 comprises updating the weight of the one or more control policies 310 model parameters based on the observed data.For example, the updating is such so as to minimize a cost function using numericaland analytical techniques.The method 200 comprises obtaining 240 a current state 340 of the communicationnetwork 100. An embodiment of the step 240 of the method 200 is provided in Figure3. The current state 340 of the communication network 100 may be a set of inputfeatures for the trained control policies 330.In an embodiment, the current state 340 comprises number of active services withinthe communication network 100, resource usage from the number of the activeservices, and / or operational cost of the number of the active services. The currentstate 340 in this embodiment may be a tuple (number of active services, resource usage, operational cost) so as to optimize the number of service instances to satisfy a quality of service, while also minimizing operational cost. In an embodiment, the step 240 of method 200 of obtaining 240 the current state 340 comprises receiving the current state 340 of the communication network 100 from the network node 120.The method 200 comprises obtaining 250, from each of the trained control policies330, a predicted action 350. Each of the predicted actions 350 is based on the currentstate 340 obtained in the step 240 of the method 200. An embodiment of the step 250of the method 200 is provided in Figure 3. Only two predicted actions (e.g., predictedaction_1, and predicted action_2) are illustrated, however the skilled person would understand that it is possible to have only one predicted action, or a plurality of predicted actions. In an embodiment, the step 250 of the method 200 obtaining 250 comprises runningeach of the trained control policies 330. The current state 340 is an input of the one ormore trained control policies 330. Each of the predicted actions 350 is an output of therespective one or more trained control policies 330. For example, the trained controlpolicy_1 is run with input the current state 340, and with output the predicted action_1, and the trained control policy_2 is run with input the current state 340, and with output the predicted action_2.The method 200 comprises obtaining 260 a current contextual information 360 relatedto the communication network 100. An embodiment of the step 260 of the method 200is provided in Figure 3. In an embodiment, the step 260 of the method 200 of obtaining260 the current contextual information 360 comprises receiving the current contextual information 360 from the network node 120. The current contextual information 360 may relate to key performance indicators of the communication network 100.The key performance indicators may be data throughput (e.g., average number of databits transmitted to the intended receiver within the communication network 100);latency (e.g., the time interval between when data was first available and the transmit buffer and when data was delivered to the received buffer within the communication network 100); reliability (e.g., the portion of data packet transmission that are successful); and / or power consumption (e.g., the amount of energy spent in delivering a data packet). The current contextual information 360 may be specific to a service type. In an embodiment, the service type may be Ultra Reliable and Low LatencyCommunications (URLLC). As URLLC requirements relate to latency within thecommunication network 100 and reliability within the communication network 100, aswell as a high priority, thus the current contextual information 360 may relate to thelatency within the communication network 100, and / or the reliability within thecommunication network 100. In another embodiment, the service type may be Massive Machine-TypeCommunications (mMTC). As mMTC requirements relate to power consumption withinthe communication network 100, as well as medium priority, thus the current contextual information 360 may relate to the power consumption within the communication network 100. In a further embodiment, the service type may be video streaming throughput. As video streaming requirement relates to throughput, thus the current contextual information 360 may relate to throughput. Alternatively, or additionally, the current contextual information 360 may relate to network load within the communication network 100. Indeed, network load within the communication network 100 may impact the actions given for a specific service type.The method 200 comprises obtaining 270, for each of the trained control policies 330,an action fusion parameter 370. Each of the action fusion parameters 370 is based onthe current contextual information 360. An embodiment of the step 270 of the method200 is provided in Figure 3. Only two action fusion parameters (e.g., action fusion parameter_1, and action fusion parameter_2) are illustrated in Figure 3, however the skilled person would understand that it is possible to have only one action fusion parameter, or a plurality of action fusion parameters. The one or more predictedactions 350 are to be combined with the one or more action fusion parameters 370.For example, the action fusion parameter_1 is to be combined with the predicted action_1, and the fusion parameter_2 is to be combined with the predicted action_2. In an embodiment, the step 270 of method 200 of obtaining 270 the action fusionparameter 370 for each of the one or more trained control policies 330 comprisesperforming a grid search. In particular, the grid search may comprises defining a rangeof action fusion parameter values for each of the one or more action fusion parameters370. Each of the ranges is subdivided into a finite number of discrete bins comprisingaction fusion parameter values. Each discrete bin may either be uniform or non- uniform, and / or overlapping or non-overlapping. The grid search may further comprise sampling a set of action fusion parameters values from each of the discrete bins. The sampling may be random. The grid search may further comprise executing with the sample fusion parameter forthe current contextual information 360 and recording the performance metrics. Theperformance metrics may be stored in a database. The performance metrics may relate to an objective for which the control policies have been trained. For example, in the event the objective is modulation and coding scheme selection, the performance metric may be the average throughput which is recorded periodically in the database. Alternatively, or additionally, the step 270 of method 200 of obtaining 270 the action fusion parameter 370 for each of the one or more trained control policies 330comprises performing a gradient descent. The gradient descent may comprisedetermining a set of initial action fusion parameters. The determination of the set of initial action fusion parameter may be random. Alternatively, or additionally, the determining of the set of initial action fusion parameter may be based on prior knowledge, such as a mathematical model of system dynamics, values derived fromsimulated or experimental data, or a Gaussian distribution. The set of initial actionfusion parameter is used to generate initial performance metrics. The initial performance metrics may be stored in the database. The initial performance metrics may relate to an objective for which the control policies have been trained. For example, in the event the objective is modulation and coding scheme selection, the initial performance metric may be the average throughput which is recordedperiodically in the database. The gradient descent may further comprise adjusting theset of initial action fusion parameters by adding, or subtracting, offset values andrecording new performance metrics. The performance metrics may be stored in thedatabase. The performance metrics also relate to an objective for which the controlpolicies have been trained. For example, in the event the objective is modulation andcoding scheme selection, the performance metric may be the average throughput which is recorded periodically in the database. The offset values are the difference between the performance metric value predicted by the trained control policies 330, and the recorded performance metric values. this difference is used to calculate a gradient that enables to updated the control policies. The gradient descent may further comprise computing a performance gradient with respect to the fusion adjusted set ofinitial action fusion parameters and using the performance gradient to determine anew offset values. The adjusting of the set of initial action fusion parameters, the recording, the computing of the performance gradient, and determination of a new offset values may be performed till a condition or criterion is reached. If the condition or criterion is reached the new offset values correspond to the one or more actionfusion parameters 370. For example, the condition or criterion is after a fixed numberof iterations or when the difference between initial action fusion parameters and the action fusion parameters 370 is smaller than a fixed number. Alternatively, or additionally, the step 270 of method 200 of obtaining 270 the action fusion parameter 370 for each of the one or more trained control policies 330comprises performing Bayesian optimisation. The Bayesian Optimisation maycomprise defining a prior function parameterized by one or more action fusion parameters 370. The prior function may be a Gaussian distribution over a range of action fusion parameter values. The Bayesian Optimisation may further comprise obtaining a set of action fusion parameters by sampling from the prior function andrecording performance metrics for the current contextual information 360. TheBayesian Optimisation may further comprise computing a second prior function by combining the prior function and the recording performance metrics. The obtaining a set of action fusion parameters, recording and computing the second prior function may be performed till a condition or criterion is reached. If the condition or criterion isreached the set of action fusion parameters from the last iteration correspond to theone or more action fusion parameters 370. For example, the condition or criterion isafter a fixed number of iterations or when the difference between initial action fusionparameters and the action fusion parameters 370 is smaller than a fixed number.Alternatively, or additionally, the step 270 of method 200 of obtaining 270 the action fusion parameter 370 for each of the one or more trained control policies 330comprises performing reinforcement learning. In particular, the action fusionparameter 370 may be learning by performing reinforcement learning. The reinforcement learning is performed iteratively using a greedy or stochastic approach. The reinforcement learning is based on the current contextual information 360 and performance metrics. The method 200 comprises obtaining 280 a singular action 380 to be performed by the network node 120. The singular action 380 is based on each of the predictedactions 350, the current contextual information 360, and each of the action fusionparameters 370. An embodiment of the step 280 of the method 200 is provided inFigure 3. The step 280 of method 200 of obtaining 280 the singular action 380comprises processing the one or more predicted actions 350 obtained in the step 250of the method 200, the one or more action fusion parameters 370 obtained in the step 270 of the method 200, and the current contextual information 360 obtained in the step 260 of the method 200. Such process may be called statistical action fusion.In an embodiment, the step 280 of the method 200 of obtaining 280 the singular action380 comprises performing a simple weighted averaging. Each of the one or moretrained control policies 330 may be assigned a weight value. For example, trainedcontrol policy_1 is assigned weight value ^^, and trained control policy_2 isassigned weight value ^^. The current state 340 is noted ^ . The predictedaction_1 is noted ^^ and the predicted action_2 is noted ^^ . Thus, thesingular action obtained by computing a simple ^∗ = ^^ ^^ ^^ In this embodiment, the weights ^^ and information 360. The current contextualinformation 360 may be encoded as a vector. Alternatively, the step 280 of the method 200 of obtaining 280 the singular action 380comprises variance weight averaging. For example, the mean reward is inverselyweighted by a standard deviation. A mean reward predicted by trained policy_1 for the predicted action_1, and the current state 340 may be noted ^^^,^ . A meanreward predicted by trained policy_2 for the predicted action_2, and the current state 340 may be noted ^^ ^^,^ . A reward standard deviation for the trained control policy_1may be noted^(^^^,^), and a reward standard deviation for the trained controlpolicy_2 may be noted^(^^^,^ ). Thus, the singular action 380 would be obtained^^^ ,^^Further, additional linear weight values may be used in additional to the^(^^^,^ ).variance weighting. Alternatively, the step 280 of the method 200 of obtaining 280 the singular action 380comprises assigning the singular action 380 as one of the predicted action of the oneor more predicted action 350, that have been obtained in the step 250 of the method200, that is most present within the one or more predicted action 350. For example, ifpredicted action_1 is the output of trained policy_1 and trained policy_2, then predictedaction_1 is assigned as the singular action 380. Alternatively, the step 280 of the method 200 of obtaining 280 the singular action 380 comprises assigning randomly the singular action 380 as one of the predicted actionof the one or more predicted action 350 that have been obtained in the step 250 of themethod 200. Alternatively, the step 280 of the method 200 of obtaining 280 the singular action 380comprises assigning, using a round robin sequential manner, the singular action 380as one of the predicted action of the one or more predicted action 350 that have been obtained in the step 250 of the method 200. For example, at a first iteration, the outputof the trained policy_1 (e.g., predicted action_1) is assigned as the singular action 380,and then at the second iteration, the output of the trained policy_2 (e.g., predicted action_2) is assigned as the singular action 380.The method 200 comprises sending 290, to the network node 120, the singular action380 to be performed. An embodiment of the step 290 of the method 200 is provided inFigure 3. In Figure 4, a further embodiment of the method 200 is provided. The method 200 comprises of performing the steps 210-290 of the method 200 as described herein. The method 200 may comprises obtaining 410, from the network node 120, aresponse 510 of the network node from performing the singular action 380. Anembodiment of the step 410 of the method 200 is provided in Figure 4. The response 510 is a reward for the singular action 380 performed by the network node 120, andthe state reached by the communication network 100. The response 510 is appendedto the observed data 320. For example, a tuple comprising [current state 340, singularaction 380, response 510]. In particular, the reward depends on what is beingoptimized within the communication network 100. In the event of modulation and coding scheme selection optimization, the reword may be the data throughput within the communication network 100. The response 510, in this same example, may comprise the reward and the state reached by the communication network 100. The state reached by the communication network 100, in this same example of modulation and coding scheme selection, may be estimated channel state information (CSI) values. The method 200 may comprise obtaining 420, for each of the one or more trainedcontrol policies 330, a second trained control policies 520 for the network node 120.Each of the second trained control policies 520 is based on the observed data 320and the response 510. An embodiment of the step 420 of the method 200 is providedin Figure 5. Only two second trained control policies (second trained policy_1 andsecond trained policy_2) are illustrated, however the skilled person would understand that it is possible to have only one second trained policy, or a plurality of second trainedpolicies. In an embodiment, the step 420 of the method 200 of obtaining 420 comprisestraining each of the one or more trained control policies 330 with the observed data320 appended with the response 510. The one or more trained control policies 320are trained to optimize a problem in the same manner as the one or more controlpolicies 310 are trained, as described herein. For example, the trained control policy_1is trained with the observed data 320 and response 510 to obtain second trainedcontrol policy_1, and the trained control policy_2 is trained with the observed data 320 and the response 510 to obtain second trained control policy_2. In an embodiment,the training of each of the one or more trained control policies 320 comprises updatingthe weight of the one or more trained control policies 320 model parameters based onthe observed data 320 and the response 510. For example, the updating is such soas to minimize a cost function using numerical and analytical techniques.The method 200 may comprise obtaining 430 a second current state 530 of thecommunication network 100. An embodiment of the step 430 of the method 200 isprovided in Figure 5. The second current state 430 of the communication network 100 may be a set of input features for the second trained control policies 520. In an embodiment, the second current state 530 comprises number of active services within the communication network 100, resource usage from the number of the active services, and / or operational cost of the number of the active services. The secondcurrent state 530 in this embodiment may be a tuple (number of active services,resource usage, operational cost) so as to optimize the number of service instances to satisfy a quality of service, while also minimizing operational cost. In an embodiment, the step 430 of method 200 of obtaining 430 the second currentstate 530 comprises receiving the second current state 530 of the communicationnetwork 100 from the network node 120. The second current state 530 is later in time compared to the current state 340.The method 200 may comprise obtaining 440, from each of the second trained controlpolicies 520, a second predicted action 540. Each of the second predicted actions 540is based on the second current state 530 of the communication network 100. An embodiment of the step 440 of the method 200 is provided in Figure 5. Only two second predicted actions (e.g., second predicted action_1, and second predicted action_2) are illustrated, however the skilled person would understand that it is possible to have only one second predicted action, or a plurality of second predicted actions. In an embodiment, the step 440 of the method 200 obtaining 440 comprises running each of the second trained control policies 520. The second current state 530 is aninput of the one or more second trained control policies 530. Each of the secondpredicted actions 540 is an output of the respective one or more second trained controlpolicies 530. For example, the second trained control policy_1 is run with input the second current state 530, and with output the second predicted action_1, and the second trained control policy_2 is run with input the second current state 530, and withoutput the second predicted action_2.The method 200 may comprise obtaining 450 a second current contextual information550 related to the communication network 100. An embodiment of the step 450 of themethod 200 is provided in Figure 5. In an embodiment, the step 450 of the method 200 of obtaining 450 the second current contextual information 550 comprises receiving the second current contextual information 550 from the network node 120. The second current contextual information 550 may relate to key performance indicators of the communication network 100. The second current contextual information 550 may be specific to a service type. In an embodiment, the service type may be Ultra Reliable and Low Latency Communications (URLLC). As URLLC requirements relate to latency within the communication network 100 and reliability within the communication network 100, as well as a high priority, thus the current contextual information 360 may relate to the latency within the communication network 100, and / or the reliability within the communication network 100. In another embodiment, the service type may be Massive Machine-Type Communications (mMTC). As mMTC requirements relate to power consumption within the communication network 100, as well as medium priority, thus the current contextual information 360 may relate to the power consumption within the communication network 100. In a further embodiment, the service type may be video streaming throughput. As video streaming requirement relates to throughput, thus the current contextual information 360 may relate to throughput. Alternatively, or additionally, the second current contextual information 550 may relate to network load within the communication network 100. Indeed, network load within the communication network 100 may impact the actions given for a specific service type. The second current contextual information 550 is later in time compared to the current contextual information 360. The method 200 may comprise obtaining 460, for each of the second trained control policies 520, a second action fusion parameter 560. Each second action fusionparameter 560 is based on the second contextual information 550. An embodiment ofthe step 460 of the method 400 is provided in Figure 5. Only two second action fusionparameters (e.g., second action fusion parameter_1, and second action fusion parameter_2) are illustrated in Figure 5, however the skilled person would understand that it is possible to have only one action fusion parameter, or a plurality of actionfusion parameters. The one or more second predicted actions 540 are to be combinedwith the one or more second action fusion parameters 560. For example, the second action fusion parameter_1 is to be combined with the second predicted action_1, and the second fusion parameter_2 is to be combined with the second predicted action_2.In an embodiment, the step 460 of method 400 of obtaining 460 the second actionfusion parameter 560 for each of the one or more second trained control policies 520comprises performing a grid search. In particular, the grid search may comprises defining a range of second action fusion parameter values for each of the one or more second action fusion parameters 560. Each of the ranges is subdivided into a finite number of discrete bins comprising action fusion parameter values. Each discrete bin may either be uniform or non-uniform, and / or overlapping or non-overlapping. The grid search may further comprise sampling a set of second action fusion parameters values from each of the discrete bins. The sampling may be random. The grid search may further comprise executing with the second sample fusionparameter for the second current contextual information 550 and recording theperformance metrics. The performance metrics may be stored in a database. Theperformance metrics may relate to an objective for which the trained control policies have been trained. For example, in the event the objective is modulation and coding scheme selection, the performance metric may be the average throughput which is recorded periodically in the database.Alternatively, or additionally, the step 460 of method 400 of obtaining 460 the secondaction fusion parameter 560 for each of the one or more second trained control policies520 comprises performing a gradient descent. The gradient descent may comprisedetermining a set of initial second action fusion parameters. The determination of the set of initial second action fusion parameter may be random. Alternatively, oradditionally, the determining of the set of initial second action fusion parameter maybe based on prior knowledge, such as a mathematical model of system dynamics, values derived from simulated or experimental data, or a Gaussian distribution. The set of initial second action fusion parameter is used to generate initial performance metrics. The initial performance metrics may be stored in the database. The initial performance metrics may relate to an objective for which the trained control policies have been trained. For example, in the event the objective is modulation and coding scheme selection, the initial performance metric may be the average throughput which is recorded periodically in the database. The gradient descent may further comprise adjusting the set of initial second action fusion parameters by adding, or subtracting,offset values and recording new performance metrics. The performance metrics maybe stored in the database. The performance metrics also relate to an objective for which the control policies have been trained. For example, in the event the objective is modulation and coding scheme selection, the performance metric may be the average throughput which is recorded periodically in the database. The offset values are the difference between the performance metric value predicted by the second trained control policies 540, and the recorded performance metric values. this difference is used to calculate a gradient that enables to updated the second trained control policies. The gradient descent may further comprise computing a performance gradient with respect to the fusion adjusted set of initial second action fusion parameters and using the performance gradient to determine a new offset values. The adjusting of the set of initial second action fusion parameters, the recording, the computing of the performance gradient, and determination of a new offset values may be performed till a condition or criterion is reached. If the condition or criterion isreached the new offset values correspond to the one or more second action fusionparameters 560. For example, the condition or criterion is after a fixed number of iterations or when the difference between initial action fusion parameters and the second action fusion parameters 560 is smaller than a fixed number.Alternatively, or additionally, the step 460 of method 400 of obtaining 460 the secondaction fusion parameter 560 for each of the one or more second trained control policies 520 comprises performing Bayesian optimisation. The Bayesian Optimisation may comprise defining a prior function parameterized by one or more second action fusion parameters 560. The prior function may be a Gaussian distribution over a range of action fusion parameter values. The Bayesian Optimisation may further comprise obtaining a set of action fusion parameters by sampling from the prior function and recording performance metrics for the second current contextual information 550. The Bayesian Optimisation may further comprise computing a second prior function by combining the prior function and the recording performance metrics. The obtaining a set of second action fusion parameters, recording and computing the second priorfunction may be performed till a condition or criterion is reached. If the condition orcriterion is reached the set of second action fusion parameters from the last iterationcorrespond to the one or more second action fusion parameters 560. For example, thecondition or criterion is after a fixed number of iterations or when the difference between initial action fusion parameters and the second action fusion parameters 560 is smaller than a fixed number.Alternatively, or additionally, the step 460 of method 400 of obtaining 460 the secondaction fusion parameter 560 for each of the one or more second trained control policies 520 comprises performing reinforcement learning. In particular, the second actionfusion parameter 560 may be learning by performing reinforcement learning. Thereinforcement learning is performed iteratively using a greedy or stochastic approach. The reinforcement learning is based on the second current contextual information 550 and performance metrics. The method 200 may comprise obtaining 470 a second singular action 570 to beperformed by the network node 120. The second singular action 570 is based on eachof the second predicted actions 540, the second current contextual information 550, and each of the second action fusion parameters 560.An embodiment of the step 470 of the method 400 is provided in Figure 5. The step470 of method 400 of obtaining 470 the second singular action 570 comprisesprocessing the one or more second predicted actions 540 obtained in the step 440 of the method 400, the one or more second action fusion parameters 560 obtained in the step 460 of the method 400, and the second current contextual information 550 obtained in the step 430 of the method 400. Such process may be called statistical action fusion. In an embodiment, the step 470 of the method 400 of obtaining 470 the second singular action 570 comprises performing a simple weighted averaging. Each of the one or more second trained control policies 520 may be assigned a weight value. Forexample, second trained control policy_1 is assigned weight value ^^, and secondtrained control policy_2 is assigned weight value ^^. The second current state 530is noted ^ . The second predicted action_1 is noted ^^ and the second^ predicted action_2 is noted ^ . Thus, the second singular action 570 would be obtained by computing a averaging:^∗ = ^ ^^ ^ ^^ In this embodiment, the weights ^^ and are based on the second current contextual information 550. The secondcurrent contextual information 550 may be encoded as a vector.Alternatively, the step 470 of the method 400 of obtaining 470 the second singularaction 570 comprises variance weight averaging. For example, the mean reward is inversely weighted by a standard deviation. A mean reward predicted by second trained policy_1 for the second predicted action_1, and the second current state 530 may be noted ^^^,^ . A mean reward predicted by second trained policy_2 for thesecond predicted action_2, and the second current state 530 may be noted^^ ^^,^ . Areward standard deviation for the second trained control policy_1 may be noted ^(^^^,^), and a reward standard deviation for the second trained control policy_2may be noted^(^^^,^ ). Thus, the second singular action 570 would be obtained by^^^ ^^^ ,^^ )^. Further, additional linear weight values may be used in additional to the^(^^,^variance weighting.Alternatively, the step 470 of the method 400 of obtaining 470 the second singularaction 570 comprises assigning the second singular action 570 as one of the secondpredicted action of the one or more second predicted action 540, that have beenobtained in the step 440 of the method 400, that is most present within the one or moresecond predicted action 540. For example, if second predicted action_1 is the output of second trained policy_1 and second trained policy_2, then second predicted action_1 is assigned as the second singular action 570.Alternatively, the step 470 of the method 400 of obtaining 470 the second singularaction 570 comprises assigning randomly the second singular action 570 as one ofthe second predicted action of the one or more second predicted action 540 that havebeen obtained in the step 440 of the method 400.Alternatively, the step 470 of the method 400 of obtaining 470 the second singularaction 570 comprises assigning, using a round robin sequential manner, the second singular action 570 as one of the second predicted action of the one or more second predicted action 540 that have been obtained in the step 440 of the method 400. For example, at a first iteration, the output of the second trained policy_1 (e.g., second predicted action_1) is assigned as the second singular action 570, and then at the second iteration, the output of the second trained policy_2 (e.g., second predicted action_2) is assigned as the second singular action 570.The method 200 may comprise sending 480 to the network node 120, the secondsingular action 570 to be performed. An embodiment of the step 480 of the method200 is provided in Figure 5.In Figure 6, a flowchart depicting an embodiment of a method 600 is provided. Themethod 600 is performed by the network node 120 in the communication network 100.The method 600 is for context-aware communication network 100 optimization. The method 600 comprises sending 610 the observed data 320. The observed data320 of the step 610 of the method 600 corresponds to the observed data 320 obtainedin the step 210 of the method 200, as described herein. The method 600 comprises sending 620 the current state 340 of communication network 100. The current state 340 of the step 620 of the method 600 corresponds to the current state 340 obtained in the step 240 of the method 200, as described herein. The method 600 comprises sending 630 the current contextual information 360. The current contextual information 360 of the step 630 of the method 600 corresponds to the current contextual information 360 of the step 260 of the method 200, as described herein. The method 600 comprises receiving 640, from the apparatus 110, the singular action 380. The singular action 380 of the step 640 of the method 600 corresponds to the singular action 380 of the step 280 and 290 of the method 200, as described herein.The method 600 comprises performing 650 the singular action 380. Performing 650the singular action 380 may comprise executing the singular action 380. In the event of modulation and coding scheme selection optimization, performing 650 the singularaction 380 may be transmitting a data packet after it has been processed using thesingular action 380. In this same example, the singular action 380 is considered complete once the data packet has been received by the receiver and processed to recover the constituent data bits. In an embodiment, the method 600 further comprises sending 660, to the apparatus 110, the response 510. The response 510 of the step 660 of the method 600 corresponds to the response 510 of the step 410 of the method 200, as described herein. In an embodiment, the method 600 further comprises sending 670 the second current state 530. The second current state 530 of the step 670 of the method 600 corresponds to the second current state 530 of the step 430 of the method 200, as described herein.In an embodiment, the method 600 further comprises sending 680 the second currentcontextual information 550. The second contextual information 550 of the step 680 of the method 600 corresponds to the second contextual information 550 of the step 450 of the method 200, as described herein. In an embodiment, the method 600 further comprises receiving 690, from the apparatus 110, the second singular action 570. The second singular action 570 of the step 690 of the method 600 corresponds to the second singular action 570 of the step 480 of the method 200, as described herein.In Figure 7, a flowchart of a method 700 performed by a system 130 is provided. Themethod 700 is performed by the system 130. The system 130 comprises the apparatus110 and the network node 120. The method 700 is for context-aware communicationnetwork 100 optimization.The method 700 comprises performing the steps 210-290 of the method 200, asdescribed herein. The method 700 further comprises performing the steps 410-480 of the method 200, as described herein. The method 700 further comprises performing the steps 610-690 of the method 600, as described herein. In Figure 8, an illustrative example of messages exchanges between the apparatus 110 and the network node 120, is provided.The apparatus 110 is configured to obtain observed data 810 from the network node120. In an embodiment, the observed data 810 corresponds to the observed data 320 obtained in the step 220 of the method 200. In an embodiment, the observed data 810corresponds to the observed data 320 sent in the step 610 of the method 600.The apparatus 110 is configured to obtain a current state 820 from the network node120. In an embodiment, the current state 820 corresponds to the current state 340 obtained in the step 240 of the method 200. In an embodiment, the current state 820 corresponds to the current state 340 sent in the step 610 of the method 600.The apparatus 110 is configured to obtain a current contextual information 830 fromthe network node 120. In an embodiment, the current contextual information 830corresponds to the current contextual information 360 obtained in the step 260 of themethod 200. In an embodiment, the current contextual information 830 correspondsto the current contextual information 360 sent in the step 630 of the method 600.The apparatus 110 is configured to send a singular action 840 to the network node120. In an embodiment, the singular action 840 corresponds to the singular action 380sent in the step 290 of the method 200. In an embodiment, the singular action 840corresponds to the singular action 380 received in the step 640 of the method 600. The apparatus 110 is configured to receive a response 850 from the network node120. In an embodiment, the response 850 corresponds to the response 510 obtainedin the step 410 of the method 200. In an embodiment, the response 850 correspondsto the response 510 sent in the step 660 of the method 600. The apparatus 110 is configured to obtain a second current state 860 from the networknode 120. In an embodiment, the second current state 860 corresponds to the secondcurrent state 530 obtained in the step 430 of the method 200. In an embodiment, thesecond current state 860 corresponds to the second current state 530 sent in the step670 of the method 600.The apparatus 110 is configured to obtain a second current contextual information 870from the network node 120. In an embodiment, the second current contextualinformation 870 corresponds to the second current contextual information 550obtained in the step 450 of the method 200. In an embodiment, the second currentcontextual information 870 corresponds to the second current contextual information 550 sent in the step 680 of the method 600. The apparatus 110 is configured to send a second singular action 880 to the networknode 120. In an embodiment, the second singular action 880 corresponds to thesecond singular action 880 sent in the step 480 of the method 200. In an embodiment,the second singular action 880 corresponds to the second singular action 880 received in the step 690 of the method 900. In Figure 9, a block diagram of the apparatus 110 is provided. The apparatus 110 comprises a defining unit 910, a first obtaining unit 915, a second obtaining unit 920, a third obtaining unit 925, a fourth obtaining unit 930, a fifth obtaining unit 935, a sixthobtaining unit 940, a seventh obtaining unit 945, and a first sending unit 950.The defining unit 910 is configured to cause the apparatus 110 to perform the step 210 of the method 200, as described herein. The first obtaining unit 915 is configured to cause the apparatus 110 to perform the step 220 of the method 200, as described herein. The first obtaining unit 915 may be a receiving unit. The second obtaining unit 920 is configured to cause the apparatus 110 to perform the step 230 of the method 220, as described herein. The second obtaining unit 920may be a training unit.The third obtaining unit 925 is configured to cause the apparatus 110 to perform the step 240 of the method 200, as described herein. The third obtaining unit 925 is a receiving unit. The fourth obtaining unit 930 is configured to cause the apparatus 110 to perform the step 250 of the method 200, as described herein. The fourth obtaining unit 930 may be an executing unit. The fifth obtaining unit 935 is configured to cause the apparatus 110 to perform the step 260 of the method 200, as described herein. The fifth obtaining unit 935 may be a receiving unit. The sixth obtaining unit 940 is configured to cause the apparatus 110 to perform the step 270 of the method 200, as described herein. The sixth obtaining unit 940 may be performing unit. The seventh obtaining unit 945 is configured to cause the apparatus 110 to perform the step 280 of the method 200, as described herein. The seventh obtaining unit 945 may be a performing unit. The first sending unit 950 is configured to cause the apparatus 110 to perform the step 290 of the method 200, as described herein.In an embodiment, the apparatus 110 comprises an eighth obtaining unit 955. Theeighth obtaining unit 955 is configured to cause the apparatus 110 to perform the step 410 of the method 200, as described herein. The eighth obtaining unit 955 may be a receiving unit.In an embodiment, the apparatus 110 comprises a ninth obtaining unit 960. The ninthobtaining unit 960 is configured to cause the apparatus 110 to perform the step 420 of the method 200, as described herein. The ninth obtaining unit 960 may be a training unit.In an embodiment, the apparatus 110 comprises a tenth obtaining unit 965. The tenthobtaining unit 965 is configured to cause the apparatus 110 to perform the step 430 of the method 200, as described herein. The tenth obtaining unit 965 may be a receiving unit.In an embodiment, the apparatus 110 comprises an eleventh obtaining unit 970. Theeleventh obtaining unit 970 is configured to cause the apparatus 110 to perform the step 440 of the method 200, as described herein. The eleventh obtaining unit 970 may be an executing unit.In an embodiment, the apparatus 110 comprises a twelfth obtaining unit 975. Thetwelfth obtaining unit 975 is configured to cause the apparatus 110 to perform the step 450 of the method 200, as described herein. The twelfth obtaining unit 975 may be a receiving unit.In an embodiment, the apparatus 110 comprises a thirteenth obtaining unit 980. Thethirteenth obtaining unit 980 is configured to cause the apparatus 110 to perform the step 460 of the method 200, as described herein. The thirteenth obtaining unit 980 may be a performing unit.In an embodiment, the apparatus 110 comprises a fourteenth obtaining unit 985. Thefourteenth obtaining unit 985 is configured to cause the apparatus 110 to perform the step 470 of the method 200, as described herein. The fourteenth obtaining unit 985 may be a performing unit.In an embodiment, the apparatus 110 comprises a second sending unit 990. Thesecond sending unit 990 is configured to cause the apparatus 110 to perform the step 480 of the method 200, as described herein. In an embodiment, the first obtaining unit 915, the third obtaining unit 925, the fifthobtaining unit 935, the eighth obtaining unit 955, the tenth obtaining unit 965, and thetwelfth obtaining unit 975 may be a single obtaining unit.In an embodiment, the second obtaining unit 920, the fourth obtaining unit 930, sixth obtaining unit 940, the seventh obtaining unit 945, the ninth obtaining unit 960,eleventh obtaining unit 970, the thirteenth obtaining unit 980, and the fourteenthobtaining unit 985 may be a single performing unit. In an embodiment, the first sending unit 950, and the second sending unit 990 are a single sending unit.The defining unit 910, the first obtaining unit 915, the second obtaining unit 920, thethird obtaining unit 925, the fourth obtaining unit 930, the fifth obtaining unit 935, the sixth obtaining unit 940, the seventh obtaining unit 945, the first sending unit 950, theeighth obtaining unit 955, the ninth obtaining unit 960, the tenth obtaining unit 965, theeleventh obtaining unit 970, the twelfth obtaining unit 975, the thirteenth obtaining unit 980, the fourteenth obtaining unit 985, and the second sending unit 990, illustrated in Figure 9, may be implemented as a hardware solution or as a combination of software and hardware, e.g., by one or more of: a processor or a micro-processor and adequate software and memory for storing of the software, a Programmable Logic Device (PLD) or other electronic component(s) or processing circuitry configured to perform the actions described above with regards to the method 200. In Figure 10, a block diagram of the network node 120 is provided. The network node120 comprises a first sending unit 1010, a second sending unit 1020, a third sendingunit 1030, a first receiving unit 1040, and a performing unit 1050. The first sending unit 1010 is configured to cause the network node 120 to perform the step 610 of the method 600, as described herein. The second sending unit 1020 is configured to cause the network node 120 to perform the step 620 of the method 600, as described herein. The third sending unit 1030 is configured to cause the network node 120 to perform the step 630 of the method 600, as described herein. The first receiving unit 1040 is configured to cause the network node 120 to perform the step 640 of the method 600, as described herein. The performing unit 1050 is configured to cause the network node 120 to perform the step 650 of the method 600, as described herein.In an embodiment, the network node 120 comprises a fourth sending unit 1060. Thefourth sending unit 1060 is configured to cause the network node 120 to perform the step 660 of the method 600, as described herein.In an embodiment, the network node 120 comprises a fifth sending unit 1070. The fifthsending unit 1070 is configured to cause the network node 120 to perform the step 670 of the method 600, as described herein.In an embodiment, the network node 120 comprises a sixth sending unit 1080. Thesixth sending unit 1080 is configured to cause the network node 120 to perform the step 680 of the method 600, as described herein.In an embodiment, the network node 120 comprises a second receiving unit 1090. Thesecond receiving unit 1090 is configured to cause the network node 120 to perform the step 690 of the method 600, as described herein. In an embodiment, the first sending unit 1010, the second sending unit 1020, the thirdsending unit 1030, the fourth sending unit 1060, the fifth sending unit 1070, and thesixth sending unit 1080, are a same sending unit. In an embodiment, the first receiving unit 1040, and the second receiving unit 1090, are a same receiving unit. In an embodiment, the first sending unit 1010, the second sending unit 1020, the third sending unit 1030, the fourth sending unit 1060, the fifth sending unit 1070, and the sixth sending unit 1080, the first receiving unit 1040, and the second receiving unit 1090 are a same transceiver unit.The first sending unit 1010, the second sending unit 1020, the third sending unit 1030,the fourth sending unit 1060, the fifth sending unit 1070, and the sixth sending unit 1080, the first receiving unit 1040, and the second receiving unit 1090, and the performing unit 1050, illustrated in Figure 10, may be implemented as a hardware solution or as a combination of software and hardware, e.g., by one or more of: a processor or a micro-processor and adequate software and memory for storing of the software, a Programmable Logic Device (PLD) or other electronic component(s) or processing circuitry configured to perform the actions described above with regards to the method 600. In Figure 11, a block diagram of the system 130 is provided. The system 130 comprises the apparatus 110 as described with relation to Figure 9. The system 130 comprises the network node 120 as described with relation to Figure 10. In Figure 12, an embodiment of the apparatus 110 is provided. The apparatus 110 comprises a processor 1210, and a computer readable storage medium 1220 in the form of a memory 1225. The memory 1625 contains a computer program 1230 comprising instructions executable by the processor 1210 whereby the apparatus 110 is operative to perform the steps of the method 200.In Figure 13, an embodiment of the network node 120 is provided. The network node120 comprises a processor 1310, and a computer readable storage medium 1320 in the form of a memory 1325. The memory 1325 contains a computer program 1330 comprising instructions executable by the processor 1310 whereby the network node 120 is operative to perform the steps of the method 600.In Figure 14, an embodiment of the system 130 is provided. The system 130comprises a processor 1410, and a computer readable storage medium 1420 in the form of a memory 1425. The memory 1425 contains a computer program 1430 comprising instructions executable by the processor 1410 whereby the system 130 is operative to perform the steps of the method 700. The (non-transitory) computer readable storage media mentioned above may be an Electrically Erasable Programmable Read-Only Memory (EEPROM), a flash memory,Field Programmable Gate Array, and a hard drive.The processor 1210 of Figure 12, the processor 1310 of Figure 13, and the processor1410 of the Figure 14 may be a single Central Processing Unit (CPU), but could alsocomprise two or more processing units. For example, the processor 1210 of Figure12, the processor 1310 of Figure 13, and the processor 1410 of the Figure 14 mayinclude general purpose microprocessors; instructions set processors and / or related chips sets and / or special purpose microprocessors such as Application Specific Circuits (ASICs). The processor 1210 of Figure 12, the processor 1310 of Figure 13,and the processor 1410 of the Figure 14 may also comprise board memory for cachingpurposes. The computer program 1230 of Figure 12, the computer program 1330 of Figure 13,and the computer program 1430 of Figure 14 may be carried by a computer program product connected to the processor 1210 of Figure 12, the processor 1310 of Figure 13, and the processor 1410 of the Figure 14. The computer program products may be or comprise a non-transitory computer readable storage medium on which computer program 1230 of Figure 12, the computer program 1330 of Figure 13, and the computer program 1430 of Figure 14 are stored. For example, the computer program products may be a flash memory, a Random-Access memory (RAM), a Read-Only Memory (ROM), or an EEPROM, and the computer programs described above could in alternative embodiments be distributed on different computer program products inthe form of memories.In Figure 15, a block diagram illustrating a virtualization environment QQ500 in which method steps implemented by some embodiments may be virtualized. In the presentcontext, virtualizing means creating virtual versions of apparatus 110 and / or thenetwork node 120 which may include virtualizing hardware platforms, storage devicesand networking resources. As used herein, virtualization can be applied to apparatus110 and / or the network node 120 described herein, or components thereof, and relatesto an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the method steps described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments QQ500 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node may be entirely virtualized. In some embodiments, the virtualization environment QQ500 includes components defined by the O-RAN Alliance, such as an O-Cloud environment orchestrated by a Service Management and Orchestration Framework via an O-2 interface. Applications QQ502 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in the virtualization environment QQ500 to implement some of the method steps, features, functions, and / or benefits of some of the embodiments disclosed herein. Hardware QQ504 includes processing circuitry, memory that stores software and / or instructions executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers QQ506 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs QQ508a and QQ508b (one or more of which may be generally referred to as VMs QQ508), and / or perform any of the method steps, functions, features and / or benefits described in relation with some embodiments described herein. The virtualization layer QQ506 may present a virtual operating platform that appears like networking hardware to the VMs QQ508. The VMs QQ508 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer QQ506. Different embodiments of the instance of a virtual appliance QQ502 may be implemented on one or more of VMs QQ508, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment. In the context of NFV, a VM QQ508 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, non-virtualized machine. Each of the VMs QQ508, and that part of hardware QQ504 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs QQ508 on top of the hardware QQ504 and corresponds to the application QQ502. Hardware QQ504 may be implemented in a standalone network node with generic or specific components. Hardware QQ504 may implement some functions via virtualization. Alternatively, hardware QQ504 may be part of a larger cluster of hardware (e.g. such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration QQ510, which, among others, oversees lifecycle management of applications QQ502. In some embodiments, hardware QQ504 is coupled to one or more radio units that each include one or more transmitters and one or more receivers that may be coupled to one or more antennas. Radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with the virtualcomponents to provide a virtual node with radio capabilities, such as a radio accessnode or a base station. In some embodiments, some signaling can be provided with the use of a control system QQ512 which may alternatively be used for communication between hardware nodes and radio units. It should be understood that although the terms “first”, “second”, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Forexample, the first trained control policy could be termed the second trained controlpolicy, and similarly, the second trained control policy could be termed the first trained control policy.As used herein, the term “and / or” includes any and all combinations of one or more ofthe associated listed terms. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limited of example embodiments. As used herein, the single forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising”, “has”, “having”, “includes”, and / or “including”, when used herein, specify the presence of stated features, elements, and / or components etc. but not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. The disclosure has been described above in reference to embodiments thereof. It should be understood that various modifications, alternatives, and additions can be made by those skilled in the art without departing from the scope of the disclosure. Therefore, the scope of the disclosure is not limited to the above particular embodiments but only defined by the claims as attached.

Claims

CLAIMS 1. An apparatus (110) for context-aware communication network (100)optimization, the apparatus (110) being configured to: -define one or more control policies (310) for a network node (120) in acommunication network (100); -obtain observed data (320);- obtain, for each of the one or more control policies (310), a trained controlpolicy for the network node (120), wherein each of the trained control policies (330) is based on the observed data (320);- obtain a current state (340) of the communication network (100);- obtain, from each of the trained control policies (330), a predicted action,wherein each of the predicted actions (350) is based on the current state(340) of the communication network (100); -obtain current contextual information (360) related to the communicationnetwork (100); -obtain, for each of the trained control policies (330), an action fusionparameter, wherein each of the action fusion parameters (370) is based onthe current contextual information (360);- obtain a singular action (380) to be performed by the network node (120),wherein the singular action (380) is based on each of the predicted actions (350), the current contextual information (360), and each of the action fusionparameters (370); and- send, to the network node (120), the singular action (380) to be performed.

2. The apparatus (110) according to claim 1, is configured to:- obtain, from the network node (120), a response (510) of the network node(120) from performing the singular action (380);- obtain, for each of the one or more trained control policies (330), a secondtrained control policy for the network node (120), wherein each of thesecond trained control policies (520) is based on the observed data (320) and the response (510);- obtain a second current state (530) of the communication network (100);- obtain, from each of the second trained control policies (520), a secondpredicted action, wherein each of the second predicted actions (540) isbased on the second current state (530) of the communication network (100); -obtain a second current contextual information (550) related to thecommunication network (100); -obtain, for each of the second trained control policies (520), a second actionfusion parameter, wherein each of the second action fusion parameters (560) is based on the second contextual information (550);- obtain a second singular action (570) to be performed by the network node(120), wherein the second singular information (570) is based on each ofthe second predicted actions (540), the second current contextualinformation (550), and each of the second action fusion parameters (560);and -send, to the network node (120), the second singular action (570) to beperformed.

3. The apparatus (110) according to any one of claims 1-2, wherein obtaining theobserved data (320) comprises receiving the observed data (320) from the network node (120).

4. The apparatus (110) according to any one of claims 1-3, wherein obtaining thetrained control policy comprises training each of the one or more control policies(310) with the observed data (320).

5. The apparatus (110) according to any one of claims 1-4, wherein obtaining thecurrent state (340) of the communication network (100) comprises receiving the current state (340) of the communication network (100) from the network node (120).

6. The apparatus (110) according to any one of claims 1-5, wherein obtaining thecontextual information (360) comprises receiving the contextual information (360) from the network node (120).

7. The apparatus (110) according to any one of claims 1-6, wherein obtaining thepredicted action comprises running each of the trained control policies (330),wherein the current state (340) is an input to the one or more trained control policies (330), and each of the predicted actions (350) is an output of therespective one or more trained control policies (330).

8. The apparatus (110) according to any one of claims 1-7, wherein the currentcontextual information (360) relates to key performance indicators of thecommunication network (100).

9. The apparatus (110) according to any one of claims 1-8, wherein obtaining theaction fusion parameter (370) for each of the one or more trained controlpolicies (330) comprises:- performing a grid search;- performing a gradient descent;- performing Bayesian optimisation; and / or- performing reinforcement learning.

10. A network node (120) in a communication network (100), for context-awarecommunication network (100) optimization, the network node (120) isconfigured to: -send observed data (320);- send a current state (340) of the communication network (100);- send current contextual information (360);- receive, from an apparatus (110), a singular action (380), wherein thesingular action (380) is based on one or more predicted actions (350) basedon the current state (340) of the communication network (100), the current contextual information (360), and one or more action fusion parameters(370) based on the current contextual information (360); and- perform the singular action (380).11.The network node (120) according to claim 10, is configured to: -send, to the apparatus (110), a response (510) of the communicationnetwork (100) from performing the singular action (380);- send a second current state (530) of the communication network (100);- send a second current contextual information (550);- receive, from the apparatus (110), a second singular action (570), whereinthe second singular action (570) is based on one or more second predictedactions (540) based on the second current state (530) of the communication network (100), the second current contextual information (550), and one ormore second action fusion parameters (560) based on the second current contextual information (550) and obtained by the apparatus (110); and -perform the second singular action (570).12.The network node (120) according to any one of claims 10-11, wherein the current contextual information (360) relates to key performance indicators ofthe communication network (100). 13.A system (130) for context-aware communication network (100) optimization, the system comprising the apparatus (110) according to any of claims 1-9 and the network node (120) according to any one of claims 10-12.14.A method (200), performed by an apparatus (110), for context-awarecommunication network (100) optimization, the method comprising: -defining (210) one or more control policies (310) for a network node (120)in a communication network (100); -obtaining (220) observed data (320);- obtaining (230), for each of the one or more control policies (310), a trainedcontrol policy for the network node (120), wherein each of the trained control policies (330) is based on the observed data (320);- obtaining (240) a current state (340) of the communication network (100);- obtaining (250), from each of the trained control policies, a predicted action,wherein each of the predicted actions (350) is based on the current state (340) of the communication network (100); -obtaining (260) current contextual information (360) related to thecommunication network (100);- obtaining (270), for each of the trained control policies (330), an actionfusion parameter, wherein each of the action fusion parameters (370) isbased on the current contextual information (360);- obtaining (280) a singular action (380) to be performed by the network node(120), wherein the singular action (380) is based on each of the predictedactions (350), the current contextual information (360), and each of theaction fusion parameters (370); and- sending (290), to the network node (120), the singular action (380) to beperformed.The method (200) according to claim 14, comprising:- obtaining (410), from the network node (120), a response (510) of thenetwork node (120) from performing the singular action (380);- obtaining, for each of the one or more trained control policies (330), asecond trained control policy for the network node (120), wherein each ofthe second trained control policies (520) is based on the observed data(320) and the response (510);- obtaining (420) a second current state (530) of the communication network(100);- obtaining (430), from each of the second trained control policies (520), asecond predicted action, wherein each of the second predicted actions (560)is based on the second current state (530) of the communication network(100);- obtaining (440) a second current contextual information (550) related to thecommunication network (100);- obtaining (450), for each of the second trained control policies (520), asecond action fusion parameter (560), wherein each of the second actionfusion parameters (560) is based on the second contextual information (550);- obtaining (460) a second singular action (570) to be performed by thenetwork node (120), wherein the second singular action (570) is based on each of the second predicted actions (540), the second current contextualinformation (550), and each of the second action fusion parameters (560);and- sending (470), to the network node (120), the second singular action (570)to be performed.The method (200) according to any one of claims 14-15, wherein obtaining(210) the observed data (320) comprises receiving the observed data (320) from the network node (120).The method (200) according to any one of claims 14-16, wherein obtaining(230) the trained control policy comprises training each of the one or morecontrol policies (310) with the observed data (320).The method (200) according to any one of claims 14-17, wherein obtaining(240) the current state (340) of the communication network (100) comprises receiving the current state (340) of the communication network (100) from the network node (120).The method (200) according to any one of claims 14-18, wherein obtaining(260) the current contextual information (360) comprises receiving the currentcontextual information (360) from the network node (120).The method (200) according to any one of claims 14-19, wherein obtaining(250) the predicted action comprises running each of the trained control policies(330), wherein the current state (340) is an input to the one or more trainedcontrol policies (330), and each of the predicted actions (350) is an output ofthe respective one or more trained control policies (330).The method (200) according to any one of claims 14-20, wherein the currentcontextual information (360) relates to key performance indicators of thecommunication network (100).The method (200) according to any one of claims 14-21, wherein obtaining(270) the action fusion parameter (370) for each of the one or more trained control policies (330) comprises:- performing a grid search;- performing a gradient descent;- performing Bayesian optimisation; and / or- performing reinforcement learning.

23. A method (600), performed by a network node (120) in a communicationnetwork (100), for context-aware communication network (100) optimization, the method (600) comprising: -sending (610) observed data (320);- sending (620) a current state (340) of the communication network (100);- sending (630) current contextual information (360);- receiving (640), from an apparatus (110), a singular action (380), whereinthe singular action (380) is based on one or more predicted actions (350) based on the current state (340) of the communication network (100), the current contextual information (360), and one or more action fusionparameters (370) based on the current contextual information (360); and- performing (650) the singular action (380).24.The method (600) according to claim 23, comprising:- sending (660), to the apparatus (110), a response (510) of thecommunication network (100) from performing the singular action (380);- sending (670) a second current state (530) of the communication network(100); -sending (680) a second current contextual information (550);- receiving (690), from the apparatus (110), a second singular action (570),wherein the second singular action (570) is based on one or more second predicted actions (540) based on the second current state (530) of the communication network (100), the second current contextual information (550), and one or more second action fusion parameters (560) based on thesecond current contextual information (550) and obtained by the apparatus (110); and -performing the second singular action (570).25.The method (600) according to any one of claims 23-24, wherein the currentcontextual information (360) relates to key performance indicators of thecommunication network (100).26.A method (700), performed by a system (130) for improving traffic performanceof a communication network (100), the method (700) performing method steps according to any one of claims 14-22, and any one of claims 23-25. 27.A computer program (1230) comprising instructions, which when executed on an apparatus (110), causes the apparatus (110) to perform the method (200) according to one or more of claims 14-22.28.A computer program (1330) comprising instructions, which when executed on a network node (120), causes the network node (120) to perform the method (600) according to one or more of claims 23-25. 29.A computer program (1430) comprising instructions, which when executed on a system (130), causes the system to perform the method (700) according to claim 26. 30.A computer readable storage medium (1225, 1325, 1425) comprising a computer program (1230, 1330, 1430) according to claim 27, claim 28, and / or claim 29.

Citation Information

Patent Citations

  • Provision of precoder selection policy for a multi-antenna transmitter

    US20220393729A1

  • Reward estimation for a target policy

    WO2022199792A1

  • Training a policy for managing a communication network environment

    WO2022253453A1

  • Controlling reinforcement learning agents using geometric policy composition

    WO2023144395A1