Systems and methods for explainable reinforcement learning for cell parameter optimization
By employing a Q-learning RL agent with a DNN to forecast KPI gains, the system addresses the lack of performance gain estimates in current RL methods, offering explainable decision-making and improved optimization of cell parameters.
Patent Information
- Application Number
- PCT/IB2024/052904
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-03-26
- Publication Date
- 2025-06-19
AI Technical Summary
Current Reinforcement Learning (RL) methods for cell parameter optimization, such as RET optimization, do not provide estimates of performance gain for Key Performance Indicators (KPIs) that form the reward, requiring waiting for actual changes to be implemented and new statistics to be collected.
The proposed solution involves using a Q-learning RL agent with a Deep Neural Network (DNN) to forecast the gain in KPIs before applying changes, allowing for explainable decision-making and enabling the prediction of KPI gains to determine whether to implement individual actions.
This approach provides explainability of RL agent decisions, allows for forecasting KPI gains before changes are made, and enables the collection of statistics on prediction errors, enhancing user trust and decision-making in cell parameter optimization.
Smart Images

Figure IB2024052904_19062025_PF_FP_ABST
Abstract
Description
[0001] SYSTEMS AND METHODS FOR EXPLAINABLE REINFORCEMENT LEARNING FOR
[0002] CELL PARAMETER OPTIMIZATION
[0003] TECHNICAL FIELD
[0004] The present disclosure relates, in general, to wireless communications and, more particularly, systems and methods for explainable Reinforcement Learning (RL) for cell parameter optimization.
[0005] BACKGROUND
[0006] Reinforcement Learning (RL) has recently become a very popular and powerful tool to tackle parameter optimization problems in wireless networks. One of the most studied parameters is Remote Electrical Tilt (RET), which defines the antenna tilt of the cell. The parameter values can be changed remotely, and by modifying it, it is possible to find a trade-off between Signal to Interference plus Noise Ratio (SINR) and coverage, in both the Uplink (UL) and the Downlink (DL). Examples of RET optimizers based on RL have been considered by IEEE Wireless Communications Networking Conference (WCNC). See, Vannella, G. lakovidis, et. al, Remote Electrical Tilt Optimization via Safe Reinforcement Learning, 2021 IEEE WCNC, 2021, pp. 1-7, doi: 10.1109 / WCNC49053.2021.9417363; Weisi Guo et. al, Spectral- and Energy-Efficient Antenna Tilting in a HetNet using Reinforcement Learning, IEEE WCNC: MAC, 2013.
[0007] A method for iterative parameter optimization based on RL has been proposed, where multiple instances of a unique agent are used by each cell separately. See, PCT International Publication Number WO2021 / 190772A1, to Mendo Mateo et. al, Policy for Optimising Cell Parameters, published September 30, 2021. This approach uses a unique and common policy, which has the benefit of accelerating the learning and, therefore, reducing the number of iterations in which the agents might take exploratory erroneous decisions. Whatever lesson is learnt by the agent instance of one cell is immediately available in the common policy for the rest of cells.
[0008] The Ericsson Mobility Report released in June 2021 includes an article that describes an optimization experience on two live networks in which an optimizer based on the RL concept defined in WO2021 / 190772A1 is executed and provides excellent results. See, Al: Enhancing Customer Experience in a Complex 5G World, Ericsson Mobility Report, June 2021, https: / / www.ericsson.com / 49ced5 / assets / local / mobility-report / documents / 2021 / ai_enhancing- customer-experience.pdf. The RL agents use a reward function that considers several Key Performance Indicators (KPI) from the cell under optimization and its neighboring cells. The KPIs are computed as the relative difference between the values measured at the actual and the previous step for each iteration. The approach and the results obtained with this approach was appreciated by operators of the two live networks where the trials were conducted.
[0009] A frequent request from customers related to the optimization of radio network parameters is the capability of providing some forecast of the expected performance improvement after each set of simultaneous changes. Currently, this is not possible since, although previous techniques are based on Q-learning that rely on produced estimations of the reward that will be obtained in the following iteration, the reward is typically defined as an unintelligible measure, which contains the aggregation of KPIs understandable by the customers but is impossible to segregate.
[0010] A solution has been proposed to decompose an environment reward into a sum of meaningful reward types. See, Z. Juozapaitis et. al, Explainable Reinforcement Learning Via Reward Decomposition, IJCAI / ECAI Workshop on Explainable Artificial Intelligence, 2019, https: / / par.nsf. gov / serylets,zpurl / l 0159391 . However, this solution is not applicable in cases where the reward is derived from a non-linear combination of meaningful metrics.
[0011] Thus, there currently exist certain challenge(s). For example, the RET optimization solution disclosed in WO2021 / 190772A1 has proven to yield exceptional results in the field, slowly steering the RET in steps of 1 dB towards the most convenient configuration without any human intervention, not only in making the decisions on the RET changes, but also in the design of the algorithm, which is entirely generated by Artificial Intelligence (Al) using data. In this respect, the only human input is the selection of the input features and the definition of the reward function. However, this solution does not provide any estimates of the performance gain for the KPIs that are part of the reward, and it is necessary to wait until the suggested changes are actually carried out and new statistics are collected to obtain an indication of such performance. An estimation of the performance gain for the KPIs that form the reward is frequently demanded by customers and could bring an added value to the product and services based on the product if available.
[0012] As another example, a solution has been proposed for decomposing an environment reward into a sum of meaningful reward components. See, Juozapaitis et al., Explainable Reinforcement Learning Via Reward Decomposition, IJCAI / ECAI Workshop on Explainable Artificial Intelligence, 2019, https: / / par.nsf.gov / servlets / purl / 10159391. The solution is focused on predicting the q-values of the reward components and assumes that the q-value of the final reward is the sum of the q-values of the reward components, but is not applicable in cases where the reward is not derived only from the sum of meaningful metrics.
[0013] SUMMARY
[0014] Certain aspects of the disclosure and their embodiments may provide solutions to these or other challenges. For example, methods and systems are provided for explainable RL for cell parameter optimization based on a Q-learning RL agent with a Deep Neural Network (DNN), which permits forecasting the gain in the KPIs that are part of the reward before applying the changes.
[0015] According to certain embodiments, a method of using an RL agent in a communication network to optimize one or more cell parameters in a cell of the communication network includes identifying a plurality of KPIs associated with a reward for performing at least one action to optimize one or more cell parameters. An associated DNN is used for each of the plurality of KPIs to estimate an associated gain for each of the plurality of KPIs. Based on the associated gain predicted for each of the plurality of KPIs, a determinization is made as to whether or not to implement at least one action to optimize the one or more cell parameters.
[0016] According to certain embodiments, a computer program product comprising a computer readable medium having computer readable code embodied therein is provided. The computer readable code is configured such that, on execution by a computer or processor, the computer or processor is caused to identify a plurality of KPIs associated with a reward for performing at least one action to optimize one or more cell parameters. The computer or processor is further caused to use an associated DNN for each of the plurality of KPIs to estimate an associated gain for each of the plurality of KPIs. Based on the associated gain predicted for each of the plurality of KPIs, the computer or processor is caused to determine whether or not to implement at least one action to optimize the one or more cell parameters. According to certain embodiments, an apparatus in a communication network includes a RL agent for optimizing one or more cell parameters in a cell of the communication network. The apparatus is configured to identify a plurality of KPIs associated with a reward for performing at least one action to optimize one or more cell parameters. The apparatus is configured to use an associated DNN for each of the plurality of KPIs to estimate an associated gain for each of the plurality of KPIs. Based on the associated gain predicted for each of the plurality of KPIs, the apparatus is configured to determine whether or not to implement at least one action to optimize the one or more cell parameters.
[0017] Certain embodiments may provide one or more of the following technical advantage(s). For example, certain embodiments may provide a technical advantage of providing “explainability” of RL agent decisions. For example, certain embodiments provide the possibility of explaining the target of every decision made by the optimizer based on RL. Rather than estimating an unintelligible reward value, certain embodiments may provide a technical advantage of providing an estimation of every KPI that composes the reward. This boosts the value offered to the customers, who in general do not feel totally comfortable by accepting the outputs of the Al-based algorithms without explanation.
[0018] As another example, unlike previous solutions and techniques that are not applicable in cases where the reward is not derived from the sum of meaningful metrics, certain embodiments may provide a technical advantage of being applicable to reward definition for RET optimization based on RL. In the considered case of RET optimization, based on the assumption of discount factor equal to zero, the reward is computed as a relative return gain with respect to the existing performance before applying an action into the environment. This implies a non-linear relation between the KPIs used to compute the reward and the reward itself. On the other hand, a discount factor of zero is not just an assumption, but a compulsory key configuration decision when optimizing RET with RL.
[0019] As another example, certain embodiments may provide a technical advantage of enabling predicted KPI gains to be used to decide whether to apply or not each individual action.
[0020] As still another example, certain embodiments provide a technical advantage of enabling statistics on predicted gain error to be collected, which will give the user additional information of the uncertainty of the predicted gains. Other advantages may be readily apparent to one having skill in the art. Certain embodiments may have none, some, or all of the recited advantages.
[0021] BRIEF DESCRIPTION OF THE DRAWINGS
[0022] For a more complete understanding of the disclosed embodiments and their features and advantages, reference is now made to the following description, taken in conjunction with the accompanying drawings, in which:
[0023] FIGURE 1 illustrates an optimizer based on RL using an unintelligible reward function, according to certain embodiments;
[0024] FIGURE 2 illustrates an example predictor utilizing four trained DNNs to predict the relative gain of four associated KPIs, according to certain embodiments;
[0025] FIGURE 3 illustrates another example predictor having an additional DNN operating as a reward predictor for simultaneously predicting KPI gains and the rewards for possible actions, according to certain embodiments;
[0026] FIGURE 4 illustrates another example predictor where the DNNs used to predict the KPI gains produce additional outputs to help in the prediction of the reward, according to certain embodiments;
[0027] FIGURE 5 illustrates an example optimizer including a predictor such as those described with regard to FIGURES 2, 3, and 4, according to certain embodiments;
[0028] FIGURE 6 illustrates the methods and systems being implemented as a single rAPP in a Non Real-Time Radio Access Network Intelligent Controller (Non-RT RIC) located in a Service Management and Orchestration (SMO) framework in communication with an Open Radio Access Network (O-RAN) based network, according to certain embodiments;
[0029] FIGURE 7 illustrates an example communication system, according to certain embodiments;
[0030] FIGURE 8 illustrates an example UE, according to certain embodiments;
[0031] FIGURE 9 illustrates an example network node, according to certain embodiments;
[0032] FIGURE 10 illustrates a block diagram of a host, according to certain embodiments;
[0033] FIGURE 11 illustrates a virtualization environment in which functions implemented by some embodiments may be virtualized, according to certain embodiments; FIGURE 12 illustrates a host communicating via a network node with a UE over a partially wireless connection, according to certain embodiments; and
[0034] FIGURE 13 illustrates a method of using a RL agent in a communication network to optimize one or more cell parameters in a cell of the communication network, according to certain embodiments.
[0035] DETAILED DESCRIPTION
[0036] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.
[0037] As used herein, ‘node’ can be a network node or a UE. Examples of network nodes are NodeB, base station (BS), multi-standard radio (MSR) radio node such as MSR BS, eNodeB (eNB), gNodeB (gNB), Master eNB (MeNB), Secondary eNB (SeNB), integrated access backhaul (IAB) node, network controller, radio network controller (RNC), base station controller (BSC), relay, donor node controlling relay, base transceiver station (BTS), Central Unit (e.g. in a gNB), Distributed Unit (e.g. in a gNB), Baseband Unit, Centralized Baseband, C-RAN, access point (AP), transmission points, transmission nodes, Remote Radio Unit (RRU), Remote Radio Head (RRH), nodes in distributed antenna system (DAS), core network node (e.g. Mobile Switching Center (MSC), Mobility Management Entity (MME), etc.), Operations & Maintenance (O&M), Operations Support System (OSS), Self Organizing Network (SON), positioning node (e.g. E- SMLC), etc. The terms network node and radio network node are used interchangeably herein.
[0038] Another example of a node is user equipment (UE), which is a non-limiting term and refers to any type of wireless device communicating with a network node and / or with another UE in a cellular or mobile communication system. Examples of UE are target device, device to device (D2D) UE, vehicular to vehicular (V2V), machine type UE, MTC UE or UE capable of machine to machine (M2M) communication, Personal Digital Assistant (PDA), Tablet, mobile terminals, smart phone, laptop embedded equipment (LEE), laptop mounted equipment (LME), Unified Serial Bus (USB) dongles, etc.
[0039] The term radio access technology (RAT), may refer to any RAT such as, for example, Universal Terrestrial Radio Access Network (UTRA), Evolved Universal Terrestrial Radio Access Network (E-UTRA), narrow band internet of things (NB-IoT), WiFi, Bluetooth, next generation RAT, NR, 4G, 5G, etc. Any of the equipment denoted by the terms node, network node or radio network node may be capable of supporting a single or multiple RATs.
[0040] The term signal or radio signal used herein can be any physical signal or physical channel. Examples of downlink (DL) physical signals are reference signal (RS) such as Primary Synchronization Signal (PSS), Secondary Synchronization Signal (SSS), Channel State Information-Reference Signal (CSLRS), Demodulation Reference Signal (DMRS) signals in SS / PBCH block (SSB), discovery reference signal (DRS), Cell Specific Reference Signal (CRS), Positioning Reference Signal (PRS), etc. RS may be periodic. For example, RS occasions carrying one or more RSs may occur with certain periodicity (e.g., 20 ms, 40 ms, etc.). The RS may also be aperiodic.
[0041] Each SSB carries New Radio-Primary Synchronization Signal (NR-PSS), New RadioSecondary Synchronization Signal (NR-SSS) and New Radio-Physical Broadcast Channel (NR- PBCH) in four successive symbols. One or multiple Synchronization Signal Blocks (SSBs) are transmitted in one SSB burst which is repeated with certain periodicity such as, for example, 5 ms, 10 ms, 20 ms, 40 ms, 80 ms, and 160 ms. The UE is configured with information about SSB on cells of certain carrier frequency by one or more SS / PBCH block measurement timing configuration (SMTC) configurations. The SMTC configuration comprising parameters such as SMTC periodicity, SMTC occasion length in time or duration, SMTC time offset with regard to reference time (e.g., serving cell’s SFN) etc. Therefore, SMTC occasion may also occur with certain periodicity (e.g., 5 ms, 10 ms, 20 ms, 40 ms, 80 ms, and 160 ms). Examples of uplink (UL) physical signals are reference signals such as Sounding Reference Signals (SRS), Demodulation Reference Signals (DMRS), etc. The term physical channel refers to any channel carrying higher layer information e.g. data, control etc. Examples of physical channels are Physical Broadcast Channel (PBCH), Physical Downlink Control Channel (PDCCH), Physical Downlink Shared Channel (PDSCH), Physical Uplink Shared Channel (PUSCH), Physical Uplink Control Channel (PUCCH), Physical Uplink Shared Channel (PUSCH), Short PUSCH (sPUCCH), Short PDSCH (sPDSCH), Short PUCCH (sPUCCH), Short PUSCH (sPUSCH), MTC PDCCH (MPDCCH), Narrowband PBCH (NPBCH), Narrowband PDCCH (NPDCCH), Narrowband PDSCH (NPDSCH), Narrowband PUSCH (NPUSCH), Enhanced PDCCH (E-PDCCH), etc.
[0042] The term time resource used herein may correspond to any type of physical resource or radio resource expressed in terms of length of time. Examples of time resources are symbol, time slot, subframe, radio frame, transmission time interval (TTI), interleaving time, slot, sub-slot, minislot, system frame number (SFN) cycle, hyper-SFN (H-SFN) cycle, etc.
[0043] According to certain embodiments, improved systems, methods, and techniques are proposed that build on top of existing cell parameter optimization approaches based on a Q- learning RL agent with a Deep Neural Network (DNN). The proposed systems, methods, and techniques permit forecasting the gain in the KPIs that form the reward before applying the changes. The proposed systems, methods, and techniques help to explain the reasons behind each change, which is an insight frequently demanded by customers. Moreover, the predictions can be used to decide whether each individual action should be applied or skipped at every step. For this decision, statistics on measured prediction errors can also be used, which may determine the uncertainty of the estimated KPI gains.
[0044] According to certain embodiments, the proposed systems, methods, and techniques rely on using multiple DNNs instead of one. For example, according to a proposed method, an associated one of a plurality of DNNs is used to perform estimation for each KPI gain. The outputs of the multiple DNNs are combined through an additional DNN to obtain a single value per action at a given state. This combined value is exactly the reward.
[0045] According to certain embodiments, the predicted KPI gains, combined with collected statistics on the gain prediction error, can be used to decide whether to apply or not each individual action. In a particular embodiment, for example, the KPI gain predictions are compared to one or more thresholds when determining if a proposed action for a particular cell at a particular time is worth applying.
[0046] As part of reward decomposition, certain embodiments disclosed herein make it possible to obtain the predicted KPI gains that form the reward for each possible action. This way, the user has the possibility to perceive the contribution of each KPI to the decision of selecting an action, but also to the non-selected actions.
[0047] Certain embodiments enable the collection of statistics of the error between the predicted KPI gains and the actual measured KPI gains after applying the proposed action. This error statistics is useful to determine the probability of future KPI gain estimates being above the user threshold used to decide whether an action should be applied. An alternative criterion to make the decision is to verify if this probability is higher than a minimum value. This probability is also information useful to the user, contributing to the explainability.
[0048] Certain embodiments disclosed herein relate to the particular case of RET optimization based on the techniques described in WO2021 / 190772A1. However, the systems, methods, and technique disclosed herein may be extended to the optimization of other cell parameters based on RL solutions applied at cell level.
[0049] Two different approaches are discussed below. First, an existing reference approach is described for RET optimization based on RL with an unintelligible reward as disclosed in WO2021 / 190772A1. Then, a modified and improved approach is described that enables providing estimations on the KPIs that compose the reward.
[0050] Reference Approach
[0051] FIGURE 1 illustrates a block diagram of an optimizer 100 based on RL using an unintelligible reward function, according to certain embodiments. In particular, FIGURE 1 relates to a reference scheme for RET optimization such as, for example, the RL solution proposed in WO2021 / 190772A1. As illustrated, optimizer 100 includes a reward predictor 105, which predicts the reward for three possible actions 140, which may also be known as q- values. Reward predictor 105 may be implemented with a neural network.
[0052] In operation, when optimizing a parameter, there are typically three different actions: keep the same value, increment the value in a fixed step, and decrement the value in a fixed step. The predictor 105 is trained continuously in block 135 with the new data coming from the network once the action is implemented. The target of the training process is to minimize the error between the predicted reward and the actual reward. Thus, Block 135 is the lost function to minimize during the training of the reward predictor.
[0053] The policy represented in FIGURE 1 is what is called an epsilon-greedy policy. Two different actions are proposed: the best action (action with highest predicted reward) in Block 145 and a random action in Block 150. More specifically, Block 145 is a selection block and operates to select the index of the input with highest value (q-value). Block 150 is a random action generator used for exploration (which occurs with probability equal to epsilon) instead of action with the best predicted q-value.
[0054] Block 155 is a decision block and operates to decide whether to explore or not. In particular, Block 155 is a random selection block, which selects between both proposals: the best action is chosen with probability 1 - epsilon and the random action is selected with probability epsilon. Epsilon is typically an hyperparameter and in many cases the value is dynamic, starting with higher values at the beginning and reducing its value progressively. Exploration happens with probability epsilon and implies selecting a random action. Otherwise, the action with the best predicted q- value is selected. As used herein the phrase unintelligible reward function means that the reward is defined as a function of several KPIs that are impossible to derive from the original reward once calculated. In this particular case, the reward at instant t + 1 is defined as where RMt+1is the Reward Metric (RM) at instant t + 1 (right after the parameter update), and RMtis the RM at instant t (right before the parameter update). The RM at instant t is computed as
[0055] RMt= GTt+ GTNt+ ARt+ ARNt, (2) where GTtand ARtare the good traffic and the accessibility rate at the cell at instant t respectively, and GTNtand ARNtare the average good traffic and the average accessibility rate measured at instant t at the closest neighboring cells, weighted by their overlapping factors with respect to the cell under consideration.
[0056] The good traffic is defined as the ratio of traffic with good coverage and good quality with respect to the total traffic.
[0057] Good coverage means having reference signal received power (RSRP) over a predefined threshold.
[0058] Good quality means having SINR over a predefined threshold.
[0059] The overlapping factor between two cells can be obtained as the periodicity in which both cells are reported simultaneously by the same UE in Cell Traffic Recording (CTR). Accessibility rate is equal to one minus the Radio Resource Control (RRC) congestion rate.
[0060] It is noted that, in a particular embodiment of FIGURE 1, a DNN 105 is trained by minimizing a loss function 135 defined as the mean square residuals between the estimated and the measured rewards. A one hot vector is used to represent the action for which every reward is obtained. By using the inner product of the one hot vector with the estimated reward, the DNN 105 may associate every reward value to learn with the action with which it was obtained. As with any other approach based on deep Q-learning, a training step of the DNN is simultaneously executed with the calculation of the best action. In this case, an epsilon greedy policy has been adopted, where an explorative random action is selected with probability s, and otherwise the action with the highest q-value (expected reward) is selected. Improved Approach
[0061] The target of a proposed approach is to deliver the same results as with the scheme depicted in FIGURE 1 while also providing estimated gain values for each KPIs that form the reward. For this reason, certain embodiments disclosed herein use multiple DNNs instead of one. For example, according to certain embodiments described below, one DNN is used per KPI to estimate KPI gains. In a particular embodiment, an additional DNN is used to obtain the final reward estimations.
[0062] The gains for each KPI used to calculate the reward, as defined in Equation (1) and Equation (2), are defined as: where gainGTis the good traffic (GT) gain, gainGTNis the gain in good traffic at neighbors (GTN), gainARis the accessibility rate (AR), and gainARNis the gain in accessibility rate at neighbors (ARN).
[0063] FIGURE 2 illustrates an example predictor 200 utilizing four trained DNNs 205A, 205B, 205C, 205D to predict the relative gain of four associated KPIs 210A, 210B, 210C, 210D , according to certain embodiments. The scheme of every DNN is very similar to the scheme of the DNN 105 in FIGURE 1 , with the difference being that, instead of learning the reward for every action at a certain state, each DNN in FIGURE 2 learns how to predict the gain for an associated one of the four KPIs that form the reward, for every possible action 215 at a certain state 220.
[0064] According to certain embodiments, the estimated reward for each action 215 can be calculated from the estimated KPI gains and the KPI values measured in the previous time instant as follows: where all the variables that contain the circumflex accent (A) are prediction vectors of three estimated values (one per possible action). It is noted that the expression in Equation (4) does not show the reward as a simple sum of the desired metrics since there are other metrics involved that complicate the relation. For that reason, it is not possible to just apply the solution proposed by Z. Juozapaitis et. al as discussed above. Computing the reward as a function of predicted values might involve higher error than directly predicting the reward. Although the Equation (4) is theoretically exact, the prediction errors of each DNN would be aggregated.
[0065] Thus, according to certain embodiments, it is proposed that an additional DNN is used to learn how to calculate the estimated reward Rt+1for each action from the estimated KPI gains and the KPI values measured in the previous time instant. FIGURE 3 illustrates another example predictor 300 having an additional DNN 305 operating as a reward predictor for simultaneously predicting KPI gains and the rewards for possible actions (i.e., the q-values), according to certain embodiments. Certain elements of FIGURE 3 are similar to those described above with regard to FIGURE 2 and are shown with like reference numerals. However, the additional DNN 305 is useful for better adjustment and get better predictions.
[0066] FIGURE 4 illustrates an alternative version of yet another example predictor 400 where the DNNs 205 A, 205B, 205C, 205D used to predict the KPI gains produce additional outputs 310A, 310B, 310C, D10D, 310E, 31 OF, 310G, and 31 OH to help in the prediction of the reward, according to certain embodiments. The training of the neurons that produce these additional outputs 310A, 310B, 310C, D10D, 310E, 31 OF, 310G, and 31 OH is focused on minimizing the loss function of the reward DNN 305. Thanks to these outputs 310A, 310B, 310C, D10D, 310E, 31 OF, 310G, and 31 OH, the prediction of the reward is also based on information from the state 220, ensuring the same accuracy in the predictions as with the reference approach described above.
[0067] It is noted that the values Rt+1[0], 7?t+1[l] and 7?t+1[2] in FIGURE 3 and FIGURE 4 exactly match the q-values q[0], q[l] and q[2] of FIGURE 1 for the reference approach. Moreover, the reward predictor of the proposed approach can also provide the estimated values of the four KPIs that form the reward: GT (GT[0], GT[1] and GT[2]), GTN (G7 [0], G7 [1] and G7 [2]), AR (AR[0], AR[1] and AR[2]), and ARN (ARA[0], ARA[1] and ARA[2]).
[0068] FIGURE 5 illustrates a high-level block diagram of an example optimizer 500, according to certain embodiments. It is noted that the diagram is very similar to that of FIGURE 1 for the reference approach; however, in this case, the reward predictor 505 illustrated in FIGURE 5 is not just a DNN but a more complex system as described with relation to FIGURE 2, FIGURE 3, and FIGURE 4. In addition to providing the expected reward 515A, 515B, and 515C for every action at a given state 520, the reward predictor 505 illustrated in FIGURE 5 also provides, for each KPI 510A, 51 OB, 510C, and 510D, the expected gain 525 A, 525B, 525C, 525D that form the reward for every action and at a given state 520.
[0069] Thought the methods, systems, and techniques disclosed herein are depicted using an example that includes four DNNs used for four associated KPIs, it is recognized that the methods, systems, and techniques may be generalized to any other case where the reward is composed of a different number of KPIs, also using a different aggregation method, as far as the discount factor is configured to zero, because otherwise, Equation (4) would not be valid, as the reward would be a function of current a previous KPI gains.
[0070] Certain embodiments disclosed herein provide additional improvements over the reference approach, thanks to having estimated the KPI gains. For example, certain embodiments provide the possibility to discard applying an action in a cell in a particular step if none of the estimated KPI gains are significant. This can be determined based on user thresholds. As another example, certain embodiments enable the collection of statistics on the prediction error for each KPI gain. This may be performed by comparing the predictions with the actual values when the changes are applied. These error statistics can be used to evaluate the uncertainty of successive predictions and use them to fine tune the decision on discarding actions: in this case based on a minimum probability that the KPI gains are over the user thresholds.
[0071] Implementation Using NWADF
[0072] According to certain embodiments, an application that implements the explainable RL concept runs on a server with access to an 5G Network Data Analytics Function (NWDAF) instance that provides the required RAN statistical inputs to execute the explainable RL application. The NWDAF instance has access to the source RAN statistical data. The application that implements the explainable RL concept registers to the NWDAF instance with the aim of obtaining the required input data. The NWDAF instance also has the required processing capabilities to produce input data expected by the application.
[0073] Implementation using the SMO Framework FIGURE 6 illustrates an Open Radio Access Network (O-RAN) architecture 600 that includes the methods and systems disclosed herein being implemented as a single rApp 605A in a Non Real-Time Radio Access Network Intelligent Controller (Non-RT RIC) 610 located in the SMO framework 615 in communication with an O-RAN based network 620. It can be used as a standalone app and / or the results can be used by other rApps 605B and 605C.
[0074] FIGURE 7 shows an example of a communication system 700 in accordance with some embodiments. In the example, the communication system 700 includes a telecommunication network 702 that includes an access network 704, such as a radio access network (RAN), and a core network 706, which includes one or more core network nodes 708. The access network 704 includes one or more access network nodes, such as network nodes 710a and 710b (one or more of which may be generally referred to as network nodes 710), or any other similar 3rdGeneration Partnership Project (3 GPP) access node or non-3GPP access point. The network nodes 710 facilitate direct or indirect connection of user equipment (UE), such as by connecting UEs 712a, 712b, 712c, and 712d (one or more of which may be generally referred to as UEs 712) to the core network 706 over one or more wireless connections.
[0075] Example wireless communications over a wireless connection include transmitting and / or receiving wireless signals using electromagnetic waves, radio waves, infrared waves, and / or other types of signals suitable for conveying information without the use of wires, cables, or other material conductors. Moreover, in different embodiments, the communication system 700 may include any number of wired or wireless networks, network nodes, UEs, and / or any other components or systems that may facilitate or participate in the communication of data and / or signals whether via wired or wireless connections. The communication system 700 may include and / or interface with any type of communication, telecommunication, data, cellular, radio network, and / or other similar type of system.
[0076] The UEs 712 may be any of a wide variety of communication devices, including wireless devices arranged, configured, and / or operable to communicate wirelessly with the network nodes 710 and other communication devices. Similarly, the network nodes 710 are arranged, capable, configured, and / or operable to communicate directly or indirectly with the UEs 712 and / or with other network nodes or equipment in the telecommunication network 702 to enable and / or provide network access, such as wireless network access, and / or to perform other functions, such as administration in the telecommunication network 702.
[0077] In the depicted example, the core network 706 connects the network nodes 710 to one or more hosts, such as host 716. These connections may be direct or indirect via one or more intermediary networks or devices. In other examples, network nodes may be directly coupled to hosts. The core network 706 includes one more core network nodes (e.g., core network node 708) that are structured with hardware and software components. Features of these components may be substantially similar to those described with respect to the UEs, network nodes, and / or hosts, such that the descriptions thereof are generally applicable to the corresponding components of the core network node 708. Example core network nodes include functions of one or more of a Mobile Switching Center (MSC), Mobility Management Entity (MME), Home Subscriber Server (HSS), Access and Mobility Management Function (AMF), Session Management Function (SMF), Authentication Server Function (AUSF), Subscription Identifier De-concealing function (SIDF), Unified Data Management (UDM), Security Edge Protection Proxy (SEPP), Network Exposure Function (NEF), and / or a User Plane Function (UPF).
[0078] The host 716 may be under the ownership or control of a service provider other than an operator or provider of the access network 704 and / or the telecommunication network 702, and may be operated by the service provider or on behalf of the service provider. The host 716 may host a variety of applications to provide one or more service. Examples of such applications include live and pre-recorded audio / video content, data collection services such as retrieving and compiling data on various ambient conditions detected by a plurality of UEs, analytics functionality, social media, functions for controlling or otherwise interacting with remote devices, functions for an alarm and surveillance center, or any other such function performed by a server.
[0079] As a whole, the communication system 700 of FIGURE 7 enables connectivity between the UEs, network nodes, and hosts. In that sense, the communication system may be configured to operate according to predefined rules or procedures, such as specific standards that include, but are not limited to: Global System for Mobile Communications (GSM); Universal Mobile Telecommunications System (UMTS); Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, 5G standards, or any applicable future generation standard (e.g., 6G); wireless local area network (WLAN) standards, such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards (WiFi); and / or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave, Near Field Communication (NFC) ZigBee, LiFi, and / or any low-power wide-area network (LPWAN) standards such as LoRa and Sigfox.
[0080] In some examples, the telecommunication network 702 is a cellular network that implements 3 GPP standardized features. Accordingly, the telecommunications network 702 may support network slicing to provide different logical networks to different devices that are connected to the telecommunication network 702. For example, the telecommunications network 702 may provide Ultra Reliable Low Latency Communication (URLLC) services to some UEs, while providing Enhanced Mobile Broadband (eMBB) services to other UEs, and / or Massive Machine Type Communication (mMTC)ZMassive loT services to yet further UEs.
[0081] In some examples, the UEs 712 are configured to transmit and / or receive information without direct human interaction. For instance, a UE may be designed to transmit information to the access network 704 on a predetermined schedule, when triggered by an internal or external event, or in response to requests from the access network 704. Additionally, a UE may be configured for operating in single- or multi-RAT or multi-standard mode. For example, a UE may operate with any one or combination of Wi-Fi, NR (New Radio) and LIE, i.e. being configured for multi-radio dual connectivity (MR-DC), such as E-UTRAN (Evolved-UMTS Terrestrial Radio Access Network) New Radio - Dual Connectivity (EN-DC).
[0082] In the example, the hub 714 communicates with the access network 704 to facilitate indirect communication between one or more UEs (e.g., UE 712c and / or 712d) and network nodes (e.g., network node 710b). In some examples, the hub 714 may be a controller, router, content source and analytics, or any of the other communication devices described herein regarding UEs. For example, the hub 714 may be a broadband router enabling access to the core network 706 for the UEs. As another example, the hub 714 may be a controller that sends commands or instructions to one or more actuators in the UEs. Commands or instructions may be received from the UEs, network nodes 710, or by executable code, script, process, or other instructions in the hub 714. As another example, the hub 714 may be a data collector that acts as temporary storage for UE data and, in some embodiments, may perform analysis or other processing of the data. As another example, the hub 714 may be a content source. For example, for a UE that is a VR headset, display, loudspeaker or other media delivery device, the hub 714 may retrieve VR assets, video, audio, or other media or data related to sensory information via a network node, which the hub 714 then provides to the UE either directly, after performing local processing, and / or after adding additional local content. In still another example, the hub 714 acts as a proxy server or orchestrator for the UEs, in particular in if one or more of the UEs are low energy loT devices.
[0083] The hub 714 may have a constant / persistent or intermittent connection to the network node 710b. The hub 714 may also allow for a different communication scheme and / or schedule between the hub 714 and UEs (e.g., UE 712c and / or 712d), and between the hub 714 and the core network 706. In other examples, the hub 714 is connected to the core network 706 and / or one or more UEs via a wired connection. Moreover, the hub 714 may be configured to connect to an M2M service provider over the access network 704 and / or to another UE over a direct connection. In some scenarios, UEs may establish a wireless connection with the network nodes 710 while still connected via the hub 714 via a wired or wireless connection. In some embodiments, the hub 714 may be a dedicated hub - that is, a hub whose primary function is to route communications to / from the UEs from / to the network node 710b. In other embodiments, the hub 714 may be a nondedicated hub - that is, a device which is capable of operating to route communications between the UEs and network node 710b, but which is additionally capable of operating as a communication start and / or end point for certain data channels.
[0084] FIGURE 8 shows a UE 800 in accordance with some embodiments. As used herein, a UE refers to a device capable, configured, arranged and / or operable to communicate wirelessly with network nodes and / or other UEs. Examples of a UE include, but are not limited to, a smart phone, mobile phone, cell phone, voice over IP (VoIP) phone, wireless local loop phone, desktop computer, personal digital assistant (PDA), wireless cameras, gaming console or device, music storage device, playback appliance, wearable terminal device, wireless endpoint, mobile station, tablet, laptop, laptop- embedded equipment (LEE), laptop-mounted equipment (LME), smart device, wireless customer-premise equipment (CPE), vehicle-mounted or vehicle embedded / integrated wireless device, etc. Other examples include any UE identified by the 3rd Generation Partnership Project (3GPP), including a narrow band internet of things (NB-IoT) UE, a machine type communication (MTC) UE, and / or an enhanced MTC (eMTC) UE.
[0085] A UE may support device-to-device (D2D) communication, for example by implementing a 3GPP standard for sidelink communication, Dedicated Short-Range Communication (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-everything (V2X). In other examples, a UE may not necessarily have a user in the sense of a human user who owns and / or operates the relevant device. Instead, a UE may represent a device that is intended for sale to, or operation by, a human user but which may not, or which may not initially, be associated with a specific human user (e.g., a smart sprinkler controller). Alternatively, a UE may represent a device that is not intended for sale to, or operation by, an end user but which may be associated with or operated for the benefit of a user (e.g., a smart power meter).
[0086] The UE 800 includes processing circuitry 802 that is operatively coupled via a bus 804 to an input / output interface 806, a power source 808, a memory 810, a communication interface 812, and / or any other component, or any combination thereof. Certain UEs may utilize all or a subset of the components shown in FIGURE 8. The level of integration between the components may vary from one UE to another UE. Further, certain UEs may contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0087] The processing circuitry 802 is configured to process instructions and data and may be configured to implement any sequential state machine operative to execute instructions stored as machine-readable computer programs in the memory 810. The processing circuitry 802 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, field- programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.); programmable logic together with appropriate firmware; one or more stored computer programs, general-purpose processors, such as a microprocessor or digital signal processor (DSP), together with appropriate software; or any combination of the above. For example, the processing circuitry 802 may include multiple central processing units (CPUs).
[0088] In the example, the input / output interface 806 may be configured to provide an interface or interfaces to an input device, output device, or one or more input and / or output devices. Examples of an output device include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. An input device may allow a user to capture information into the UE 800. Examples of an input device include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, and the like. The presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. A sensor may be, for instance, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as an input device. For example, a Universal Serial Bus (USB) port may be used to provide an input device and an output device.
[0089] In some embodiments, the power source 808 is structured as a battery or battery pack. Other types of power sources, such as an external power source (e.g., an electricity outlet), photovoltaic device, or power cell, may be used. The power source 808 may further include power circuitry for delivering power from the power source 808 itself, and / or an external power source, to the various parts of the UE 800 via input circuitry or an interface such as an electrical power cable. Delivering power may be, for example, for charging of the power source 808. Power circuitry may perform any formatting, converting, or other modification to the power from the power source 808 to make the power suitable for the respective components of the UE 800 to which power is supplied.
[0090] The memory 810 may be or be configured to include memory such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, and so forth. In one example, the memory 810 includes one or more application programs 814, such as an operating system, web browser application, a widget, gadget engine, or other application, and corresponding data 816. The memory 810 may store, for use by the UE 800, any of a variety of various operating systems or combinations of operating systems.
[0091] The memory 810 may be configured to include a number of physical drive units, such as redundant array of independent disks (RAID), flash memory, USB flash drive, external hard disk drive, thumb drive, pen drive, key drive, high-density digital versatile disc (HD-DVD) optical disc drive, internal hard disk drive, Blu-Ray optical disc drive, holographic digital data storage (HDDS) optical disc drive, external mini-dual in-line memory module (DIMM), synchronous dynamic random access memory (SDRAM), external micro-DIMM SDRAM, smartcard memory such as tamper resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs), such as a USIM and / or ISIM, other memory, or any combination thereof. The UICC may for example be an embedded UICC (eUICC), integrated UICC (iUICC) or a removable UICC commonly known as ‘SIM card.’ The memory 810 may allow the UE 800 to access instructions, application programs and the like, stored on transitory or non-transitory memory media, to off-load data, or to upload data. An article of manufacture, such as one utilizing a communication system may be tangibly embodied as or in the memory 810, which may be or comprise a device-readable storage medium.
[0092] The processing circuitry 802 may be configured to communicate with an access network or other network using the communication interface 812. The communication interface 812 may comprise one or more communication subsystems and may include or be communicatively coupled to an antenna 822. The communication interface 812 may include one or more transceivers used to communicate, such as by communicating with one or more remote transceivers of another device capable of wireless communication (e.g., another UE or a network node in an access network). Each transceiver may include a transmitter 818 and / or a receiver 820 appropriate to provide network communications (e.g., optical, electrical, frequency allocations, and so forth). Moreover, the transmitter 818 and receiver 820 may be coupled to one or more antennas (e.g., antenna 822) and may share circuit components, software or firmware, or alternatively be implemented separately.
[0093] In the illustrated embodiment, communication functions of the communication interface 812 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof. Communications may be implemented in according to one or more communication protocols and / or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / internet protocol (TCP / IP), synchronous optical networking (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), and so forth.
[0094] Regardless of the type of sensor, a UE may provide an output of data captured by its sensors, through its communication interface 812, via a wireless connection to a network node. Data captured by sensors of a UE can be communicated through a wireless connection to a network node via another UE. The output may be periodic (e.g., once every 15 minutes if it reports the sensed temperature), random (e.g., to even out the load from reporting from several sensors), in response to a triggering event (e.g., when moisture is detected an alert is sent), in response to a request (e.g., a user initiated request), or a continuous stream (e.g., a live video feed of a patient). As another example, a UE comprises an actuator, a motor, or a switch, related to a communication interface configured to receive wireless input from a network node via a wireless connection. In response to the received wireless input the states of the actuator, the motor, or the switch may change. For example, the UE may comprise a motor that adjusts the control surfaces or rotors of a drone in flight according to the received input or to a robotic arm performing a medical procedure according to the received input.
[0095] A UE, when in the form of an Internet of Things (loT) device, may be a device for use in one or more application domains, these domains comprising, but not limited to, city wearable technology, extended industrial application and healthcare. Non-limiting examples of such an loT device are a device which is or which is embedded in: a connected refrigerator or freezer, a TV, a connected lighting device, an electricity meter, a robot vacuum cleaner, a voice controlled smart speaker, a home security camera, a motion detector, a thermostat, a smoke detector, a door / window sensor, a flood / moisture sensor, an electrical door lock, a connected doorbell, an air conditioning system like a heat pump, an autonomous vehicle, a surveillance system, a weather monitoring device, a vehicle parking monitoring device, an electric vehicle charging station, a smart watch, a fitness tracker, a head-mounted display for Augmented Reality (AR) or Virtual Reality (VR), a wearable for tactile augmentation or sensory enhancement, a water sprinkler, an animal- or itemtracking device, a sensor for monitoring a plant or animal, an industrial robot, an Unmanned Aerial Vehicle (UAV), and any kind of medical device, like a heart rate monitor or a remote controlled surgical robot. A UE in the form of an loT device comprises circuitry and / or software in dependence of the intended application of the loT device in addition to other components as described in relation to the UE 800 shown in FIGURE 8.
[0096] As yet another specific example, in an loT scenario, a UE may represent a machine or other device that performs monitoring and / or measurements, and transmits the results of such monitoring and / or measurements to another UE and / or a network node. The UE may in this case be an M2M device, which may in a 3 GPP context be referred to as an MTC device. As one particular example, the UE may implement the 3 GPP NB-IoT standard. In other scenarios, a UE may represent a vehicle, such as a car, a bus, a truck, a ship and an airplane, or other equipment that is capable of monitoring and / or reporting on its operational status or other functions associated with its operation. In practice, any number of UEs may be used together with respect to a single use case. For example, a first UE might be or be integrated in a drone and provide the drone’s speed information (obtained through a speed sensor) to a second UE that is a remote controller operating the drone. When the user makes changes from the remote controller, the first UE may adjust the throttle on the drone (e.g. by controlling an actuator) to increase or decrease the drone’s speed. The first and / or the second UE can also include more than one of the functionalities described above. For example, a UE might comprise the sensor and the actuator, and handle communication of data for both the speed sensor and the actuators.
[0097] FIGURE 9 shows a network node 900 in accordance with some embodiments. As used herein, network node refers to equipment capable, configured, arranged and / or operable to communicate directly or indirectly with a UE and / or with other network nodes or equipment, in a telecommunication network. Examples of network nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, Node Bs, evolved Node Bs (eNBs) and NR NodeBs (gNBs)).
[0098] Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and so, depending on the provided amount of coverage, may be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be a relay node or a relay donor node controlling a relay. A network node may also include one or more (or all) parts of a distributed radio base station such as centralized digital units and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS).
[0099] Other examples of network nodes include multiple transmission point (multi-TRP) 5G access nodes, multi-standard radio (MSR) equipment such as MSRBSs, network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi-cell / multicast coordination entities (MCEs), Operation and Maintenance (O&M) nodes, Operations Support System (OSS) nodes, Self-Organizing Network (SON) nodes, positioning nodes (e.g., Evolved Serving Mobile Location Centers (E-SMLCs)), and / or Minimization of Drive Tests (MDTs). The network node 900 includes a processing circuitry 902, a memory 904, a communication interface 906, and a power source 908. The network node 900 may be composed of multiple physically separate components (e.g., a NodeB component and a RNC component, or a BTS component and a BSC component, etc.), which may each have their own respective components. In certain scenarios in which the network node 900 comprises multiple separate components (e.g., BTS and BSC components), one or more of the separate components may be shared among several network nodes. For example, a single RNC may control multiple NodeBs. In such a scenario, each unique NodeB and RNC pair, may in some instances be considered a single separate network node. In some embodiments, the network node 900 may be configured to support multiple radio access technologies (RATs). In such embodiments, some components may be duplicated (e.g., separate memory 904 for different RATs) and some components may be reused (e.g., a same antenna 910 may be shared by different RATs). The network node 900 may also include multiple sets of the various illustrated components for different wireless technologies integrated into network node 900, for example GSM, WCDMA, LTE, NR, WiFi, Zigbee, Z-wave, LoRaWAN, Radio Frequency Identification (RFID) or Bluetooth wireless technologies. These wireless technologies may be integrated into the same or different chip or set of chips and other components within network node 900.
[0100] The processing circuitry 902 may comprise a combination of one or more of a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application-specific integrated circuit, field programmable gate array, or any other suitable computing device, resource, or combination of hardware, software and / or encoded logic operable to provide, either alone or in conjunction with other network node 900 components, such as the memory 904, to provide network node 900 functionality.
[0101] In some embodiments, the processing circuitry 902 includes a system on a chip (SOC). In some embodiments, the processing circuitry 902 includes one or more of radio frequency (RF) transceiver circuitry 912 and baseband processing circuitry 914. In some embodiments, the radio frequency (RF) transceiver circuitry 912 and the baseband processing circuitry 914 may be on separate chips (or sets of chips), boards, or units, such as radio units and digital units. In alternative embodiments, part or all of RF transceiver circuitry 912 and baseband processing circuitry 914 may be on the same chip or set of chips, boards, or units. The memory 904 may comprise any form of volatile or non-volatile computer-readable memory including, without limitation, persistent storage, solid-state memory, remotely mounted memory, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), mass storage media (for example, a hard disk), removable storage media (for example, a flash drive, a Compact Disk (CD) or a Digital Video Disk (DVD)), and / or any other volatile or non-volatile, non-transitory device-readable and / or computer-executable memory devices that store information, data, and / or instructions that may be used by the processing circuitry 902. The memory 904 may store any suitable instructions, data, or information, including a computer program, software, an application including one or more of logic, rules, code, tables, and / or other instructions capable of being executed by the processing circuitry 902 and utilized by the network node 900. The memory 904 may be used to store any calculations made by the processing circuitry 902 and / or any data received via the communication interface 906. In some embodiments, the processing circuitry 902 and memory 904 is integrated.
[0102] The communication interface 906 is used in wired or wireless communication of signaling and / or data between a network node, access network, and / or UE. As illustrated, the communication interface 906 comprises port(s) / terminal(s) 916 to send and receive data, for example to and from a network over a wired connection. The communication interface 906 also includes radio frontend circuitry 918 that may be coupled to, or in certain embodiments a part of, the antenna 910. Radio front-end circuitry 918 comprises filters 920 and amplifiers 922. The radio front-end circuitry 918 may be connected to an antenna 910 and processing circuitry 902. The radio frontend circuitry may be configured to condition signals communicated between antenna 910 and processing circuitry 902. The radio front-end circuitry 918 may receive digital data that is to be sent out to other network nodes or UEs via a wireless connection. The radio front-end circuitry 918 may convert the digital data into a radio signal having the appropriate channel and bandwidth parameters using a combination of filters 920 and / or amplifiers 922. The radio signal may then be transmitted via the antenna 910. Similarly, when receiving data, the antenna 910 may collect radio signals which are then converted into digital data by the radio front-end circuitry 918. The digital data may be passed to the processing circuitry 902. In other embodiments, the communication interface may comprise different components and / or different combinations of components.
[0103] In certain alternative embodiments, the network node 900 does not include separate radio front-end circuitry 918, instead, the processing circuitry 902 includes radio front-end circuitry and is connected to the antenna 910. Similarly, in some embodiments, all or some of the RF transceiver circuitry 912 is part of the communication interface 906. In still other embodiments, the communication interface 906 includes one or more ports or terminals 916, the radio front-end circuitry 918, and the RF transceiver circuitry 912, as part of a radio unit (not shown), and the communication interface 906 communicates with the baseband processing circuitry 914, which is part of a digital unit (not shown).
[0104] The antenna 910 may include one or more antennas, or antenna arrays, configured to send and / or receive wireless signals. The antenna 910 may be coupled to the radio front-end circuitry 918 and may be any type of antenna capable of transmitting and receiving data and / or signals wirelessly. In certain embodiments, the antenna 910 is separate from the network node 900 and connectable to the network node 900 through an interface or port.
[0105] The antenna 910, communication interface 906, and / or the processing circuitry 902 may be configured to perform any receiving operations and / or certain obtaining operations described herein as being performed by the network node. Any information, data and / or signals may be received from a UE, another network node and / or any other network equipment. Similarly, the antenna 910, the communication interface 906, and / or the processing circuitry 902 may be configured to perform any transmitting operations described herein as being performed by the network node. Any information, data and / or signals may be transmitted to a UE, another network node and / or any other network equipment.
[0106] The power source 908 provides power to the various components of network node 900 in a form suitable for the respective components (e.g., at a voltage and current level needed for each respective component). The power source 908 may further comprise, or be coupled to, power management circuitry to supply the components of the network node 900 with power for performing the functionality described herein. For example, the network node 900 may be connectable to an external power source (e.g., the power grid, an electricity outlet) via an input circuitry or interface such as an electrical cable, whereby the external power source supplies power to power circuitry of the power source 908. As a further example, the power source 908 may comprise a source of power in the form of a battery or battery pack which is connected to, or integrated in, power circuitry. The battery may provide backup power should the external power source fail. Embodiments of the network node 900 may include additional components beyond those shown in FIGURE 9 for providing certain aspects of the network node’s functionality, including any of the functionality described herein and / or any functionality necessary to support the subject matter described herein. For example, the network node 900 may include user interface equipment to allow input of information into the network node 900 and to allow output of information from the network node 900. This may allow a user to perform diagnostic, maintenance, repair, and other administrative functions for the network node 900.
[0107] FIGURE 10 is a block diagram of a host 1000, which may be an embodiment of the host 716 of FIGURE 7, in accordance with various aspects described herein. As used herein, the host 1000 may be or comprise various combinations hardware and / or software, including a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, container, or processing resources in a server farm. The host 1000 may provide one or more services to one or more UEs.
[0108] The host 1000 includes processing circuitry 1002 that is operatively coupled via a bus 1004 to an input / output interface 1006, a network interface 1008, a power source 1010, and a memory 1012. Other components may be included in other embodiments. Features of these components may be substantially similar to those described with respect to the devices of previous figures, such as Figures 8 and 9, such that the descriptions thereof are generally applicable to the corresponding components of host 1000.
[0109] The memory 1012 may include one or more computer programs including one or more host application programs 1014 and data 1016, which may include user data, e.g., data generated by a UE for the host 1000 or data generated by the host 1000 for a UE. Embodiments of the host 1000 may utilize only a subset or all of the components shown. The host application programs 1014 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Coding (WC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG, VP9) and audio codecs (e.g., FLAC, Advanced Audio Coding (AAC), MPEG, G.711), including transcoding for multiple different classes, types, or implementations of UEs (e.g., handsets, desktop computers, wearable display systems, heads-up display systems). The host application programs 1014 may also provide for user authentication and licensing checks and may periodically report health, routes, and content availability to a central node, such as a device in or on the edge of a core network. Accordingly, the host 1000 may select and / or indicate a different host for over-the-top services for a UE. The host application programs 1014 may support various protocols, such as the HTTP Live Streaming (HLS) protocol, Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Dynamic Adaptive Streaming over HTTP (MPEG-DASH), etc.
[0110] FIGURE 11 is a block diagram illustrating a virtualization environment 1100 in which functions implemented by some embodiments may be virtualized. In the present context, virtualizing means creating virtual versions of apparatuses or devices which may include virtualizing hardware platforms, storage devices and networking resources. As used herein, virtualization can be applied to any device described herein, or components thereof, and relates to an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1100 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node may be entirely virtualized.
[0111] Applications 1102 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in the virtualization environment Q400 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.
[0112] Hardware 1104 includes processing circuitry, memory that stores software and / or instructions executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 1106 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 1108a and 1108b (one or more of which may be generally referred to as VMs 1108), and / or perform any of the functions, features and / or benefits described in relation with some embodiments described herein. The virtualization layer 1106 may present a virtual operating platform that appears like networking hardware to the VMs 1108.
[0113] The VMs 1108 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer 1106. Different embodiments of the instance of a virtual appliance 1102 may be implemented on one or more of VMs 1108, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment.
[0114] In the context of NFV, a VM 1108 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, non-virtualized machine. Each of the VMs 1108, and that part of hardware 1104 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs 1108 on top of the hardware 1104 and corresponds to the application 1102.
[0115] Hardware 1104 may be implemented in a standalone network node with generic or specific components. Hardware 1104 may implement some functions via virtualization. Alternatively, hardware 1104 may be part of a larger cluster of hardware (e.g. such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration 1110, which, among others, oversees lifecycle management of applications 1102. In some embodiments, hardware 1104 is coupled to one or more radio units that each include one or more transmitters and one or more receivers that may be coupled to one or more antennas. Radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with the virtual components to provide a virtual node with radio capabilities, such as a radio access node or a base station. In some embodiments, some signaling can be provided with the use of a control system 1112 which may alternatively be used for communication between hardware nodes and radio units.
[0116] FIGURE 12 shows a communication diagram of a host 1202 communicating via a network node 1204 with a UE 1206 over a partially wireless connection in accordance with some embodiments. Example implementations, in accordance with various embodiments, of the UE (such as a UE 712a of FIGURE 7 and / or UE 800 of FIGURE 8), network node (such as network node 710a of FIGURE 7 and / or network node 900 of FIGURE 9), and host (such as host 716 of FIGURE 7 and / or host 1000 of FIGURE 10) discussed in the preceding paragraphs will now be described with reference to FIGURE 12.
[0117] Like host 1000, embodiments of host 1202 include hardware, such as a communication interface, processing circuitry, and memory. The host 1202 also includes software, which is stored in or accessible by the host 1202 and executable by the processing circuitry. The software includes a host application that may be operable to provide a service to a remote user, such as the UE 1206 connecting via an over-the-top (OTT) connection 1250 extending between the UE 1206 and host 1202. In providing the service to the remote user, a host application may provide user data which is transmitted using the OTT connection 1250.
[0118] The network node 1204 includes hardware enabling it to communicate with the host 1202 and UE 1206. The connection 1260 may be direct or pass through a core network (like core network 706 of FIGURE 7) and / or one or more other intermediate networks, such as one or more public, private, or hosted networks. For example, an intermediate network may be a backbone network or the Internet.
[0119] The UE 1206 includes hardware and software, which is stored in or accessible by UE 1206 and executable by the UE’s processing circuitry. The software includes a client application, such as a web browser or operator-specific “app” that may be operable to provide a service to a human or non-human user via UE 1206 with the support of the host 1202. In the host 1202, an executing host application may communicate with the executing client application via the OTT connection 1250 terminating at the UE 1206 and host 1202. In providing the service to the user, the UE's client application may receive request data from the host's host application and provide user data in response to the request data. The OTT connection 1250 may transfer both the request data and the user data. The UE's client application may interact with the user to generate the user data that it provides to the host application through the OTT connection 1250.
[0120] The OTT connection 1250 may extend via a connection 1260 between the host 1202 and the network node 1204 and via a wireless connection 1270 between the network node 1204 and the UE 1206 to provide the connection between the host 1202 and the UE 1206. The connection 1260 and wireless connection 1270, over which the OTT connection 1250 may be provided, have been drawn abstractly to illustrate the communication between the host 1202 and the UE 1206 via the network node 1204, without explicit reference to any intermediary devices and the precise routing of messages via these devices. As an example of transmitting data via the OTT connection 1250, in step 1208, the host 1202 provides user data, which may be performed by executing a host application. In some embodiments, the user data is associated with a particular human user interacting with the UE 1206. In other embodiments, the user data is associated with a UE 1206 that shares data with the host 1202 without explicit human interaction. In step 1210, the host 1202 initiates a transmission carrying the user data towards the UE 1206. The host 1202 may initiate the transmission responsive to a request transmitted by the UE 1206. The request may be caused by human interaction with the UE 1206 or by operation of the client application executing on the UE 1206. The transmission may pass via the network node 1204, in accordance with the teachings of the embodiments described throughout this disclosure. Accordingly, in step 1212, the network node 1204 transmits to the UE 1206 the user data that was carried in the transmission that the host 1202 initiated, in accordance with the teachings of the embodiments described throughout this disclosure. In step 1214, the UE 1206 receives the user data carried in the transmission, which may be performed by a client application executed on the UE 1206 associated with the host application executed by the host 1202.
[0121] In some examples, the UE 1206 executes a client application which provides user data to the host 1202. The user data may be provided in reaction or response to the data received from the host 1202. Accordingly, in step 1216, theUE 1206 may provide user data, which may be performed by executing the client application. In providing the user data, the client application may further consider user input received from the user via an input / output interface of the UE 1206. Regardless of the specific manner in which the user data was provided, the UE 1206 initiates, in step 1218, transmission of the user data towards the host 1202 via the network node 1204. In step 1220, in accordance with the teachings of the embodiments described throughout this disclosure, the network node 1204 receives user data from the UE 1206 and initiates transmission of the received user data towards the host 1202. In step 1222, the host 1202 receives the user data carried in the transmission initiated by the UE 1206.
[0122] One or more of the various embodiments improve the performance of OTT services provided to the UE 1206 using the OTT connection 1250, in which the wireless connection 1270 forms the last segment. More precisely, the teachings of these embodiments may improve one or more of, for example, data rate, latency, and / or power consumption and, thereby, provide benefits such as, for example, reduced user waiting time, relaxed restriction on file size, improved content resolution, better responsiveness, and / or extended battery lifetime.
[0123] In an example scenario, factory status information may be collected and analyzed by the host 1202. As another example, the host 1202 may process audio and video data which may have been retrieved from a UE for use in creating maps. As another example, the host 1202 may collect and analyze real-time data to assist in controlling vehicle congestion (e.g., controlling traffic lights). As another example, the host 1202 may store surveillance video uploaded by a UE. As another example, the host 1202 may store or control access to media content such as video, audio, VR or AR which it can broadcast, multicast or unicast to UEs. As other examples, the host 1202 may be used for energy pricing, remote control of non-time critical electrical load to balance power generation needs, location services, presentation services (such as compiling diagrams etc. from data collected from remote devices), or any other function of collecting, retrieving, storing, analyzing and / or transmitting data.
[0124] In some examples, a measurement procedure may be provided for the purpose of monitoring data rate, latency and other factors on which the one or more embodiments improve. There may further be an optional network functionality for reconfiguring the OTT connection 1250 between the host 1202 and UE 1206, in response to variations in the measurement results. The measurement procedure and / or the network functionality for reconfiguring the OTT connection may be implemented in software and hardware of the host 1202 and / or UE 1206. In some embodiments, sensors (not shown) may be deployed in or in association with other devices through which the OTT connection 1250 passes; the sensors may participate in the measurement procedure by supplying values of the monitored quantities exemplified above, or supplying values of other physical quantities from which software may compute or estimate the monitored quantities. The reconfiguring of the OTT connection 1250 may include message format, retransmission settings, preferred routing etc.; the reconfiguring need not directly alter the operation of the network node 1204. Such procedures and functionalities may be known and practiced in the art. In certain embodiments, measurements may involve proprietary UE signaling that facilitates measurements of throughput, propagation times, latency and the like, by the host 1202. The measurements may be implemented in that software causes messages to be transmitted, in particular empty or ‘dummy’ messages, using the OTT connection 1250 while monitoring propagation times, errors, etc. FIGURE 13 illustrates an example method 1300 of using a RL agent 100, 500 in a communication network to optimize one or more cell parameters in a cell of the communication network, according to certain embodiments. The method begins at step 1302 when the RL agent 100, 500 identifies a plurality of KPIs associated with a reward for performing at least one action to optimize one or more cell parameters. At step 1304, the RL agent 100, 500 uses a plurality of DNNs to estimate an associated gain for each of the plurality of KPIs. An associated DNN is used for each of the plurality of KPIs. Based on the associated gain predicted for each of the plurality of KPIs, the RL agent 100, 500 determines whether to implement the at least one action to optimize the one or more cell parameters, at step 1306.
[0125] In a particular embodiment, the RL agent 100, 500 determines, based on the associated gain for each of the plurality of KPIs, an estimated reward associated with implementing the at least one action at a state.
[0126] In a further particular embodiment, when determining the estimated reward, the RL agent 100, 500 uses a reward DNN to determine the estimated reward based on each of the estimated associated gains for the plurality of KPIs.
[0127] In a further particular embodiment, the step of determining whether to implement the at least one action to optimize the one or more cell parameters is based at least in part on the estimated reward determined using the reward DNN.
[0128] In a further particular embodiment, when determining whether to implement the at least one action to optimize the one or more cell parameters, the RL agent 100, 500 determines to implement the at least one action to optimize the one or more cell parameters if a value associated with the estimated reward is greater than a threshold. Alternatively, the RL agent 100, 500 determines not to implement the at least one action to optimize the one or more cell parameters if a value associated with the estimated reward is less than a threshold.
[0129] In a particular embodiment, the associated gain is estimated for each of the plurality of KPIs based on at least a first value of a respective KPI at a time before the implementation of the at least one action to optimize the one or more cell parameters and a second value of a respective KPI at a time after the implementation of the at least one action to optimize the one or more cell parameters.
[0130] In a particular embodiment, the plurality of KPIs comprise at least one of: GT; an average GTN; AR; and an average ARN. In a particular embodiment, estimating the associated gain for each of the plurality of KPIs comprises computing a Good Traffic gain for the cell, gainor, as a relative difference between GTt+i and GTt, wherein GTt, is a first value associated with a GT measure before implementation of the at least one action to optimize the one or more cell parameters and GTt+i is a second value associated with a GT measure after implementation of the at least one action to optimize the one or more cell parameters.
[0131] In a particular embodiment, estimating the associated gain for each of the plurality of KPIs comprises estimating an average Good Traffic gain for at least one neighboring cell, gainer , wherein the gainer is computed as a relative difference between GTNt+1and GTNt, and where GTNt, is a first value associated with a GTN measure before implementation of the at least one action to optimize the one or more cell parameters and GTNt+i is a second value associated with a GTN measure after implementation of the at least one action to optimize the one or more cell parameters.
[0132] In a particular embodiment, estimating the associated gain for each of the plurality of KPIs comprises estimating a Accessibility Rate gain for the cell, gainAR,, wherein the gainAR, is computed as a relative difference of ARt+1and ARt, and wherein ARt, is a first value associated with a AR measure before implementation of the at least one action to optimize the one or more cell parameters and ARt+i is a second value associated with a AR measure after implementation of the at least one action to optimize the one or more cell parameters.
[0133] In a particular embodiment, estimating the associated gain for each of the plurality of KPIs comprises estimating an average Accessibility Rate gain for at least one neighboring cell, gainARN, and wherein the gainARN s computed as a relative difference of ARNt+1and ARNt, and wherein ARNt, is a first value associated with a ARN measure before implementation of the at least one action to optimize the one or more cell parameters and ARNt+i is a second value associated with a ARN measure after implementation of the at least one action to optimize the one or more cell parameters.
[0134] In a further particular embodiment, determining whether to implement the at least one action to optimize the one or more cell parameters comprises determining an estimated reward, Rt+1, associated with implementing the at least one action based on at least one of: the gainer, gainGTN, gainAR, and gainARN. In a particular embodiment, determining the estimated reward, RL+1, associated with implementing the at least one action to optimize the one or more cell parameters is based on:
[0135] In a particular embodiment, the method further includes at least one of: optimizing the one or more cell parameters; determining a value associated with an actual reward for implementing the at least one action to optimize the one or more cell parameters; comparing the value associated with the actual reward to a value associated with the estimated reward; and determining at least one error statistic based on the comparing.
[0136] In a particular embodiment, the method further includes using the at least one error statistic for training of an ML model.
[0137] In a particular embodiment, the method is performed on a network node 710, wireless device 712, core network node 708, or other node.
[0138] In a particular embodiment, the method is performed on a server associated with or in communication with a 5GNWDAF.
[0139] In a particular embodiment, the method is performed by an rApp in a Non Real-Time Radio Access Network Intelligent Controller in communication with an O-RAN based network.
[0140] In a particular embodiment, the one or more cell parameters relate to at least one downlink transmission to one or more wireless devices in the cell and / or at least one uplink transmission from one or more wireless devices in the cell.
[0141] In a particular embodiment, the one or more cell parameters comprise at least one of: an antenna tilt of at least one antenna associated for the cell, and a target power level expected for at least one uplink transmission and / or downlink transmission.
[0142] In a particular embodiment, the method further includes at least one of: receiving values associated with measurements performed by at least one radio node in the cell, wherein the values and / or measurements are associated with one or more of the plurality of KPIs; and receiving values associated with measurements performed by at least one radio node in at least one neighboring cell, wherein the values and / or measurements are associated with one or more of the plurality of KPIs. In a particular embodiment, the associated gain for each of the plurality of KPIs provide explainability and / or are intelligible and / or comprehensible to an end user.
[0143] Although the computing devices described herein (e.g., UEs, network nodes, hosts) may include the illustrated combination of hardware components, other embodiments may comprise computing devices with different combinations of components. It is to be understood that these computing devices may comprise any suitable combination of hardware and / or software needed to perform the tasks, features, functions and methods disclosed herein. Determining, calculating, obtaining or similar operations described herein may be performed by processing circuitry, which may process information by, for example, converting the obtained information into other information, comparing the obtained information or converted information to information stored in the network node, and / or performing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Moreover, while components are depicted as single boxes located within a larger box, or nested within multiple boxes, in practice, computing devices may comprise multiple different physical components that make up a single illustrated component, and functionality may be partitioned between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of the components may be partitioned between the processing circuitry and the communication interface. In another example, non-computationally intensive functions of any of such components may be implemented in software or firmware and computationally intensive functions may be implemented in hardware.
[0144] In certain embodiments, some or all of the functionality described herein may be provided by processing circuitry executing instructions stored on in memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuitry without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hard-wired manner. In any of those particular embodiments, whether executing instructions stored on a non-transitory computer-readable storage medium or not, the processing circuitry can be configured to perform the described functionality. The benefits provided by such functionality are not limited to the processing circuitry alone or to other components of the computing device, but are enjoyed by the computing device as a whole, and / or by end users and a wireless network generally.
Claims
CLAIMS1. A method (1300) of using a reinforcement learning, RL, agent (100, 500) in a communication network to optimize one or more cell parameters in a cell of the communication network, the method comprising: identifying (1302) a plurality of key performance indicators, KPIs, associated with a reward for performing at least one action to optimize one or more cell parameters; using (1304), by the RL agent, a plurality of Deep Neural Networks, DNNs, to estimate an associated gain for each of the plurality of KPIs, wherein an associated DNN is used for each of the plurality of KPIs; and based on the associated gain predicted for each of the plurality of KPIs, determining (1306) whether to implement the at least one action to optimize the one or more cell parameters.
2. The method of Claim 1, comprising determining, based on the associated gain for each of the plurality of KPIs, an estimated reward associated with implementing the at least one action at a state.
3. The method of Claim 2, wherein determining the estimated reward comprises using a reward DNN to determine the estimated reward based on each of the estimated associated gains for the plurality of KPIs.
4. The method of any one of Claims 2 to 3, wherein determining whether to implement the at least one action to optimize the one or more cell parameters is based at least in part on the estimated reward determined using the reward DNN.
5. The method of any one of Claims 2 to 4, wherein determining whether to implement the at least one action to optimize the one or more cell parameters comprises: determining to implement the at least one action to optimize the one or more cell parameters if a value associated with the estimated reward is greater than a threshold; or determining not to implement the at least one action to optimize the one or more cell parameters if a value associated with the estimated reward is less than a threshold.
6. The method of any one of Claims 1 to 5, wherein the associated gain is estimated for each of the plurality of KPIs based on at least a first value of a respective KPI at a time before the implementation of the at least one action to optimize the one or more cell parameters and a secondvalue of a respective KPI at a time after the implementation of the at least one action to optimize the one or more cell parameters.
7. The method of any one of Claims 1 to 6, wherein the plurality of KPIs comprise at least one of:Good traffic in the cell, GT; an average GT in a plurality of neighboring cells, GTN;Accessibility Rate in the cell, AR; and an average AR in the plurality of neighboring cells, ARN.
8. The method of any one of Claims 1 to 7, wherein estimating the associated gain for each of the plurality of KPIs comprises computing a Good Traffic gain for the cell, gainor, as a relative difference between GTt+i and GTt, wherein GTt, is a first value associated with a GT measure before implementation of the at least one action to optimize the one or more cell parameters and GTt+i is a second value associated with a GT measure after implementation of the at least one action to optimize the one or more cell parameters.
9. The method of any one of Claims 1 to 8, wherein estimating the associated gain for each of the plurality of KPIs comprises estimating an average Good Traffic gain for at least one neighboring cell, gainGTN, wherein the gainGTN is computed as a relative difference between GTNt+1and GTNt, and where GTNt, is a first value associated with a GTN measure before implementation of the at least one action to optimize the one or more cell parameters and GTNt+i is a second value associated with a GTN measure after implementation of the at least one action to optimize the one or more cell parameters.
10. The method of any one of Claims 1 to 9, wherein estimating the associated gain for each of the plurality of KPIs comprises estimating a Accessibility Rate gain for the cell, gainAR, , wherein the gainAR, is computed as a relative difference of ARt+1and ARt, and wherein ARt, is a first value associated with a AR measure before implementation of the at least one action to optimize the one or more cell parameters and ARt+i is a second value associated with a AR measure after implementation of the at least one action to optimize the one or more cell parameters.
11. The method of any one of Claims 1 to 10, wherein estimating the associated gain for each of the plurality of KPIs comprises estimating an average Accessibility Rate gain for at least oneneighboring cell, gainARN, and wherein the gainARN is computed as a relative difference of ARNt+1and ARNt, and wherein ARNt, is a first value associated with a ARN measure before implementation of the at least one action to optimize the one or more cell parameters and ARNt+i is a second value associated with a ARN measure after implementation of the at least one action to optimize the one or more cell parameters.
12. The method of any one of Claims 8 to 11, wherein determining whether to implement the at least one action to optimize the one or more cell parameters comprises determining an estimated reward, Rt+1, associated with implementing the at least one action based on at least one of: the gainGT, gairiGTN, gairiAR, and gainARN.
13. The method of Claim 12, wherein determining the estimated reward, Rt+, associated with implementing the at least one action to optimize the one or more cell parameters is based on:
14. The method of any one of Claims 1 to 13, comprising at least one of: optimizing the one or more cell parameters; determining a value associated with an actual reward for implementing the at least one action to optimize the one or more cell parameters; comparing the value associated with the actual reward to a value associated with the estimated reward; determining at least one error statistic based on the comparing.
15. The method of Claim 14, comprising using the at least one error statistic for training of a Machine Learning, ML, model.
16. The method of any one of Claims 1 to 15, wherein the method is performed on a network node (710), wireless device (712), core network node (708), or other node.
17. The method of any one of Claims 1 to 16, wherein the method is performed on a server associated with or in communication with a 5G Network Data Analytics Function, NWDAF.
18. The method of any one of Claims 1 to 16, wherein the method is performed by an rApp in a Non Real-Time Radio Access Network Intelligent Controller in communication with an O-RAN based network.
19. The method of any one of Claims 1 to 18, wherein the one or more cell parameters relate to at least one downlink transmission to one or more wireless devices in the cell and / or at least one uplink transmission from one or more wireless devices in the cell.
20. The method of any one of Claims 1 to 19, wherein the one or more cell parameters comprise at least one of: an antenna tilt of at least one antenna associated for the cell, and a target power level expected for at least one uplink transmission and / or downlink transmission.
21. The method of any one of Claims 1 to 20, comprising at least one of: receiving values associated with measurements performed by at least one radio node in the cell, wherein the values and / or measurements are associated with one or more of the plurality of KPIs; and receiving values associated with measurements performed by at least one radio node in at least one neighboring cell, wherein the values and / or measurements are associated with one or more of the plurality of KPIs.
22. The method of any one of Claims 1 to 21, wherein the associated gain for each of the plurality of KPIs provide explainability and / or are intelligible and / or comprehensible to an end user.
23. A computer program product comprising a computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a computer or processor, the computer or processor is caused to: identify a plurality of key performance indicators, KPIs, associated with a reward for performing at least one action to optimize one or more cell parameters; use, by the RL agent, a plurality of Deep Neural Networks, DNNs, to estimate an associated gain for each of the plurality of KPIs, wherein an associated DNN is used for each of the plurality of KPIs; and based on the associated gain predicted for each of the plurality of KPIs, determine whether to implement the at least one action to optimize the one or more cell parameters.
24. The computer program product of Claim 23, wherein the computer or processor is caused to perform the method of any of Claims 2 to 22.
25. An apparatus (600, 708, 710, 712) in a communication network, the apparatus comprising a reinforcement learning, RL, agent (100, 500) for optimizing one or more cell parameters in a cell of the communication network, the apparatus configured to: identify a plurality of key performance indicators, KPIs, associated with a reward for performing at least one action to optimize one or more cell parameters; use, by the RL agent, a plurality of Deep Neural Networks, DNNs, to estimate an associated gain for each of the plurality of KPIs, wherein an associated DNN is used for each of the plurality of KPIs; and based on the associated gain predicted for each of the plurality of KPIs, determine whether to implement the at least one action to optimize the one or more cell parameters.
26. The apparatus of Claim 25, wherein the apparatus is configured to perform the method of any of Claims 2 to 22.
Citation Information
Patent Citations
Policy for optimising cell parameters
WO2021190772A1
System and method for deep learning and wireless network optimization using deep learning
EP3635505B1
Improving operation of a communication network
WO2021244765A1