Event-based reinforcement learning for RRM parameter optimization

A reinforcement learning model optimizes RRM parameters using event-based strategies to overcome latency challenges in 5G NR, enhancing adaptability and responsiveness in dynamic radio environments.

WO2025248476A1PCT designated stage Publication Date: 2025-12-04NOKIA TECHNOLOGIES OY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055537
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2025-05-28
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

The latency of radio resource control (RRC) configuration in 5G NR limits the tuning of layer-3 handover parameters, such as cell individual offset (CIO) and time-to-trigger (TTT), which is also a challenge in uplink power control and antenna panel selection, particularly in dynamic and high-speed mobility scenarios.

Method used

Implementing a reinforcement learning (RL) model to dynamically optimize radio resource management (RRM) parameters based on radio measurement metrics, using event triggering and exit conditions to enhance adaptability and responsiveness.

Benefits of technology

The RL model enables agile and precise tuning of RRM parameters, addressing latency issues and ensuring quality of service requirements in dynamic radio environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055537_04122025_PF_FP_ABST
    Figure IB2025055537_04122025_PF_FP_ABST
Patent Text Reader

Abstract

According to an aspect, there is provided an apparatus for performing the following. The apparatus transmits, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model. The apparatus receives, from the network entity, at least one configuration message comprising the one or more RL strategies which comprise one or more RL event conditions for entering and / or exiting exploration and / or exploitation events. The apparatus performs exploration and / or exploitation using the RL model. The apparatus evaluates the one or more RL event conditions and transmits, to the network entity, an evaluation report comprising results of the evaluating. The apparatus receives, from the network entity, a positive or negative acknowledgement. Based on the reception of the positive acknowledgment and the results, the apparatus performs triggering an exploration or exploitation event and / or exiting an exploration or the exploitation event.
Need to check novelty before this filing date? Find Prior Art

Description

EVENT-BASED REINFORCEMENT LEARNING FOR RRM PARAMETEROPTIMIZATIONTECHNICAL FIELD

[0001] Various example embodiments relate to wireless communications.BACKGROUND

[0002] In Fifth Generation New Radio (5G NR), layer-3 (L3) handover (HO) parameters such as the cell individual offset (CIO) and time-to-trigger (TTT) are determined conventionally per terminal device via radio resource control (RRC) (re)configuration for measurement events verification. Specifically, the L3 HO parameters may be configured in a static manner, e.g., in “MeasObjectNR” and “ReportConfigNR” information elements transmitted from the access node to the terminal device. However, the latency of RRC configuration may place considerable limits to the tuning of these parameters. Similar challenges relating to the latency of the RRC configuration may be identified also in the context of other use cases such as in uplink power control and antenna panel selection.SUMMARY

[0003] According to an aspect, there is provided the subject matter of the independent claims. Embodiments are defined in the dependent claims.

[0004] According to a first aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: maintaining, in the at least one memory, a reinforcement learning, RL, model trained to determine one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; receiving, from a network entity, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model; and carrying out optimization of the one or more RRM parameters using the RL model according to the one or more RL strategies.

[0005] According to a second aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting, to a terminal device storing an RL model trained to evaluate one or more RRM parameters of the terminal device based on one or more radio measurement metrics, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model.

[0006] According to a third aspect, there is provided a method comprising: maintaining, in at least one memory, a reinforcement learning, RL, model trained to determine one or more radio resource management, RRM, parameters based on one or more radio measurement metrics; receiving, from a network entity, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model; and carrying out optimization of the one or more RRM parameters using the RL model according to the one or more RL strategies.

[0007] According to a fourth aspect, there is provided a method comprising: transmitting, to a terminal device storing an RL model trained to evaluate one or more RRM parameters of the terminal device based on one or more radio measurement metrics, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model.

[0008] According to a fifth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: receiving, from a network entity, at least one configuration message comprising one or more reinforcement learning, RL, strategies for an RL model trained to determine one or more radio resource management, RRM, parameters based on one or more radio measurement metrics, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model; and carrying out optimization of the one or more RRM parameters using the RL model according to the one or more RL strategies.

[0009] According to a sixth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: transmitting, to a terminal device storing an RL model trained to evaluate one or more RRM parameters of the terminal device based on one or more radio measurement metrics, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model.

[0010] According to a seventh aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model;receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating; receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0011] According to an eighth aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; receiving, from the terminal device, an evaluation report comprising results of evaluating of the one or more RL event conditions; determining whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event; and based on the determining, transmitting, to the terminal device, a positive or negative acknowledgement for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event.

[0012] According to a ninth aspect, there is provided a method comprising: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating;receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0013] According to a tenth aspect, there is provided a method comprising: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; receiving, from the terminal device, an evaluation report comprising results of evaluating of the one or more RL event conditions; determining whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event; and based on the determining, transmitting, to the terminal device, a positive or negative acknowledgement for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event.

[0014] According to an eleventh aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following:transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating; receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0015] According to a twelfth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; receiving, from the terminal device, an evaluation report comprising results of evaluating of the one or more RL event conditions; determining whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event; and based on the determining, transmitting, to the terminal device, a positive or negative acknowledgement for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event.

[0016] According to a thirteenth aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event;performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; and based on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0017] According to a fourteenth aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; and transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event.

[0018] According to a fifteenth aspect, there is provided a method comprising: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters based on one or more radio measurement metrics; evaluating the one or more RL event conditions; and based on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0019] According to a sixteenth aspect, there is provided a method comprising: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; and transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event.

[0020] According to seventeenth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model;receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; and based on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0021] According to an eighteenth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; and transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event.

[0022] One or more examples of implementations are set forth in more detail in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG. 1 illustrates a system to which some embodiments may be applied;

[0024] FIG. 2 illustrates a basic operating principle of reinforcement learning;

[0025] FIGs. 3 to 4 illustrate processes according to some embodiments;

[0026] FIGs. 5 & 6 illustrate examples of evaluating exploration triggering and exit conditions and exploitation triggering and exit conditions, respectively;

[0027] FIGs. 7 to 9 illustrate processes according to some embodiments;

[0028] FIGs. 10 and 11 illustrate signaling between a terminal device and an access node according to some embodiments; and

[0029] FIG. 12 illustrates an apparatus according to some embodiments.DETAILED DESCRIPTION OF SOME EMBODIMENTS

[0030] The following embodiments are only presented as examples. Although the specification may refer to “an”, “one”, or “some” embodiment(s) and / or example(s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment(s) or example(s), or that a particular feature only applies to a single embodiment and / or example. Single features of different embodiments and / or examples may also be combined to provide other embodiments and / or examples.

[0031] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.

[0032] In the following, different exemplifying embodiments will be described using, as an example of an access architecture to which the embodiments may be applied, a radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR, 5G), without restricting the embodiments to such an architecture, however. It is obvious for a person skilled in the art that the embodiments may also be applied to other kinds of communications networks having suitable means by adjusting parameters and procedures appropriately. Some examples of other options for suitable systems are the universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), wireless local area network (WLAN or WiFi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad- hoc networks (MANETs), Internet Protocol multimedia subsystems (IMS), rebel SIM (R-SIM) for code division multiple access (CDMA) technologies such as lx and lx evolution data optimized (IxEV-DO), global system for mobile communications (GSM), open radio access network (O-RAN) or any combination thereof.

[0033] FIG. 1 depicts examples of simplified system architectures only showing some elements and functional entities, all being logical units, whose implementation may differ from what is shown. The connections shown in FIG. 1 are logical connections; the actual physical connections may be different. It is apparent to a person skilled in the art that the system typically comprises also other functions and structures than those shown in FIG. 1.

[0034] The embodiments are not, however, restricted to the system given as an example but a person skilled in the art may apply the solution to other communication systems provided with necessary properties.

[0035] The example of FIG. 1 shows a part of an exemplifying radio access network.

[0036] A communications system typically comprises more than one (e / g)NodeB 104 in which case the (e / g)NodeBs may also be configured to communicate with one another over links, wired or wireless, designed for the purpose. These links may be used for signaling purposes. The (e / g)NodeB is a computing device configured to control the radio resources of communication system it is coupled to. The NodeB may also be referred to as a base station, an access point or any other type of interfacing device including a relay station capable of operating in a wireless environment. The (e / g)NodeB includes or is coupled to transceivers.From the transceivers of the (e / g)NodeB, a connection is provided to an antenna unit that establishes bi-directional radio links to user devices. The antenna unit may comprise a plurality of antennas or antenna elements. The (e / g)NodeB is further connected to core network 110 (CN or next generation core NGC). Depending on the system, the counterpart on the CN side can be a serving gateway (S-GW, routing and forwarding user data packets), packet data network gateway (P-GW), for providing connectivity of user devices (UEs) to external packet data networks, or mobile management entity (MME), etc.

[0037] The user device 100, 102 (also called UE, user equipment, user terminal, terminal device, etc.) illustrates one type of an apparatus to which resources on the air interface are allocated and assigned, and thus any feature described herein with a user device may be implemented with a corresponding apparatus, such as a relay node. An example of such a relay node is a layer 3 relay (self-backhauling relay) towards the base station. The user equipment may comprise a mobile equipment and at least one universal integrated circuit card (UICC).

[0038] The user device 100, 102 typically refers to a portable computing device that includes wireless mobile communication devices operating with or without a subscriber identity (or identification) module (SIM) or UICC, including, but not limited to, the following types of devices: a mobile station (mobile phone), smartphone, personal digital assistant (PDA), handset, device using a wireless modem (alarm or measurement device, etc.), laptop and / or touch screen computer, tablet, game console, notebook, and multimedia device. Here, the SIM may be a physical SIM which may be removable by a user or an embedded SIM (eSIM) embedded directly into the user device 100, 102 (and thus not being removable by a user). It should be appreciated that a user device may also be a nearly exclusive uplink only device, of which an example is a camera or video camera loading images or video clips to a network. A user device may also be a device having capability to operate in Internet of Things (loT) network which is a scenario in which objects are provided with the ability to transfer data over a network without requiring human-to-human or human-to-computer interaction. Thus, the user devices may not enable direct user interaction or may enable only limited user interaction (e.g., during setup). The user device (or in some embodiments a layer 3 relay node) is configured to perform one or more of user equipment functionalities. The user device may also be called a terminal device, a subscriber unit, mobile station, remote terminal, access terminal, user terminal or user equipment (UE) just to mention but a few names or apparatuses. Each user device 100, 102 may comprise one or more antennas.

[0039] Various techniques described herein may also be applied to a cyber-physical system (CPS) (a system of collaborating computational elements controlling physical entities). CPS may enable the implementation and exploitation of massive amounts of interconnected ICT devices (sensors, actuators, processors microcontrollers, etc.) embedded in physical objects at different locations. Mobile cyber physical systems, in which the physical system in question has inherent mobility, are a subcategory of cyber-physical systems. Examples of mobile physical systems include mobile robotics and electronics transported by humans or animals.

[0040] Additionally, although the apparatuses have been depicted as single entities, different units, processors and / or memory units (not all shown in FIG. 1) may be implemented.

[0041] 5G enables using MIMO antennas, many more base stations or nodes than the LTE (a so-called small cell concept), including macro sites operating in co-operation with smaller stations and employing a variety of radio technologies depending on service needs, use cases and / or spectrum available. 5G mobile communications supports a wide range of use cases and related applications including video streaming, augmented reality, different ways of data sharing and various forms of machine type applications, including vehicular safety, different sensors and real-time control. 5G is expected to have multiple radio interfaces, namely below 6GHz, cmWave and mmWave, and also being integrable with existing legacy radio access technologies, such as the LTE. Integration with the LTE may be implemented, at least in the early phase, as a system, where macro coverage is provided by the LTE and 5G radio interface access comes from small cells by aggregation to the LTE. In other words, 5G is planned to support both inter-RAT operability (such as LTE-5G) and inter-RI operability (inter-radio interface operability, such as below 6GHz - cmWave, below 6GHz - cmWave - mmWave). One of the concepts considered to be used in 5G networks is network slicing in which multiple independent and dedicated virtual sub-networks (network instances) may be created within the same infrastructure to run services that have different requirements on latency, reliability, throughput and mobility.

[0042] The current architecture in LTE networks is fully distributed in the radio and fully centralized in the core network. The low latency applications and services in 5G require to bring the content close to the radio which leads to local break out and multi-access edge computing (MEC). 5G enables analytics and knowledge generation to occur at the source of the data. This approach requires leveraging resources that may not be continuously connectedto a network such as laptops, smartphones, tablets and sensors. MEC provides a distributed computing environment for application and service hosting. It also has the ability to store and process content in close proximity to cellular subscribers for faster response time. Edge computing covers a wide range of technologies such as wireless sensor networks, mobile data acquisition, mobile signature analysis, cooperative distributed peer-to-peer ad hoc networking and processing also classifiable as local cloud / fog computing and grid / mesh computing, dew computing, mobile edge computing, cloudlet, distributed data storage and retrieval, autonomic self-healing networks, remote cloud services, augmented and virtual reality, data caching, Internet of Things (massive connectivity and / or latency critical), critical communications (autonomous vehicles, traffic safety, real-time analytics, time-critical control, healthcare applications).

[0043] The communication system is also able to communicate with other networks, such as a public switched telephone network or the Internet 112, or utilize services provided by them. The communication network may also be able to support the usage of cloud services, for example at least part of core network operations may be carried out as a cloud service (this is depicted in FIG. 1 by “cloud” 114). The communication system may also comprise a central control entity, or a like, providing facilities for networks of different operators to cooperate for example in spectrum sharing.

[0044] Edge cloud may be brought into the RAN by utilizing network function virtualization (NVF) and software defined networking (SDN). Using edge cloud may mean access node operations to be carried out, at least partly, in a server, host or node operationally coupled to a remote radio head or unit (RU) 116, 118 or base station comprising radio parts. It is also possible that node operations will be distributed among a plurality of servers, nodes or hosts. Application of cloudRAN architecture enables RAN real time functions being carried out at the RAN side (in a distributed unit, DU 104) and non-real time functions being carried out in a centralized manner (in a central or centralized unit, CU 108). Thus, in summary, the RAN may comprise, in some embodiments, at least one distributed access node comprising a central unit 108, one or more distributed units 104 communicatively connected to the central unit 108 and one or more (remote) radio heads or units 116, 118, each of which is communicatively connected to at least one of the one or more distributed units 104.

[0045] In some embodiments, element 104 is a non-distributed access node. In such embodiments, RUs 116, 118 are omitted.

[0046] It should also be understood that the distribution of labor between core network operations and base station operations may differ from that of the LTE or even be non-existent. Some other technology advancements probably to be used are Big Data and all-IP, which may change the way networks are being constructed and managed. 5G (or new radio, NR) networks are being designed to support multiple hierarchies, where MEC servers can be placed between the core and the base station or nodeB (gNB). It should be appreciated that MEC can be applied in 4G networks as well.

[0047] 5G may also utilize satellite communication to enhance or complement the coverage of 5G service, for example by providing backhauling. Possible use cases are providing service continuity for machine-to-machine (M2M) or Internet of Things (loT) devices or for passengers on board of vehicles, or ensuring service availability for critical communications, and future rail-way / maritime / aeronautical communications. Satellite communication may utilize geostationary earth orbit (GEO) satellite systems, but also low earth orbit (LEO) satellite systems, in particular mega-constellations (systems in which hundreds of (nano)satellites are deployed). Each satellite 106 in the mega-constellation may cover several satellite-enabled network entities that create on-ground cells. The on-ground cells may be created through an on-ground relay node 104 or by a gNB located on-ground or in a satellite.

[0048] It is obvious for a person skilled in the art that the depicted system is only an example of a part of a radio access system and in practice, the system may comprise a plurality of (e / g)NodeBs, the user device may have an access to a plurality of radio cells and the system may comprise also other apparatuses, such as physical layer relay nodes or other network elements, etc. At least one of the (e / g)NodeBs or may be a Home(e / g)nodeB. Additionally, in a geographical area of a radio communication system a plurality of different kinds of radio cells as well as a plurality of radio cells may be provided. Radio cells may be macro cells (or umbrella cells) which are large cells, usually having a diameter of up to tens of kilometers, or smaller cells such as micro-, femto- or picocells. The (e / g)NodeBs of FIG. 1 may provide any kind of these cells. A cellular radio system may be implemented as a multilayer network including several kinds of cells. Typically, in multilayer networks, one access node provides one kind of a cell or cells, and thus a plurality of (e / g)NodeBs are required to provide such a network structure.

[0049] For fulfilling the need for improving the deployment and performance of communication systems, the concept of “plug-and-play” (e / g)NodeBs has been introduced.Typically, a network which is able to use “plug-and-play” (e / g)NodeBs, includes, in addition to Home (e / g)NodeBs (H(e / g)nodeBs), a home node B gateway, or HNB-GW (not shown in FIG. 1). A HNB Gateway (HNB-GW), which is typically installed within an operator’s network may aggregate traffic from a large number of HNBs back to a core network.

[0050] 6G architecture is targeted to enable easy integration of everything, such as a network of networks, joint communication and sensing, non-terrestrial networks and terrestrial communication. 6G systems are envisioned to encompass machine learning algorithms as well as local and distributed computing capabilities, where virtualized network functions can be distributed over core and edge computing resources. Far edge computing, where computing resources are pushed to the very edge of the network, will be part of the distributed computing environment, for example in “zero-delay” scenarios. Some 5G systems may also employ such capabilities. More generally, the actual (radio) communication system is envisaged to be comprised of one or more computer programs executed within a programmable infrastructure, such as general-purpose computing entities (servers, processors, and like).

[0051] 6G networks are expected to adopt flexible decentralized and / or distributed computing systems and architecture and ubiquitous computing, with local spectrum licensing, spectrum sharing, infrastructure sharing, and intelligent automated management underpinned by mobile edge computing, artificial intelligence, short-packet communication, distributed ledgers and blockchain technologies. Key features of 6G will include intelligent connected management and control functions, programmability, integrated sensing and communication, reduction of energy footprint, trustworthy infrastructure, scalability and affordability. In addition to these, 6G is also targeting new use cases covering the integration of localization and sensing capabilities into system definition to unifying user experience across physical and digital worlds.

[0052] At least some of the terminal devices 100, 102 may be maintaining or hosting reinforcement learning models according to embodiments. To provide context for the more detailed discussion on embodiments relating to use of reinforcement learning, reinforcement learning is discussed in the following in general in reference to FIG. 2 showing a typical functional block 200 of reinforcement learning (equally called a reinforcement learning model 200).

[0053] Reinforcement learning (RL) is a field of machine learning where an (intelligent) agent 201 learns by trial and error a policy to maximize the given performance objective. Theagent 201 selects, at a given discrete time step t, an action Atbased on the current state Stof the environment 202. The action influences the environment 202 and leads to a new state St+1and reward Rt(goodness of the action in a given state). The reward may be calculated using a reward function. The objective of the agent 201 is to learn which actions with given state leads to highest cumulative future reward. In other words, the reinforcement learning model is based on a Markov decision process defining: a set of states S, a set of actions A of the agent 201 from each of the set of states S, probabilities of transitioning from state Stto all possible states St+1under respective actions Atand a set of (immediate) rewards R after transitioning from state Stto state St+1.

[0054] The purpose of the trial-and-error is to gain knowledge about what is working and what is not. This includes a concept of exploration vs. exploitation where exploration refers to occasional random actions, while exploitation means that the agent exploits the learned policy, i.e., the information received so far. In practice, RL algorithms must balance between the exploration and exploitation. Typically, before converging to a good policy, the RL algorithms start by exploring with a high probability. This probability of exploration typically decreases as the agent starts to learn. Agent is typically considered to be learning, when the received rewards increase over time and stabilize to a certain level.

[0055] A classical approach to any RL problem is to explore and to exploit, that is, to explore the most rewarding way that reaches the target and keep on exploiting a certain action. In general, exploration is hard. Without proper reward functions, the algorithms can end up chasing their own tails to eternity. Over the years, many exploration strategies have been formulated by incorporating mathematical approaches. Examples of the exploration strategies used in reinforcement learning models comprise:Epsilon-greedy: The agent does random exploration occasionally with probability e and takes the optimal action most of the time with probability 1 — 6.Upper Confidence Bound (UCB): The agent selects the greediest action to maximize the upper confidence bound Qt(u) + I / t(u), where Qt(a) is the average rewards associated with action a up to time t and I / t(u) is a function reversely proportional to how many times action a has been taken.Thompson Sampling: The agent keeps track of a belief over the probability of optimal actions and samples from this distribution.Boltzmann Exploration: The agent draws actions from a Boltzmann distribution (softmax) over the learned Q values, regulated by a temperature parameter T.

[0056] The aforementioned policies are inherently reliant on algorithms designed to strike a balance between exploration and exploitation. However, aligning these strategies with specific radio use cases may necessitate the incorporation of new criteria. These criteria should be intricately tailored to consider the individual impact of radio environments of terminal devices and access nodes and key performance indicators (KPIs) or quality of service (QoS) requirements such as mobility performance, user throughput, and latency. These additional constraints remain open for development, poised to be crafted as supplementary triggers or criteria. They are intended to augment the existing legacy RL exploration policies as listed above, dynamically regulating exploration policies in accordance with the specified parameters.

[0057] The embodiments to be discussed below seek to introduce RL strategies enabling meaningful balancing between the exploration and exploitation in various wireless communications use cases. Namely, the embodiments seek to uniquely harness measures such as reference signal received power (RSRP), speed and / or one or more mobility counters, not merely as inputs of the used RL model for radio resource management (RRM) parameter optimization, but as critical components in establishing a sophisticated event-based configuration for an RL agent of the RL model within a terminal device. More specifically, the embodiments provide an event-based management solution enabling definition of a novel configuration of various conditions (e.g., thresholds and / or timers) that dictate the entry and exit of exploration and exploitation events. These parameters are configured to align with the needs of RL-based radio resource management, providing unprecedented precision and control. The strategic use of these conditions (e.g., thresholds and / or timers) facilitates a level of policy adaptation and responsiveness that represents a significant advancement over the prior art.

[0058] One exemplary use case benefitting from the embodiments is terminal devicedriven RL-based HO parameter tuning. In 5G NR, layer-3 (L3) HO parameters such as the cell individual offset (CIO) and time-to-trigger (TTT) are determined conventionally per terminal device via radio resource control (RRC) (re)configuration for measurement events verification. Specifically, the layer-3 (L3) HO parameters may be configured in a static manner, e.g., in“MeasObjectNR” and “ReportConfigNR” information elements transmitted from the access node to the terminal device. If a machine-learning (ML) solution (e.g., an RL solution) were to be implemented for adapting HO parameters, such a solution would need to overcome the following problems.Problem 1: Re-configuring HO parameters is time consuming for some delay sensitive services such as Ultra Reliable Low Latency Communications (URLLC) and extended reality (XR). Therefore, frequent updating or tuning parameters may not fit the QoS requirement and ML algorithm design.Problem 2: On the other hand, designing layer- 1 (Ll) / layer-2 (L2) signaling to convey the parameters adjustment is possible to fit the delay sensitive service, however, the signaling cost might be unacceptably high.Problem 3: The ML algorithm and relevant implementation methods are up to terminal device or access node (e.g., gNB) vendor. However, the parameter search space is usually large to ensure the terminal device optimality. To have reduced dimension of search space is desired especially for RL exploration.

[0059] The challenges highlighted above are intricately tied to a specific use case, exemplified above by the intricate task of HO parameter tuning. However, these hurdles are not confined solely to this scenario; they permeate across diverse radio communication landscapes. In these contexts, the refinement and optimization of configuration parameters heavily rely on the often sluggish mechanisms of RRC procedures.

[0060] The application of RL presents a promising avenue for dynamically managing these parameters with enhanced adaptability. However, a consideration lies in the fact it is an operational speed or pace of an exploration process of the RL model is typically rather slow. In dynamic radio communication landscapes, particularly those entailing high-speed mobility scenarios or harnessing the frequencies within the 5G NR frequency range 2 (FR2), rapid and constant fluctuations of the radio environment necessitate an agile responsiveness. In the context of the RL, this demand often calls for frequent re-exploration - a requirement for the system to continuously adapt and fine-tune its strategies to cope effectively with the dynamic nature of these environments.

[0061] FIG. 3 illustrates a process for obtaining one or more RL strategies for operating an RL model and using said RL model according to said one or more RL strategies accordingto embodiments. The process of FIG. 3 may be carried out by an apparatus. The apparatus may be a terminal device or a part thereof. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. In the following, the entity carrying out the process of FIG. 3 is called simply an apparatus.

[0062] Referring to FIG. 3, the apparatus maintains, in block 301 , in at least one memory, a RL model. The RL model is assumed to have been (pre-)trained to determine one or more RRM parameters of the apparatus based on one or more radio measurement metrics.

[0063] A state, an action and a reward of the RL model may be defined as follows. The state may define one or more RRM parameters (i.e., values of one or more RRM parameters). The action taken by an agent (equally called an RL agent or an intelligent agent) of the RL model from a given state is defined as a modification of at least one of the one or more RRM parameters (i.e., a modification of at least one value of at least one of the one or more RRM parameters). The reward of taking a given action in a given state is calculated based on the one or more (radio) measurement metrics so as to optimize one or more KPIs of the apparatus or a serving cell. The reward may be calculated using a reward function dependent on the one or more radio measurement metrics.

[0064] The one or more measurement metrics may comprise any metrics or parameters whose values are derivable based on measurements (e.g., radio measurements and / or sensor measurements) carried out by the apparatus. The one or more measurement metrics may comprise one or more radio measurement metrics. The one or more radio measurement metrics may comprise at least one of RSRP, reference signal received quality (RSRQ), SINR ratio, received signal strength or channel (or signal) quality. More generally, the one or more measurement metrics may comprise, e.g., at least one of RSRP, RSRQ, signal-to-interference- plus-noise (SINR) ratio, received signal strength, channel quality, speed of the apparatus, one or more mobility stability related counters, a geolocation-based metric or an RSRP difference between the serving cell and the strongest neighbor cell.

[0065] How the one or more RRM parameters and the one or more KPIs are defined may depend on the specific use case. In the following, some examples are provided.

[0066] In some embodiments for handover parameter tuning, the one or more RRM parameters comprise one or more handover parameters (e.g., one or more layer-3 handover parameters) and / or the one or more KPIs comprise one or more mobility-specific cell-levelKPIs relating, e.g., to outage and / or failure. The one or more handover parameters may comprise, e.g., at least a cell individual offset (CIO) and / or a time-to-trigger (TTT). The one or more mobility-specific cell-level KPIs may comprise, e.g., at least one of a number of handovers, a number of handover failures, a number of ping-pong (PP) handovers, a number of unnecessary handovers, a number of radio link failures, a handover success rate, a handover failure rate, outage rate, or a ping-pong handover rate. In some embodiments, the one or more mobility-specific cell-level KPIs may comprise a cell throughput.

[0067] In some embodiments for uplink or sidelink power control, the one or more RRM parameters comprise one or more uplink power control parameters and / or the one or more KPIs comprise an uplink or sidelink throughput (of the apparatus). The one or more uplink power control parameters may comprise any parameters having an effect on the transmit power of the apparatus.

[0068] In some embodiments antenna panel or beam selection, the one or more RRM parameters comprise a parameter (e.g., a panel index or a beam index) defining a selection of at least one beam or at least one antenna panel of the apparatus and / or the one or more KPIs comprise link throughput (associated with the apparatus). The link throughput may correspond to downlink or uplink throughput (or a combination thereof).

[0069] Additionally or alternative, the one or more RRM parameters may comprise one or more time domain packet scheduling parameters, one or more frequency domain packet scheduling parameters, one or more spatial domain packet scheduling parameters and / or one or more link adaptation parameters. The one or more link adaptation parameters may comprise, e.g., at least one of a modulation and coding scheme (MCS) index, a channel quality indicator (QCI), a rank indicator (RI), a transport block size (TBS), a transmission time interval (TTI), one or more adaptive beamforming parameters, a precoding matric indicator (PMI) or one or more hybrid automatic repeat request (HARQ) parameters.

[0070] In some alternative embodiments, the RL model may be maintained in a remoter server (e.g., a cloud server) accessible by the apparatus, instead of being maintained locally in the at least one memory of the apparatus.

[0071] The apparatus receives, in block 302, from a network entity, at least one configuration message comprising one or more RL strategies. The network entity may be, e.g., an access node. The one or more RL strategies comprise one or more RL event conditionsrelating to one or more (different) RL events. Namely, the one or more RL strategies (or said one or more RL event conditions defined by the one or more RL strategies) comprise at least one (or both) of: one or more RL event triggering conditions for triggering an event of the RL model, or one or more RL event exit conditions for exiting an event of the RL model. In some embodiments, each of the one or more RL strategies (or said one or more RL event conditions defined by the one or more RL strategies) may comprise at least one of: one or more RL event triggering conditions for triggering an event of the RL model or one or more RL event exit conditions for exiting an event of the RL model

[0072] In some embodiments, the one or more RL strategies (or said one or more RL event conditions defined by the one or more RL strategies) may comprise at least one (or all) of:- one or more first RL event triggering conditions for triggering a first event of the RL model,- one or more first RL event exit conditions for exiting the first event of the RL model,- one or more second RL event triggering conditions for triggering a second event of the RL model, or- one or more second RL event exit conditions for exiting the second event of the RL model.In some such embodiments, the first and second events may be defined to be mutually exclusive. In some such cases, the one or more first RL event triggering conditions (or some of them) may be comprised also in the one or more second RL event exit conditions and / or the one or more second RL event triggering conditions (or some of them) may be comprised also in the one or more first RL event exit conditions. Thus, the first event may be triggered when the second event is exited and vice versa. In other embodiments, the triggering of any RL event (e.g., exploration or exploitation event) may automatically cause exiting of any (currently active) other RL events, irrespective of whether any associated RL event exit conditions have been satisfied.

[0073] In some embodiments, the first and second events mentioned above may be, respectively, an exploration event and an exploitation event. Here and in the following, the exploitation event may be specifically an exploration-excluding exploitation event (i.e., an exploitation event during which exploration is fully disabled). In other words, in someembodiments, the one or more RL strategies (or said one or more RL event conditions defined by the one or more RL strategies) may comprise at least one (or all) of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event.

[0074] In some embodiments, each of the one or more RL event conditions (e.g., each of the exploration and / or exploitation event triggering and / or exit conditions) may be defined based on at least one of:1) an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus (or other radio coupling gain based metric),2) an average speed of the apparatus,3) one or more mobility stability related counters of the apparatus,4) a geolocation-based metric indicative of location or distance of the apparatus relative to a serving network entity (e.g., a serving access node),5) time,6) a 5G NR RRM measurement relaxation condition,7) one or more capabilities of the apparatus, or8) a QoS-based metric of the apparatus.Namely, each RL event condition may be defined based on one of the above listed quantities or based on a combination of multiple of the above listed quantities (or other quantities). The combination may correspond, e.g., to a combination formed using one or more Boolean operators (e.g., any of OR, AND, XOR and NOT). As an example of the use of the AND operation, an exploration event triggering condition may define that the condition is satisfied if the average speed of the apparatus is smaller than a (pre-defined) speed threshold and the geolocation-based metric indicates that the apparatus is within a certain distance range from the serving network entity.

[0075] In general, an exploration event may be triggered, for example, in a situation where the radio environment has become unstable (e.g., when the apparatus is moving towards the cell edge, the variation of signal quality is high or the speed of the apparatus is high orincreasing fast). Under these circumstances, the agent of the RL model will be encouraged to perform exploration (so as to adapt to the changing radio environment). On the other hand, the exploitation event may be triggered when the apparatus is in a relatively stable radio environment so that exploration does not necessarily have to be carried out.

[0076] In the following, each of the options 1) to 7) listed above is discussed separately in more detail.

[0077] In option 1), the one or more RL event triggering conditions (or one or more RL event triggering conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus exceeding or falling below a first RSRP difference threshold. Additionally or alternatively, the one or more RL event exit conditions (or one or more RL event exit conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on an RSRP difference between the serving cell of the apparatus and the strongest neighbor cell of the apparatus exceeding or falling below a second RSRP difference threshold. Each of the first and second RSRP difference thresholds may be defined separately per RL event (e.g., separately for exploration and exploitation events). In general, a high RSRP difference may be indicative of a need for triggering an exploration event (and exiting an exploitation event) while a low RSRP difference may be indicative of a need for triggering an exploitation event (and exiting an exploration event.

[0078] More specifically, in some embodiments, the one or more exploration event triggering conditions may comprise at least: a condition triggered when an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus exceeds a first exploration RSRP difference threshold and / or the one or more exploration event exit conditions may comprise at least: a condition triggered when an RSRP difference between the serving cell of the apparatus and the strongest neighbor cell of the apparatus falls below a second exploration RSRP difference threshold (being, e.g., equal to or lower than the first exploration RSRP difference threshold). An example of this embodiment is discussed below in connection with FIG. 5. Additionally or alternatively, the one or more exploitation event triggering conditions may comprise at least: a condition triggered when an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus falls below a first exploitation RSRP difference threshold, and / or the one or more exploitation event exit conditions may comprise at least: a condition triggered when the RSRP difference betweenthe serving cell of the apparatus and the strongest neighbor cell of the apparatus exceeds a second exploitation RSRP difference threshold (being, e.g., equal to or higher than the first exploitation RSRP difference threshold). Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments. An example of this embodiment is discussed below in connection with FIG. 6.

[0079] In option 2), the one or more RL event triggering conditions (or one or more RL event triggering conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on an average speed of the apparatus exceeding or falling below a first speed threshold and / or the one or more RL event exit conditions (or one or more RL event exit conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on an average speed of the apparatus exceeding or falling below a second speed threshold (which may be the same or different compared to the first speed threshold). Each of the first and second speed thresholds may be defined separately per RL event (i.e., separately for exploration and exploitation events). The first and second speed thresholds may be defined, e.g., in km / h or m / s. In general, high speeds may be indicative of a need for triggering an exploration event (and exiting an exploitation event) while low speeds may be indicative of a need for triggering an exploitation event (and exiting an exploration event).

[0080] More specifically, in some embodiments, the one or more exploration event triggering conditions may comprise at least: a condition triggered when an average speed of the apparatus exceeds a first exploration speed threshold, and / or the one or more exploration event exit conditions may comprise at least: a condition triggered when the average speed of the apparatus falls below a second exploration speed threshold (being, e.g., equal to or lower than the first exploitation speed difference threshold). Additionally or alternatively, the one or more exploitation event triggering conditions may comprise at least: a condition triggered when an average speed of the apparatus falls below a first exploitation speed threshold, and / or the one or more exploitation event exit conditions may comprise at least: a condition triggered when the average speed of the apparatus exceeds a second exploitation speed threshold (being, e.g., equal to or lower than the first exploitation RSRP difference threshold). Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments.

[0081] In some alternative embodiments, some other speed-based metric (e.g., a median speed, a median acceleration or an average acceleration) may be employed instead of the average speed.

[0082] Referring to option 3), the one or more mobility stability related counters may comprise, e.g., at least one of a counter for a number of handovers, a counter for a number of handover failures, a counter for a number of ping-pong (PP) handovers, a counter for a number of unnecessary handovers, a counter for a number of radio link failures, a counter for a number of beam switches, a counter for a number of beam failure indicators, or a counter for a number of panel switches. An unnecessary handover may be defined as a handover which does not lead to an improvement (or significant improvement) in signal strength or quality or QoS. An unnecessary handover may be followed by a handover back to the original cell.

[0083] In option 3), the one or more RL event triggering conditions (or one or more RL event triggering conditions defined for a particular RL event or particular RL events) may comprise at least: a condition based at least on at least one of one or more mobility stability related counters exceeding or falling below respective at least one of one or more first mobility stability thresholds, and / or the one or more RL event exit conditions (or one or more RL event exit conditions defined for a particular RL event or particular RL events) may comprise at least: a condition based at least on at least one of the one or more mobility stability related counters exceeding or falling below respective at least one of one or more second mobility stability thresholds. Each of the one or more first mobility stability thresholds and the one or more second mobility stability thresholds may be defined separately per RL event (i.e., separately for exploration and exploitation events). In general, a high value of any of the listed mobility stability related counters may be indicative of a need for triggering an exploration event (and exiting an exploitation event) while low value of any of the listed mobility stability related counters may be indicative of a need for triggering an exploitation event (and exiting an exploration event).

[0084] More specifically, in some embodiments, the one or more exploration event triggering conditions may comprise at least: a condition triggered when at least one of one or more mobility stability related counters exceeds respective at least one of one or more first exploitation mobility stability thresholds, and / or the one or more exploration event exit conditions may comprise at least: a condition triggered when at least one (or all) of the one or more mobility stability related counters falls below respective at least one (or all) of one ormore second exploitation mobility stability thresholds. Additionally or alternatively, the one or more exploitation event triggering conditions comprising at least: a condition triggered when at least one (or all) of one or more mobility stability related counters for the apparatus falls below respective at least one (or all) of one or more first exploitation mobility stability thresholds, and / or the one or more exploitation event exit conditions may comprise at least: a condition triggered when at least one (or all) of the one or more mobility stability related counters exceeds below respective at least one (or all) of one or more second exploitation mobility stability thresholds. Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments.

[0085] Referring to option 4), the geolocation-based metric (or measure) may be, for example, a (reference) distance with respect to the serving access node or cell or a round trip time (RTT) based on which said reference distance may be derived using timing advance (TA) estimation. Alternatively, the geolocation-based metric may be a metric indicating a position of the apparatus relative to (or within) a pre-defined area defined by a set of coordinates (e.g., an area encompassing a particular building). Alternatively, the geolocation metric may be a metric indicating whether the apparatus is within coverage area(s) of one or more access nodes (e.g., as specified by physical cell identity, PCI) (i.e., is being served by any of the one or more access nodes) and optionally whether the signal strength from any of the one or more access nodes is above a pre-defined threshold (indicating that the associated cells are detectable). By specifying serving or detectable PCIs, control of exploration impact is directly linked to base stations which might be affected.

[0086] In option 4), the one or more RL event triggering conditions (or one or more RL event triggering conditions defined for a particular RL event or particular RL events) may comprise, e.g., at least: a condition based at least on a geolocation-based metric for the apparatus indicating that the apparatus is within or outside of a first distance range defined, e.g., relative to a serving access node or within or outside of a first area, and / or the one or more RL event exit conditions (or one or more RL event exit conditions defined for a particular RL event or particular RL events) may compris at least: a condition based at least on a geolocationbased metric for the apparatus indicating that the apparatus is within or outside of a second distance range defined, e.g., relative to the serving access node or within or outside of a second area. The first and second distance ranges or areas may be the same or different. Each of the first and second distance ranges or areas may be defined separately per RL event (e.g., separately for exploration and exploitation events). In general, an exploration event may betriggered and an exploitation event may be exited when the apparatus is farther from the serving access node (i.e., closer to the cell edge) while an exploitation event may be triggered and an exploration event may be exited when the apparatus is close to the serving access node (i.e., farther from the cell edge). Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments.

[0087] Referring to option 5), the one or more RL event triggering conditions (or specifically one or more exploration event exit conditions and / or one or more exploitation event triggering conditions thereof) may comprise a condition based at least on a time range or a time of day. For instance, during a duration of an emergency event, exploration may be disabled. To give another example, exploration may be disabled or at least limited to avoid overload during a scheduled sport event in a sport stadium. In general, the condition based on the time range or a time of day may be used for exiting an exploration event and / or triggering an exploitation event during known busy hours. Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments.

[0088] Referring to option 6), the 5G NR RRM measurement relaxation conditions may comprise, e.g., cellEdgeEvaluation and / or lowMobilityEvaluation. The legacy threshold of RRM measurement relaxation and associated conditions may be reused at least for the exploration event triggering and exiting. Also these conditions may be combined with one or more other conditions via Boolean operation(s), in some embodiments

[0089] Referring to option 7), at least one or more exploration event exit conditions and / or one or more exploitation event triggering conditions may be defined based on the one or more capabilities of the apparatus, in some embodiments. In other words, the capabilities of the apparatus may limit the performing of the exploration. The one or more capabilities of the apparatus may comprise, e.g., at least one of a maximum allowed exploration time duration (e.g., in ms), maximum allowed power consumption in exploration (e.g., in mW / dBm), a maximum allowed delay budget (e.g., in ms), a minimum allowed user throughput requirement (e.g., in kbps), or a maximum and / or minimum allowed action range (depends on the use case). To give an example of the last option, the network may configure the apparatus to explore the full action space of handover parameters, e.g., 0 ms - 400 ms with resolution of 10 ms. However, the apparatus may be only capable of exploration in the 0 ms - 120 ms range.

[0090] In option 7), the or more exploration event exit conditions and / or the one or more exploitation event triggering conditions may comprise one or more apparatus capability conditions comprising, e.g., at least one of:- a condition triggered when a duration of an exploration event exceeds a predefined maximum allowed duration of the exploration event for the apparatus,- a condition triggered when a power consumption of the apparatus exceeds a predefined maximum allowed power consumption during the exploration event,- a condition triggered when a pre-defined maximum allowed delay budget of the apparatus is exceeded during the exploration event,- a condition triggered when a user throughput of the apparatus falls below a predefined minimum allowed user throughput requirement of the apparatus during the exploration event,- a condition triggered when one or more pre-defined maximum allowed action ranges of the one or more RRM parameters is exceed during the exploration event, or- a condition triggered when a value of any of the one or more RRM parameters falls below a respective minimum allowed action range of the one or more RRM parameters during the exploration event.Any of the conditions of this paragraph may be combined with each other or with one or more other conditions via Boolean operation(s), in some embodiments.

[0091] In option 8), a QoS-based metric may be, for example, uplink throughput of the apparatus. Here, the one or more RL event triggering conditions (or one or more RL event triggering conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on a value of the QoS-based metric exceeding or falling below a first QoS threshold and / or the one or more RL event exit conditions (or one or more RL event exit conditions defined for a particular RL event or particular RL events) may comprise a condition based at least on a value of the QoS-based metric exceeding or falling below a second QoS threshold (which may be the same or different compared to the first QoS threshold). In general, low value of the QoS-based metric (indicating high QoS) may be indicative of a need for triggering an exploration event (and exiting an exploitation event) while high value of the QoSbased metric may be indicative of a need for triggering an exploitation event (and exiting an exploration event).

[0092] More specifically, in some embodiments, the one or more exploration event triggering conditions may comprise at least: a condition triggered when a value of the QoS-based metric for the apparatus falls below a first exploration QoS threshold, and / or the one or more exploration event exit conditions may comprise at least: a condition triggered when a value of the QoS-based metric for the apparatus exceeds a second exploration QoS threshold. Additionally or alternatively, the one or more exploitation event triggering conditions may comprise at least: a condition triggered when a value of the QoS-based metric for the apparatus exceeds a first exploitation QoS threshold, and / or the one or more exploitation event exit conditions may comprise at least: a condition triggered when a value of the QoS-based metric for the apparatus falls below a second exploitation QoS threshold. Any of the conditions of this paragraph may be combined with one or more other conditions via Boolean operation(s), in some embodiments.

[0093] In some embodiments, the at least one configuration message further comprises one or more configuration parameters for at least one event (e.g., at least one of an exploration event or an exploitation event) of the RL model defined in the one or more RL strategies. The one or more configuration parameters may form a part of the one or more RL strategies. The one or more configuration parameters for each of the at least one event may comprise at least one of:- a time window for evaluating the one or more RL event exit conditions for an event (the time window having a start time corresponding to a start of the evaluation of the event),- a timer for evaluating the one or more RL event exit conditions for the event,- a first time-to-trigger for satisfying the one or more RL event triggering conditions for the event, or- a second time-to-trigger for satisfying the one or more RL event exit conditions for the event.Here, the time-to-trigger enables avoiding frequent switching between starting and stopping of exploration or exploitation. The first / second time-to-trigger may trigger the timer. The time window, the timer and the first and second times-to-trigger may be equally called the (event) evaluation time window, the (event) evaluation timer and the first and second (event) evaluation times-to-trigger. In some embodiments, at least one of the time window, the timer and the first and second time-to-triggers as listed above may be defined separately for both exploration and exploitation events. In some embodiments, the one or more configuration parameters for each of the at least one event of RL model may comprise one of the time window or the timer, the first time-to-trigger and the second time-to-trigger.

[0094] The apparatus carries out, in block 303, optimization of the one or more RRM parameters using the RL model according to the one or more RL strategies (i.e., according to one or more RL event triggering and / or exit conditions and optionally one or more configuration parameters). As indicated above, the one or more RL event triggering and / or exit conditions may comprise at least one (or all) of one or more exploration triggering conditions, one or more exploration exit conditions, one or more exploitation triggering conditions or one or more exploitation exit conditions. In other words, the apparatus may perform exploration during exploitation events and / or exploitation (at least) during exploitation events while simultaneously evaluating the one or more RL event triggering and / or exit conditions and transitioning between the exploration and exploitation events accordingly. The operation pertaining to block 303 is discussed in more detail in connection with further embodiments.

[0095] In some embodiments, the apparatus may notify any exploration event status (i.e., active / inactive) changes to the network entity (e.g., a (serving) access node). This notification may be carried out directly following the change or periodically or regularly. Alternatively, the apparatus may transmit the information on the exploration event status change upon receiving an associated request from the network entity. In some embodiments, similar functionality may be implemented for the exploitation event status changes, in addition or alternative to the exploration event status changes.

[0096] FIG. 4 illustrates a process for configuring a terminal device to use an RL model using according to one or more RL strategies according to embodiments. The process of FIG. 4 may be carried out by an apparatus. The apparatus may be a network entity or node or a part thereof. The network entity or node may be a (non-distributed) access node or a part thereof, a distributed access node or an RU, DU or CU of a distributed access node or a part thereof. The apparatus may be, e.g., a non-distributed access node 104 of FIG. 1 or an RU 116, 118 of FIG. 1, a DU 104 of FIG. 1, a CU 108 of FIG. 1 or a combination thereof. In the following, the entity carrying out the process of FIG. 4 is called simply an apparatus.

[0097] The process of FIG. 4 corresponds to a process carried out at the network side in parallel with the carrying out of the process of FIG. 3 at the terminal device. Thus, any of the features, descriptions and definitions provided in connection with the process of FIG. 3 apply, mutatis mutandis, also for the process of FIG. 4. These features, descriptions and definitions are not repeated here in full merely for brevity.

[0098] Referring to FIG. 4, the apparatus transmits, in block 401, to a terminal device, at least one configuration message comprising one or more RL strategies. The terminal device may be storing an RL model trained to evaluate one or more RRM parameters of the apparatus based on one or more radio measurement metrics or at least have access to such an RL model, as described in connection with FIG. 3. The one or more RL strategies (or one or more RL event conditions defined by the one or more RL strategies) comprise at least one of: one or more RL event triggering conditions for triggering an event of the RL model, or one or more RL event exit conditions for exiting an event of the RL model. In some embodiments, the one or more RL strategies (or said one or more RL event conditions defined by the one or more RL strategies) may comprise at least one (or all) of: one or more first RL event triggering conditions for triggering a first event (e.g., an exploration event) of the RL model, one or more first RL event exit conditions for exiting the first event of the RL model, one or more second RL event triggering conditions for triggering a second event (e.g., an exploitation event) of the RL model, or one or more second RL event exit conditions for exiting the second event of the RL model. Any of the further definitions for the at least one configuration message provided in connection with the process of FIG. 3 may apply also here.

[0099] In some embodiments, the apparatus may determine the one or more RL strategies, before block 401, based on historical statistics of operation of the terminal device (e.g., measurements carried out, speed variations and / or applied RRM parameters).

[0100] In some embodiments, the transmission of the at least one configuration message is preceded by reception of a configuration request from the terminal device, as will be described in further detail in connection with FIGs. 10 & 11. In some such embodiments, the apparatus may determine the one or more RL strategies, before block 401, based on the configuration request received from a terminal device.

[0101] FIG. 5 shows an example of applying an exploration event triggering condition and an exploration event exit condition according to the option 1) described in connection with block 302 of FIG. 3. In other words, FIG. 5 corresponds to the case where the one or more exploration event triggering conditions consists of a condition triggered when an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus exceeds a first exploration RSRP difference threshold and the one or more exploration event exit conditions consists of a condition triggered when an RSRP difference between the serving cell of the apparatus and the strongest neighbor cell of the apparatus falls below a secondexploration RSRP difference threshold (being lower than the first exploration RSRP difference threshold). FIG. 5 shows the measured RSRP difference (“RSRPdiff’) in dB plotted against time in ms. The left and right dotted lines illustrate the first and second exploration RSRP difference thresholds. The rectangular box illustrates a time range when the exploitation event is active (including also the initial time-to-trigger time when the event evaluation timer is not running).

[0102] In FIG. 5, when the RSRP difference exceeds the first exploration RSRP difference threshold for triggering the exploration event, the event evaluation timer is started following a first time-to-trigger time (optionally also configured via a configuration parameter of the at least one configuration message). After some time, the RSRP difference falls below the second exploration RSRP difference threshold for exiting the exploration event. After a second time-to-trigger time (optionally also configured via a configuration parameter of the at least one configuration message), the exploration is ended.

[0103] FIG. 6 shows an example of applying an exploitation event triggering condition and an exploitation event exit condition according to the option 1) described in connection with block 302 of FIG. 3. In other words, FIG. 6 corresponds to the case where the one or more exploitation event triggering conditions consists of a condition triggered when an RSRP difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus falls below a first exploitation RSRP difference threshold and the one or more exploitation event exit conditions consists of a condition triggered when an RSRP difference between the serving cell of the apparatus and the strongest neighbor cell of the apparatus exceeds a second exploitation RSRP difference threshold (being lower than the first exploitation RSRP difference threshold). FIG. 6 shows the measured RSRP difference (“RSRPdiff’) in dB plotted against time in ms. The left and right dotted lines illustrate the first and second exploitation RSRP difference thresholds. The rectangular box illustrates a time range when the exploitation event is active (including also the initial time-to-trigger time when the event evaluation timer is not running).

[0104] In FIG. 6, when the RSRP difference falls below the first exploitation RSRP difference threshold for triggering the exploitation event, the event evaluation timer is started following a first exploitation time-to-trigger time (optionally also configured via a configuration parameter of the at least one configuration message). After some time, the RSRP difference exceeds the second exploitation RSRP difference threshold for exiting theexploitation event. After a second exploitation time-to-trigger time (optionally also configured via a configuration parameter of the at least one configuration message), the exploitation is ended.

[0105] FIG. 7 illustrates a process for triggering an exploration event, operating during the exploration event and exiting the exploration event according to embodiments. The process of FIG. 7 may be carried out by an apparatus. The apparatus may be a terminal device or a part thereof. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. In the following, the entity carrying out the process of FIG. 7 is called simply an apparatus.

[0106] The process of FIG. 7 may correspond to one more detailed implementation of block 303 of FIG. 3. Alternatively, the process of FIG. 7 may form a part of the operation associated block 303 of FIG. 3. Accordingly, any of the features, descriptions and definitions provided in connection with FIG. 3 may apply, mutatis mutandis, also here.

[0107] In FIG. 7, it is initially assumed that the apparatus is maintaining, in at least one memory, an RL model (or at least has access to such a model) and has received at least one configuration message comprising one or more RL strategies, as described in connection with blocks 301, 302 of FIG. 3. It is assumed here that the one or more RL strategies comprise at least one or more exploration event triggering conditions and one or more exploration event exit conditions. At the beginning of the process of FIG. 7, no exploration or exploitation event may be active.

[0108] The apparatus initially evaluates, in block 701, the one or more exploration event triggering conditions. In other words, the apparatus monitors whether at least one of the one or more exploration event triggering conditions has been satisfied. The evaluation in block 701 may be carried out continuously, periodically or regularly.

[0109] In response to none of the one or more exploration event triggering conditions being satisfied in block 702, the evaluation of the one or more exploration event triggering conditions may continue, i.e., the process proceeds back to block 701.

[0110] In response to at least one of the one or more exploration event triggering conditions being satisfied in block 702, the apparatus triggers an exploration event. Blocks 703 to 706 relate to operation carried out during the exploration event.

[0111] The apparatus obtains, in block 703, values of the one or more radio measurement metrics. Here, the obtaining may correspond to measuring, directly or indirectly, values of the one or more radio measurement metrics.

[0112] The apparatus performs, in block 704, exploration using the RL model based at least on the obtained values of the one or more radio measurement metrics according to a predefined exploration strategy. The pre-defined exploration strategy may be any strategy described above in connection with the general description of reinforcement learning (e.g., epsilon-greedy).

[0113] In some alternative embodiments, the one or more radio measurement metrics in block 703, 704 may be replaced with one or more measurement metrics which may comprise one or more radio measurement metrics and / or one or more non-radio measurement metrics (e.g., speed of the apparatus).

[0114] The apparatus evaluates, in block 705, the one or more exploration event exit conditions. In other words, the apparatus monitors whether at least one of the one or more exploration event exit conditions has been satisfied.

[0115] In response to none of the one or more exploration event exit conditions being satisfied in block 706, the exploration may continue, i.e., the process proceeds back to block 703.

[0116] In response to at least one of the one or more exploration event exit conditions being satisfied in block 706, the apparatus triggers exiting of the exploration event. Thus, the process may proceed back to block 701.

[0117] In some embodiments, the one or more exploration triggering and exit conditions may define a so-called exploration region. Specifically, the exploration region may be defined based on the parameters associated with the one or more exploitation triggering and exit conditions (including but not limited to a location of the apparatus). Within this region, no matter how the RL model is executing, the RL agent at the apparatus will be enforced to perform exploration.

[0118] FIG. 8 illustrates a process for triggering an exploitation event, operating during the exploitation event and exiting the exploitation event according to embodiments. The process of FIG. 8 may be carried out by an apparatus. The apparatus may be a terminal deviceor a part thereof. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. In the following, the entity carrying out the process of FIG. 8 is called simply an apparatus.

[0119] The process of FIG. 8 may correspond to one more detailed implementation of block 303 of FIG. 3. Alternatively, the process of FIG. 8 may form a part of the operation associated block 303 of FIG. 3 (e.g., along with the process of FIG. 7). Accordingly, any of the features, descriptions and definitions provided in connection with FIG. 3 may apply, mutatis mutandis, also here.

[0120] In FIG. 8, it is initially assumed that the apparatus is maintaining, in at least one memory, an RL model (or at least has access to such a model) and has received at least one configuration message comprising one or more RL strategies, as described in connection with blocks 301, 302 of FIG. 3. It is assumed here the one or more RL strategies comprise at least one or more exploitation event triggering conditions and one or more exploitation event exit conditions. At the beginning of the process of FIG. 8, no exploration or exploitation event may be active.

[0121] The apparatus initially evaluates, in block 801, the one or more exploitation event triggering conditions. In other words, the apparatus monitors whether at least one of the one or more exploitation event triggering conditions have been satisfied.

[0122] In response to none of the one or more exploitation event triggering conditions being satisfied in block 802, the evaluation of the one or more exploitation event triggering conditions may continue, i.e., the process proceeds back to block 801

[0123] In response to at least one of the one or more exploitation event triggering conditions being satisfied in block 802, the apparatus triggers an exploitation event. Blocks 703 to 706 relate to operation carried out during the exploitation event

[0124] The apparatus obtains, in block 803, values of the one or more radio measurement metrics. Here, the obtaining may correspond to measuring, directly or indirectly, values of the one or more radio measurement metrics.

[0125] The apparatus performs, in block 804, exploitation using the RL model based at least on the obtained values of the one or more radio measurement metrics. In other words, theapparatus at least calculates, in block 804, using the RL model, values of the one or more RRM parameters based on the collected values of the one or more radio measurement metrics.

[0126] In some alternative embodiments, the one or more radio measurement metrics in block 803, 804 may be replaced with one or more measurement metrics which may comprise one or more radio measurement metrics and / or one or more non-radio measurement metrics (e.g., speed of the apparatus).

[0127] The apparatus evaluates, in block 805, the one or more exploitation event exit conditions. In other words, the apparatus monitors whether at least one of the one or more exploitation event exit conditions has been satisfied.

[0128] In response to none of the one or more exploitation event exit conditions being satisfied in block 806, the exploration may continue, i.e., the process proceeds back to block 803.

[0129] In response to at least one of the one or more exploitation event exit conditions being satisfied in block 806, the apparatus triggers exiting of the exploitation event. Thus, the process may proceed back to block 801.

[0130] The one or more exploitation triggering and exit conditions may define a so- called exploitation region. Specifically, the exploration region may be defined based on the parameters associated with the one or more exploitation triggering and exit conditions (including but not limited to a location of the apparatus). Within this region, no matter how the RL model is executing, the RL agent at the apparatus will be enforced to perform exploitation.

[0131] FIG. 9 illustrates a process for triggering and exiting exploration and exploitation events and operating during the exploration and exploitation events according to embodiments. The process of FIG. 9 may be carried out by an apparatus. The apparatus may be a terminal device or a part thereof. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. In the following, the entity carrying out the process of FIG. 9 is called simply an apparatus.

[0132] The process of FIG. 9 may correspond to one more detailed implementation of block 303 of FIG. 3. Accordingly, any of the features, descriptions and definitions provided in connection with FIG. 3 may apply, mutatis mutandis, also here. The process of FIG. 9 corresponds effectively to a combination of processes of FIGs. 7 and 8 and, thus, any of thefeatures, descriptions and definitions provided in connection with FIGs. 7 and / or 8 may apply, mutatis mutandis, also here.

[0133] In FIG. 9, it is initially assumed that the apparatus is maintaining, in at least one memory, an RL model (or at least has access to such a model) and has received at least one configuration message comprising one or more RL strategies, as described in connection with blocks 301, 302 of FIG. 3. It is assumed here the one or more RL strategies comprise at least one or more exploration event triggering conditions and one or more exploration event exit conditions. At the beginning of the process of FIG. 9, no exploration or exploitation event may be active

[0134] The apparatus initially evaluates, in block 901, the one or more exploration event triggering conditions and the one or more exploitation event triggering conditions. In other words, the apparatus monitors whether at least one of the one or more exploration event triggering conditions or the one or more exploitation event triggering conditions has been satisfied. The evaluation in block 901 may be carried out continuously, periodically or regularly.

[0135] In response to none of the one or more exploration event triggering conditions being satisfied in block 902, the process proceeds to block 905 where it is checked whether any of the one or more exploitation event triggering conditions are satisfied. In response to none of the one or more exploitation event triggering conditions being satisfied in block 905, the evaluation of the exploration and exploitation event triggering conditions may continue, i.e., the process goes back to block 901.

[0136] In response to at least one of the one or more exploration or exploitation event triggering conditions being satisfied in block 902, the apparatus triggers an exploration event. The apparatus performs, in block 903, exploration while the exploration event is active. The exploration during the exploration event carried out in block 903 may correspond to operation discussed in connection with blocks 703 to 705 of FIG. 7.

[0137] During the exploration in block 903, the apparatus evaluates the one or more exploration event exit conditions, as discussed in connection with block 705 of FIG. 7. In response to none of the one or more exploration event exit conditions being satisfied in block 904, the exploration may continue, i.e., the process proceeds back to block 903. In response to at least one of the one or more exploration event exit conditions being satisfied in block 904,the apparatus triggers exiting of the exploration event. In this case, the process proceeds to evaluate and check, in block 905, whether any of the one or more exploitation event triggering conditions have been satisfied.

[0138] In response to at least one of the one or more exploitation event triggering conditions being satisfied in block 905, the apparatus triggers an exploitation event. The apparatus performs, in block 906, exploitation while the exploitation event is active. The exploitation during the exploitation event carried out in block 906 may correspond to operation discussed in connection with blocks 803 to 805 of FIG. 8.

[0139] During the exploitation in block 906, the apparatus evaluates the one or more exploitation event exit conditions, as discussed in connection with block 805 of FIG. 8. In response to none of the one or more exploitation event exit conditions being satisfied in block 907, the exploitation may continue, i.e., the process goes back to block 906. In response to at least one of the one or more exploration event exit conditions being satisfied in block 907, the apparatus triggers exiting of the exploitation event. In this case, the process proceeds to back to block 901.

[0140] Based on established exploration event & (exploration-exclusive) exploitation event (set up by the network and indicated to a terminal device), a terminal device may perform additional steps in order to verify these newly configured constraints and report related monitoring status to the network. Here, two cases may be considered. In a network-centric solution, the terminal device may perform exploration & exploitation transitions for the followup RL actions update (e.g., adapting one or more RRM parameters such as one or more handover parameters) only when a positive acknowledgment (ACK) is received from the network. On other hand, in a terminal device centric solution, a terminal device may autonomously carry out the exploration & exploitation transitions and then execute the followup RL actions to adapt the one or more RRM parameters without dedicated network instruction.

[0141] FIG. 10 illustrates signaling between a terminal device and a network entity or node according to the network centric solution of some embodiments. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. The network entity or node may be an access node or a part thereof (serving the terminal device). The network entity or node may be a non-distributed access node, a distributed access node or an RU, DU or CU of a distributed access node. The network entity or node may be, e.g., a non-distributed access node104 of FIG. 1 or an RU 116, 118 of FIG. 1, a DU 104 of FIG. 1, a CU 108 of FIG. 1 or a combination thereof. In FIG. 10, dashed lines are used for indicating optional features.

[0142] The process of FIG. 10 corresponds to one more detailed implementation of the processes of FIGs. 3 & 4. Thus, any of the features, descriptions and definitions provided in connection with FIG. 3 and / or 4 (and / or any of FIGs. 5 to 8) may apply, mutatis mutandis, also here.

[0143] Referring to FIG. 10, the terminal device may initially maintain, in at least one memory, (or at least have access to) an RL model though it may not yet have been configured and trained for determining one or more RRM parameters based on one or more radio measurement metrics.

[0144] The terminal device-network entity capability exchange process may be initially carried out in message(s) 1001 via RRC (re)configuration. The terminal device- network entity capability exchange process may comprise at least transmitting a first capability exchange message from the network entity to the terminal device and receiving the first capability exchange message at the terminal device, transmitting a second capability exchange message from the terminal device to the network entity and receiving the second capability exchange message at the network entity and transmitting a third capability exchange message from the network entity to the terminal device and receiving the third capability exchange message at the terminal device. The first capability exchange message may be a request for capability information of the terminal device. The second capability exchange message may comprise information on one or more capabilities of the terminal device such as maximum allowed exploration time duration, maximum allowed power consumption during exploration event, a maximum allowed delay budget, a minimum allowed user throughput requirement and / or maximum and / or minimum allowed action range. The third capability exchange message may comprise, e.g., one or more default RL model settings to be used by the terminal device. The one or more default RL model settings may, in some embodiments, be determined based on contents of the second capability exchange message. The one or more default RL model settings may comprise, e.g., at least one of CIO, TTT, timing granularity, one or more RL hyperparameters, one or more timing constraints, one or more used exploration policies or RL use case information (defining, e.g., state, action, reward of the RL model). The one or more RL hyperparameters may comprise, e.g., at least one of learning rate, exploration probability or discounting factor.

[0145] The terminal device transmits, in message 1002, to the network entity, a configuration request requesting configuration of one or more RL strategies for an RL model. As in previous embodiments, the one or more RL strategies may comprise at least one or more RL event conditions. The RL model may for determining one or more RRM parameters of the terminal device based on one or more radio measurement metrics. The configuration request may be an RRC message.

[0146] The configuration request may request configuration of one or more particular RL event conditions (e.g., configuration of one or more exploration event triggering conditions, one or more exploration event exit conditions, one or more exploitation event triggering conditions and / or one or more exploitation event exit conditions). In some embodiments, the configuration request may further request one or more configuration parameters (e.g., a time window, a timer and / or a TTT) for at least one event (e.g., an exploration event and / or exploitation event) of the RL model defined in the one or more RL strategies (as described above). In other embodiments, the configuration request may be a more generic request not explicitly specifying, e.g., the specific RL events to be configured. In such cases, the network entity may decide on the specifics of the configuration (e.g., RL events, RL event triggering / exit conditions and / or configuration parameters to be configured).

[0147] In general, the RL model and the one or more RL strategies (and the one or more RL event conditions defined therein) may be defined similar to any of the embodiments discussed above (e.g., as discussed in connection with FIG. 3). Thus, a state, an action and a reward of the RL model may be defined as follows: the state is defined by one or more RRM parameters of the terminal device, the action taken by an agent of the RL model from a given state is defined as a modification of at least one of the one or more RRM parameters, and the reward of taking a given action in a given state is calculated based on the one or more radio measurement metrics for optimizing values of one or more KPIs of the terminal device or of a serving cell (provided by the network entity of FIG. 10). Additionally or alternatively, each of the one or more RL event conditions may be defined based on at least one of: RSRP difference between a serving cell of the terminal device and a strongest neighbor cell of the terminal device, an average speed of the terminal device, one or more mobility stability related counters of the terminal device, a geolocation-based metric indicative of location or distance of the terminal device (e.g., relative to the serving network entity), time, a 5G New Radio RRM measurement relaxation condition, or one or more capabilities of the terminal device. Here,one or more RL event conditions based at least on the one or more capabilities of the terminal device may relate at least to an (exploration-excluding) exploitation event (or exiting thereof).

[0148] The network entity receives, in block 1003, the configuration request. Moreover, the network entity transmits, in message 1004, to the terminal device, at least one configuration message comprising the requested one or more RL strategies. The one or more RL strategies may have been defined, e.g., based on the configuration request and / or information obtained during the terminal device-network entity capability exchange process (messages 1001). The terminal device receives, in block 1005, the at least one configuration message. Each of the at least one configuration message may be an RRC message. The one or more RL event conditions may comprise all or at least one of: one or more exploration event triggering conditions for triggering an exploration event of the RL model, one or more exploration event exit conditions for exiting the exploration event of the RL model, one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or one or more exploitation event exit conditions for exiting the exploitation event. In some embodiments, the configuration message may also comprise one or more configuration parameters for the at least one event of the RL model, as in connection with FIG. 3.

[0149] In some embodiments, the at least one configuration message comprises information defining an evaluation time window (or an evaluation timer).

[0150] The terminal device may initialize, in block 1006, the default RL model settings for the performing of the exploration and / or the exploitation. The default RL model settings may have been configured via the terminal device-network entity capability exchange process (messages 1001).

[0151] The terminal device starts performing, in block 1007, exploration and / or exploitation using the RL model for determining, using the RL model, the one or more RRM parameters of the terminal device based on the one or more radio measurement metrics. The default RL model settings may be used here initially.

[0152] The terminal device evaluates, in block 1008, the one or more RL event conditions as defined in the one or more RL strategies. In other words, the terminal device may monitor the one or more RL event conditions (dependent, e.g., on RSRP difference between the serving and strongest non-serving cells, mobility stability related counters and / or location of the terminal device) configured via message 1004.

[0153] In some embodiments, the evaluating of block 1008 may be performed within the evaluation time window defined via message 1004 (or according to an evaluation timer). Thus, values of parameters associated with the one or more RL event conditions may be accumulated within this evaluation time window (or while the evaluation timer is running) and, then, compared to corresponding RL event conditions (e.g., thresholds) for triggering entering or exiting an RL event.

[0154] Following the evaluating in block 1008, the terminal device transmits, message 1009, to the network entity, an evaluation report comprising results of the evaluation. The message 1009 may be a layer- 1 / layer-2 message for minimizing signaling cost. The evaluation report may comprise, e.g., information on whether the one or more RL event conditions have been satisfied (during the evaluation time window or before expiration of the evaluation timer) and optionally for how long (in time) have they been satisfied. Said information on whether the one or more RL event conditions have been satisfied (during the evaluation time window or before expiration of the evaluation timer) may be provided, e.g., as binary information. For example, ‘1’ & ‘0’ may indicate, respectively, that a given RL event condition has been satisfied or not satisfied (or has been satisfied for a particular time slot or range). Additionally or alternatively, the evaluation report may comprise values of the parameters associated with the one or more RL event conditions (e.g., values of RSRP difference, one or more mobility stability related counters, a geolocation-based metric of the terminal device, an average speed of the terminal device and / or time).

[0155] In some embodiments, if a plurality of RL event conditions are jointly evaluated in block 1008, the evaluation report may comprise A- bit information in a compound format, where A is a positive integer.

[0156] In some embodiments, the evaluation report may be generated (by an RL agent) according to one or more reporting signal criteria.

[0157] The network entity receives, in block 1010, from the terminal device, the evaluation report comprising the results of the evaluation of the one or more RL event conditions. The network entity determines, in block 1011, whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event. For example, the network entity may check or verify whether at least one of the one or more RL event conditions has been satisfied. The network entity may be assumed to have knowledge of the current exploration or exploitation event (ifany) active in the terminal device (as all RL events may be activated via the network entity). Thus, if the exploration event is active at the terminal device, the network entity may check or verify at least whether the one or more exploration event exit conditions and / or the one or more exploitation event triggering conditions has been satisfied. If the exploitation event is active at the terminal device, the network entity may check or verify at least whether the one or more exploration event triggering conditions and / or the one or more exploitation event exit conditions has been satisfied.

[0158] In some embodiments, said least one of the triggering or the exiting may comprise one of the exploration event or the exploitation event and exiting other of the exploration event or the exploitation event. Thus, the exploration event and the exploitation may be defined as mutually exclusive states of the terminal device.

[0159] Based on the determining in block 1011, the network entity transmits, in message 1012, to the terminal device, a positive or negative acknowledgement (an ACK or a NACK) for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event. The ACK or NACK may be a layer- 1 / layer-2 message. For example, the ACK or NACK may a medium access control (MAC) or downlink control information (DO) message. In FIG. 10, the case where an ACK is transmitted has been illustrated.

[0160] The terminal device receives, in block 1013, from the network entity, the positive acknowledgement. Based on the reception of the positive acknowledgment in block 1010 and on the results of the evaluation performed in block 1008, the terminal device performs, in block 1014, at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event. Thereafter, the terminal device may continue exploration and / or exploration in the new state.

[0161] In some embodiments, the terminal device may perform, in block 1015, RL policy update on the RL model in response to the received positive acknowledgment. In other words, the terminal device may redefine how an action taken when in a given state is derived. The RL policy update may be based on results of the exploration and / or the exploitation carried out by the terminal device.

[0162] The terminal device may transmit, in message 1016, to the network entity, a report comprising a plurality of values of the one or more RRM parameters determined, by theterminal device, using the RL model (and subsequently used by the apparatus). Message 1016 may be a layer- 1 / layer-2 message (e.g., a MAC control element, CE, or uplink control information, UCI). Said report is received in block 1017 by the network entity. The network entity may store the plurality of values of the one or more RRM parameters to at least one memory. The plurality of values of the one or more RRM parameters may be used subsequently, e.g., for evaluation decision-making. For example, as mentioned above, the network entity may employ the plurality of values of the one or more RRM parameters for determining one or more further RL strategies for the terminal device.

[0163] Following the execution of actions pertaining to block 1014 (and optionally block 1015 and / or message 1016), the terminal device may repeat actions pertaining to blocks 1007, 1008, 1009, 1013, 1014 (and optionally block 1015 and / or message 1016). In other words, the actions pertaining to blocks 1007, 1008, 1009, 1013, 1014 (and optionally block 1015 and / or message 1016) may be carried out continuously, periodically or regularly.

[0164] As was mentioned above, in some cases (not shown in FIG. 10), the network entity may transmit, in message 1012, a negative acknowledgment (instead of the positive acknowledgment) if no exploration and / or exploitation event entry or exit is to be triggered. In such cases, the terminal device may, upon receiving the negative acknowledgment, repeat execution of blocks 1007 to 1009.

[0165] FIG. 11 illustrates signaling between a terminal device and a network entity or node according to the terminal device centric solution of some embodiments. Said terminal device may be, e.g., any of the terminal devices 100, 102 of FIG. 1. The network entity or node may be an access node or a part thereof (serving the terminal device). The network entity or node may be a non-distributed access node, a distributed access node or an RU, DU or CU of a distributed access node. The network entity or node may be, e.g., a non-distributed access node 104 of FIG. 1 or an RU 116, 118 of FIG. 1, a DU 104 of FIG. 1, a CU 108 of FIG. 1 or a combination thereof. In FIG. 11, dashed lines are used for indicating optional features.

[0166] The process of FIG. 11 corresponds to one more detailed implementation of the processes of FIGs. 3 & 4. Thus, any of the features, descriptions and definitions provided in connection with FIG. 3 and / or 4 (and / or any of FIGs. 5 to 8) may apply, mutatis mutandis, also here.

[0167] Many of the steps depicted in FIG. 11 correspond to steps of FIG. 10. Thus, the process of FIG. 11 is discussed in the following only briefly.

[0168] As described also in connection with FIG. 10, the terminal device may initially maintain, in at least one memory, an RL model though it may not yet have been configured and trained for determining one or more RRM parameters based on one or more radio measurement metrics. A terminal device-network entity capability exchange process may be initially carried out in message(s) 1101 via RRC (re)configuration. These message(s) may correspond fully to messages 1001 discussed above.

[0169] The terminal device transmits, in message 1102, to the network entity, a configuration request requesting configuration of one or more RL strategies for an RL model. As in previous embodiments, the one or more RL strategies may comprise at least one or more RL event conditions, and the RL model is for determining one or more RRM parameters of the terminal device based on one or more radio measurement metrics. The configuration request may be an RRC message. The configuration request is received, in block 1103, by the network entity. Message 1102 and block 1103 may correspond, respectively, fully to message 1002 and block 1003 discussed above.

[0170] The network entity transmits, in message 1104, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions. Each of the at least one configuration message may be an RRC message. The one or more RL event conditions may comprise all or at least one of: one or more exploration event triggering conditions for triggering an exploration event of the RL model, one or more exploration event exit conditions for exiting the exploration event of the RL model, one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or one or more exploitation event exit conditions for exiting the exploitation event. The terminal device receives, in block 1105, from the network entity, at least one configuration message comprising the one or more RL strategies comprising the one or more RL event conditions. Message 1104 and block 1105 may correspond, respectively, fully to message 1004 and block 1005 discussed above.

[0171] The terminal device may initialize, in block 1106, the default RL model settings for the performing of the exploration and / or the exploitation, similar to block 1006 of FIG. 10.

[0172] The terminal device performs, in block 1107, exploration and / or exploitation using the RL model for determining, using the RL model, one or more RRM parameters of the apparatus based on one or more radio measurement metrics. Moreover, the terminal device evaluates, in block 1108, the one or more RL event conditions as defined in the one or more RL strategies. Blocks 1107, 1108 may correspond, respectively, fully to blocks 1007, 1008 of FIG. 10.

[0173] In some embodiments, the evaluating of block 1108 may be performed within the evaluation time window (or according to an evaluation timer) defined via message 1104. Thus, values of parameters associated with the one or more RL event conditions may be accumulated within this evaluation time window (or while the evaluation timer is running) and, then, compared to corresponding RL event conditions (e.g., thresholds) for triggering entering or exiting an RL event.

[0174] In the terminal device centric solution of FIG. 11 , the terminal device may not seek verification from the network for triggering entering or exiting of an RL event (i.e., an exploration or exploitation event). Instead, the terminal device may directly proceed to triggering of an exploration or exploitation event and / or exiting an exploitation or exploration event when corresponding exploration and / or exploitation event triggering and / or exiting condition(s) are satisfied. Accordingly, the terminal device performs, in block 1109, at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event based on the results of the evaluation in block 1108. Apart from the fact that block 1109 is not based on (or triggered by) reception of a positive acknowledgment from the network entity, block 1109 may correspond fully to block 1014 of FIG. 10.

[0175] In some embodiments, the performing of the at least one of the triggering or the exiting in block 1109 comprises, based on the results of the evaluating, triggering one of the exploration event or the exploitation event and exiting other of the exploration event or the exploitation event.

[0176] The terminal device may perform, in block 1110, policy update on the RL model, similar to block 1015 of FIG. 10.

[0177] Additionally or alternatively, the terminal device may transmit, in message 1111, to the network entity, a report comprising results of the evaluation (determined in block 1108) and / or a plurality of values of the one or more RRM parameters determined, by the apparatus,using the RL model (and subsequently used by the apparatus). The results of the evaluation may correspond to the results of the evaluation as discussed above in connection with the evaluation report of message 1009 of FIG. 10. Message 1111 may be a layer- 1 / layer-2 message (e.g., a MAC CE or UCI). In some embodiments, the transmission of the report of message 1111 may be triggered in response to an end of the evaluation time window or in response to expiring of an evaluation timer. The report is received, in block 1112, by the network entity (and optionally stored to at least one memory). Elements 1111, 1112 may correspond, respectively, mutatis mutandis, to blocks 1016, 1016 of FIG. 10.

[0178] Following the execution of actions pertaining to block 1109 (and optionally block 1110 and / or message 1111), the terminal device may repeat actions pertaining to blocks 1107 to 1109 (and optionally block 1110 and / or message 1111). In other words, the actions pertaining to blocks 1107 to 1109 (and optionally block 1110 and / or message 1111) may be carried out continuously, periodically or regularly

[0179] In some embodiments, the terminal device and / or the network entity may be configured to carry both of the processes of FIGs. 10 and 11. In such cases, the network entity may be able to control or enforce any transitions between the network-centric operation (of FIG. 10) of the terminal device and the terminal device centric operation (of FIG. 11) of the terminal device.

[0180] In the following, three specific use cases are discussed in detail. These use cases may be applied to any of the embodiments discussed above.

[0181] Example use case - 1: UE RL-based HO parameters tuning

[0182] The embodiments may be applied to performing the tuning of handover parameters of a terminal device in either of baseline handover (BHO) and layer- 1 / layer-2 triggered mobility (LTM) scenarios. In such a use case, the objective is to determine the optimal combinations of CIOs and TTTs per terminal device such that one or more mobility-specific (overall) cell-level handover KPIs in terms of outage and failure can be reduced. The RL agent may be deployed at the terminal device side to adjust the handover parameters based on local radio conditions (as quantified by values of one or more radio measurement metrics such as signal strength or RSRP). Thus, in these embodiments, the one or more RRM parameters may comprise one or more handover parameters (e.g., one or more layer-3 handover parameters) and the one or more KPIs may comprise one or more mobility-specific cell-level KPIs relatingto outage and / or failure. The one or more handover parameters may comprise, e.g., at least a CIO and / or a TTT.

[0183] The one or more exploration event triggering conditions, the one or more exploration event exit conditions, the one or more exploitation event triggering conditions and the one or more exploitation event exit conditions may be defined in this use case as follows.

[0184] The one or more exploration event triggering conditions may be based at least on one or more mobility stability related scalar counters exceeding or falling below respective at least one of one or more first mobility stability thresholds. Similarly, the one or more exploration event exit conditions may be based at least on one or more mobility stability related scalar counters exceeding or falling below respective at least one of one or more second mobility stability thresholds. The possible options for a parameter or metric using which the one or more exploration event triggering conditions and / or the one or more exploration event exit conditions are defined comprise at least one of:- number of handovers, a number of handover failures, a number of ping-pong handovers, a number of unnecessary handovers, a number of radio link failures vs. respective network-configured thresholds,- number of beam switches vs. a network-configured threshold,- a beam failure indicator (BFI) counter vs. a network-configured threshold,- number of terminal device panel switch counters vs. a network-configured threshold.As the value of the one or more mobility stability related scalar counters (or at least one of them) exceeds a respective first mobility stability threshold (indicating that the radio environment is relatively unstable), the exploration event may be triggered. On the other hand, as the value of the one or more mobility stability related scalar counters (or at least one of them) falls below a respective second mobility stability threshold (indicating that the radio environment is relatively stable), the exploration event may be exited. As the mobility failure counters are more critical for reward / cost function, a rather aggressive action searching may be required here.

[0185] The one or more exploitation event triggering conditions may be based at least on a (reference) distance with respect to the serving cell exceeding or falling below a first distance threshold. Similarly, the one or more exploitation event exit conditions may be based at least on a (reference) distance with respect to the serving cell exceeding or falling below asecond distance threshold. As the value of the (reference) distance falls below a respective first distance threshold (i.e., the terminal device is relatively close to a center of the serving cell and far from the cell edge meaning that the radio environment may be relatively stable), the exploitation event may be triggered. On the other hand, as the value of the (reference) distance exceeds a respective second distance threshold (i.e., the terminal device is farther away from the center of the serving cell and closer to the cell edge meaning that the radio environment is likely unstable), the exploitation event may be exited.

[0186] Example use case - 2; UE RL-based uplink power control

[0187] The embodiments may be applied to performing uplink transmit power control (TPC) parameter optimization, where the objective is to determine an optimal transmit power value (or, more generally, values of one or more uplink power control parameters) per terminal device such that the uplink or sidelink terminal device throughput is maximized without significant impact on the overall uplink (or sidelink) cell performance (or the mobility-specific uplink or sidelink cell performance). The RL agent may be deployed at the terminal device side to adjust the uplink (or sidelink) transmit power of the terminal device based on local radio conditions (as quantified by values of one or more radio measurement metrics such as signal strength or RSRP). These power adjustments may be assumed to be combined with conventional uplink open loop power control (OLPC) and / or uplink closed loop power control (CLPC) mechanisms. Thus, in these embodiments, the one or more RRM parameters may comprise one or more uplink power control parameters and the one or more KPIs may comprise an uplink or sidelink throughput.

[0188] The one or more exploration event triggering conditions may be based at least on a radio coupling gain based metric exceeding or falling below respective at least one of one or more first threshold for the radio coupling gain based metric. Similarly, the one or more exploration event exit conditions may be based at least on a radio coupling gain based metric exceeding or falling below a second thresholds for the radio coupling gain based metric. The radio coupling gain based metric may be, e.g., an RSRP difference between the serving cell and the strongest neighbor cell (i.e., the strongest non-serving cell). As the value of the radio coupling gain based metric (e.g., the RSRP difference) falls below the first threshold (indicating that the terminal device is likely at the cell edge where the radio environment may be relatively unstable), the exploration event may be triggered. On the other hand, as the value of the radio coupling gain based metric (e.g., the RSRP difference) exceeds the secondthreshold (indicating that the terminal device is likely not near the cell edge but at a more central region where the radio environment is relatively stable), the exploration event may be exited.

[0189] The one or more exploitation event triggering conditions may be based at least on QoS-based metric (e.g., the uplink throughput) of the terminal device exceeding or falling below respective at least one of one or more first threshold for the QoS metric. Similarly, the one or more exploitation event exit conditions may be based at least on the QoS-based metric of the terminal device exceeding or falling below a second threshold for the QoS-based metric. As the value of the QoS-based metric exceeds the first threshold (indicating that the radio environment of the terminal device is likely relatively stable), the exploitation event may be triggered. On the other hand, as the value of the QoS-based metric falls below the second threshold (indicating that the radio environment of the terminal device is likely relatively unstable), the exploitation event may be exited.

[0190] Example use case - 3: UE RL-based antenna panel selection

[0191] The embodiments may be applied to performing the terminal device antenna panel (or beam) selection optimization, where the objective is to determine an optimal transmission / reception antenna panel (or beam) per terminal device such that the link throughput (e.g., uplink or downlink throughput) of the terminal device is maximized. The RL agent may be deployed at the terminal device side to adjust the UE antenna panel (or beam) to be used based on local radio conditions (as quantified by values of one or more radio measurement metrics such as signal strength or RSRP). In these use cases, another RL agent (or other machine-learning model) may be deployed at the network entity to track the global performance in the cell related to terminal device panel selections. This RL architecture allows for collaboration with the access node and finetuning of the terminal device antenna selection, while maintaining terminal device autonomy. Thus, in these embodiments, the one or more RRM parameters may comprise a parameter defining a selection of at least one beam or at least one antenna panel of the terminal device (e.g., a panel index or a beam index) and the one or more KPIs may comprise a link throughput (e.g., an uplink or downlink throughput).

[0192] The one or more exploration event triggering conditions may be based at least on one or more mobility stability related (scalar) counters exceeding or falling below respective at least one of one or more first mobility stability thresholds. Similarly, the one or more exploration event exit conditions may be based at least on one or more mobility stability related(scalar) counters exceeding or falling below respective at least one of one or more second mobility stability thresholds. Here, the one or more mobility stability related (scalar) counter comprise, e.g., a counter for the number of terminal device panel switches. As the value of the one or more mobility stability related scalar counters (or at least one of them) exceeds a respective first mobility stability threshold (indicating that the radio environment is relatively unstable), the exploration event may be triggered. On the other hand, as the value of the one or more mobility stability related scalar counters (or at least one of them) falls below a respective second mobility stability threshold (indicating that the radio environment is relatively stable), the exploration event may be exited. As frequent terminal device antenna panel switches are risky for uplink / downlink data transmission according to the operation mode of “Assumption 1”, this necessitates a rather aggressive action searching. For this reason, the terminal device may not be able to support multiple panel concurrent transmission and receiving.

[0193] The one or more exploitation event triggering conditions may be based at least on a speed of the terminal device exceeding or falling below a first speed threshold. Similarly, the one or more exploitation event exit conditions may be based at least on a speed of the terminal device exceeding or falling below a second speed threshold. As the value of the speed of the terminal device falls below a first speed threshold (i.e., the terminal device is moving so slowly that the radio environment may be assumed to be relatively stable), the exploitation event may be triggered. On the other hand, as the value of the speed of the terminal device exceeds a respective second speed threshold (i.e., the terminal device is moving fast meaning that the radio environment may likely be unstable), the exploitation event may be exited.

[0194] The blocks, related functions, and information exchanges described above by means of FIGs. 3, 4 and 7 to 11 are in no absolute chronological order, and some of them may be performed simultaneously or in an order differing from the given one. Other functions can also be executed between them or within them, and other information may be sent, and / or other rules applied. Some of the blocks or part of the blocks or one or more pieces of information can also be left out or replaced by a corresponding block or part of the block or one or more pieces of information.

[0195] FIG. 12 provides an apparatus 1201 according to some embodiments. Specifically, FIG. 12 may illustrate an apparatus configured to carry out at least some of the functions described above. The apparatus 1201 may be or form a part of a terminal device. Theapparatus 1201 may be or form a part of a (non-distributed) access node or at least one of an RU, a DU or a CU of a distributed access node.

[0196] The apparatus 1201 may comprise one or more communication control circuitry 1220, such as at least one processor, and at least one memory 1230, including one or more algorithms 1231, such as a computer program code (software) wherein the at least one memory and the computer program code (software) are configured, with the at least one processor, to cause the apparatus 1201 to carry out any one of the exemplified functionalities of the apparatus described above in connection with any of FIGs. 3 to 11. Said at least one memory 1230 may also comprise at least one database 1232.

[0197] When the one or more communication control circuitry 1220 comprises more than one processor, the apparatus 1201 may be a distributed device wherein processing of tasks takes place in more than one physical unit. Each of the at least one processor may comprise one or more processor cores. A processing core may comprise, for example, a Cortex-A12 processing core manufactured by ARM Holdings or a Zen processing core designed by Advanced Micro Devices Corporation. The one or more control circuitry 1220 may comprise at least one Qualcomm Snapdragon and / or Intel Atom processor.

[0198] Referring to FIG. 12, the one or more communication control circuitry 1220 of the apparatus 1201 is configured to carry out functionalities described above by means of any of elements of FIGs. 2 to 11 using one or more individual circuitries. It may also be feasible to use specific integrated circuits, such as DSP block, digital signal processor, ASIC or FPGA, or other components and devices for implementing said functionalities in accordance with different embodiments.

[0199] Referring to FIG. 12, the apparatus 1201 may further comprise different interfaces 1210 such as one or more communication interfaces comprising hardware and / or software for realizing communication connectivity according to one or more communication protocols. For example, the one or more communication interfaces 1210 may comprise at least one interface enabling communication between the apparatus and one or more terminal devices (if the apparatus is a terminal device or a part thereof) or between the apparatus and one or more terminal device (if the apparatus is an access node or a part thereof or a terminal device for sidelink communication). Additionally, if the apparatus is an access node, the one or more communication interfaces 1210 may comprise at least one interface enabling communication between the apparatus and at least one core network node and / or at least one interface enablingcommunication between the apparatus and one or more (other) access nodes. If the apparatus 1201 forms a part of a distributed access node, the one or more communication interfaces 1210 may comprise a one or more interfaces providing one or more connections between the apparatus and other parts of the distributed access node.

[0200] Referring to FIG. 12, the memory 1230 may be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.

[0201] As used in this application, the term ‘circuitry’ may refer to one or more or all of the following: (a) hardware-only circuit implementations, such as implementations in only analog and / or digital circuitry, and (b) combinations of hardware circuits and software (and / or firmware), such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software, including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus, such as a terminal device or an access node, to perform various functions, and (c) hardware circuit(s) and processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software may not be present when it is not needed for operation. This definition of ‘circuitry’ applies to all uses of this term in this application, including any claims. As a further example, as used in this application, the term ‘circuitry’ also covers an implementation of merely a hardware circuit or processor (or multiple processors) or a portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware.

[0202] In an embodiment, at least some of the processes described in connection with FIGs. 2 to 11 may be carried out by an apparatus comprising corresponding means for carrying out at least some of the described processes. Some example means for carrying out the processes may include at least one of the following: detector, processor (including dual-core and multiple-core processors), digital signal processor, controller, receiver, transmitter, encoder, decoder, memory, register, multiply-accumulate (MAC) unit, delay element, RAM, ROM, software, firmware, display, user interface, display circuitry, user interface circuitry, user interface software, display software, circuit, filter (low-pass, high-pass, bandpass and / or bandstop), sensor, circuitry, inverter, capacitor, inductor, resistor, operational amplifier, diode and transistor. In some embodiments, at least some of the processes may be implemented usingdiscrete components. In an embodiment, at least some of the processes described in connection with FIGs. 2 to 11 may be carried out by an apparatus comprising corresponding hardware means for carrying out at least some of the described processes. Said hardware means may comprise at least one of: a multiplier, an adder, a MAC unit, a barrel shifter, a register, a shift register, a memory unit, a control logic, a clocking circuitry or a finite state machine.

[0203] According to an embodiment, there is provided an apparatus (e.g., a terminal device) comprising means for performing: maintaining, in at least one memory, a reinforcement learning, RL, model trained to determine one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; receiving, from a network entity, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model; and carrying out optimization of the one or more RRM parameters using the RL model according to the one or more RL strategies.

[0204] According to an embodiment, there is provided an apparatus (e.g., a network entity or node or an access node) comprising means for performing: transmitting, to a terminal device storing an RL model trained to evaluate one or more RRM parameters of the terminal device based on one or more radio measurement metrics, at least one configuration message comprising one or more RL strategies, wherein the one or more RL strategies comprise at least one of:- one or more RL event triggering conditions for triggering an event of the RL model, or- one or more RL event exit conditions for exiting an event of the RL model.

[0205] According to an embodiment, there is provided an apparatus (e.g., a terminal device) comprising means for performing: transmitting, to a network entity, a configuration request requesting configuration of one or more RL strategies for a reinforcement learning, RL, model;receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating; receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0206] According to an embodiment, there is provided an apparatus (e.g., a network entity or node or an access node) comprising means for performing: receiving, from a terminal device, a configuration request requesting configuration of one or more RL strategies for a reinforcement learning, RL, model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; receiving, from the terminal device, an evaluation report comprising results of evaluating of the one or more RL event conditions; determining whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event; and based on the determining, transmitting, to the terminal device, a positive or negative acknowledgement for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event.

[0207] According to an embodiment, there is provided an apparatus (e.g., a terminal device) comprising means for performing: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions according to the one or more RL strategies; and based on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

[0208] According to an embodiment, there is provided an apparatus (e.g., a network entity or node or an access node) comprising means for performing: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics; and transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event

[0209] Embodiments as described above may also be carried out, fully or at least in part, in the form of a computer process defined by a computer program or portions thereof. Embodiments of the methods described in connection with FIGs. 2 to 11 may be carried out by executing at least one portion of a computer program comprising corresponding instructions. The computer program may be provided as a computer readable medium comprising program instructions stored thereon or as a non-transitory computer readable medium comprising program instructions stored thereon. The computer program may be in source code form, object code form, or in some intermediate form, and it may be stored in some sort of carrier, which may be any entity or device capable of carrying the program. For example, the computer program may be stored on a computer program distribution medium readable by a computer or a processor. The computer program medium may be, for example but not limited to, a record medium, computer memory, read-only memory, electrical carrier signal, tele-communications signal, and software distribution package, for example. The computer program medium may be a non-transitory medium. Coding of software for carrying out the embodiments as shown and described is well within the scope of a person of ordinary skill in the art.

[0210] The term “non-transitory”, as used herein, is a limitation of the medium itself (that is, tangible, not a signal) as opposed to a limitation on data storage persistency (for example, RAM vs. ROM).

[0211] Reference throughout this specification to one embodiment or an embodiment means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present solution. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0212] As used herein, a plurality of items, structural elements, compositional elements, and / or materials may be presented in a common list for convenience. However, these lists should be construed as though each member of the list is individually identified as a separate and unique member. Thus, no individual member of such list should be construed as a de facto equivalent of any other member of the same list solely based on their presentation in a common group without indications to the contrary. In addition, various embodiments and example of the present solution may be referred to herein along with alternatives for the various components thereof. It is understood that such embodiments, examples, and alternatives are not to be construed as de facto equivalents of one another, but are to be considered as separate and autonomous representations of the present solution.

[0213] Even though embodiments have been described above with reference to examples according to the accompanying drawings, it is clear that the embodiments are not restricted thereto but can be modified in several ways within the scope of the appended claims. Therefore, all words and expressions should be interpreted broadly and they are intended to illustrate, not to restrict, the embodiment. It will be obvious to a person skilled in the art that, as technology advances, the inventive concept can be implemented in various ways. Further, it is clear to a person skilled in the art that the described embodiments may, but are not required to, be combined with other embodiments in various ways.INDUSTRIAL APPLICABILITY

[0214] At least some embodiments find industrial application in wireless communications.

Claims

CLAIMS1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting, to a network entity, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters of the apparatus based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating; receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

2. The apparatus of claim 1, wherein the performing of the at least one of the triggering or the exiting comprises:based on the reception of the positive acknowledgment and the results of the evaluating, triggering one of the exploration event or the exploitation event and exiting other of the exploration event or the exploitation event.

3. The apparatus of claim 1 or 2, wherein the at least one memory further storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: performing policy update on the RL model in response to the received positive acknowledgement.

4. The apparatus according to any preceding claim, wherein the at least one memory further storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting, to the network entity, a report comprising a plurality of values of the one or more RRM parameters determined, by the apparatus, using the RL model.

5. The apparatus according to any preceding claim, wherein the at least one memory further storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform, following the reception of the at least one configuration message and before the performing of the exploration and / or the exploitation, at least the following: initializing default RL model settings for the performing of the exploration and / or the exploitation.

6. The apparatus according to any preceding claim, wherein the at least one memory further storing instructions that, when executed by the at least one processor, cause the apparatus at least to, following the performing of the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or the exploitation event or following the reception of the negative acknowledgment, repeat the performing of the exploration and / or the exploitation, the evaluating, the transmitting of the evaluation report comprising results of the evaluating, the receiving of the positive or negative acknowledgement and, based on the reception of the positiveacknowledgment and on the results of the evaluating, the performing of the at least one of the triggering of an event and the exiting of an event.

7. The apparatus according to any preceding claim, wherein each of the one or more RL event conditions is defined based on at least one of: reference signal received power, RSRP, difference between a serving cell of the apparatus and a strongest neighbor cell of the apparatus, an average speed of the apparatus, one or more mobility stability related counters of the apparatus, a geolocation-based metric indicative of location or distance of the apparatus relative to a serving network entity, time, a 5G New Radio RRM measurement relaxation condition, one or more capabilities of the apparatus, or a quality of service, QoS, based metric of the apparatus.

8. The apparatus according to any preceding claim, wherein a state, an action and a reward of the RL model are defined as follows: the state defines one or more radio resource management, RRM, parameters, the action taken by an agent of the RL model from a given state is defined as a modification of at least one of the one or more RRM parameters, and the reward of taking a given action in a given state is calculated based on the one or more radio measurement metrics for optimizing values of one or more key performance indicators, KPIs, of the apparatus or of a serving cell.

9. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, from a terminal device, a configuration request requesting configuration of one or more reinforcement learning, RL, strategies for an RL model, wherein the RL model is for determining one or more radio resource management, RRM, parameters of the terminal device based on one or more radio measurement metrics;transmitting, to the terminal device, at least one configuration message comprising one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; receiving, from the terminal device, an evaluation report comprising results of evaluating of the one or more RL event conditions; determining whether the evaluation report is indicative of a need for at least one of triggering the exploration or exploitation event or exiting the exploration or exploitation event; and based on the determining, transmitting, to the terminal device, a positive or negative acknowledgement for, respectively, verifying or not verifying the at least one of the triggering of the exploration or exploitation event or the exiting of the exploration or exploitation event.

10. The apparatus of claim 9, wherein the at least one of the triggering or the exiting comprises: triggering one of the exploration event or the exploitation event and exiting other of the exploration event or the exploitation event.

11. The apparatus of claim 9 or 10, wherein the at least one memory further storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, from the terminal device, a report comprising a plurality of values of the one or more RRM parameters determined, by the terminal device, using the RL model.

12. The apparatus according to any preceding claim, wherein at least one of the at least one configuration message comprises information defining an evaluation timewindow or an evaluation timer, the evaluating being performed within the evaluation time window or before expiry of the evaluation timer.

13. The apparatus according to any preceding claim, wherein the configuration request and / or the at least one configuration message is a radio resource control, RRC, message.

14. The apparatus according to any preceding claim, wherein each of the evaluation report and / or the positive or negative acknowledgment is a layer- l / layer-2 message.

15. A method comprising: transmitting, to a network entity, a configuration request requesting configuration of one or more RL strategies for a reinforcement learning, RL, model; receiving, from the network entity, at least one configuration message comprising the one or more RL strategies comprising one or more RL event conditions, wherein the one or more RL event conditions comprise all or at least one of:- one or more exploration event triggering conditions for triggering an exploration event of the RL model,- one or more exploration event exit conditions for exiting the exploration event of the RL model,- one or more exploitation event triggering conditions for triggering an exploitation event for performing exploitation using the RL model, or- one or more exploitation event exit conditions for exiting the exploitation event; performing exploration and / or exploitation using the RL model for determining, using the RL model, one or more radio resource management, RRM, parameters based on one or more radio measurement metrics; evaluating the one or more RL event conditions; transmitting, to the network entity, an evaluation report comprising results of the evaluating; receiving, from the network entity, a positive or negative acknowledgement; and based on the reception of the positive acknowledgment and on the results of the evaluating, performing at least one of triggering the exploration or exploitation event or exiting the exploration or the exploitation event.

Citation Information

Patent Citations

  • Deep reinforcement learning exploration method and system based on extreme novelty search

    CN117150927A

  • Method and unit for radio resource management using reinforcement learning

    US20190239238A1

  • Methods and devices for determination of an update timescale for radio resource management algorithms

    US20240098575A1

  • A packet data unit session for machine learning exploration for wireless communication network optimization

    WO2022253414A1

  • Model request method, model request processing method and related device

    WO2023066288A1