Reward signal requirement parameter

By integrating delay requirement parameters for reward signals in AIML functionalities, the challenge of delayed feedback in RL-based AIML is addressed, resulting in improved performance and efficiency of wireless communication systems.

WO2026092949A1PCT designated stage Publication Date: 2026-05-07NOKIA TECHNOLOGIES OY
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2025-10-02
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing wireless communication systems face challenges in efficiently incorporating reinforcement learning (RL) due to the impact of delayed or invalid reward signals on artificial intelligence and machine learning (AIML) functionalities, particularly when stringent delay requirements are not met, leading to potential performance degradation in user equipment (UE) and network nodes.

Method used

Incorporating requirement parameters, such as delay requirements, for reward signals in RL-based AIML functionalities to ensure timely feedback, thereby enhancing the accuracy and performance of AIML models by ensuring that reward signals are received within specified time thresholds.

Benefits of technology

This approach improves the efficiency and effectiveness of AIML model training and operation by ensuring that reward signals are provided in a timely manner, thereby maintaining high-quality service and reducing resource wastage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025078342_07052026_PF_FP_ABST
    Figure EP2025078342_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A user device may transmit to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request may include a requirement parameter associated with the at least one reward signal of the supported AIML functionality. The user device may receive the at least one reward signal based on the requirement parameter.
Need to check novelty before this filing date? Find Prior Art

Description

REWARD SIGNAL REQUIREMENT PARAMETERTECHNICAL FIELD

[0001] This description relates to wireless communications.BACKGROUND

[0002] A communication system may be a facility that enables communication between two or more nodes or devices, such as fixed or mobile communication devices. Signals can be carried on wired or wireless carriers.

[0003] An example of a cellular communication system is an architecture that is being standardized by the 3rd Generation Partnership Project (3GPP). Long-term evolution (LTE) is referred to as 4G radio-access technology of the Universal Mobile Telecommunications System (UMTS). EUTRA (evolved UMTS Terrestrial Radio Access) is the air interface of 3GPP’s Long Term Evolution (LTE) upgrade path for mobile networks. In LTE, base stations or access points (APs), which are referred to as enhanced Node AP (eNBs), provide wireless access within a coverage area or cell. In LTE, mobile devices, or mobile stations are referred to as user equipments (UE). LTE has included a number of improvements or developments. Aspects of LTE are also continuing to improve.

[0004] 5G New Radio (NR) development is part of a continued mobile broadband evolution process to meet the requirements of 5G, similar to earlier evolution of 3G and 4G wireless networks. In addition, 5G is also targeted at the new emerging use cases in addition to mobile broadband. A goal of 5G is to provide significant improvement in wireless performance, which may include new levels of data rate, latency, reliability, and security. 5G NR may also scale to efficiently connect the massive Internet of Things (loT) and may offer new types of mission-critical services. For example, ultra-reliable and low-latency communications (URLLC) devices may require high reliability and very low latency. 6G and other networks are also being developed.SUMMARY

[0005] In some aspects, the techniques described herein relate to a method including: transmitting, by a user device to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request including a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receiving the at least one reward signal based on the requirement parameter.

[0006] In some aspects, the techniques described herein relate to a method including: receiving, by a network node from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request including a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmitting the at least one reward signal based on the requirement parameter.

[0007] In some aspects, the techniques described herein relate to an apparatus including: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request including a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receiving the at least one reward signal based on the requirement parameter.

[0008] In some aspects, the techniques described herein relate to an apparatus including: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request including a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmitting the at least one reward signal based on the requirement parameter.

[0009] In some aspects, the techniques described herein relate to a method including: transmitting, by a user device to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication including a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; transmitting a request for the at least one reward signal associated with the supported AIML model; and receiving the at least one reward signal based on the indicated requirement parameter.

[0010] In some aspects, the techniques described herein relate to a method including: receiving, by a network node from a user device, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication including a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; receiving a request for the at least one reward signal associated with the supported AIML model; and transmitting the at least one reward signal based on the indicated requirement parameter.

[0011] In some aspects, the techniques described herein relate to an apparatus including: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication including a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; transmitting a request for the at least one reward signal associated with the supported AIML model; and receiving the at least one reward signal based on the indicated requirement parameter.

[0012] In some aspects, the techniques described herein relate to an apparatus including: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication including a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; receiving a request for the at least one reward signal associated with the supported AIML model; and transmitting the at least one reward signal based on the indicated requirement parameter.

[0013] Other example embodiments are provided or described for each of the example methods, including: means for performing any of the example methods; a non-transitory computer- readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform any of the example methods; and an apparatus including at least one processor, and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform any of the example methods.

[0014] The details of one or more examples of embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] FIG. 1 is a block diagram of a wireless network 130.

[0016] FIG. 2 is a diagram illustrating functional framework for radio access network intelligence based on AIML.

[0017] FIG. 3 is a diagram illustrating an example embodiment.

[0018] FIG. 4 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment.

[0019] FIG. 5 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment.

[0020] FIG. 6 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment.

[0021] FIG. 7 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment.

[0022] FIG. 8 is a block diagram of a wireless station or node (e.g., UE, user device, AP, BS, eNB, gNB, RAN node, network node, TRP, or other node) 1300 according to an example embodiment.DETAILED DESCRIPTION

[0023] It shall be understood that although the terms “first,” “second,”. . ., etc., in front of noun(s) and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another and they do not limit the order of the noun(s). For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0024] As used herein, unless stated explicitly, performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.

[0025] FIG. 1 is a block diagram of a wireless network 130. In the wireless network 130 of FIG. 1 , user devices 131, 132, 133 and 135, which may also be referred to as mobile stations (MSs) or user equipment (UEs), may be connected (and in communication) with a base station (BS) 134, which may also be referred to as an access point (AP), an enhanced Node B (eNB), a gNB or a network node. The terms user device and user equipment (UE) may be used interchangeably. A BS may also include or may be referred to as a RAN (radio access network) node, and may include a portion of a BS or a portion of a RAN node, such as e.g., such as a centralized unit (CU) and / or a distributed unit (DU) in the case of a split BS or split gNB. At least part of the functionalities of a BS (e.g., access point (AP), base station (BS) or (e)Node B (eNB), gNB, RAN node) may also be carried out by any node, server or host which may be operably coupled to a transceiver, such as a remote radio head. BS (or AP) 134 provides wireless coverage within a cell 136, including to user devices (or UEs) 131 , 132, 133 and 135. Although only four user devices (or UEs) are shown as being connected or attached to BS 134, any number of user devices may be provided. BS 134 is alsoconnected to a core network 150 via a S1 interface 151 . This is merely one simple example of a wireless network, and others may be used.

[0026] A base station (e.g., such as BS 134) is an example of a radio access network (RAN) node within a wireless network. A BS (or a RAN node) may be or may include (or may alternatively be referred to as), e.g., an access point (AP), a gNB, an eNB, or portion thereof (such as a centralized unit (CU) and / or a distributed unit (DU) in the case of a split BS or split gNB), or other network node.

[0027] Some functionalities of the communication network may be carried out, at least partly, in a central / centralized unit, CU, (e.g., server, host or node) operationally coupled to distributed unit, DU, (e.g., a radio head / node). Thus, 5G networks architecture may be based on a so-called CU-DU split. The gNB-CU (central node) may control a plurality of spatially separated gNB-DUs, acting at least as transmit / receive (Tx / Rx) nodes. In some embodiments, however, the gNB-DUs (also called DU) may comprise e.g., a radio link control (RLC), medium access control (MAC) layer and a physical (PHY) layer, whereas the gNB-CU (also called a CU) may comprise the layers above RLC layer, such as a packet data convergence protocol (PDCP) layer, a radio resource control (RRC) and an internet protocol (IP) layer. Other functional splits are possible too.

[0028] According to an illustrative example, a BS node (e.g., BS, eNB, gNB, CU / DU, ...) or a radio access network (RAN) may be part of a mobile telecommunication system. A RAN (radio access network) may include one or more BSs or RAN nodes that implement a radio access technology, e.g., to allow one or more UEs to have access to a network or core network (CN). Thus, for example, the RAN (RAN nodes, such as BSs or gNBs) may reside between one or more user devices or UEs and a core network. According to an example embodiment, each RAN node (e.g., BS, eNB, gNB, CU / DU, ...) or BS may provide one or more wireless communication services for one or more UEs or user devices, e.g., to allow the UEs to have wireless access to a network, via the RAN node. Each RAN node or BS may perform or provide wireless communication services, e.g., such as allowing UEs or user devices to establish a wireless connection to the RAN node, and sending data to and / or receiving data from one or more of the UEs. For example, after establishing a connection to a UE, a RAN node or network node (e.g., BS, eNB, gNB, CU / DU, ...) may forward data to the UE that is received from a network or the core network, and / or forward data received from the UE to the network or core network. RAN nodes or network nodes (e.g., BS, eNB, gNB, CU / DU, . . .) may perform a wide variety of other wireless functions or services, e.g., such as broadcasting control information (e.g., such as system information or on-demand system information) to UEs, paging UEs when there is data to be delivered to the UE, assisting in handoverof a UE between cells, scheduling of resources for uplink data transmission from the UE(s) and downlink data transmission to UE(s), sending control information to configure one or more UEs, and the like. These are a few examples of one or more functions that a RAN node or BS may perform.

[0029] A user device or user node (user terminal, user equipment (UE), mobile terminal, handheld wireless device, etc.) may refer to a portable computing device that includes wireless mobile communication devices operating either with or without a subscriber identification module (SIM), including, but not limited to, the following types of devices: a mobile station (MS), a mobile phone, a cell phone, a smartphone, a personal digital assistant (PDA), a handset, a device using a wireless modem (alarm or measurement device, etc.), a laptop and / or touch screen computer, a tablet, a phablet, a game console, a notebook, a vehicle, a sensor, and a multimedia device, as examples, or any other wireless device. It should be appreciated that a user device may also be (or may include) a nearly exclusive uplink only device, of which an example is a camera or video camera loading images or video clips to a network. Also, a user node may include a user equipment (UE), a user device, a user terminal, a mobile terminal, a mobile station, a mobile node, a subscriber device, a subscriber node, a subscriber terminal, or other user node. For example, a user node may be used for wireless communications with one or more network nodes (e.g., gNB, eNB, BS, AP, CU, DU, CU / DU) and / or with one or more other user nodes, regardless of the technology or radio access technology (RAT). In LTE (as an illustrative example), core network 150 may be referred to as Evolved Packet Core (EPC), which may include a mobility management entity (MME) which may handle or assist with mobility / handover of user devices between BSs, one or more gateways that may forward data and control signals between the BSs and packet data networks or the Internet, and other control functions or blocks. Other types of wireless networks, such as 5G (which may be referred to as New Radio (NR)) may also include a core network.

[0030] In addition, the techniques described herein may be applied to various types of user devices or data service types, or may apply to user devices that may have multiple applications running thereon that may be of different data service types. New Radio (5G) development may support a number of different applications or a number of different data service types, such as for example: machine type communications (MTC), enhanced machine type communication (eMTC), Internet of Things (loT), and / or narrowband loT user devices, enhanced mobile broadband (eMBB), and ultra-reliable and low-latency communications (URLLC). Many of these new 5G (NR) - related applications may require generally higher performance than previous wireless networks.

[0031] loT may refer to an ever-growing group of objects that may have Internet or network connectivity, so that these objects may send information to and receive information from othernetwork devices. For example, many sensor type applications or devices may monitor a physical condition or a status and may send a report to a server or other network device, e.g., when an event occurs. Machine Type Communications (MTC, or Machine to Machine communications) may, for example, be characterized by fully automatic data generation, exchange, processing and actuation among intelligent machines, with or without intervention of humans. Enhanced mobile broadband (eMBB) may support much higher data rates than currently available in LTE.

[0032] Ultra-reliable and low-latency communications (URLLC) is a new data service type, or new usage scenario, which may be supported for New Radio (5G) systems. This enables emerging new applications and services, such as industrial automations, autonomous driving, vehicular safety, e-health services, and so on. 3GPP targets in providing connectivity with reliability corresponding to block error rate (BLER) of 10-5 and up to 1 ms U-Plane (user / data plane) latency, by way of illustrative example. Thus, for example, URLLC user devices / UEs may require a significantly lower block error rate than other types of user devices / UEs as well as low latency (with or without requirement for simultaneous high reliability). Thus, for example, a URLLC UE (or URLLC application on a UE) may require much shorter latency, as compared to an eMBB UE (or an eMBB application running on a UE).

[0033] The techniques described herein may be applied to a wide variety of wireless technologies or wireless networks, such as 5G (New Radio (NR)), cmWave, and / or mmWave band networks, loT, MTC, eMTC, eMBB, URLLC, 6G, etc., or any other wireless network or wireless technology. These example networks, technologies or data service types are provided only as illustrative examples.

[0034] A user device (or UE) may measure various signals and may transmit one or more measurement reports to the network. For example, a UE may measure reference signals received from one or more network nodes (e.g., gNBs or DUs), including channel state informationreference signals (CSI-RSs) and / or synchronization signal block (SSB) reference signals, demodulation references signals, and / or other reference signals. Based on received reference signals, the UE may measure various signal parameters, e.g., such as reference signal received power (RSRP), reference signal received quality (RSRQ), signal to interference plus noise ratio (SINR), received signal strength indicator (RSSI), or other signal parameter.

[0035] The PHY (physical) layer may refer to layer 1 (L1) and MAC (media access control) may refer to layer 2 (L2). RSRP, RSRQ, SINR and RSSI are signal quantities measured at layer 1 (L1). The UE may send L1 measurement reports (e.g., CSI-RS reports, which include measurements of one or more signal parameters for one or more cells) to a gNB, source DU or serving cell. These L1measurement reports may be sent periodically, for example, or aperiodically. L1 / L2 measurement reports may include no averaging or filtering of measurement values or may include less averaging or filtering than what is performed for L3 measurement reports. L1 (or L1 / L2) measurement reports may be transmitted by a UE to a serving network node or source DU and may cause the network node to trigger or initiate a L1 / L2 triggered mobility (LTM) handover of the UE to another cell. L1 measurements (e.g., RSRP RSRQ, RSSI) may be provided or reported periodically to the DU (MAC / PHY).

[0036] A machine learning (ML) model may be used within a wireless network to perform (or assist with performing) one or more tasks. In general, one or more nodes (e.g., BS, gNB, eNB, RAN node, user node, UE, user device, relay node, or other wireless node) within a wireless network may use or employ a ML model, e.g., such as, for example a neural network model (e.g., which may be referred to as a neural network, an artificial intelligence (Al) neural network, an Al neural network model, an Al model, a machine learning (ML) model or algorithm, a model, or other term) to perform, or assist in performing, one or more ML-enabled tasks. Other types of models may also be used. A ML-enabled task may include tasks that may be performed (or assisted in performing) by a ML model, or a task for which a ML model has been trained to perform or assist in performing).

[0037] ML-based algorithms or ML models may be used to perform and / or assist with performing a variety of wireless and / or radio resource management (RRM) and / or RAN-related functions or tasks to improve network performance, such as, e.g., in the UE for beam prediction (e.g., predicting a best beam or best beam pair based on measured reference signals), antenna panel or beam control, RRM (radio resource measurement) measurements and feedback (channel state information (CSI) feedback), link monitoring, Transmit Power Control (TPC), etc. In some cases, ML models may be used to improve performance of a wireless network in one or more aspects or as measured by one or more performance indicators or performance criteria.

[0038] Models (e.g., neural networks or ML models) may be or may include, for example, computational models used in machine learning made up of nodes organized in layers. The nodes are also referred to as artificial neurons, or simply neurons, and perform a function on provided input to produce some output value. A neural network or ML model may typically require a training period to learn the parameters, i.e., weights, used to map the input to a desired output. The mapping may occur via the function that is learned from a given data for the problem in question. Thus, the weights are weights for the mapping function of the neural network. Each neural network model or ML model may be trained for a particular task.

[0039] To provide the output given the input, the ML functionality of a neural network model or ML model should be trained, which may involve learning the proper value for a large number of parameters (e.g., weights and / or biases) for the mapping function (or of the ML functionality of the ML model). For example, the parameters may be used to weight and / or adjust terms in the mapping function. This training may be an iterative process, with the values of the weights and / or biases being tweaked over many (e.g., tens, hundreds and / or thousands) of rounds of training episodes or training iterations until arriving at the optimal, or most accurate, values (or weights and / or biases). In the context of neural networks (neural network models) or ML models, the parameters may be initialized, often with random values, and a training optimizer iteratively updates the parameters (e.g., weights) of the neural network to minimize error in the mapping function. In other words, during each round, or step, of iterative training the network updates the values of the parameters so that the values of the parameters eventually converge to the optimal values.

[0040] ML models may be trained in either a supervised or unsupervised manner, as examples. In supervised learning, training examples are provided to the ML model or other machine learning algorithm. A training example includes the inputs and a desired or previously observed output. Training examples are also referred to as labeled data because the input is labeled with the desired or observed output. In the case of a neural network (which may be a specific case of ML model), the network (or ML model) learns the values for the weights used in the mapping function or ML functionality of the ML model that most often result in the desired output when given the training inputs. In unsupervised training, the ML model learns to identify a structure or pattern in the provided input. In other words, the model identifies implicit relationships in the data. Unsupervised learning is used in many machine learning problems and typically requires a large set of unlabeled data.

[0041] According to an example embodiment, a ML model may be classified into (or may include) two broad categories (supervised and unsupervised), depending on whether there is a learning “signal” or “feedback” available to a model. Thus, for example, within the field of machine learning, there may be two main types of learning or training of a model: supervised, and unsupervised. The main difference between the two types is that supervised learning is done using known or prior knowledge of what the output values for certain samples of data should be. Therefore, a goal of supervised learning may be to learn a function that, given a sample of data and desired outputs, best approximates the relationship between input and output observable in the data. Unsupervised learning, on the other hand, does not have labeled outputs, so its goal is to infer the natural structure present within a set of data points.

[0042] Supervised learning: The computer is presented with example inputs and their desired outputs, and the goal may be to learn a general rule that maps inputs to outputs. Supervised learning may, for example, be performed in the context of classification, where a computer or learning algorithm attempts to map input to output labels, or regression, where the computer or algorithm may map input(s) to a continuous output(s). Common algorithms in supervised learning may include, e.g., logistic regression, naive Bayes, support vector machines, artificial neural networks, and random forests. In both regression and classification, a goal may include finding specific relationships or structure in the input data that allow us to effectively produce correct output data. In some example cases, the input signal may be only partially available, or restricted to special feedback. Semisupervised learning: the computer may be given only an incomplete training signal; a training set with some (often many) of the target outputs missing. Active learning: the computer can only obtain training labels for a limited set of instances (based on a budget), and also may optimize its choice of objects for which to acquire labels. When used interactively, these can be presented to the user for labeling.

[0043] Unsupervised learning: No labels are given to the learning algorithm, leaving it on its own to find structure in its input. Some example tasks within unsupervised learning may include clustering, representation learning, and density estimation. In these cases, the computer or learning algorithm is attempting to learn the inherent structure of the data without using explicitly-provided labels. Some common algorithms include k-means clustering, principal component analysis, and auto-encoders. Since no labels are provided, there may be no specific way to compare model performance in most unsupervised learning methods.

[0044] Continual Learning (CL) may refer to or may include a capability of the ML model to adapt to ever-changing (or continuously changing, or periodically changing) surrounding environment or data by learning or adapting the ML model continually based on incoming data (or new or updated data), e.g., without forgetting original or previous knowledge or ML model settings, and, e.g., which may be based on less than a full or complete set of data. For example, given a (e.g., potentially unlimited or continuous) stream of data (e.g., data reflecting changing or updated conditions or environment upon which the ML model should be updated), a continual learning (CL) algorithm may (or should) learn, e.g., by updating or adapting weights or other parameters of the ML model, based on a sequence of partial experiences or partial data (e.g., a most recent set of data) where all data may not be available at once, since new or updated data will be received later (thus, the new data potentially renders the weights or parameter settings of the ML model obsolete or inaccurate). Thus, a full or complete set of data may not be considered available at that time of ML model updating oradaptation, since the data or environment may be continuously or continually changing over time. Thus, at any given point or moment in time, data (upon which the ML model may be updated or adapted) may be considered incomplete because there may be a continuous stream of data. Thus, a CL algorithm may include or may refer to iteratively updating or adapting weights or other parameters of the ML model based on an updated set of data, and then repeating the learning or adaptation process for the ML model when a second (or later) set of updated data is received subsequently.

[0045] Reinforcement learning (RL) may include may be or may include an interdisciplinary area of machine learning and optimal control concerned with how an intelligent agent should perform actions in a dynamic environment in order to maximize a reward. RL may be a goal-based optimization approach where an agent performs an action (e.g., an action performed by a ML model) based on the observed state / context (or inputs) and then receives a reward to learn the optimal policy or train the ML model. To distinguish the good (or preferred) actions from the bad (or non-preferred) actions, the agent may explore the action space (e.g., performing various actions) by performing various actions and observing or obtaining a reward (feedback that may be used to train the ML model). Due to its ability to optimize radio functions, e.g., such as various radio resource management (RRM) functions, based on a reward, RL may be used in future wireless networks. There are many radio functions for which a ML model and / or RL may be used to assist and / or improve performance of the radio function, e.g., such as beam selection (or beam management), power control, assisting in performing handovers, and many others.

[0046] In an example, the ML based algorithm may include an artificial intelligence and / or machine learning (AIML) algorithm.

[0047] In an example, artificial intelligence and / or machine learning (AIML) techniques may be implemented to improve the performance of wireless communication systems. The implementation of the AIML may include implementation of mechanisms at the network side and the UE side. For example, the AIML techniques may enhance data collection for NR and dual connectivity scenarios. The AIML model, herein, is exchangeable with Al and ML model, Al or ML model, Al model, or ML model.

[0048] In an example, artificial intelligence and machine learning (AIML) based methods may be employed for enhancement of positioning accuracy. For example, positioning accuracy enhancements may include direct AIML positioning, or AIML assisted positioning. For example, for positioning enhancement, training data may be generated by a UE, a gNB, a location management function (LMF), and / or the like for model training. For example, for a LMF-side model inference, input data may be generated by the UE or the gNB and terminated at the LMF. In an example, for a gNB-side model inference, input data may be internally available at the gNB. In an example, for UE-side model inference, input data may be internally available at the UE. In an example, for performance monitoring at the LMF side, calculated performance metrics (if needed) or data needed for performance metric calculation (if needed) may be generated by the UE or the gNB and terminated at LMF. In an example, for performance monitoring at the gNB side, calculated performance metrics (if needed) or data needed for performance metric calculation (if needed) may be generated by at least the gNB.

[0049] In an example, artificial intelligence and machine learning (AIML) based methods for beam management may be employed in wireless communication systems. In an example, the beam management may include spatial domain beam prediction (e.g., beam management BM-Case1) and time domain beam prediction (e.g., BM-Case2). In an example, the scope of spatial beam prediction (BM-Case1) may be to predict the best TX / RX beams in different spatial locations. In an example, time-domain beam predictions (BM-Case2) may include methods to predict the most likely beam to use for next time instants.

[0050] FIG. 2 is a diagram illustrating functional framework for radio access network intelligence based on AIML. In an example, data collection may be a function that provides input data to model training and model inference functions. AIML algorithm specific data preparation (e.g., data pre-processing and cleaning, formatting, and transformation) may not be carried out in the data collection function. In an example, input data may include measurements from UEs or different network entities, feedback from actor, output from an AIML model, and / or the like. In an example, training data may include the data needed as input for the AIML model training function.

[0051] In an example, inference data may include the data needed as input for the AIML model inference function. In an example, model training may be a function that performs the AIML model training, validation, and testing which may generate model performance metrics as part of the model testing procedure. The model training function may perform data preparation (e.g., data preprocessing and cleaning, formatting, and transformation) based on training data delivered by a data collection function. In an example, model deployment / update may be employed to initially deploy a trained, validated, and tested AIML model to the model inference function or to deliver / provide an updated model to the model inference function. In an example, the model inference may be a function that provides AIML model inference output (e.g., predictions or decisions). In an example, the model inference function may provide model performance feedback to model training function. The model inference function may perform data preparation (e.g., data pre-processing and cleaning, formatting, and transformation) based on inference data delivered / provided by a data collectionfunction. In an example, the output may include the inference output of the AIML model produced by a model inference function. In an example, the actor may include a function that receives the output from the model inference function and triggers or performs corresponding actions. The actor may trigger actions directed to other entities (of the AIML model, network, and / or the like) or to itself. In an example, the feedback may include information that may be needed to derive training data, inference data or to monitor the performance of the AIML model and its impact to the network through updating of key performance indicators (KPIs) and performance counters.

[0052] In an example, reinforcement learning (RL) may include a process of training an AI / ML model from input (a.k.a. state) and a feedback signal (a.k.a. reward signal) resulting from the model’s output (a.k.a. action) in an environment the model is interacting with. RL may involve / include a model (usually referred to as an agent) that makes decisions - referred to as actions, in an environment, based on its observations of the environment - referred to as the state of the environment. The environment state may consist of several features depending on the task the agent is tackling. For example, for an RL agent that makes handover decisions, channel measurements from candidate cells or beams may be a feature that the agent may take as an input. The decisions of the agent may be based on the agent’s policy, e.g., represented by a deep neural network, which is trained with the use of the reward signal provided upon each action. Reward signal may indicate how good (or bad) the selected action was, for example reflecting the impacted QoS during handover.

[0053] After training the RL agent, which can be conducted offline, e.g., without actually using the actions of the agent in the network, the agent may be deployed in the network for real-time inference. During inference, the agent only requires the state information as input (e.g., no reward required), and outputs the selected actions. Yet, it is also possible to further train an agent, e.g., for fine tuning purposes in a new environment, or monitor the performance of the model, which then again requires the reward signal.

[0054] In an example, a reward signal may be a feedback mechanism used in reinforcement learning that may indicate the success or failure of an agent’s actions. The reward signal may assist the agent to understand which behaviors lead to positive outcomes and which do not, guiding the learning process.

[0055] RL models may require a reward signal for training. For example, the reward signal may indicate how good or bad is the action taken by the model at each time period or instance. For example, when the RL is performed on the UE side, e.g., UE-side models, the network may provide the reward signal to the UE, as the network may be in a better position to evaluate the actions taken by the UE.

[0056] In existing technologies, a problem may arise when RL is employed with stringent delay requirements for a reward signal. For example, a challenge for RL may be the impact of validity of a reward signal on the performance of the AIML functionality and / or the AIML model. For example, in some AIML functionalities, when RL is performed on the UE side, a reward signal may be rendered invalid (or less effective) if received after an extended period of time beyond an acceptable delay threshold. For example, a reward signal that is received with significant delay may not provide accurate information for the AIML model or AIML functionality and corresponding algorithms. In some cases, the UE and / or the gNB performance may be negatively impacted if a bad or poor (or nonpreferred) action is selected based on the (missing or delayed) reward signal by the agent or AIML model and performed by the UE or network (e.g., gNB) as part of exploration. For example, if a bad or non-preferred action is selected and performed by the UE or the gNB, the UE may experience lower QoS and / or the gNB may experience a higher resource usage due to transmission errors.

[0057] In other words, each RL model may have a different requirement in terms of a reward signal. In an example, the reward signal needs to be received within a certain time after an action takes place, depending on its use, e.g., methodology used for training the AIML model. As an example, if a training is conducted in which the AIML model is updated after each and every action, then the reward signal should be provided immediately after an action takes place and before a next action is performed. As a result, the reward signal may be time sensitive and should be received within a required time period. For example, a delay between performing an action and receiving the corresponding reward signal may not exceed a threshold. In an example implementation, one or more actions may be performed frequently (such as a resource allocation decision) every few milliseconds. If an AIML model trained for such a use case needs to be updated frequently, then the reward signal should be received within the few milliseconds of time. Additionally, or alternatively, if a training or monitoring is conducted in batches, e.g., the AIML model gets updated only after performing a sequence of actions, then the corresponding set of reward signals may be provided with a less restricted delay requirement, e.g., depending on the length of the batch. In another example, in a case of an offline training, delay requirement of the reward signal may be less restrictive (less stringent) since the actions may not be required to be taken in real-time.

[0058] It is therefore advantageous to incorporate requirement parameters (such as delay requirements) of the reward signal to ensure accuracy and better performance of the RL. By incorporating the requirement parameters of the reward signal, RL algorithms and AIML models / functionalities may result in a better performance and / or outcome.

[0059] Example embodiments are directed to incorporation of requirement parameters associated with a reward signal or at least one reward signal. The at least one reward signal may be associated with an AIML functionality and / or an AIML model. In an example, a UE may transmit to a network node (e.g., a base station, gNB, a core network node, and / or the like), a request (or at least one request) for a reward signal (or at least one reward signal) associated with a supported artificial intelligence and machine learning (AIML) functionality. In an example, the request may include a requirement parameter associated with the at least one reward signal of the supported AIML functionality. In an example, the UE may receive the at least one reward signal based on the requirement parameter. In an example, the request may include an identifier of an AIML model associated with the supported AIML functionality.

[0060] Therefore, when example embodiments are implemented, the requirement parameter (such as a delay requirement) of the reward signal may be incorporated into operations pertaining the AIML functionalities (at the UE side, the RAN side, core network side, and / or the like) for RL based AIML functionalities or AIML models. By incorporating the requirement parameter, training of the AIML functionalities or AIML models may be performed more efficiently and may result in a better performance of the system.

[0061] In an example embodiment, the UE may transmit to the network node, the indication of support for one or more artificial intelligence and machine learning (AIML) models based on at least one of a medium access control (MAC), a radio resource control (RRC), a non-access stratum (NAS), and / or the like. The indication may be a capability indication information element (IE). The network node may be a base station, an eNB, a gNB, a gNB-CU, a gNB-DU, a core network node, an access and mobility management function (AMF), a location management function (LMF), a network data analytics function (NWDAF), and / or the like. In an example, transmission from the UE to the core network node may be via the base station (or gNB, NG-RAN, and / or the like).

[0062] In an example embodiment, the requirement parameter associated with the at least one reward signal may include a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality. For example, the time parameter may include at least one of a time value, a minimum time value, a maximum time value, a range of time values, and / or the like. In an example, the requirement parameter may include at least one of a period associated with (transmission or reception of) the at least one reward signal of the supported AIML functionality, a starting time associated with the at least one reward signal of the supported AIML functionality, an ending time associated with the at least one reward signal of the supported AIML functionality, a time duration associated with the at least one reward signal of the supported AIMLfunctionality, a triggering condition associated with the at least one reward signal of the supported AIML functionality, a total number of the at least one reward signal (or reward signals), a batch size of the at least one reward signal, wherein the batch size corresponds to a number of instances of each reward signal, and / or the like.

[0063] In an example, the UE may transmit to the network node, an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities. For example, the indication may include one or more supported AIML functionalities, such as identifiers of the one or more AIML functionalities.

[0064] In an example, the one or more AIML functionalities may include at least one of one or more functionalities related to beam management, one or more functionalities related to beamforming, one or more functionalities related to positioning, one or more functionalities related to allocation of radio resources, one or more functionalities related to scheduling, one or more functionalities related to power control, one or more functionalities related to link adaptation, one or more functionalities related to mobility, one or more functionalities related to selection of a modulation and coding (MCS) scheme, and / or the like.

[0065] In an example, the UE may receive from the network node, a request to activate one of the one or more AIML functionalities. In an example, the UE may select the one of the one or more AIML functionalities. In an example, the UE may transmit an indication of the selected one of the one or more AIML functionalities.

[0066] In an example, the supported AIML functionality may be associated with an AIML model. For example, the AIML model may employ an AIML algorithm that is based on a reinforcement learning (RL).

[0067] In an example, the UE may transmit to the network node, an outcome of the supported AIML functionality. In an example, the UE may receive the at least one reward signal based on the outcome of the supported AIML functionality. For example, the outcome may indicate a success or failure, or whether a minimum performance requirement has been met. In an example, the outcome may correspond to an output of the AIML model, e.g., the actions of the RL. For example, the actions may include determining a selected anchor for a position estimate, a selected cell for handover, a selected resource for transmission, and / or the like. For example, the outcome may impact a quality of service (QoS) of a service. In an example, the outcome may relate to determining an intermediate metric towards achieving a goal or an objective of the service. For example, in case of a positioning estimation, the outcome may include higher / lower accuracy of the positioning estimate or an intermediate feature, e.g., line of sight (LOS) or non-line of sight ( / NLOS) indicator. In anotherexample, for a case of measurement prediction, beam prediction, and / or the like, the outcome may impact an accuracy of the prediction. For example, in a handover scenario, the outcome may impact success / failure of a handover. As another example, for resource allocation, the impact of the outcome may be reflected in terms of increased / degraded QoS, e.g., lower throughput, or the like.

[0068] In an example, the UE may receive from the network node, an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality. In other words, the network node may indicate to the UE that the network may not be able to transmit the at least one reward signal within the required time or periodicity. This may be for example due to current network load, number of users, and / or the like.

[0069] In an example, the UE may receive from the network node, an indication of deactivating a first supported AIML functionality. For example, the UE may receive from the network node, an indication to activate a second supported AIML functionality (e.g., instead of the deactivated first supported AIML functionality). In an example, the UE may determine to deactivate the first supported AIML functionality. Then, the UE may transmit to the network node, an indication of deactivating the first supported AIML functionality. In an example, the UE may determine to activate the second supported AIML functionality. The UE may then transmit an indication of activating the second supported AIML functionality.

[0070] In an example, the UE may receive from the network node, an indication of suspending (or pausing) the supported AIML functionality. For example, the UE may receive from the network node, an indication to resume the supported AIML functionality. In an example, the UE may determine to suspend (or pause) the supported AIML functionality. Then, the UE may transmit to the network node, an indication of suspending (or pausing) the supported AIML functionality. In an example, the UE may determine to resume the supported AIML functionality. The UE may then transmit an indication of resuming the supported AIML functionality.

[0071] In an example, the UE may determine to suspend (or pause) receiving (from the network node) of the at least one reward signal. Then, the UE may transmit to the network node, a request to suspend (or pause) transmission of the at least one reward signal. Then the network node may suspend (or pause) transmitting the at least one reward signal to the UE. In an example, the UE may determine to resume receiving the at least one reward signal (from the network node). The UE may then transmit a request to the network node indicating resume of transmission (by the network node) of the at least one reward signal. Then the network node may resume transmitting the at least one reward signal to the UE.

[0072] In an example, the UE may transmit an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality. For example, the update may be based on changes of UE battery condition, UE mobility, and / or the like.

[0073] In an example, the UE may transmit to the network node, the request for the at least one reward signal. The transmitting may be based on (or in response to) at least one of determining that an outcome of an AIML functionality (or the supported AIML functionality) does not meet a performance requirement, a triggering event associated with the supported AIML functionality, and / or the like.

[0074] In an example, the UE may employ the received reward signals e.g., the at least one reward signal, to perform at least one of a training, re-training, monitoring, or update for a reinforcement learning model (e.g., based on the at least one reward signal).

[0075] In an example, the UE may transmit a first request indicating start of transmitting reward signals. For example, the first request may include a time interval between transmission of reward signals. In an example, the UE may transmit a second request indicating stop of transmitting reward signals. In other words, the first request and / or the second request may be a first command (to the network node) indicating start of transmitting reward signals, or a second command indicating stop of transmitting reward signals. In another example, the UE may transmit a first request indicating start of transmitting reward signals for a time duration. For example, the first request may include a value of the time duration.

[0076] In an example, each reward signal of the at least one reward signal may be associated with a corresponding action of the AIML functionality and a corresponding requirement parameter. In other words, for each single action performed by an AIML functionality, a single reward signal may be required. In an example, the UE may request a batch of reward signals, after performing a batch of actions. For example, for a batch length of 5, the corresponding request for the at least one reward signal may may include a list such as [r1 ,r2,r2,r4,r5] corresponding to [a1 ,a2,a3,a4,a5], where rk is reward signal k, and ak is action k. Therefore, the request for the at least one reward signal may include (indication of) one or more reward signals.

[0077] In an example, the request for the at least one reward signal may include one request for the at least one reward signal per AIML functionality. For example, the UE may send the reward request only once, e.g., which indicates the batch size, likely with a fixed delay requirement that would apply to every reward message for the corresponding AIML functionality.

[0078] In an example, the request for the at least one reward signal may include one request for one reward signal. For example, the one reward signal may correspond to one action of the AIMLfunctionality. In other words, the UE may send the request for the at least one reward signal after every action that is taken by the UE, e.g., where the UE may dynamically update and indicate the delay requirement for each reward signal per action.

[0079] In an example, the request for the at least one reward signal may include one request for at least one reward signal, wherein each reward signal of the at least one reward signal may correspond to at least one action of the AIML functionality. In other words, the request for the at least one reward signal may include one request per a number of reward signals [r_t, r_t+1 , ...] corresponding to multiple actions [a_t, a_t+1 , ...] of an AIML functionality / model. For example, the UE may send the request after a batch of actions it takes, e.g., after performing multiple actions, and the UE may dynamically adjust the delay requirement for each batch.

[0080] In an example, the request for the at least one reward signal may include a request for one or more instances of the at least one reward signal. For example, the one or more instances may correspond to one or more actions performed in the past. In an example, the request for the at least one reward signal may include a request for a list of reward signals ordered based on a time of the one or more actions. In other words, based on the reward signal requirement, provided message may include one or more instances of reward signal corresponding to one or more actions taken in the past, e.g., in batches. The reward signals within a message may be ordered with respect to the order of the taken actions, or may be associated with actions based on timestamps of the actions or identifiers of actions, e.g., transaction ID of a message containing an action signaling. In an example, the receiving the at least one reward signal may include receiving at least one of: at least one reward signal per AIML model / functionality, one reward signal, wherein the one reward signal corresponds to one action of the AIML model, or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model / functionality. In an example, the receiving the at least one reward signal may include receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past, or a list of reward signals ordered based on a time of the one or more actions.

[0081] Example embodiments are also directed to indication of requirement parameters associated with a reward signal or at least one reward signal. The indication may be transmitted via initial messaging exchange of the UE with the network. In an example embodiment, a UE may transmit to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models. In an example, the indication may include a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIMLmodels. In an example, the UE may transit to the network node, a request for the at least one reward signal associated with the supported AIML model. In an example, the request for the at least one reward signal may include the requirement parameter. The UE may receive from the network node, the at least one reward signal based on the indicated requirement parameter.

[0082] Therefore, when example embodiments are implemented, the outcome of a RL algorithm may be improved. The improvements may be achieved by incorporating the requirement parameter (such as a delay requirement) of the reward signal into operations pertaining the AIML models (at the UE side and / or the RAN side) for RL based AIML functionalities or AIML models. Therefore, training of the AIML functionalities or AIML models may be performed more efficiently, and result in a better performance of the system.

[0083] In an example embodiment, the UE may transmit to the network node, the indication of support for one or more artificial intelligence and machine learning (AIML) models via a radio resource control (RRC) message, a non-access stratum (NAS) message, and / or the like. The indication may be a capability indication information element (IE). The network node may be a base station, an eNB, a gNB, a gNB-CU, a gNB-DU, a core network node, an access and mobility management function (AMF), a location management function (LMF), a network data analytics function (NWDAF), and / or the like. In an example, transmission from the UE to the core network node may be via the base station (or gNB, NG-RAN, and / or the like).

[0084] In an example embodiment, the requirement parameter associated with the at least one reward signal may include a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML model. For example, the time parameter may include at least one of a time value, a minimum time value, a maximum time value, a range of time values, and / or the like. In an example, the requirement parameter may include at least one of a period associated with (transmission or reception of) the at least one reward signal of the supported AIML model, a starting time associated with the at least one reward signal of the supported AIML model, an ending time associated with the at least one reward signal of the supported AIML model, a time duration associated with the at least one reward signal of the supported AIML model, a triggering condition associated with the at least one reward signal of the supported AIML model, a total number of the at least one reward signal (or reward signals), a batch size of the at least one reward signal, wherein the batch size corresponds to a number of instances of each reward signal, and / or the like.

[0085] In an example, the UE may transmit to the network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models. For example, the indicationmay include one or more supported AIML models, such as identifiers of the one or more AIML models.

[0086] In an example, the UE may receive from the network node, the at least one reward signal, in response to (or based on) transmitting the request for the at least one reward signal. In an example, the UE may determine that a requirement indicated by the requirement parameter may not be met. For example, the at least one reward signal may arrive with a delay beyond a threshold value. In an example, the UE may determine to discard the at least one rewards signal based on the determining. In an example, the UE may send to the network node an indication of selecting a different AIML model. In another example, the UE may determine to update the requirement parameters, in another example, the UE may revert to a default action such as determining not to use an AIML model or an AIML functionality.

[0087] In an example, the request for the at least one reward signal may include the requirement parameter. In an example, the request may include an identifier of the supported AIML model.

[0088] In an example, the UE may receive an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML model. For example, the network node may indicate to the UE that at least one requirement of the requirement parameter may not be met, e.g., due to a power saving state of the network node, a load condition of the network node, and / or the like.

[0089] In an example, the UE may receive a request to activate one of the one or more AIML models. In an example, the UE may select the one of the one or more AIML models. In an example, the UE may transmit to the network node, an indication of the selected one of the one or more AIML models.

[0090] In an example, the UE may transmit an update of the requirement parameter associated with the at least one reward signal of the supported AIML model. For example, the update may be based on changes of UE battery level, power saving state, UE mobility, an accuracy requirement of an AIML functionality, and / or the like.

[0091] In an example, the indication of support for the one or more artificial intelligence and machine learning (AIML) models may include identifiers of the one or more AIML models. In an example, the one or more AIML models may employ or use AIML algorithms that are based on a reinforcement learning (RL).

[0092] In an example, the transmitting the request for the at least one reward signal may be based on at least one of a determining that an outcome of an AIML functionality (or the supportedAIML functionality) does not meet a performance requirement, a triggering event associated with an AIML functionality (or the supported AIML functionality), and / or the like.

[0093] In an example, the UE may employ the received reward signals e.g., the at least one reward signal, to perform at least one of a training, re-training, monitoring, or update for a reinforcement learning model (e.g., based on the at least one reward signal).

[0094] In an example, the UE may transmit a first request indicating start of transmitting reward signals. For example, the first request may include a time interval between transmission of reward signals. In an example, the UE may transmit a second request indicating stop of transmitting reward signals. In other words, the first request and / or the second request may be a first command (to the network node) indicating start of transmitting reward signals, or a second command indicating stop of transmitting reward signals. In another example, the UE may transmit a first request indicating start of transmitting reward signals for a time duration. For example, the first request may include a value of the time duration.

[0095] In an example, the UE may receive an indication of deactivating a first supported AIML model. In an example, the indication may also indicate activating a second supported AIML model. In another example, the UE may receive another indication indicating activating the second supported AIML model.

[0096] In an example, the UE may determine to deactivate a first supported AIML model and the UE may then transmit to the network node, an indication of deactivating the first supported AIML model. In an example, the UE may determine to activate a second supported AIML model and the UE may then transmit to the network node, an indication of activating the second supported AIML model.

[0097] In an example, the UE may transmit an outcome of the supported AIML model. In an example, the UE may receive the at least one reward signal based on the outcome of the supported AIML model.

[0098] In an example, each reward signal of the at least one reward signal may be associated with a corresponding action of the AIML model or the AIML functionality and a corresponding requirement parameter. In other words, for each single action performed by an AIML model or an AIML functionality, a single reward signal may be required. In an example, the UE may request a batch of reward signals, after performing a batch of actions. For example, for a batch length of 5, the corresponding request for the at least one reward signal may may include a list such as [r1 ,r2,r2,r4,r5] corresponding to [a1 ,a2,a3,a4,a5], where rk is reward signal k, and ak is action k.Therefore, the request for the at least one reward signal may include (indication of) one or more reward signals.

[0099] In an example, the request for the at least one reward signal may include one request for the at least one reward signal per AIML functionality. For example, the UE may send the reward request only once, e.g., which indicates the batch size, likely with a fixed delay requirement that would apply to every reward message for the corresponding AIML functionality.

[0100] In an example, the request for the at least one reward signal may include one request for one reward signal. For example, the one reward signal may correspond to one action of the AIML functionality. In other words, the UE may send the request for the at least one reward signal after every action that is taken by the UE, e.g., where the UE may dynamically update and indicate the delay requirement for each reward signal per action.

[0101] In an example, the request for the at least one reward signal may include one request for at least one reward signal, wherein each reward signal of the at least one reward signal may correspond to at least one action of the AIML functionality. In other words, the request for the at least one reward signal may include one request per a number of reward signals [r_t, r_t+1 , ...] corresponding to multiple actions [a_t, a_t+1 , ...] of an AIML functionality / model. For example, the UE may send the request after a batch of actions it takes, e.g., after performing multiple actions, and the UE may dynamically adjust the delay requirement for each batch.

[0102] In an example, the request for the at least one reward signal may include a request for one or more instances of the at least one reward signal. For example, the one or more instances may correspond to one or more actions performed in the past. In an example, the request for the at least one reward signal may include a request for a list of reward signals ordered based on a time of the one or more actions. In other words, based on the reward signal requirement, provided message may include one or more instances of reward signal corresponding to one or more actions taken in the past, e.g., in batches. The reward signals within a message may be ordered with respect to the order of the taken actions, or may be associated with actions based on timestamps of the actions or identifiers of actions, e.g., transaction ID of a message containing an action signaling. In an example, the receiving the at least one reward signal may include receiving at least one of: at least one reward signal per AIML model, one reward signal, wherein the one reward signal corresponds to one action of the AIML model, or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model. In an example, the receiving the at least one reward signal may include receiving at least one of: one or more instances of the at leastone reward signal, wherein the one or more instances correspond to one or more actions performed in the past, or a list of reward signals ordered based on a time of the one or more actions.

[0103] FIG. 3 is a diagram illustrating an example embodiment. The network node may be a base station, an eNB, a gNB, a gNB-CU, a gNB-DU, or core network node such as a LMF. At step 1 , the UE may (pre-)train one or more AIML models or RL model(s) for one or more use cases. In an example, the UE may obtain (or receive), the AIML models from another network entity, e.g., from another network of a same or different vendor. In an example, each AIML model or RL model may be (pre-)trained using a specific reward signal. At step 2, the UE may inform the network node about its support for one or more AIML models or AIML functionalities or supported features that may use the RL. The UE may also indicate a request for at least one reward signal and at least one requirement parameter (or delay requirement) corresponding to the at least one reward signal. In an example embodiment, the AIML functionality or the AIML model may be associated with positioning. In an example, for positioning use case, an RL model may be trained to select positioning anchors by determining a set of positioning measurements to be performed and used to estimate the UE location. The AIML model may be trained using at least one reward signal that indicates the resulting accuracy of estimating the UE position (e.g., error between the estimated UE location and the actual (or ground truth) UE location). In an example, the UE may determine to monitor the performance of the AIML model in a new area that the UE entered first time. Therefore, the UE may determine (or prefer) that the at least one reward signal to be provided to UE after each anchor selection the UE performs, and before performing subsequent selections. In an example, when positioning requests may arrive one after another in the order of minutes, then the UE may determine a delay requirement for the reward signal to be in the order of seconds.

[0104] In another example, the AIML model or the AIML functionality may be associated with a mobility condition / procedure of the UE. In an example, the UE may have a pretrained RL model that is trained to select a new cell for handover, which may use variations in downlink throughput during handover as the at least one reward signal. In an example, the UE may determine to fine-tune an AIML model in a new environment, yet the UE may perform the AIML model update in batches, e.g., after performing Y actions. Thus, the delay requirement of at least one reward signal corresponding to an action may be required to be provided to the UE within next Y actions.

[0105] In yet another example, the UE may have two RL models that the UE may use to select a MCS for UL transmissions: a first RL model may be trained using uplink throughput as the at least one reward signal, which may be periodically measured in the order of milliseconds, and a second RL model may be trained using packet jitter as the at least one reward signal. For the first RL model,that UE may determine to frequently update the corresponding AIML model using the at least one reward signal, thus the UE may determine a periodic reward signal requirement with X ms of periodicity. For the second RL model, the UE’s training algorithm may use a longer horizon to update the model, and the UE may determine to collect the at least one reward signal for every Y number of packets received by gNB before performing each model update. The following table shows an example of indication of reward requirement per functionality that uses RL model:

[0106] The following table shows an example of indication of reward requirement per RL model:

[0107] In an example, the requirement parameter (e.g., indication of the delay requirement) may be expressed in terms of min., max., or a range of time values in terms of a time period (e.g., in the case of periodic reward signals). In an example, the requirement parameter may include a start and / or end time if the at least one reward signal is periodic. In an example, the requirement parameter that indicates a delay requirement may be based on a reference time (e.g., with respect to the reception time of the signal indicating the action), or with respect to (or conditioned upon) a (repetitive or periodic) event / action / outcome or a state change in the environment (e.g., radio channel measurements reaching a certain threshold, RSRP, and / or the like).

[0108] In an example, the information transmitted by the UE to the network node at step 1 of Fl. 3 may be in response to a request from the network node, or in response to a determination by the UE to transmit the information (e.g., unsolicited).

[0109] At step 3 of FIG. 3, the network node may determine necessary configuration for one or more AIML functionalities or AIML models. For example, based on the received information from the UE (and other requirements, e.g., QoS, or the like), the network node may determine to configureone or more AIML functionalities, AIML models, AIML operations, and / or the like at the UE side. For example, the network node may determine to select and / or activate a certain AIML functionality and / or AIML model that uses RL, for inference, monitoring, or training purposes. The network may determine (prefer) to select an AIML functionality and / or an AIML model that supports relaxed (less restrictive) delay requirements for the at least one reward signal over the ones that have more stringent delay requirements. In an example, the network may determine (or have a preference over) a model that has a certain type of reward signal.

[0110] At step 4 of FIG. 3, the network node may request and / or activate the selected RL functionality / model supporting RL (e.g., AIML functionalities, AIML models, or AIML operations). The network node may further provide necessary configurations for the indicated AIML functionalities, AIML models, or AIML operations. For example, the network node may configure (and transmit) necessary reference signals depending on the AIML functionality, e.g., for the case of positioning related functionality, the network node may configure and / or transmit positioning reference signals for positioning estimation, as well as the configurations for UE measurement and reporting.

[0111] At step 5 of FIG. 3, the UE may perform RL inference, e.g., by using any configuration provided by network node as provided in step 4, e.g., reference signal configuration. At step 6, the UE may provide or transmit to the network node, an outcome of the RL inference in order to get associated reward signal.

[0112] At step 7 of FIG. 3, the UE may transmit a request for the at least one reward signal associated with the supported AIML functionality. The request may include a requirement parameter associated with the at least one reward signal of the supported AIML functionality. In other words, the UE may request the reward signal associated with an AIML model / functionality / feature that used RL, e.g., the one that the network has configured for the AI / ML functionality / operation or associated with the RL inference outcome that was provided by the UE. The UE may also indicate the associated delay requirement with the reward signal, if not provided in Step 2. For example, the UE may indicate (in the request for the at least one reward signal) the total duration as well as a starting and / or ending time or a threshold event during which it requires the at least one reward signal, or the total number of the at least one reward signal (or reward signals) the UE requires. In an example, this step may be performed before previous steps, e.g., before step 3 or step 4.

[0113] At step 8 of FIG. 3, the network node may determine the at least one reward signal based on the RL inference outcome. Depending on the use case and the RL task, the network node may perform various procedures and calculations to determine the reward signal. For example, in the positioning use case of anchor selection, a LMF may calculate the resulting positioning accuracy. For example, if the UE is a positioning reference unit (PRU) with a known location, the UE may be used to collect data, e.g., tuples of state / action / rewards. Other examples may include various other positioning methods, e.g., global navigation satellite system (GNSS)-based positioning may be used to obtain at least an approximate ground truth. In another example, the network node may calculate the throughput and / or jitter to determine the at least one reward signal.

[0114] At step 9 of FIG. 3, the network may transmit (provide) the at least one reward signal to the UE based on the indicated requirement parameter such as a delay requirement of the at least one reward signal. In an example, the receiving the at least one reward signal may include receiving at least one of: at least one reward signal per AIML model, one reward signal, wherein the one reward signal corresponds to one action of the AIML model, or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model. In an example, the receiving the at least one reward signal may include receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past, or a list of reward signals ordered based on a time of the one or more actions. In other words, based on the requirement parameter, the message may include one or more instances of reward signal corresponding to one or more actions taken in the past, e.g., in batches. The reward signals within the message may be ordered with respect to the order of the taken actions (or associated with actions) based on timestamps of the actions or IDs of actions, e.g., transaction ID of a message containing an action signaling. Step 9 may be performed multiple times (e.g., as many as the requested (batches of) rewards), but drawn only once in the figure for simplicity. In an example embodiment, the network may not be able to provide the requested reward based on the requirement parameter (e.g., within the delay requirement), e.g., due to high processing or signaling load or congestion at the NW side. In this case, at least one or more of the following may take place: The network node may provide an indication to UE that the network node may not be able to (and / or it will not be able to, e.g., for a certain duration) provide the requested at least one reward signal within the delay requirement. In an example, the network node may indicate to the UE to deactivate / suspend / pause the previously indicated AI / ML model / functionality / feature, and / or select / switch / activate another AI / ML model / functionality / feature or fallback (or revert) to a non-AI / ML functionality / feature since it is not able to provide the rewardassociated with the original model / functionality / feature as requested. In another implementation, although the network node may transmit the requested at least one reward signal within the delay requirement, the UE may not receive it, such as due to channel conditions, handover being performed, and / or the like. As a result, the UE may perform at least one of the following actions: The UE may provide an indication to the network node, e.g., when the channel condition gets better, that the UE was not able to receive the requested reward within the delay requirement, or the UE may deactivate / suspend / pause the previously indicated AI / ML model / functionality / feature, and / or select / switch / activate another AI / ML model / functionality / feature, or fallback (or revert) to a non-AI / ML functionality / feature since it was not able to receive the at least one reward signal associated with the original AIML model / functionality / feature as requested. The UE may then inform the network node about the decision / action it has made, e.g., indicates that it has switched to a different AIML model / functionality.

[0115] At step 10 of FIG. 3, the UE may perform update of the AIML functionality or AIML model (e.g., training / fine tuning). The UE may monitor and / or utilize the at least one reward signal it has obtained. If UE is / was not able to receive the at least one reward signal, e.g., within the required delay time, the UE may perform switching / fallback of AIML model / functionality / feature.

[0116] At step 11 of FIG. 3, if UE identifies or determines any necessary changes to the requirement parameter (or the delay requirement) for the at least one reward signal (e.g., due to a performance monitoring outcome), and if the UE determines / decides to update an AIML model after performing each action instead of every X actions, the UE may informs the network node about an update of the requirement parameter (e.g., a new delay requirement associated with the RL reward signal) used for the AIML model / functionality / feature.

[0117] At least one aspect of example embodiments may be performed by other network nodes, or a core network node / entity (e.g., LMF, AMF) or a RAN node, e.g., a gNB, or OAM entity / function as described.

[0118] Example embodiments are also applicable to implementation on the network node side e.g., on the gNB-side or in general RAN-side models. For example, similar methods performed by the UE may be performed by other apparatus such as a gNB, a LMF, OAM, an AMF, a core network node, and / or the like. Thus, implementation of example embodiments may include utilization of different AIML use cases, that may employ different reference signals for positioning, channel sounding or channel state information, synchronization, phase tracking, demodulation, as (equivalent of) PRS / SRS, CSI-RS, DMRS, PTRS, SRS, PSS / SSS, and / or the like in 5G NR. For example, different reference signals may need to be configured and transmitted for different AI / ML use cases.In an example embodiment, signaling between the network node and the UE may be performed via different protocols, such as (equivalent of) LTE positioning protocol (LPP) (between LMF and UE), RRC (between gNB and UE), MAC (between gNB and UE), and / or the like.

[0119] Some examples will now be described, based on the description and figures provided herein.

[0120] FIG. 4 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment.

[0121] Example 1 . FIG. 4 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment. Operation 410 includes transmitting, by a user device to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality. Operation 420 includes receiving the at least one reward signal based on the requirement parameter.

[0122] Example 2. The method of Example 1 , wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality; a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of the at least one reward signal, wherein the batch size corresponds to a number of instances of each reward signal.

[0123] Example 3. The method of Example 1 or 2, further comprising transmitting an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities.

[0124] Example 4. The method of any of Example 3, wherein the one or more AIML functionalities comprise at least one of: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities relatedto link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

[0125] Example 5. The method of Example 4, further comprising receiving a request to activate one of the one or more AIML functionalities.

[0126] Example 6. The method of Example 5, further comprising: selecting the one of the one or more AIML functionalities; and transmitting an indication of the selected one of the one or more AIML functionalities.

[0127] Example 7. The method of any of Examples 1 to 6, wherein the supported AIML functionality is associated with an AIML model, and wherein the AIML model uses an AIML algorithm that is based on a reinforcement learning (RL).

[0128] Example 8. The method of any of Examples 1 to 7, further comprising: transmitting an outcome of the supported AIML functionality; and receiving the at least one reward signal based on the outcome of the supported AIML functionality.

[0129] Example 9. The method of any of Examples 1 to 8, further comprising receiving an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0130] Example 10. The method of any of Examples 1 to 9, further comprising receiving an indication of deactivating a first supported AIML functionality and activating a second supported AIML functionality.

[0131] Example 11. The method of any of Examples 1 to 10, further comprising: determining to deactivate a first supported AIML functionality and transmitting an indication of deactivating the first supported AIML functionality; and determining to activate a second supported AIML functionality and transmitting an indication of activating the second supported AIML functionality.

[0132] Example 12. The method of any of Examples 1 to 11 , further comprising transmitting an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0133] Example 13. The method of any of Examples 1 to 12, wherein the transmitting the request for the at least one reward signal is based on at least one of: determining that an outcome of the supported AIML functionality does not meet a performance requirement; or a triggering event associated with the supported AIML functionality.

[0134] Example 14. The method of any of Examples 1 to 13, wherein the request comprises an identifier of an AIML model associated with the supported AIML functionality.

[0135] Example 15. The method of any of Examples 1 to 14, further comprising performing at least one of a training, re-training, monitoring, or update for a reinforcement learning model based on the at least one reward signal.

[0136] Example 16. The method of any of Examples 1 to 15, further comprising: transmitting a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of the reward signals; and transmitting a second request indicating stop of transmitting the reward signals.

[0137] Example 17. The method of any of Examples 1 to 16, further comprising transmitting a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0138] Example 18. The method of any of Examples 1 to 17, wherein each reward signal of the at least one reward signal is associated with a corresponding action of the AIML functionality and a corresponding requirement parameter.

[0139] Example 19. The method of any of Examples 1 to 18, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML functionality; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0140] Example 20. The method of any of Examples 1 to 19, wherein the receiving the at least one reward signal comprises receiving at least one of: at least one reward signal per AIML functionality; one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0141] Example 21. The method of any of Examples 1 to 20, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0142] Example 22. The method of any of Examples 1 to 21 , wherein the receiving the at least one reward signal comprises receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0143] FIG. 5 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment.

[0144] Example 23. FIG. 5 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment. Operation 510 includes receiving, by a network node from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality. Operation 520 includes transmitting the at least one reward signal based on the requirement parameter.

[0145] Example 24. The method of Example 23, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality; a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0146] Example 25. The method of Example 23 or 24, further comprising receiving an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities.

[0147] Example 26. The method of any of Example 25, wherein the one or more AIML functionalities comprise at least: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities related to link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

[0148] Example 27. The method of Example 26, further comprising transmitting a request to activate one of the one or more AIML functionalities.

[0149] Example 28. The method of any of Examples 23 to 27, wherein the supported AIML functionality is associated with an AIML model, and wherein the AIML model uses an AIML algorithm that is based on a reinforcement learning (RL).

[0150] Example 29. The method of any of Examples 23 to 28, further comprising: receiving an outcome of the supported AIML functionality; and transmitting the at least one reward signal based on the outcome of the supported AIML functionality.

[0151] Example 30. The method of any of Examples 23 to 29, further comprising transmitting an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0152] Example 31 . The method of any of Examples 23 to 30, further comprising transmitting an indication of deactivating a first supported AIML functionality and activating a second supported AIML functionality.

[0153] Example 32. The method of any of Examples 23 to 31 , further comprising receiving an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0154] Example 33. The method of any of Examples 23 to 32, wherein the request comprises an identifier of an AIML model associated with the supported AIML functionality.

[0155] Example 34. The method of any of Examples 23 to 33, further comprising: receiving a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of reward signals; and receiving a second request indicating stop of transmitting reward signals.

[0156] Example 35. The method of any of Examples 23 to 34, further comprising receiving a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0157] Example 36. The method of any of Examples 23 to 35, wherein each reward signal of the at least one reward signal is associated with a corresponding action of the AIML functionality and a corresponding requirement parameter.

[0158] Example 37. The method of any of Examples 23 to 36, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML functionality; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0159] Example 38. The method of any of Examples 23 to 37, wherein the transmitting the at least one reward signal comprises transmitting at least one of: at least one reward signal per AIML functionality; one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0160] Example 39. The method of any of Examples 23 to 38, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0161] Example 40. The method of any of Examples 23 to 39, wherein the transmitting the at least one reward signal comprises transmitting at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0162] Example 41 . An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receiving the at least one reward signal based on the requirement parameter.

[0163] Example 42. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmitting the at least one reward signal based on the requirement parameter.

[0164] FIG. 6 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment.

[0165] Example 43. FIG. 6 is a flow chart illustrating operation of an apparatus (e.g., which may be a UE or user device, or other apparatus) according to an example embodiment. Operation 610 includes transmitting, by a user device to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models. Operation 620 includes transmitting a request for the at least one reward signalassociated with the supported AIML model. Operation 630 includes receiving the at least one reward signal based on the indicated requirement parameter.

[0166] Example 44. The method of Example 43, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML model, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML model; a starting time associated with the at least one reward signal of the supported AIML model; an ending time associated with the at least one reward signal of the supported AIML model; a time duration associated with the at least one reward signal of the supported AIML model; or a triggering condition associated with the at least one reward signal of the supported AIML model; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0167] Example 45. The method of Example 43 or 44, further comprising: receiving the at least one reward signal, in response to transmitting the request for the at least one reward signal; determining that a requirement indicated by the requirement parameter is not met; and discarding, based on the determining, the at least one rewards signal.

[0168] Example 46. The method of any of Examples 43 to 45, wherein the request for the at least one reward signal further comprises the requirement parameter.

[0169] Example 47. The method of any of Examples 43 to 46, further comprising receiving an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0170] Example 48. The method of any of Examples 44 to 47, further comprising receiving a request to activate one of the one or more AIML models.

[0171] Example 49. The method of Example 48, further comprising: selecting the one of the one or more AIML models; and transmitting, an indication of the selected one of the one or more AIML models.

[0172] Example 50. The method of any of Examples 43 to 49, further comprising transmitting an update of the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0173] Example 51 . The method of any of Examples 43 to 50, wherein the indication of support for the one or more artificial intelligence and machine learning (AIML) models comprises identifiers of the one or more AIML models.

[0174] Example 52. The method of any of Examples 43 to 51 , wherein the one or more AIML models use AIML algorithms that are based on a reinforcement learning (RL).

[0175] Example 53. The method of any of Examples 43 to 52, wherein the transmitting the request for the at least one reward signal is based on at least one of: determining that an outcome of the supported AIML functionality does not meet a performance requirement; or a triggering event associated with the supported AIML functionality.

[0176] Example 54. The method of any of Examples 43 to 53, wherein the request comprises an identifier of the supported AIML model.

[0177] Example 55. The method of any of Examples 43 to 53, further comprising performing at least one of a training, re-training, monitoring, or update for a reinforcement learning model based on the at least one reward signal.

[0178] Example 56. The method of any of Examples 43 to 55, further comprising: transmitting a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of the reward signals; and transmitting a second request indicating stop of transmitting the reward signals.

[0179] Example 57. The method of any of Examples 43 to 56, further comprising transmitting a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0180] Example 58. The method of any of Examples 43 to 57, wherein each reward signal of the at least one reward signal is associated with a corresponding action of an AIML functionality and a corresponding requirement parameter, and wherein the corresponding action of the AIML functionality is based on at least one AIML model.

[0181] Example 59. The method of any of Examples 43 to 58, further comprising receiving an indication of deactivating a first supported AIML model and activating a second supported AIML model.

[0182] Example 60. The method of any of Examples 43 to 59, further comprising: determining to deactivate a first supported AIML model and transmitting an indication of deactivating the first supported AIML model; and determining to activate a second supported AIML model and transmitting an indication of activating the second supported AIML model.

[0183] Example 61 . The method of any of Examples 43 to 60, further comprising: transmitting an outcome of the supported AIML model; and receiving the at least one reward signal based on the outcome of the supported AIML model.

[0184] Example 62. The method of any of Examples 43 to 61 , wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML model; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0185] Example 63. The method of any of Examples 43 to 62, wherein the receiving the at least one reward signal comprises receiving at least one of: at least one reward signal per AIML model; one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0186] Example 64. The method of any of Examples 43 to 63, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0187] Example 65. The method of any of Examples 43 to 64, wherein the receiving the at least one reward signal comprises receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0188] FIG. 7 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment.

[0189] Example 66. FIG. 7 is a flow chart illustrating operation of an apparatus (e.g., which may be a network node or a gNB, or other apparatus) according to an example embodiment. Operation 710 includes receiving, by a network node from a user device, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models. Operation 720 includes receiving a request for the at least one reward signal associated with the supported AIML model. Operation 730 includes transmitting the at least one reward signal based on the indicated requirement parameter.

[0190] Example 67. The method of Example 66, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML model, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of thesupported AIML model; a starting time associated with the at least one reward signal of the supported AIML model; an ending time associated with the at least one reward signal of the supported AIML model; a time duration associated with the at least one reward signal of the supported AIML model; or a triggering condition associated with the at least one reward signal of the supported AIML model; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0191] Example 68. The method of Example 66 or 67, wherein the request for the at least one reward signal further comprises the requirement parameter.

[0192] Example 69. The method of any of Examples 66 to 68, further comprising transmitting an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0193] Example 70. The method of any of Examples 67 to 69, further comprising transmitting a request to activate one of the one or more AIML models.

[0194] Example 71 . The method of any of Examples 66 to 70, further comprising receiving an update of the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0195] Example 72. The method of any of Examples 66 to 71 , wherein the indication of support for the one or more artificial intelligence and machine learning (AIML) models comprises identifiers of the one or more AIML models.

[0196] Example 73. The method of any of Examples 66 to 72, wherein the one or more AIML models use AIML algorithms that are based on a reinforcement learning (RL).

[0197] Example 74. The method of any of Examples 66 to 73, wherein the request comprises an identifier of the supported AIML model.

[0198] Example 75. The method of any of Examples 66 to 74, further comprising: receiving a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of reward signals; and receiving a second request indicating stop of transmitting reward signals.

[0199] Example 76. The method of any of Examples 66 to 75, further comprising receiving a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0200] Example 77. The method of any of Examples 66 to 76, wherein each reward signal of the at least one reward signal is associated with a corresponding action of an AIML functionality anda corresponding requirement parameter, and wherein the corresponding action of the AIML functionality is based on at least one AIML model.

[0201] Example 78. The method of any of Examples 66 to 77, further comprising transmitting an indication of deactivating a first supported AIML model and activating a second supported AIML model.

[0202] Example 79. The method of any of Examples 66 to 78, further comprising: receiving an outcome of the supported AIML model; and transmitting the at least one reward signal based on the outcome of the supported AIML model.

[0203] Example 80. The method of any of Examples 66 to 79, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML model; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0204] Example 81 . The method of any of Examples 66 to 85, wherein the transmitting the at least one reward signal comprises transmitting at least one of: at least one reward signal per AIML model; one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0205] Example 82. The method of any of Examples 66 to 81 , wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0206] Example 83. The method of any of Examples 66 to 82, wherein the transmitting the at least one reward signal comprises transmitting at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0207] Example 84. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; transmitting a request for the at least one reward signal associated with the supported AIML model; and receiving the at least one reward signal based on the indicated requirement parameter.

[0208] Example 85. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; receiving a request for the at least one reward signal associated with the supported AIML model; and transmitting the at least one reward signal based on the indicated requirement parameter.

[0209] Clause 1 . An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receiving the at least one reward signal based on the requirement parameter.

[0210] Clause 2. The apparatus of Clause 1 , wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality; a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of the at least one reward signal, wherein the batch size corresponds to a number of instances of each reward signal.

[0211] Clause 3. The apparatus of Clause 1 or 2, wherein the apparatus is further caused to perform transmitting an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities.

[0212] Clause 4. The apparatus of any of Clause 3, wherein the one or more AIML functionalities comprise at least one of: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities relatedto link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

[0213] Clause 5. The apparatus of Clause 4, wherein the apparatus is further caused to perform receiving a request to activate one of the one or more AIML functionalities.

[0214] Clause 6. The apparatus of Clause 5, wherein the apparatus is further caused to perform: selecting the one of the one or more AIML functionalities; and transmitting an indication of the selected one of the one or more AIML functionalities.

[0215] Clause 7. The apparatus of any of Clauses 1 to 6, wherein the supported AIML functionality is associated with an AIML model, and wherein the AIML model uses an AIML algorithm that is based on a reinforcement learning (RL).

[0216] Clause 8. The apparatus of any of Clauses 1 to 7, wherein the apparatus is further caused to perform: transmitting an outcome of the supported AIML functionality; and receiving the at least one reward signal based on the outcome of the supported AIML functionality.

[0217] Clause 9. The apparatus of any of Clauses 1 to 8, wherein the apparatus is further caused to perform receiving an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0218] Clause 10. The apparatus of any of Clauses 1 to 9, wherein the apparatus is further caused to perform receiving an indication of deactivating a first supported AIML functionality and activating a second supported AIML functionality.

[0219] Clause 11 . The apparatus of any of Clauses 1 to 10, wherein the apparatus is further caused to perform: determining to deactivate a first supported AIML functionality and transmitting an indication of deactivating the first supported AIML functionality; and determining to activate a second supported AIML functionality and transmitting an indication of activating the second supported AIML functionality.

[0220] Clause 12. The apparatus of any of Clauses 1 to 11 , wherein the apparatus is further caused to perform transmitting an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0221] Clause 13. The apparatus of any of Clauses 1 to 12, wherein the transmitting the request for the at least one reward signal is based on at least one of: determining that an outcome of the supported AIML functionality does not meet a performance requirement; or a triggering event associated with the supported AIML functionality.

[0222] Clause 14. The apparatus of any of Clauses 1 to 13, wherein the request comprises an identifier of an AIML model associated with the supported AIML functionality.

[0223] Clause 15. The apparatus of any of Clauses 1 to 14, wherein the apparatus is further caused to perform performing at least one of a training, re-training, monitoring, or update for a reinforcement learning model based on the at least one reward signal.

[0224] Clause 16. The apparatus of any of Clauses 1 to 15, wherein the apparatus is further caused to perform: transmitting a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of the reward signals; and transmitting a second request indicating stop of transmitting the reward signals.

[0225] Clause 17. The apparatus of any of Clauses 1 to 16, wherein the apparatus is further caused to perform transmitting a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0226] Clause 18. The apparatus of any of Clauses 1 to 17, wherein each reward signal of the at least one reward signal is associated with a corresponding action of the AIML functionality and a corresponding requirement parameter.

[0227] Clause 19. The apparatus of any of Clauses 1 to 18, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML functionality; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0228] Clause 20. The apparatus of any of Clauses 1 to 19, wherein the receiving the at least one reward signal comprises receiving at least one of: at least one reward signal per AIML functionality; one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0229] Clause 21 . The apparatus of any of Clauses 1 to 20, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0230] Clause 22. The apparatus of any of Clauses 1 to 21 , wherein the receiving the at least one reward signal comprises receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0231] Clause 23. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmitting the at least one reward signal based on the requirement parameter.

[0232] Clause 24. The apparatus of Clause 23, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality; a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0233] Clause 25. The apparatus of Clause 23 or 24, wherein the apparatus is further caused to perform receiving an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities.

[0234] Clause 26. The apparatus of any of Clause 25, wherein the one or more AIML functionalities comprise at least: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities related to link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

[0235] Clause 27. The apparatus of Clause 26, wherein the apparatus is further caused to perform transmitting a request to activate one of the one or more AIML functionalities.

[0236] Clause 28. The apparatus of any of Clauses 23 to 27, wherein the supported AIML functionality is associated with an AIML model, and wherein the AIML model uses an AIML algorithm that is based on a reinforcement learning (RL).

[0237] Clause 29. The apparatus of any of Clauses 23 to 28, wherein the apparatus is further caused to perform: receiving an outcome of the supported AIML functionality; and transmitting the at least one reward signal based on the outcome of the supported AIML functionality.

[0238] Clause 30. The apparatus of any of Clauses 23 to 29, wherein the apparatus is further caused to perform transmitting an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0239] Clause 31 . The apparatus of any of Clauses 23 to 30, wherein the apparatus is further caused to perform transmitting an indication of deactivating a first supported AIML functionality and activating a second supported AIML functionality.

[0240] Clause 32. The apparatus of any of Clauses 23 to 31 , wherein the apparatus is further caused to perform receiving an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

[0241] Clause 33. The apparatus of any of Clauses 23 to 32, wherein the request comprises an identifier of an AIML model associated with the supported AIML functionality.

[0242] Clause 34. The apparatus of any of Clauses 23 to 33, wherein the apparatus is further caused to perform: receiving a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of reward signals; and receiving a second request indicating stop of transmitting reward signals.

[0243] Clause 35. The apparatus of any of Clauses 23 to 34, wherein the apparatus is further caused to perform receiving a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0244] Clause 36. The apparatus of any of Clauses 23 to 35, wherein each reward signal of the at least one reward signal is associated with a corresponding action of the AIML functionality and a corresponding requirement parameter.

[0245] Clause 37. The apparatus of any of Clauses 23 to 36, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML functionality; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0246] Clause 38. The apparatus of any of Clauses 23 to 37, wherein the transmitting the at least one reward signal comprises transmitting at least one of: at least one reward signal per AIMLfunctionality; one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

[0247] Clause 39. The apparatus of any of Clauses 23 to 38, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0248] Clause 40. The apparatus of any of Clauses 23 to 39, wherein the transmitting the at least one reward signal comprises transmitting at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0249] Clause 41 . An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: transmitting to a network node, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; transmitting a request for the at least one reward signal associated with the supported AIML model; and receiving the at least one reward signal based on the indicated requirement parameter.

[0250] Clause 42. The apparatus of Clause 41 , wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML model, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML model; a starting time associated with the at least one reward signal of the supported AIML model; an ending time associated with the at least one reward signal of the supported AIML model; a time duration associated with the at least one reward signal of the supported AIML model; or a triggering condition associated with the at least one reward signal of the supported AIML model; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0251] Clause 43. The apparatus of Clause 41 or 42, wherein the apparatus is further caused to perform: receiving the at least one reward signal, in response to transmitting the request for the at least one reward signal; determining that a requirement indicated by the requirement parameter is not met; and discarding, based on the determining, the at least one rewards signal.

[0252] Clause 44. The apparatus of any of Clauses 41 to 43, wherein the request for the at least one reward signal further comprises the requirement parameter.

[0253] Clause 45. The apparatus of any of Clauses 41 to 44, further comprising receiving an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0254] Clause 46. The apparatus of any of Clauses 42 to 45, further comprising receiving a request to activate one of the one or more AIML models.

[0255] Clause 47. The apparatus of Clause 46, wherein the apparatus is further caused to perform: selecting the one of the one or more AIML models; and transmitting, an indication of the selected one of the one or more AIML models.

[0256] Clause 48. The apparatus of any of Clauses 41 to 47, wherein the apparatus is further caused to perform transmitting an update of the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0257] Clause 49. The apparatus of any of Clauses 41 to 48, wherein the indication of support for the one or more artificial intelligence and machine learning (AIML) models comprises identifiers of the one or more AIML models.

[0258] Clause 50. The apparatus of any of Clauses 41 to 49, wherein the one or more AIML models use AIML algorithms that are based on a reinforcement learning (RL).

[0259] Clause 51 . The apparatus of any of Clauses 41 to 50, wherein the transmitting the request for the at least one reward signal is based on at least one of: determining that an outcome of the supported AIML functionality does not meet a performance requirement; or a triggering event associated with the supported AIML functionality.

[0260] Clause 52. The apparatus of any of Clauses 41 to 51 , wherein the request comprises an identifier of the supported AIML model.

[0261] Clause 53. The apparatus of any of Clauses 41 to 51 , wherein the apparatus is further caused to perform performing at least one of a training, re-training, monitoring, or update for a reinforcement learning model based on the at least one reward signal.

[0262] Clause 54. The apparatus of any of Clauses 41 to 53, wherein the apparatus is further caused to perform: transmitting a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of the reward signals; and transmitting a second request indicating stop of transmitting the reward signals.

[0263] Clause 55. The apparatus of any of Clauses 41 to 54, wherein the apparatus is further caused to perform transmitting a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0264] Clause 56. The apparatus of any of Clauses 41 to 55, wherein each reward signal of the at least one reward signal is associated with a corresponding action of an AIML functionality and a corresponding requirement parameter, and wherein the corresponding action of the AIML functionality is based on at least one AIML model.

[0265] Clause 57. The apparatus of any of Clauses 41 to 56, wherein the apparatus is further caused to perform receiving an indication of deactivating a first supported AIML model and activating a second supported AIML model.

[0266] Clause 58. The apparatus of any of Clauses 41 to 57, wherein the apparatus is further caused to perform: determining to deactivate a first supported AIML model and transmitting an indication of deactivating the first supported AIML model; and determining to activate a second supported AIML model and transmitting an indication of activating the second supported AIML model.

[0267] Clause 59. The apparatus of any of Clauses 41 to 58, wherein the apparatus is further caused to perform: transmitting an outcome of the supported AIML model; and receiving the at least one reward signal based on the outcome of the supported AIML model.

[0268] Clause 60. The apparatus of any of Clauses 41 to 59, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML model; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0269] Clause 61 . The apparatus of any of Clauses 41 to 60, wherein the receiving the at least one reward signal comprises receiving at least one of: at least one reward signal per AIML model; one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0270] Clause 62. The apparatus of any of Clauses 41 to 61 , wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0271] Clause 63. The apparatus of any of Clauses 41 to 62, wherein the receiving the at least one reward signal comprises receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0272] Clause 64. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving from a user device, an indication of support for one or more artificial intelligence and machine learning (AIML) models, the indication comprising a requirement parameter associated with at least one reward signal of a supported AIML model of the one or more AIML models; receiving a request for the at least one reward signal associated with the supported AIML model; and transmitting the at least one reward signal based on the indicated requirement parameter.

[0273] Clause 65. The apparatus of Clause 64, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML model, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; a range of time values; a period associated with the at least one reward signal of the supported AIML model; a starting time associated with the at least one reward signal of the supported AIML model; an ending time associated with the at least one reward signal of the supported AIML model; a time duration associated with the at least one reward signal of the supported AIML model; or a triggering condition associated with the at least one reward signal of the supported AIML model; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

[0274] Clause 66. The apparatus of Clause 64 or 65, wherein the request for the at least one reward signal further comprises the requirement parameter.

[0275] Clause 67. The apparatus of any of Clauses 64 to 66, wherein the apparatus is further caused to perform transmitting an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0276] Clause 68. The apparatus of any of Clauses 65 to 67, wherein the apparatus is further caused to perform transmitting a request to activate one of the one or more AIML models.

[0277] Clause 69. The apparatus of any of Clauses 64 to 68, further comprising receiving an update of the requirement parameter associated with the at least one reward signal of the supported AIML model.

[0278] Clause 70. The apparatus of any of Clauses 64 to 69, wherein the indication of support for the one or more artificial intelligence and machine learning (AIML) models comprises identifiers of the one or more AIML models.

[0279] Clause 71 . The apparatus of any of Clauses 64 to 70, wherein the one or more AIML models use AIML algorithms that are based on a reinforcement learning (RL).

[0280] Clause 72. The apparatus of any of Clauses 64 to 71 , wherein the request comprises an identifier of the supported AIML model.

[0281] Clause 73. The apparatus of any of Clauses 64 to 72, wherein the apparatus is further caused to perform: receiving a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of reward signals; and receiving a second request indicating stop of transmitting reward signals.

[0282] Clause 74. The apparatus of any of Clauses 64 to 73, wherein the apparatus is further caused to perform receiving a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

[0283] Clause 75. The apparatus of any of Clauses 64 to 74, wherein each reward signal of the at least one reward signal is associated with a corresponding action of an AIML functionality and a corresponding requirement parameter, and wherein the corresponding action of the AIML functionality is based on at least one AIML model.

[0284] Clause 76. The apparatus of any of Clauses 64 to 75, wherein the apparatus is further caused to perform transmitting an indication of deactivating a first supported AIML model and activating a second supported AIML model.

[0285] Clause 77. The apparatus of any of Clauses 64 to 76, wherein the apparatus is further caused to perform: receiving an outcome of the supported AIML model; and transmitting the at least one reward signal based on the outcome of the supported AIML model.

[0286] Clause 78. The apparatus of any of Clauses 64 to 77, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML model; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML model; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0287] Clause 79. The apparatus of any of Clauses 64 to 78, wherein the transmitting the at least one reward signal comprises transmitting at least one of: at least one reward signal per AIML model; one reward signal, wherein the one reward signal corresponds to one action of the AIMLmodel; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML model.

[0288] Clause 80. The apparatus of any of Clauses 64 to 79, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a request for a list of reward signals ordered based on a time of the one or more actions.

[0289] Clause 81 . The apparatus of any of Clauses 64 to 80, wherein the transmitting the at least one reward signal comprises transmitting at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

[0290] FIG. 8 is a block diagram of a wireless station or node (e.g., UE, user device, AP, BS, eNB, gNB, RAN node, network node, TRP, or other node) 1300 according to an example embodiment. The wireless station 1300 may include, for example, one or more (e.g., two as shown in FIG. 8) RF (radio frequency) or wireless transceivers 1302A, 1302B, where each wireless transceiver includes a transmitter to transmit signals and a receiver to receive signals. The wireless station also includes a processor or control unit / entity (controller) 1304 to execute instructions or software and control transmission and receptions of signals, and a memory 1306 to store data and / or instructions.

[0291] Processor 1304 may also make decisions or determinations, generate frames, packets or messages for transmission, decode received frames or messages for further processing, and other tasks or functions described herein. Processor 1304, which may be a baseband processor, for example, may generate messages, packets, frames or other signals for transmission via wireless transceiver 1302 (1302A or 1302B). Processor 1304 may control transmission of signals or messages over a wireless network, and may control the reception of signals or messages, etc., via a wireless network (e.g., after being down-converted by wireless transceiver 1302, for example). Processor 1304 may be programmable and capable of executing software or other instructions stored in memory or on other computer media to perform the various tasks and functions described above, such as one or more of the tasks or methods described above. Processor 1304 may be (or may include), for example, hardware, programmable logic, a programmable processor that executes software or firmware, and / or any combination of these. Using other terminology, processor 1304 and transceiver 1302 together may be considered as a wireless transmitter / receiver system, for example.

[0292] In addition, referring to FIG. 8, a controller (or processor) 1308 may execute software and instructions, and may provide overall control for the station 1300, and may provide control for other systems not shown in FIG. 8, such as controlling input / output devices (e.g., display, keypad), and / or may execute software for one or more applications that may be provided on wireless station 1300, such as, for example, an email program, audio / video applications, a word processor, a Voice over IP application, or other application or software.

[0293] In addition, a storage medium may be provided that includes stored instructions, which when executed by a controller or processor may result in the processor 1304, or other controller or processor, performing one or more of the functions or tasks described above.

[0294] According to another example embodiment, RF or wireless transceiver(s) 1302A / 1302B may receive signals or data and / or transmit or send signals or data. Processor 1304 (and possibly transceivers 1302A / 1302B) may control the RF or wireless transceiver 1302A or 1302B to receive, send, broadcast or transmit signals or data.

[0295] Example embodiments are provided or described for each of the example methods, including: An apparatus (e.g., 1300, FIG. 8) including means (e.g., processor 1304, RF transceivers 1302A and / or 1302B, and / or memory 1306, in FIG. 8) for carrying out any of the methods; a non- transitory computer-readable storage medium (e.g., memory 1306, FIG. 8) comprising instructions stored thereon that, when executed by at least one processor (processor 1304, FIG. 8), are configured to cause a computing system (e.g., 1300, FIG. 8) to perform any of the example methods; and an apparatus (e.g., 1300, FIG. 8) including at least one processor (e.g., processor 1304, FIG. 8), and at least one memory (e.g., memory 1306, FIG. 8) including computer program code, the at least one memory (1306) and the computer program code configured to, with the at least one processor (1304), cause the apparatus (e.g., 1300) at least to perform any of the example methods.

[0296] Embodiments of the various techniques described herein may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Embodiments may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, e.g., in a machine-readable storage device or in a propagated signal, for execution by, or to control the operation of, a data processing apparatus, e.g., a programmable processor, a computer, or multiple computers. Embodiments may also be provided on a computer readable medium or computer readable storage medium, which may be a non- transitory medium. Embodiments of the various techniques may also include embodiments provided via transitory signals or media, and / or programs and / or software embodiments that are downloadable via the Internet or other network(s), either wired networks and / or wireless networks. Inaddition, embodiments may be provided via machine type communications (MTC), and also via an Internet of Things (IOT).

[0297] As used in this application, the term “circuitry” or “circuit” refers to all of the following: (a) hardware-only circuit implementations, such as implementations in only analog and / or digital circuitry, and (b) combinations of circuits and soft-ware (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present. This definition of “circuitry” applies to all uses of this term in this application. As a further example, as used in this application, the term “circuitry” would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term “circuitry” would also cover, for example and if applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0298] The computer program may be in source code form, object code form, or in some intermediate form, and it may be stored in some sort of carrier, distribution medium, or computer readable medium, which may be any entity or device capable of carrying the program. Such carriers include a record medium, computer memory, read-only memory, photoelectrical and / or electrical carrier signal, telecommunications signal, and software distribution package, for example. Depending on the processing power needed, the computer program may be executed in a single electronic digital computer, or it may be distributed amongst a number of computers.

[0299] Furthermore, embodiments of the various techniques described herein may use a cyberphysical system (CPS) (a system of collaborating computational elements controlling physical entities). CPS may enable the embodiment and exploitation of massive amounts of interconnected ICT devices (sensors, actuators, processors microcontrollers, ...) embedded in physical objects at different locations. Mobile cyber physical systems, in which the physical system in question has inherent mobility, are a subcategory of cyber-physical systems. Examples of mobile physical systems include mobile robotics and electronics transported by humans or animals. The rise in popularity of smartphones has increased interest in the area of mobile cyber-physical systems. Therefore, various embodiments of techniques described herein may be provided via one or more of these technologies.

[0300] A computer program, such as the computer program(s) described above, can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit or part of it suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0301] Method steps may be performed by one or more programmable processors executing a computer program or computer program portions to perform functions by operating on input data and generating output. Method steps also may be performed by, and an apparatus may be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0302] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer, chip or chipset. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. Elements of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also may include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magnetooptical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0303] To provide for interaction with a user, embodiments may be implemented on a computer having a display device, e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, for displaying information to the user and a user interface, such as a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0304] Embodiments may be implemented in a computing system that includes a backend component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an embodiment, or any combination of such backend, middleware, or frontend components. Components may be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0305] While certain features of the described embodiments have been illustrated as described herein, many modifications, substitutions, changes and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the various embodiments.

Claims

WHAT IS CLAIMED IS:1 . An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: transmit to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receive the at least one reward signal based on the requirement parameter.

2. The apparatus of claim 1 , wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; or a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality; a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of the at least one reward signal, wherein the batch size corresponds to a number of instances of each reward signal.

553. The apparatus of claim 1 or 2, the apparatus is further caused to transmit an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities.

4. The apparatus of any of claims 1 to 3, wherein the supported AIML functionality comprises at least one of: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities related to link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

5. The apparatus of claim 4, the apparatus is further caused to receive a request to activate one of the one or more AIML functionalities.

6. The apparatus of claim 5, the apparatus is further caused to: select the one of the one or more AIML functionalities; and transmit an indication of the selected one of the one or more AIML functionalities.

7. The apparatus of any of claims 1 to 6, wherein the supported AIML functionality is associated with an AIML model, and wherein the AIML model uses an AIML algorithm that is based on a reinforcement learning (RL).

8. The apparatus of any of claims 1 to 7, the apparatus is further caused to: transmit an outcome of the supported AIML functionality; and receive the at least one reward signal based on the outcome of the supported AIML functionality.

9. The apparatus of any of claims 1 to 8, the apparatus is further caused to receive an indication of whether a requirement can be met based on the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

10. The apparatus of any of claims 1 to 9, the apparatus is further caused to receive an indication of deactivating a first supported AIML functionality and activating a second supported AIML functionality.

11. The apparatus of any of claims 1 to 10, the apparatus is further caused to: determine to deactivate a first supported AIML functionality and transmitting an indication of deactivating the first supported AIML functionality; and determine to activate a second supported AIML functionality and transmitting an indication of activating the second supported AIML functionality.

12. The apparatus of any of claims 1 to 11 , the apparatus is further caused to transmit an update of the requirement parameter associated with the at least one reward signal of the supported AIML functionality.

13. The apparatus of any of claims 1 to 12, wherein the transmitting the request for the at least one reward signal is based on at least one of: determining that an outcome of the supported AIML functionality does not meet a performance requirement; or a triggering event associated with the supported AIML functionality.

14. The apparatus of any of claims 1 to 13, wherein the request comprises an identifier of an AIML model associated with the supported AIML functionality.

15. The apparatus of any of claims 1 to 14, the apparatus is further caused to performe at least one of a training, re-training, monitoring, or update for a reinforcement learning model based on the at least one reward signal.

16. The apparatus of any of claims 1 to 15, the apparatus is further caused to:transmit a first request indicating start of transmitting reward signals, wherein the first request comprises a time interval between transmission of the reward signals; and transmit a second request indicating stop of transmitting the reward signals.

17. The apparatus of any of claims 1 to 16, the apparatus is further caused to transmit a first request indicating start of transmitting reward signals for a time duration, wherein the first request comprises a value of the time duration.

18. The apparatus of any of claims 1 to 17, wherein each reward signal of the at least one reward signal is associated with a corresponding action of the AIML functionality and a corresponding requirement parameter.

19. The apparatus of any of claims 1 to 18, wherein the request for the at least one reward signal comprises at least one of: one request for the at least one reward signal per AIML functionality; one request for one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or one request for at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.

20. The apparatus of any of claims 1 to 19, wherein the receiving the at least one reward signal comprises receiving at least one of: at least one reward signal per AIML functionality; one reward signal, wherein the one reward signal corresponds to one action of the AIML functionality; or at least one reward signal, wherein each reward signal of the at least one reward signal corresponds to at least one action of the AIML functionality.21 . The apparatus of any of claims 1 to 20, wherein the request for the at least one reward signal comprises at least one of: a request for one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; ora request for a list of reward signals ordered based on a time of the one or more actions.

22. The apparatus of any of claims 1 to 21, wherein the receiving the at least one reward signal comprises receiving at least one of: one or more instances of the at least one reward signal, wherein the one or more instances correspond to one or more actions performed in past; or a list of reward signals ordered based on a time of the one or more actions.

23. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmit the at least one reward signal based on the requirement parameter.

24. The apparatus of claim 23, wherein the requirement parameter associated with the at least one reward signal comprises at least one of: a time parameter of a delay requirement associated with the at least one reward signal of the supported AIML functionality, wherein the time parameter comprises at least one of: a time value; a minimum time value; a maximum time value; or a range of time values; a period associated with the at least one reward signal of the supported AIML functionality; a starting time associated with the at least one reward signal of the supported AIML functionality; an ending time associated with the at least one reward signal of the supported AIML functionality; a time duration associated with the at least one reward signal of the supported AIML functionality;a triggering condition associated with the at least one reward signal of the supported AIML functionality; a total number of the at least one reward signal; or a batch size of reward signals, wherein the batch size corresponds to a number of instances of each reward signal.

25. The apparatus of claim 23 or 24, the apparatus is further caused to: receive an indication of support for one or more artificial intelligence and machine learning (AIML) functionalities, and wherein the one or more AIML functionalities comprise at least: one or more functionalities related to beam management; one or more functionalities related to beamforming; one or more functionalities related to positioning; one or more functionalities related to allocation of radio resources; one or more functionalities related to scheduling; one or more functionalities related to power control; one or more functionalities related to link adaptation; one or more functionalities related to mobility; or one or more functionalities related to selection of a modulation and coding (MCS) scheme.

26. An apparatus comprising: means for transmitting to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and means for receiving the at least one reward signal based on the requirement parameter.

27. An apparatus comprising: means for receiving from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; andmeans for transmitting the at least one reward signal based on the requirement parameter.

28. A method comprising: transmitting, by a user device to a network node, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and receiving the at least one reward signal based on the requirement parameter.

29. A method comprising: receiving, by a network node from a user device, a request for at least one reward signal associated with a supported artificial intelligence and machine learning (AIML) functionality, the request comprising a requirement parameter associated with the at least one reward signal of the supported AIML functionality; and transmitting the at least one reward signal based on the requirement parameter.

30. A non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform a method of any of claims 1 to 25.31 . A computer program comprising instructions stored thereon for performing a method of any of claims 1 to 25.

Citation Information

Patent Citations

  • Transmission method and device, terminal and network side equipment

    CN118158655A

  • Method, apparatus and computer program

    US20240007884A1

  • Method and apparatus for using artificial intelligence / machine learning model in wireless communication network

    US20240214840A1

  • Structure of ML model information and its usage

    US20240281708A1

  • CSI processing mode switching method and apparatus, and medium, product and chip

    US20250219698A1