Methods and apparatuses for controlling a plurality of wireless devices to utilze a respective plurality of discontinuous reception cycles.

A reinforcement learning model coordinates DRX cycles to extend non-active phases, optimizing energy savings at RAN nodes by allowing deeper sleep modes without compromising communication quality.

WO2025216674A1PCT designated stage Publication Date: 2025-10-16TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/050339
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing approaches struggle to leverage deeper sleep modes at radio access network (RAN) nodes due to unpredictable traffic patterns, limiting energy savings and increasing communication latency.

Method used

Implementing a reinforcement learning model to coordinate discontinuous reception (DRX) cycles for wireless devices, extending the common non-active phase to allow RAN nodes to enter deeper sleep modes more frequently.

Benefits of technology

Enhances energy savings at RAN nodes by enabling longer periods of deep sleep while maintaining communication quality and service requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2024050339_16102025_PF_FP_ABST
    Figure SE2024050339_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments described herein relate to methods and apparatuses for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. A method performed by a controller network node comprises applying an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period; and initiating use of the respective plurality of DRX cycles by the plurality of wireless devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS AND APPARATUSES FOR CONTROLLING A PLURALITY OF WIRELESS DEVICES TO UTILZE A RESPECTIVE PLURALITY OF DISCONTINUOUS RECEPTION CYCLES.

[0002] TECHNICAL FIELD

[0003] Embodiments described herein relate to methods and apparatuses for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. In particular, some methods and apparatuses described herein utilize an optimization process to act, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period.

[0004] BACKGROUND

[0005] Discontinuous reception (DRX) is an approach to reduce the energy consumption of a wireless device (e.g. a user equipment (UE)) by regulating the time at which the wireless devices radio receiver can go to sleep.

[0006] Figure 1 illustrates an example of power consumption during a DRX cycle process. At the beginning of a DRX cycle (or with a specified offset from it), the wireless device wakes up to monitor the channel for a period of time. This period of time is referred to the ON period 101. At the beginning of the ON period, the wireless device may set an inactivity timer and upon its expiration, it may enter a dormancy phase (also referred to as a non-active phase or an OFF phase) 102 which lasts for a specific time period. If a wireless device starts receiving data during the ON period, it consumes more energy (as illustrated by the period 103 in Figure 1 ) and after the finish of the data reception phase, it may restart the inactivity timer (e.g. during phase 104 in Figure 1 ) and wait for the expiration of the inactivity timer before entering the non-active phase again (e.g. in phase 105 of Figure 1 ). As can be seen in Figure 1 , a wireless device consumes much less energy (for the radio components) in the dormancy phase.

[0007] F. Moradi, E. Fitzgerald, M. Pioro and B. Landfeldt, “Flexible DRX Optimization for LTE and 5G,” in IEEE Transactions on Vehicular Technology, vol. 69, no. 1 , pp. 607-621 , Jan. 2020 discloses improving a UE’s energy saving by using optimized DRX configuration for video streaming service. The core of the solution is the prediction of the channel quality to activate the UE whenever the channel is good and put is on the dormancy otherwise. This way, the UE can be in an ON phase only when good channel conditions are occurring and can therefore serve the traffic faster, increasing its sleeping time opportunities.

[0008] Base station (BS) sleeping is a common approach to reduce energy consumption in radio access network (RAN) node. BSs (or any RAN node, e.g. O-RAN nodes) may temporarily turn off some of their radio hardware (e.g., power amplifier, mixer, digital to analog converter, etc.) or the entire RF chain(s) (e.g. transmission chains) to save some energy. BS sleeping often comes in various modes that are characterized by the depths of sleeping, and the resulting energy saving and latency to activate back the hardware components. Table 1 illustrates various examples of sleeping modes (SMs) (e.g. transmission sleep modes), their relative power consumptions with respect to a Deep sleep mode, as well as the wake-up time required for each SM. Each SM is also associated with a ramp-down / ramp-up energy, which increases with the depth of the sleeping mode.

[0009] Table 1 : Descriptions of various SMs [Table 5.1 -3, 5.1 -4, and 5.1 -5, 3GPP TR 38.864],

[0010] From this table, we can see the inherent trade-off in BS sleeping. The deeper the SM, the more energy saved at the RAN node, but the more delay is experienced to switch the RAN node back on, which may cause some extra delay in the communication with the UEs.

[0011] Almost all SM levels are defined on the near-real-time level. These decisions are also often local, in the sense that if a BS is in some SM, it can turn on again based on its prediction of new incoming traffic. Due to random nature of the cell traffic, however, it is difficult to see large enough traffic idle times to activate a deep sleep mode (which is responsible for most of the energy saving gains).

[0012] Vereecken, Willem, et al. "The effect of variable wake up time on the utilization of sleep modes in femtocell mobile access networks." 2012 9th Annual Conference on Wireless On-Demand Network Systems and Services (WONS). IEEE, 2012 introduced several SM levels that can vary on the set of hardware they switch off and wakeup time. Jang, Gunhee, et al. "Base station switching and sleep mode optimization with LSTM-based user prediction." IEEE Access 8 (2020): 22271 1 -222723 used a machine learning (ML) approach to find optimal BS switching and SM design based on the network traffic predictions for the subsequent time slots.

[0013] SUMMARY

[0014] The existing approaches discussed above uses various models, including ML models, to predict the traffic for near-future and apply some appropriate RAN network node (e.g. BS) sleeping. However, when the incoming traffic is random and un-controlled, there is simply not much opportunity for sleeping, even if the traffic prediction algorithm is error- free.

[0015] Embodiments described herein therefore make use of DRX cycles (or similar energy saving cycles) to coordinate or control the DRX of the UEs to account for the possibility of BS sleeping. In other words, by providing deeper sleep opportunities for the BS, the potential for energy savings at the RAN network node can be leveraged.

[0016] According to some embodiments there is provided a method performed by a controller network node for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. The method comprises applying an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period. The method further comprises initiating use of the respective plurality of DRX cycles by the plurality of wireless devices.

[0017] According to some embodiments there is provided a computer implemented method for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a radio access network, RAN, node, a respective plurality of DRX cycles. The method comprises performing a training loop process. The training loop processes comprises: obtaining first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period; inputting the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles; applying the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices; and obtaining second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period; determining a reward, Rt+i, for the second state information utilizing a reward function, and adjusting the policy according to the reward.

[0018] According to some embodiments there is provided a controller network node for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. The controller network node comprises processing and a memory, the memory containing instructions executable by the processing circuitry whereby the controller network node is operable to: apply an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period; and initiate use of the respective plurality of DRX cycles by the plurality of wireless devices.

[0019] According to some embodiments there is provided a training network node for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a radio access network, RAN, node, a respective plurality of DRX cycles. The training network node comprises processing and a memory, the memory containing instructions executable by the processing circuitry whereby the training network node is operable to perform a training loop process. The training loop process comprises obtaining first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period; inputting the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles; applying the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices; and obtaining second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period; determining a reward, Rt+i, for the second state information utilizing a reward function, and adjusting the policy according to the reward.

[0020] According to some embodiments there is provided a computer program, comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out any of the methods described above.

[0021] According to some embodiments there is provided a carrier containing the computer program as described above, wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.

[0022] According to some embodiments there is provided a computer-readable medium comprising instructions that, when executed on at least one processor, cause the at least one processor to perform a method as described above.

[0023] According to some embodiments there is provided a computer program product comprising non transitory computer readable media having stored thereon a computer program as described above.

[0024] An advantage of the claimed technology is that the RAN node may turn off its transmitter chain for longer periods, by the coordination of the DRX cycles of the UE, such that there will be a common DRX inactive phase. In particular energy savings are made when the PA of the RAN node can be turned off.

[0025] BRIEF DESCRIPTION OF THE DRAWINGS

[0026] For a better understanding of the embodiments of the present disclosure, and to show how it may be put into effect, reference will now be made, by way of example only, to the accompanying drawings, in which:

[0027] Figure 1 illustrates an example of power consumption during a DRX cycle process;

[0028] Figure 2a illustrates an example of the energy saving that can be made by a RAN node by without utilising the embodiments described herein;

[0029] Figure 2b illustrates an example of the energy saving that can be made by a RAN node by utilising the embodiments described herein. Figure 3 is a flowchart illustrating a method performed by a controller network node;

[0030] Figure 4 illustrates an example implementation of step 301 of Figure 3;

[0031] Figure 5 illustrates an example of a training network node 500 configured to train a reinforcement learning model to determine for a plurality of wireless devices in communication with a RAN node, a respective plurality of discontinuous reception, DRX, cycles;

[0032] Figure 6 illustrates a computer-implemented method for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a RAN node, a respective plurality of discontinuous reception, DRX, cycles;

[0033] Figure 7 is a signalling diagram illustrating an example implementation of the method of Figure 6 being performed by the training network node 500 of Figure 5.

[0034] Figure 8 illustrates a Complementary cumulative distribution function (CCDF) of the cell idle time with and without joint DRX optimization;

[0035] Figure 9 illustrates RAN energy saving performance when embodiments described herein are applied;

[0036] Figure 10 illustrates an apparatus comprising processing circuitry (or logic);

[0037] Figure 11 is a block diagram illustrating a controller network node according to some embodiments;

[0038] Figure 12 is a block diagram illustrating a training network node according to some embodiments.

[0039] DETAILED DESCRIPTION

[0040] Generally, all terms used herein are to be interpreted according to their ordinary meaning in the relevant technical field, unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any methods disclosed herein do not have to be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or where it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, wherever appropriate. Likewise, any advantage of any of the embodiments may apply to any other embodiments, and vice versa. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following description.

[0041] The following sets forth specific details, such as particular embodiments or examples for purposes of explanation and not limitation. It will be appreciated by one skilled in the art that other examples may be employed apart from these specific details. In some instances, detailed descriptions of well-known methods, nodes, interfaces, circuits, and devices are omitted so as not obscure the description with unnecessary detail. Those skilled in the art will appreciate that the functions described may be implemented in one or more nodes using hardware circuitry (e.g., analog and / or discrete logic gates interconnected to perform a specialized function, ASICs, PLAs, etc.) and / or using software programs and data in conjunction with one or more digital microprocessors or general purpose computers. Nodes that communicate using the air interface may have suitable radio communications circuitry. Moreover, where appropriate the technology can additionally be considered to be embodied entirely within any form of computer- readable memory, such as (ROM, EEPROM, Flash memory, a memory disc, RAM etc.) solid-state memory, magnetic disk, or optical disk containing an appropriate set of computer instructions that would cause a processor to carry out the techniques described herein.

[0042] Hardware implementation may include or encompass, without limitation, digital signal processor (DSP) hardware, a reduced instruction set processor, hardware (e.g., digital or analogue) circuitry including but not limited to application specific integrated circuit(s) (ASIC) and / or field programmable gate array(s) (FPGA(s)), and (where appropriate) state machines capable of performing such functions.

[0043] Certain aspects of the present disclosure and their embodiments may provide solutions to these or other challenges. Particular embodiments are described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject matter disclosed herein. The disclosed subject matter should not be construed as limited to only the embodiments set forth herein; rather, these embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.

[0044] For the purposes of the present disclosure, the term “ML model” or reinforcement learning (RL) model encompasses within its scope the following concepts: machine Learning algorithms, for example RL algorithms, comprising processes or instructions through which data may be used in a training process to generate a model artefact for performing a given task, or for representing a real world process or system; the model artefact that is created by such a training process, and which comprises the computational architecture that performs the task; and the process performed by the model artefact in order to complete the task.

[0045] References to “ML model”, “model”, model parameters”, “model information”, etc., may thus be understood as relating to any one or more of the above concepts encompassed within the scope of “ML model”.

[0046] It will be appreciated that a representation of a ML model, e.g. a policy of a reinforcement learning model, may comprise parameters of the ML model, the ML model itself, or any other suitable representation of the ML model that may be utilised by the node receiving the representation appropriately. A representation of an ML model may for example include details of the architecture of the model and values for its trainable parameters, for example number of hidden layers, number of neurons per layer, an activation function or particular neurons and / or weights of neurons.

[0047] Herein the term DRX cycle is utilised to describe the energy saving cycles implemented at a UE. It will be appreciated that any form of energy cycle implemented at a wireless device may be encompassed by this term.

[0048] Embodiments described herein coordinate and / or control DRX cycles of a plurality of wireless devices being served by a RAN node (e.g. a base station). The objective the control is to allow for a longer common non-active DRX time among the plurality of wireless devices thereby leaving more traffic idle time within the cell allowing the RAN node to utilize its deeper sleeping functions more frequently and efficiently. The length and frequency of the common non-active DRX time may therefore be balanced with the services that are provided to the UEs, and to the traffic load in the cell, such that the requirements on the communication service quality can be met.

[0049] Figure 2a illustrates an example of the energy saving that can be made by a RAN node without utilising the embodiments described herein. In contrast, Figure 2b illustrates an example of the energy saving that can be made by a RAN node while utilising the embodiments described herein.

[0050] In this example, a base station (BS) 200 is serving two wireless devices 201 and 202. In this particular example the wireless devices are illustrated as having common DRX cycle lengths. However, it will be appreciated that not all wireless devices served by a RAN network node (e.g. a base station) will utilise a common DRX cycle length. The DRX length and cycle may for example be adapted to a predicted DL data arrival for the specific UE, and to a latency requirement for specific type of communication service to the UE.

[0051] In Figure 2a the active periods 203 and 204 of the wireless devices 201 and 202 are not overlapping at all. This only leaves very small common DRX non-active phases 205. A common DRX non-active phase may be defined as a length of time during which none of the wireless devices served by a RAN node are in an active phase of their respective DRX cycle.

[0052] In contrast, in Figure 2b, the active periods 203 and 204 of the wireless devices 201 and 202 are overlapping. Therefore, the common DRX non-active phase 205 is much longer, allowing for the BS to enter a longer sleep period. It will be appreciated that the active periods do not need to be entirely overlapping in order to improve the length of the common DRX non-active phase 205.

[0053] It will be appreciated that the amount of energy that either a wireless device or a RAN node is able to save during a DRX non-active phase may depend on the sleep modes available at the wireless device or the RAN node, and on how the wireless device or RAN node is configured to transition between these modes. How either a wireless device or RAN node is configured to transition between sleeping modes may be referred to as a sleeping mode process. Consider an example in which the plurality of wireless devices are in a connected mode of operation, e.g. RRC_CONNECTED mode, and so they may follow the timing of the Physical Random Access Channel (PRACH) configuration period of the cell (e.g. as specified in Release 17 with a maximum of 160 ms [Table 8.1 -1 , 3GPP TS 38.213]). At any time, a wireless device m may be in an ON state (e.g. DRX active phase) or OFF state (e.g. DRX non-active phase) as shown in Table 2 below.

[0054] In this example, the OFF state has three potential sleeping modes (SMs). Each SM has three associated parameters: a relative power consumption to deep sleep mode, a transition energy, and a transition time to go to and come out of that sleeping mode. The specific values in table 2 are examples, and the values may be wireless device or implementation specific. Some generic example numbers and models may be found in 3GPP TR 38.840 v 16.0.0. These examples are utilised to exemplify the embodiments described herein.

[0055] Table 2

[0056] Similarly, table 3 below illustrates the possible SMs at a RAN node (e.g. a base station). Again, the OFF state is the non-active phase of the RAN node, and this may only be entered during a common non-active phase among the plurality of wireless devices. Table 3

[0057] Again, the OFF states for the RAN node may be characterized by ramp up energy and transition times (to active the components to operational mode) (e.g. as described in 3GPP TR 38.864 V18.0.1 ).

[0058] Figure 3 is a flowchart illustrating a method performed by a controller network node. The method may be for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. In some examples, the controller network node comprises or is comprised within the RAN node.

[0059] The controller network node may comprise a physical or virtual node, and may be implemented in a computing device or server apparatus and / or in a virtualized environment, for example in a cloud, edge cloud or fog deployment.

[0060] In step 301 the method comprises applying an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period.

[0061] The optimization time period may comprise a time period during which the DRX cycles of the plurality of wireless devices are fixed. The optimisation time period may be of a length of K consecutive DRX cycles, or at least a length of a longest DRX cycle used among the plurality of wireless devices.

[0062] For example, the optimization process may comprise for at least a subset of the plurality of DRX cycles, coordinating a start or a stop time of on durations within the subset of the plurality of DRX cycles. This optimization process may for example result in the scenario illustrated in Figure 2b in which the stop times of the active phases of the DRX cycles in wireless devices 201 and 202 are aligned. It will be appreciated that the start or stop times may not be entirely aligned, for example it may be ensured that all active phases fall within the longest active phase among the plurality of DRX cycles.

[0063] It will also be appreciated that, for example, where a RAN node is serving a large number of wireless devices, coordinating all of the DRX cycles of these wireless devices may result in performance issues due to too many wireless devices being active at once. Therefore, the wireless devices served by a RAN node may be grouped and then coordinated within their groups (or subsets). This may result in less potential energy saving at the RAN node, but may account for some of the performance issues that could be caused by coordinating all of the DRX cycles at the wireless devices served by the RAN node.

[0064] In other examples, the optimization process may comprise utilizing a policy, for example, a policy that has been determined utilizing reinforcement learning, RL. The use of such a policy will be described in more detail with reference to Figures 4 to 6. It will be appreciated that various RL training processes may be utilized update a policy including but not limited to Q-learning, policy gradient methods, and actor-critic methods and the family of algorithms derived from them.

[0065] In step 302 the method of Figure 3 comprises initiating use of the respective plurality of DRX cycles by the plurality of wireless devices.

[0066] For example, step 302 may comprise, where the controller network node is in communication with the RAN node, transmitting one or more parameters of the plurality of DRX cycles for the plurality of wireless devices to the RAN node. In other examples, where the controller network node comprises or is comprised within the RAN node, step 302 may comprise transmitting, for each wireless device, one or more parameters of the respective DRX cycle to the wireless device.

[0067] The one or more parameters may for a first DRX cycle for a first wireless device, m, may comprise one or more of:

[0068] •ma the Long DRX cycle and drx-StartOffset which defines the subframe where the first DRX cycle starts (e.g. drx-LongCylcleStartOffset) • vman active phase duration for the first DRX cycle (e.g. drx-onDurationTimer)

[0069] • xman offset time for the start of the active phase in the first DRX cycle (e.g. drx- SlotOffset)

[0070] • uman inactivity timer value for the first DRX cycle (e.g. a duration within which wireless device m remains in the active phase after receiving a PDCCH (e.g. drx- InactivityTimer)

[0071] As mentioned above, the optimisation process of step 301 may comprise a utilising a policy determined by machine learning, ML, for example reinforcement learning, RL. It will be appreciated that the actual RL process may be performed outside of the controller network node, and the policy, or a representation of the policy, may be communicated to the controller network node for use in performing the method of Figure 3. The training of the policy of the RL model will be described in more detail with reference to Figures 4 to 6.

[0072] In examples in which a ML policy, e.g. a RL policy, is utilised, step 301 may be implemented as illustrated in Figure 4.

[0073] Figure 4 comprises an example implementation of step 301 of Figure 3.

[0074] Step 401 comprises obtaining first state information for a first time period. The first state information may comprise an indication of energy saving made by the plurality of wireless devices and / or by the RAN node during the first time period.

[0075] The first state information may further comprise one or more of: an indication of performance metrics associated with the plurality of wireless devices (for example KPIs of the plurality of wireless devices); an indication of traffic levels associated with the plurality of wireless devices; an indication of a traffic level associated with the RAN node.

[0076] Step 402 comprises inputting the first state information into a policy to determine a first action, wherein the first action updates the respective plurality of DRX cycles. As described above, the policy may be determined using a reinforcement learning, RL, model. In some examples, step 401 comprises receiving from a first wireless device of the plurality of wireless devices an indication of the energy saving made by the first wireless device in the first time period. For example, each of the plurality of wireless devices may transmit to the controller network node an indication of the energy saving made by the respective wireless device (e.g. due to the adoption of the DRX configuration set by the BS for the first time period).

[0077] For example, step 401 may comprise a wireless device may report at the end of its DRX cycle a list of the sleeping modes utilised (e.g. as specified in Table 2) as well the timestamps on activation and deactivation of those modes.

[0078] In another example, step 401 may comprise a first wireless device reporting at the end of its DRX cycle an indication of a value of the energy saving due to utilised radio sleeping features during that DRX cycle.

[0079] In some examples, step 401 may comprise for a first wireless device of the plurality of wireless devices, predicting an energy saving made by the wireless device in the first time period. For example, the first wireless device and the RAN node may agree on a wireless device sleeping mode process (e.g. how and when to switch between the SMs available at the first wireless device). Based on the agreed wireless device sleeping mode process, the BS may locally predict when the various sleeping modes of the first wireless device will be utilised based on its respective DRX cycle and the scheduled traffics in downlink and uplink. It will then be appreciated, that based on whether the various sleeping modes of the firs wireless device are to be used, the base station may be able to locally predict the energy saving at the first wireless device due to radio sleeping.

[0080] In another embodiment, if no relevant energy saving data is available or being reported by the first wireless device, step 401 may comprise the RAN node using a duration of the active phase (ON state) at the first wireless device divided by the DRX cycle length (e.g. vm / m) as a measure of the energy saving of a first wireless device.

[0081] In some examples, the energy saving the of a wireless device may be represented by a distribution of traffic idle time (e.g. when wireless device does not have any UL or DL traffic including user and control signals) as well as the non-active phase of the DRX cycles. The benefit of this example may be that explicit feedback from the wireless device to predict energy saving may not be required, and the distribution of the idle times at the wireless device may be derived based on PDCCH, PDSCH, PUCCH PUSCH, and the UE DRX configuration, which are all available locally at the RAN node (and may be easily obtained by the controller network node).

[0082] In some examples, the optimization process may be further configured to determine one or more sleeping mode process parameters for the RAN node and / or one or more sleeping mode process parameters for at least one of the plurality of wireless devices. In other words, the optimization process may be able to update how the plurality of wireless devices or the RAN node transition between various SMs when in the DRX nonactive phase.

[0083] It will be appreciated that, once the policy is trained, the policy for the RAN node may remain fixed until a change in the environment of the RAN node triggers a retraining of the policy. For example, the method of Figure 3 may further comprise updating the policy responsive to a change in one or more of: a number of wireless devices in the plurality of wireless devices being served by the RAN node, the traffic being served by the RAN node or, one or more performance requirements at the plurality of wireless devices; and link capacities between the RAN node and the plurality of wireless devices. It will be appreciated that the policy may be re-trained by a training network node (e.g. as will be described later in reference to Figure 6) and the updated policy may be communicated to the controller network node.

[0084] Figure 5 illustrates an example of a training network node 500 configured to train a reinforcement learning model to determine for a plurality of wireless devices in communication with a RAN node, a respective plurality of discontinuous reception, DRX, cycles.

[0085] In this example, the training network node 500 may comprise a reinforcement learning orchestrator, RLO, 501 . The RLO 501 may determine one or more actions (e.g. updates to be applied to the DRX cycles utilised by the plurality of wireless devices and / or the RAN node).

[0086] The training network node 500 may further comprise a data processing unit 502. The data processing unit 502 may retrieve observations (e.g. data) from the environment (e.g. from the plurality of wireless devices) and may provide a reward to the RLO. The data processing unit 502 may also provide inputs to a translation unit 503 in the training node 500. The translation unit 503 may generate state information for the environment from the data provided by the data processing unit 502 and / or indicate one or more potential constraint violations to the RLO 501 . It will be appreciated that the data obtained from the environment may be enriched with additional side information relevant for, for example, the traffic experienced by the wireless devices. For example, if the traffic is periodic, this information can improve any traffic modeling performed by the translation unit 503.

[0087] In some examples, the translation unit 503 and / or the data processing unit 502 may be deployed at the RAN node, the controller network node, or the edge cloud.

[0088] The data processing unit 502 may also receive an indication of the configuration of a sleeping mode process for the RAN node and / or for the plurality of wireless devices from an Intent Management Function (IMF), Operations Support System (OSS) or other relevant network function.

[0089] It will be appreciated that one or more of the components of the training network node may be stored within a Network Data Analytics Function (NWDAF).

[0090] Figure 6 illustrates a computer-implemented method for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a RAN node, a respective plurality of discontinuous reception, DRX, cycles.

[0091] The method of Figure 6 may be performed by a training network node, for example the training network node 500 illustrated in Figure 5. It will be appreciated that the training network node may comprise a physical or virtual node, and may be implemented in a computing device or server apparatus and / or in a virtualized environment, for example in a cloud, edge cloud or fog deployment.

[0092] The method of Figure 6 comprises performing a training loop process. The training loop process may be performed over an optimisation period. The optimisation period may have a length of K, where K is an integer value, consecutive DRX cycles in at least one of the wireless devices. It will be appreciated that the length of the optimisation period may be set to a fixed value, where the fixed value is at least as big as a maximum DRX cycle length. During an optimisation period, the configuration of DRX cycles at the plurality of wireless devices remains the same (e.g. is not updated by the policy). At the end of the optimisation period, the training loop process may be performed and both the policy and the configuration of DRX cycles at the plurality of wireless devices may be updated.

[0093] The training loop process comprises the steps 601 to 606 of Figure 6.

[0094] Step 601 comprises obtaining first state information for a first time period. The first time period may comprise an optimization period as described above. The first state information may comprise an indication of an energy saving made by the plurality of wireless devices and by the RAN node during the first time period. The first state information may be obtained by the data processing unit 501 collecting information from the plurality of wireless devices and / or the RAN node. The translation unit 503 may then determine the first state information.

[0095] It will be appreciated that, as described with reference to step 401 of Figure 4, the indication of the energy saving made by the plurality of wireless devices may be predicted (e.g. by the translation unit 503) or may be reported by the respective wireless devices. Alternatively or additionally, the indication of the energy may comprise an indication of the traffic idle time that may be considered representative of the energy saving made by a wireless device.

[0096] The first state information may further comprise one or more of: an indication of performance metrics associated with the plurality of wireless devices (e.g. Key Performance Indicators (KPIs) that may be reported by the wireless devices to the RAN node); an indication of traffic levels associated with the plurality of wireless devices (e.g. an average traffic level associated with each wireless device); and an indication of a traffic level associated with the RAN node (e.g. average traffic level at the cell).

[0097] In addition to the first state information, the method of Figure 6 may obtain additional information relating to: one or more parameters describing service constraints, for example, minimum KPI requirements for the plurality of wireless devices; one or more parameters describing constraints on the actions, for example, feasible configurations of DRX cycles as well as maximum feasible lengths for a DRX cycle. For the latter, for example, the DRX cycle may be constrained to be less than or equal to the PRACH configuration period, otherwise the UE may not stay in the RRC_CONNECTED mode; one or more parameters describing sleeping mode processes of the plurality of wireless devices and / or the RAN node as well as (either algorithms describing seeping mode selection or availability of measurement / reporting the energy saving due to sleeping modes).

[0098] For example, the additional information may be obtained at the translation unit 503 And / or the data processing unit from the IMF / OSS.

[0099] In step 602 the method comprises inputting the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles. It will be appreciated that step 602 may be performed by the RLO of the training network node 500.

[0100] Considering [M] ■■= {1,2, as a set of M of the plurality of wireless devices.

[0101] The one or more parameters for a first DRX cycle for a first wireless device, m may then comprises one or more of:

[0102] •ma Long DRX cycle and drx-StartOffset which defines the subframe where the first DRX cycle starts (e.g. drx-LongCylcleStartOffset)

[0103]

[0104] • vm- an active phase duration for the first DRX cycle (drx-onDurationTimer)

[0105] • xman offset time for the start of the active phase in the first DRX cycle (drx- SlotOffset)

[0106] • uman inactivity timer value for the first DRX cycle (e.g. a duration within which wireless device m remains in the active phase after receiving a PDCCH (drx- InactivityTimer)

[0107] It will therefore be appreciated that the first action may comprise updates to any of the above parameters for any of the plurality of DRX cycles. Not all DRX cycles may be adjusted by the first action.

[0108] As previously noted, in some examples the RAN node sleeping mode process and / or the individual wireless devices sleeping mode processes may not be optimised by the method of Figure 6. It may be assumed that the RAN node sleeping model process is configured by, for example, a network data analytics function (NWDAF) or another relevant network function, or the final sleeping decisions timing and sleeping depths are available for observation and processing by the data processing unit 502. Nevertheless, the optimized DRX configurations (i.e., tm, vm, xm, um) may help sleep mode processes make informed decisions for their traffic prediction.

[0109] In other examples, other parameters may be included within the action space, including for example, parameters relating to the RAN node sleeping mode process, or indeed the wireless device sleeping mode processes.

[0110] It will be appreciated that the policy may be constrained such that the first action cannot cause the one or more updated parameters to violate one or more constraints. As described above the one or more constraints may be obtained from an IMF / OSS 504.

[0111] For example, the one or more constraints may comprise one or more of the constraints C1 to C4 listed below:

[0112] • C1 : rm, vm, um, Vm e [M] be in the set of specified values by 3GPP for long cycle, ON-duration, and inactivity time, respectively,

[0113] • C2: vm< rm, Vm e [M],

[0114] • C3: xm< vm-m, Vm e [M],

[0115] • C4: service requirements of UE m, Vm e [M]

[0116] In step 603 the method comprises applying the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices. It will be appreciated that step 603 may comprise transmitting an indication of the one or more updated parameters to the RAN node. The indication may comprise the one or more updated parameters or may indicate the change to apply to the previous one or more parameters.

[0117] In step 604 the method comprises obtaining second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period. Again, the second time period may comprise another optimization time period. It will be appreciated that the second state information may comprise similar parameters to the first state information.

[0118] In step 605 the method comprises determining a reward for the second state information utilizing a reward function. The reward function may comprise a plurality of component parts. For example, the reward function may comprise a first component relating to energy saving made by the RAN node.

[0119] For example, the first component may be defined as: fb= Ub([rm, as the utility function of the RAN node, which may be defined as the energy saving possible given the plurality of DRX cycles, for example, given the BS sleeping modes followed by 3GPP models (e.g. in 3GPP TR 38.864).

[0120] The reward function may in some examples comprise for a wireless device in the plurality of wireless devices, a second component relating to energy saving made by the wireless device. For example, the reward function may comprise a second component for each wireless device in the plurality of wireless devices. For example, the collective second components for the plurality of wireless devices may be defined as:

[0121] / UE where Umis the utility function of wireless device m over the optimisation period.

[0122] Without loss of generality, the energy model of [3GPP TR 38.840] may be used to define Umas the energy saving when activating UE sleeping modes.

[0123] In some examples, the reward function may comprise for a wireless device in the plurality of wireless devices, a third component relating to whether a performance metric at the wireless device is meeting a criterion. For example, the energy saving at the wireless device may only be taken into account if the wireless device is meeting one or more KPI constraints. For example, the collective second components and third components for the plurality of wireless devices may be defined as: if all other KPIs of all UEs are met > if some of other KPIs of all UEs are not met where c may be defined as a large number to guide the RLO to take actions that lead to maintaining the KPI constraints. Here other KPIs may be latency, throughput, etc. for various UEs. For example, c may be set to infinity, to ensure whenever the “if some of other KPIs of all UEs are not met” condition is triggered, fUEwill be negative infinity, so the RL agent learns that is a very very bad action that should avoided. The infinity can be any big number of course as long as it is much bigger (say 10x) than the value of fUEwhen the other condition is met (i.e., “if all other KPIs of all UEs are met”).

[0124] In examples in which the energy saving of a wireless device is represented by the distribution of traffic idle time (e.g. when UE does not have any UL or DL traffic including user and control signals). A reward function may comprise a fourth component, for example a probability that the traffic idle time of the wireless device is greater than a given positive constant, T. For example, the fourth component may be defined as:

[0125] Pr(Idle Time of UE m > T), for a given positive constant T, where T is typically selected based on the sleep mode transition times specified in Table 2. For example, T may be selected as one of the transition time values from Table 2, e.g. the transition time value for light sleep or deep sleep modes. Defining the fourth component as Pr(Idle Time of UE m > T) therefore indicates the probability of having enough idle time for wireless device m to ensure light sleep (if we set T as transition time of the light sleep) or deep sleep (if we set T as the transition time of the deep sleep).

[0126] The reward function, R may comprise a convex combination of the at least the first component and the combination of second components, for example:

[0127] R =afuE + (1 -a)fb foragiven constant a e [0,1],

[0128] For example, the IMF may be able to provide information relating to the multiplier, a, in the reward function, R as well as other components (e.g. third or fourth components) to be included in the reward function, R.

[0129] As an alternative embodiment, multi-objective RL may be utilised to make a trade-off between fUEand fb.

[0130] In step 606 the method comprises adjusting the policy according to the reward. It will be appreciated that there are many different methods available for performing RL to update and adjust a policy based on a reward function.

[0131] It will be appreciated that the method of Figure 6 may then further comprise repeating the training loop process until the policy converges. In this repeat, the second state information may be set as the first state information for inputting into the policy in step 602. Once the policy has converged, the training network node may provide an indication of the policy to a controller network node (e.g, the RAN node) for performing the method of Figure 3.

[0132] Figure 7 is a signalling diagram illustrating an example implementation of the method of Figure 6 being performed by the training network node 500 of Figure 5.

[0133] In step 701 , the data processing unit 502 obtains an indication of the reward function and indicating of at least the RAN node sleeping mode process from the OSS 504.

[0134] In step 702, the translation unit obtains an indication of objectives and constraint descriptions from the OSS 504.

[0135] In step 703, the data processing unit obtains information form the environment (e.g. the plurality of wireless devices and / or the RAN node), and logs this information with the translation unit 503.

[0136] In step 704, the data processing unit utilises the reward function obtained in step 701 to determine a reward. Step 704 comprises an example implementation of step 605 of Figure 6.

[0137] In response to receiving the reward in step 605 the RLO may update the policy of the RLO (e.g. an example implementation of step 606 of Figure 6).

[0138] In step 705, the RLO receives state information and optionally an indication of any constraint violations by any of the wireless devices. The RLO 501 may then utilise the state information and the policy to determine the first action (e.g. an example implementation of step 602)

[0139] In step 706, the RLO transmits an indication of the first action to the RAN node. Step 706 comprises an example implementation of step 603 of Figure 6. In step 707, the RAN node updates the parameters of the DRX cycles for the plurality of wireless devices according to the first action.

[0140] At the end of the optimisation period, the RAN node may then report, in step 708, one or more metrics associated with the previous optimisation period to the data processing unit 502. For example the RAN node may report energy saving made by the RAN node during the optimisation period. The information provided in step 708 may then form the basis of the state information for the next training loop process.

[0141] Experimental Results

[0142] Embodiments described herein have been implement on a proprietary network-level symbol-based simulator. Sleeping mode processes described in 3GPP for both a BS and wireless devices have been assumed. Table 5 illustrates a summary of the simulation parameters utilised.

[0143] Table 5: main simulation parameters.

[0144] A cell idle time may be defined as a the time during which the cell has no incoming or outgoing traffic (in uplink or downlink). We also define r as the ratio between ON duration and DRX cycle.

[0145] Figure 8 is a Complementary cumulative distribution function (CCDF) of the cell idle time with and without joint DRX optimization. The 6 ms and 50 ms lines are the minimum durations required to activate light and deep sleep, respectively (see Table 1 ). In particular, Figure 8 illustrates the distribution of the cell idle time with and without performing embodiments described herein. Instances in which the embodiments described herein were utilised are labelled “coor”, instances in which no optimisation was performed are labelled and “un-coor”.

[0146] According to Table 1 and noting the values of minimum 6ms to use light sleep and 50ms to use deep sleep, and for r=1 / 32, in comparison to the uncoordinated baseline, optimising the DRX configuration for all UEs enhances the opportunities for light sleep and deep sleep by over 2% and 36%, respectively.

[0147] Figure 9 illustrates RAN energy saving performance when embodiments described herein are applied, in this example, where the plurality of wireless devices comprises 20 wireless devices. In particular Figure 9 illustrates that the actions taken according to embodiment described herein lead to a substantial energy saving gain compared (more than 9 times) to the uncoordinated DRX baseline without violating the delay requirements of the service.

[0148] Figure 10 illustrates an apparatus 1000 comprising processing circuitry (or logic) 1001. The processing circuitry 1001 controls the operation of the apparatus 1000 and can implement the method described herein in relation to an apparatus 1000. The processing circuitry 1001 can comprise one or more processors, processing units, multicore processors or modules that are configured or programmed to control the apparatus 1000 in the manner described herein. In particular implementations, the processing circuitry 1001 can comprise a plurality of software and / or hardware modules that are each configured to perform, or are for performing, individual or multiple steps of the method described herein in relation to the apparatus 1000. It will be appreciated that the apparatus 1000 may comprise one or more virtual machines running different software and / or processes. The apparatus 1000 may therefore comprise, or be implemented in or as one or more servers, switches and / or storage devices and / or may comprise cloud computing infrastructure that runs the software and / or processes.

[0149] Optionally, the apparatus 1000 may comprise a memory 1003. In some embodiments, the memory 1003 of the apparatus 1000 can be configured to store instructions (e.g. program code) executable by the processing circuitry 1001 of the apparatus 1000 whereby the apparatus is operable to perform the method as described with reference to any one of more of Figures 3, 4 and 6.

[0150] Alternatively or in addition, the memory 1003 of the apparatus 1000, can be configured to store any requests, resources, information, data, signals, or similar that are described herein. The processing circuitry 1001 of the apparatus 1000 may be configured to control the memory 1003 of the apparatus 1000 to store any requests, resources, information, data, signals, or similar that are described herein.

[0151] In some embodiments, the apparatus 1000 may optionally comprise a communications interface 1002. The communications interface 1002 of the apparatus 1000 can be for use in communicating with other nodes, such as other virtual nodes. For example, the communications interface 1002 of the apparatus 1000 can be configured to transmit to and / or receive from other nodes requests, resources, information, data, signals, or similar. The processing circuitry 1001 of apparatus 1000 may be configured to control the communications interface 1002 of the apparatus 1000 to transmit to and / or receive from other nodes requests, resources, information, data, signals, or similar. The communications interface 1002 can use any suitable communication technology.

[0152] The apparatus 1000 may be configured operate in the manner described herein in respect of a training network node or a controller network node.

[0153] Figure 11 is a block diagram illustrating a controller network node 1 100 according to some embodiments. The controller network node 1100 can control a plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node. The control network node 1 100 comprises an applying module 1102 configured to apply an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period. The controller network node 1 100 comprises an initiating module 1 104 configured to initiate use of the respective plurality of DRX cycles by the plurality of wireless devices. The controller network node 1100 may operate in the manner described herein in respect of a controller network node. Figure 12 is a block diagram illustrating a training network node 1200 according to some embodiments. The training network node 1200 can train a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a radio access network, RAN, node, a respective plurality of DRX cycles. The training network node 1200 may be configured to perform a training loop process utilising the modules listed below. The training network node comprises a first obtaining module 1202 configured to obtain first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period. The training network node 1200 comprises an inputting module 1204 configured to input the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles. The training network node 1200 comprises can applying module 1206 configured to apply the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices. The training network node comprises a second obtaining module 1208 configured to obtain second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period. The training network node 1200 comprises a determining module 1210 configured to determine a reward, Rt+i, for the second state information utilizing a reward function. The training network node 1200 comprises an adjusting module 1212 configured to adjust the policy according to the reward. The training network node 1200 may operate in the manner described herein in respect of a training network node.

[0154] There is also provided a computer program comprising instructions which, when executed on a least one processor (such as the processing circuitry 1001 of the apparatus 1000 described earlier), cause the processor to carry out at least part of the method(s) described herein. According to some embodiments there is provided a carrier containing the computer program. In some embodiments, the carrier can be any one of an electronic signal, an optical signal, an electromagnetic signal, an electrical signal, a radio signal, a microwave signal, or a computer-readable medium. There is also provided a (for example, tangible and / or non-transient) computer-readable medium comprising instructions which, when executed by at least one processor, cause the at least one processor to perform at least part of the method(s) described herein.

[0155] It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word “comprising” does not exclude the presence of elements or steps other than those listed in a claim, “a” or “an” does not exclude a plurality, and a single processor or other unit may fulfil the functions of several units recited in the claims. Any reference signs in the claims shall not be construed so as to limit their scope.

Claims

CLAIMS1. A method performed by a controller network node for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node, the method comprising: applying (301 ) an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period; and initiating (302) use of the respective plurality of DRX cycles by the plurality of wireless devices.

2. The method as claimed claim 1 , wherein initiating use of the respective plurality of DRX cycles by the plurality of wireless devices comprises transmitting one or more parameters of the plurality of DRX cycles for the plurality of wireless devices to the RAN node.

3. The method as claimed in claim 1 wherein initiating use of the respective plurality of DRX cycles by the plurality of wireless devices comprises: for each wireless device, transmitting one or more parameters of the respective DRX cycle to the wireless device.

4. The method as claimed in claim 2 or 3, wherein the one or more parameters for a first DRX cycle for a first wireless device, m, comprise one or more of: an offset time for the start of the first DRX cycle rm, an active phase duration for the first DRX cycle, vman offset time for the start of the active phase in the first DRX cycle xm, and an inactivity timer for the first DRX cycle um5. The method as claimed in claim 1 to 3, wherein applying the optimization process comprises: obtaining (401 ) first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period;inputting (402) the first state information into a policy to determine a first action, wherein the first action updates the respective plurality of DRX cycles, and wherein the policy is determined using a reinforcement learning, RL, model.

6. The method as claimed in claim 4 wherein the first state information further comprises one or more of: an indication of performance metrics associated with the plurality of wireless devices; an indication of traffic levels associated with the plurality of wireless devices; an indication of a traffic level associated with the network node.

7. The method of claim 4 or 5, further comprising: training the reinforcement leaning model according to any one of claims 13 to 19.

8. The method as claimed in any one of claims 4 or 5, comprising updating the policy responsive to a change in one or more of: a number of wireless devices in the plurality of wireless devices being served by the RAN node, traffic being served by the RAN node; one or more performance requirements at the plurality of wireless devices; and link capacities between the RAN node and the plurality of wireless devices.

9. The method as claimed in any one of claims 4 to 7, wherein obtaining the first state information comprises: receiving from a wireless device of the plurality of wireless devices an indication of the energy saving made by the wireless device in the first time period.

10. The method as claimed in any one of claims 4 to 7, herein obtaining the first state information comprises: for a wireless device of the plurality of wireless devices, predicting the energy saving made by the wireless device in the first time period.11 . The method as claimed in any one of claims 1 to 3, wherein the optimization process comprises: for at least a subset of the plurality of DRX cycles, coordinating a start time or a stop time of active phases within the subset of the plurality of DRX cycles.

12. The method of claim 1 to 10, wherein the optimization process further determines one or more sleeping mode process parameters for the RAN node.

13. The method of any one of claims 1 to 1 1 , wherein the optimization process further determines one or more sleeping mode process parameters for at least one of the plurality of wireless devices.

14. A computer implemented method for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a radio access network, RAN, node, a respective plurality of DRX cycles, the method comprising: performing a training loop process comprising: obtaining (601 ) first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period; inputting (602) the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles; applying (603) the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices; and obtaining (604) second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period; determining (605) a reward, Rt+i, for the second state information utilizing a reward function, and adjusting (606) the policy according to the reward.

15. The method as claimed in claim 13, wherein the first state information and / or second state information further comprises one or more of: an indication of performance metrics associated with the plurality of wireless devices; an indication of traffic levels associated with the plurality of wireless devices; an indication of a traffic level associated with the RAN node.

16. The method as claimed in claim 13 or 14, wherein the reward function comprises: a first component relating to energy saving made by the RAN node.

17. The method as claimed in claim 13 to 15, wherein the reward function comprises: for a wireless device in the plurality of wireless devices, a second component relating to energy saving made by the wireless device.

18. The method as claimed in any one of claims 13 to 16, wherein the reward function comprises: for a wireless device in the plurality of wireless devices, a third component relating to whether a performance metric at the wireless device is meeting a criterion.

19. The method as claimed in any one of claims 13 to 17, wherein the policy is constrained such that the first action cannot cause the one or more updated parameters to violate one or more constraints.

20. The method as claimed in any one of claims 13 to 18, further comprising repeating the training loop process until the policy converges.

21. The method as claimed in any one of claims 13 to 19 wherein the method is performed by a training network node.

22. The method as claimed in any one of claims 13 to 20 further comprising transmitting an indication of the policy to a controller network node.

23. A controller network node (1 100) for controlling plurality of wireless devices to utilize a respective plurality of discontinuous reception, DRX, cycles for communication with a radio access network, RAN, node, the controller networknode comprising processing circuitry (1 101 ) and a memory (1103), the memory containing instructions executable by the processing circuitry whereby the controller network node is operable to: apply (301 ) an optimization process to determine the plurality of discontinuous reception, DRX, cycles, wherein the optimization process acts, at least in part, to increase a length of a common DRX non-active phase among the plurality of DRX cycles during an optimization time period; and initiate (302) use of the respective plurality of DRX cycles by the plurality of wireless devices.

24. The controller network node as claimed in claim 23 wherein the memory further contains instructions executable by the processing circuitry whereby the controller network node is operable to perform the method as claimed in any one of claims 2 to 13.

25. A training network node (1 100) for training a reinforcement learning, RL, model to determine, for a plurality of wireless devices in communication with a radio access network, RAN, node, a respective plurality of DRX cycles, the training network node comprising processing circuitry (1101 ) and a memory (1 103), the memory containing instructions executable by the processing circuitry whereby the training network node is operable to: perform a training loop process comprising: obtaining (601 ) first state information for a first time period, the first state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the first time period; inputting (602) the first state information into a policy of the RL model to determine a first action, wherein the first action updates one or more parameters of the plurality of DRX cycles; applying (603) the action by initiating indication of the one or more updated parameters to one or more relevant wireless devices in the plurality of wireless devices; and obtaining (604) second state information for a second time period, the second state information comprising an indication of energy saving made by the plurality of wireless devices and by the RAN node during the second time period;determining (605) a reward, Rt+i, for the second state information utilizing a reward function, and adjusting (606) the policy according to the reward.

26. The training network node as claimed in claim 25 wherein the memory further contains instructions executable by the processing circuitry whereby the controller network node is operable to perform the method as claimed in any one of claims 15 to 22.

27. A computer program, comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out a method according to any of claims 1 to 22.

28. A carrier containing the computer program according to claim 27, wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.

29. A computer-readable medium comprising instructions that, when executed on at least one processor, cause the at least one processor to perform the method according to any of claims 1 to 22.

30. A computer program product comprising non transitory computer readable media having stored thereon a computer program according to claim 27.

Citation Information

Patent Citations

  • 5G terminal power consumption optimization method and device based on discontinuous reception, and medium

    CN117768982A

  • DRX configuration method, terminal device, network device and communication system

    EP3609243A1

  • Base station discontinuous reception and transmission design for energy saving network

    EP4346147A1

  • Simultaneous active time modification for a plurality of ue

    WO2022023123A1

  • Methods and apparatuses for wireless communication in connected discontinuous reception mode

    WO2023278026A1