Intelligent hardware control for ran energy efficiency

By using RL and heuristic-based mechanisms to control RAN hardware energy savings, the solution addresses inefficiencies in existing RAN energy management, balancing energy consumption and user experience through adaptive and intelligent control.

WO2025224476A1PCT designated stage Publication Date: 2025-10-30TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/053914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing RAN energy efficiency solutions prioritize performance over energy savings, leading to reduced user experience and inefficiencies due to complex interactions between hardware and software states, and lack of adaptability to dynamic network conditions.

Method used

Implementing a reinforcement learning (RL) policy and heuristic-based mechanisms to control RAN hardware energy savings features, considering both application software and hardware states, with configurable objectives to balance energy consumption and user experience.

Benefits of technology

Achieves an optimal tradeoff between energy efficiency and user experience by learning from network conditions and adapting to dynamic environments, minimizing energy consumption without compromising performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024053914_30102025_PF_FP_ABST
    Figure IB2024053914_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for intelligent control of energy savings features of compute hardware of a Radio Access Network (RAN) node are disclosed. In one embodiment, a method is disclosed which is performed by an evaluation entity for controlling one or more energy savings features of a hardware component of a RAN node where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively. The method comprises obtaining a first set of inputs related to a RAN application state of the RAN application software components and obtaining a second set of inputs related to a hardware state of a hardware component of the RAN node. The method further comprises determining outputs for control of the energy saving features of the hardware component, based on the first and second sets of inputs.
Need to check novelty before this filing date? Find Prior Art

Description

INTELLIGENT HARDWARE CONTROL FOR RAN ENERGY EFFICIENCYTechnical Field

[0001] The present disclosure relates to a Radio Access Network (RAN) of a cellular communications system and, more specifically, to control of hardware energy savings of a RAN node.Background

[0002] In a cellular communications system such as, e.g., a 3rdGeneration Partnership Project (3GPP) 5thGeneration (5G) system, energy efficiency of the Radio Access Network (RAN) is of great concern to Mobile Network Operators (MNOs) and equipment vendors, given the operational cost and push for sustainable technology. To achieve a high performing RAN service, a combination of complex software (SW) and powerful compute hardware (HW) is required. Often, the compute hardware can be controlled to reduce the energy consumption via what are hereinafter referred to as energy saving features, but this reduction in energy consumption is at the cost of simultaneously reducing the processing capacity of the compute hardware. Therefore, to deliver excellent user-experience in the network, it is conventional to prioritize performance over energy efficiency.Summary

[0003] Systems and methods for intelligent control of energy savings features of compute hardware of a Radio Access Network (RAN) node are disclosed. In one embodiment, a method performed by an evaluation entity is disclosed for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node. The method comprises obtaining a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node, and obtaining a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node. The method further comprises determining a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs. In this manner, the one or more energy saving features of the hardware component can be controlled in an intelligent manner, e.g., to achieve a desired objective.

[0004] In one embodiment, the method further comprises providing the set of outputs to the hardware component of the RAN node for control of the one or more energy savings features of the hardware component of the RAN node.

[0005] In one embodiment, the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following: (a) Packet Data Convergence Protocol (PDCP) traffic load, (b) Medium Access Control (MAC) throughput, (c) MAC layer and / or Radio Link Control (RLC) layer latency, (d) Physical Resource Block (PRB) utilization, (e) RLC buffer size, (f) one or more inputs about a state of one or more User Equipments (UEs) scheduled for a transmit time interval, (g) average or combined state of multiple UEs scheduled for a transmit time interval, and (h) a predicted value for any one or more of (a)-(g).

[0006] In one embodiment, the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following: one or more configurations of the RAN node, one or more configurations of a cell operated by the RAN node, a carrier frequency of a cell operated by the RAN node, an indication of whether carrier aggregation is being utilized, a Multiple Input Multiple Output (MIMO) configuration of the RAN node, a number of cells operated by the RAN node, and a cell bandwidth of each of one or more cells operated by the RAN node.

[0007] In one embodiment, the second set of inputs related to the hardware state of the hardware component of the RAN node comprise any one or more of the following: core utilization of one or more multi-core processors comprised on the hardware component, core frequency of one or more cores of the one or more multi-core processors comprised in the hardware component, power consumption of the hardware component, power consumption of the one or more cores of the one or more multi-core processors comprised in the hardware component, a residence time in a particular state of the hardware component, and power consumption of one or more fans used for cooling the one or more multi-core processors comprised in the hardware component.

[0008] In one embodiment, the set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node comprise any one or more of the following: an output for controlling a core sleep state of one or more cores of one or more multi-core processors comprised in the hardware component, an output for controlling a clock frequency of one or more cores of the one or more multi-core processors comprised in the hardware component, and an output for controlling an uncore clock frequency of the one or more multi-core processors comprised in the hardware component.

[0009] In one embodiment, determining the set of outputs for control of the one or more energy saving features of the hardware component comprises applying the first set of inputs and the second set of inputs to a heuristic model that outputs the set of outputs for control of the one or more energy savings features of the hardware component based on the first and second sets of inputs. In one embodiment, the heuristic model comprises one or more rules that are configurable by a network operation of the wireless communication system. In one embodiment, the one or more rules comprise one or more rules that define a direct relationship between an amount of traffic or expected amount of traffic in one or more cells operated by the RAN node and which of the one or more energy savings features are to be enabled. In one embodiment, the one or more rules comprise any one or more of the following: one or more rules that define a mapping between different traffic levels in a cell operated by the RAN node and different sets of outputs that provide different setting of the one or more energy savings features, one or more rules that selects the set of outputs based on a capacity of the hardware component and expected traffic demand in the cell operated by the RAN node, and one or more rules that select the set of outputs based on a user experience metric.

[0010] In one embodiment, determining the set of outputs for control of the one or more energy saving features of the hardware component comprises applying a reinforcement learning (RL) policy to a combined state formed by the first set of inputs and the second set of inputs to thereby obtain a probability distribution for a plurality of different sets of possible outputs for the combined state achieving a desired result and selecting, based on the probability distribution, one of the plurality of different sets of possible outputs having a highest probability of achieving the desired result as the set of outputs for control of the one or more energy savings features of the hardware component. In one embodiment, the desired result is an optimal tradeoff between user experience of users of a plurality of UEs served by the RAN node and power consumption of the RAN node.

[0011] In one embodiment, the RL policy is trained to achieve an objective that is configurable by a network operator of the wireless communication system.

[0012] In one embodiment, the RL policy is trained based on a reward function that is configurable by a network operator of the wireless communication system. In one embodiment, the reward function is a function of one or more metrics indicative of user experience of users associated to UEs served by the RAN node and one or more metrics indicative of power consumption of the hardware component. In another embodiment, the reward function is a weighted sum of one or more metrics indicative of user experience of users associated to UEs served by the RAN node and one or more metrics indicative of power consumption of the hardware component. In another embodiment, the reward function is a discontinuous function of one ormore metrics indicative of user experience of users associated to UEs served by the RAN node. In another embodiment, the reward function is a state-adjusted reward function.

[0013] In one embodiment, the method further comprises obtaining a third set of inputs related to an environment of the hardware component of the RAN node and / or an environment of the RAN node, wherein determining the set of outputs for control of the one or more energy saving features of the hardware component is further based on the third set of inputs. In one embodiment, the third set of inputs comprises any one or more of the following: a location of the hardware component, a location of the RAN node, a current time of day, and current weather information for geographic area in which the hardware component and / or RAN node is located.

[0014] In one embodiment, the steps of obtaining the first set if inputs, obtaining the second set of inputs, and determining the set of outputs is repeated at a predefined or configured frequency.

[0015] In one embodiment, the evaluation entity is implemented at the RAN node.

[0016] In one embodiment, the evaluation entity is implemented at computing node that is separate from the RAN node.

[0017] Corresponding embodiments of a node for implementing an evaluation entity for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system are disclosed, where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node. In one embodiment, the node for implementing the evaluation entity is adapted to obtain a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node, obtain a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node, and determine a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs. In one embodiment, the node is the RAN node. In another embodiment, the node is a computing node that is separate from the RAN node.

[0018] In another embodiment, a node for implementing an evaluation entity for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system, where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, comprises processing circuitry configured to cause the node to obtain a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node, obtain a second set of inputs related to ahardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node, and determine a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs. In one embodiment, the node is the RAN node. In another embodiment, the node is a computing node that is separate from the RAN node.

[0019] Embodiments of a method performed by a training entity for training a RL policy for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system are also disclosed, where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node. In one embodiment, the method comprises obtaining a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node and obtaining a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node. The method further comprises calculating, based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node. The method further comprises selecting a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node and setting the one or more energy savings features of the hardware component based on the selected set of outputs.

[0020] In one embodiment, the method further comprises updating the RL policy at a previous combined state and the previous selected set of outputs based on the calculated reward and the current combined state and applying the RL policy to the combined state to thereby obtain, for each of a plurality of different sets of possible outputs for the combined state, a probability of that set of possible outputs providing a desired result. Further, selecting the new set of outputs comprises selecting one of the plurality of different sets of possible outputs as the new set of outputs for controlling the one or more energy savings features of the hardware component.

[0021] In one embodiment, the method further comprises sending the reward, the current combined state, and the previous selected set of outputs to a policy update entity.

[0022] In one embodiment, the method further comprises performing a plurality of iterations of the steps to thereby train the RL policy.

[0023] In one embodiment, the training entity is implemented at the RAN node.

[0024] In another embodiment, the training entity is implemented at computing node that is separate from the RAN node.

[0025] Embodiments of a node for implementing a training entity for training a RL policy for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system are also disclosed, where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node. In one embodiment, the node for implementing the training entity is adapted to obtain a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node and obtain a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node. The node is further adapted to calculate, based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node. The node is further adapted to select a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, and set the one or more energy savings features of the hardware component based on the selected set of outputs. In one embodiment, the node for implementing the training entity is the RAN node. In another embodiment, the node for implementing the training entity is a computing node that is separate from the RAN node.

[0026] In another embodiment, a node for implementing a training entity for training a RL policy for controlling one or more energy savings features of a hardware component of a RAN node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, comprises processing circuitry configured to cause the node to obtain a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node and obtain a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node. The processing circuitry is further configured to cause the node to calculate, based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energysavings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node. The processing circuitry is further configured to cause the node to select a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, and set the one or more energy savings features of the hardware component based on the selected set of outputs. In one embodiment, the node for implementing the training entity is the RAN node. In another embodiment, the node for implementing the training entity is a computing node that is separate from the RAN node.Brief Description of the Drawings

[0027] The accompanying drawing figures incorporated in and forming a part of this specification illustrate several aspects of the disclosure, and together with the description serve to explain the principles of the disclosure.

[0028] Figure 1 illustrates one example embodiment of a Radio Access Network (RAN) node;

[0029] Figure 2 illustrates the general structure of Markov Decision Process (MDP);

[0030] Figure 3 illustrates a MDP model for applying Reinforcement Learning (RL) to optimize control of one or more energy saving features available on compute hardware of a RAN node, in accordance with an embodiment of the present disclosure;

[0031] Figure 4 is a flow chart that illustrates the operation of an energy efficiency agent for online training of a RL policy for control of one or more energy saving features of compute hardware of a RAN node, in accordance with an embodiment of the present disclosure;

[0032] Figure 5 is a flow chart that illustrates the operation of the energy efficiency agent for offline training of the RL policy for control of one or more energy saving features of compute hardware of a RAN node, in accordance with an embodiment of the present disclosure;

[0033] Figure 6 a flow chart that illustrates the operation of the energy efficiency agent for decision making (i.e., inference), in accordance with an embodiment of the present disclosure;

[0034] Figure 7 illustrates one example implementation of the emergency efficiency agent and the RAN node in an Open RAN (O-RAN) architecture, in accordance with an embodiment of the present disclosure; and

[0035] Figure 8 is a flow chart that illustrates the operation of a control entity (e.g., a control function implemented, e.g., at the RAN node) for controlling one or more energy saving featuresof compute hardware of the RAN node using a heuristics based control scheme, in accordance with one example embodiment of the present disclosure.Detailed Description

[0036] The embodiments set forth below represent information to enable those skilled in the art to practice the embodiments and illustrate the best mode of practicing the embodiments. Upon reading the following description in light of the accompanying drawing figures, those skilled in the art will understand the concepts of the disclosure and will recognize applications of these concepts not particularly addressed herein. It should be understood that these concepts and applications fall within the scope of the disclosure.

[0037] There exist certain challenges in regard to energy efficiency in a Radio Access Network (RAN) of a cellular communications system. Compute hardware for RAN applications must be high performing to meet demanding processing and timing requirements. Generally, two types of hardware are used, either tailor made processors for RAN or Commercial Off-The-Shelf (COTS) processors, but both have a multitude of processing elements including general purpose Central Processing Units (CPUs), Digital Signal Processors (DSPs), hardware accelerators, and Graphic Processing Units (GPUs).

[0038] In the context of RAN compute and energy efficiency, it is important to distinguish between the capacity of the hardware and the load on the network. Often, the hardware is designed to meet requirements for the maximum capacity (i.e., peak traffic load). However, the traffic load on the RAN is often less than the peak traffic load. For example, low traffic load levels can often occur during the night or weekends. Whenever the traffic load on the system is below the peak, there is an opportunity to reduce the hardware processing power and thus save energy.

[0039] Energy saving in RAN compute hardware is accomplished though energy saving features available from the hardware supplier. For example, a COTS server with an Intel multicore processor has several energy saving features including core sleep states (C-states), core frequency states (P-states), and uncore frequency (i.e., the clock frequency for non-core components such as shared cache and memory). The risk in using these energy saving features is that any node in the RAN can experience a burst of traffic so the user experience can still suffer if not enough processing resources are available in time for the demand. Therefore, mobile network operators (MNOs) are hesitant or unable to fully utilize energy efficiency features in RAN compute hardware due to the risk of reduced performance and consequently poor user experience.

[0040] At a technical level, there are three main difficulties that need to be managed by entity or procedure for governing the use of RAN hardware energy efficiency features. These three main difficulties are:1. The time for hardware to change from a low-energy state to an active state, and vice versa. Generally, the more the energy savings in a particular state, the longer it takes to become active again, and therefore increases the risk on RAN performance.2. The interaction between multiple energy saving features. The overall impact of various features on the RAN application can be very complex and difficult to optimize. For example, on an Intel® processor, trying to optimize the energy performance while simultaneously manipulating C-state, P-state, and uncore frequency, each of which affect the power consumption and processing capability of the hardware in different ways, is a complex task.3. The ability to generalize to different deployment situations and adapt to advances in RAN technology. A RAN can be deployed in highly dynamic and complex environments and undergoes rapid advances in compute and software features. Therefore, energy performance optimization is a constantly moving target.

[0041] Existing solutions for managing RAN hardware efficiency are commonly implemented at the hardware or operating system level and do not have enough information about the RAN application to effectively balance energy consumption with performance. It is particularly difficult given that RAN is a real-time application with hard timing deadlines. A common issue with these solutions is adding unnecessary latency to user data flows.

[0042] Systems and methods that provide solutions to the aforementioned and / or other challenges are disclosed herein. Embodiments of systems and methods are disclosed herein for controlling one or more energy savings features of a hardware component (i.e., a compute hardware component such as, e.g., a CPU(s), DSP(s), GPU(s), hardware accelerator(s), or the like) of a RAN node based on a combined state of the hardware component of the RAN node and one or more RAN application software components of the RAN node. In some embodiments, the one or more energy saving features of the hardware component of the RAN node are controlled based on the combined state using a Reinforcement Learning (RL) policy. The RL policy is trained based on an objective (e.g., a reward function). In some embodiments, the objective (e.g., the reward function) takes into account one or more metrics related to energy efficiency of the RAN node and one or more metrics related to user experience (i.e., user experience of users of User Equipments (UEs) served by the RAN node, which may be measured via metric(s) such as, e.g., latency (e.g., MAC and / or RLC latency for a site, cell, or UE(s)), throughput (e.g., MAC and / or RLC layerthroughput for a site, cell, or UE(s)), or time to content). In some embodiments, the objective (e.g., the reward function) is configurable, e.g., by the MNO. Note that RL is a sub-field of Machine Learning (ML), where optimization problems can be modelled as Markov Decision Processes (MDPs). The design of MDPs allows a decision-making agent to continuously learn from new experiences, to maximize a reward from taking particular actions.

[0043] In some other embodiments, the one or more energy savings features of the hardware component of the RAN node are controlled based on the combined state using heuristics (e.g., one or more rules that are, in some embodiments, configurable by, e.g., the MNO). In some embodiments, heuristics take into account one or more configurable inputs (e.g., one or more configurable rules) that are configured by, e.g., the MNO.

[0044] Embodiments of the present disclosure may include any one or more of the following aspects: a heuristic procedure or RL based procedure for controlling one or more energy savings features of a hardware component of a RAN node in a manner that takes into account both energy efficiency and user experience, the heuristic procedure or RL-based procedure having inputs from both the RAN application software domain and the hardware domain, a configurable objective supplied by the MNO.

[0045] Embodiments of the present disclosure may provide a number of advantages over existing solutions. For example, in some embodiments, control of the energy saving feature(s) of the hardware component of the RAN node can easily be tuned, e.g., by the MNO to be more or less aggressive with energy savings by adjusting the inputs or objective. This enables the MNO to control the risk level associated with the use of energy saving feature(s) and the resulting impact on user experience. As another example, by using inputs from both the RAN application software domain (e.g., incoming traffic load) and the compute hardware domain (e.g., CPU / core utilization), embodiments of the present disclosure can know ahead of time the amount of compute hardware resources that will be needed to process incoming traffic and therefore make more relevant decisions than existing solutions that do not consider the RAN application software domain.

[0046] Further, embodiments of the present disclosure that utilize a RL-based mechanism enable an optimal balance of energy consumption and performance. The RL-based mechanism learns how controlling the energy efficiency features in the hardware impacts the desired outcome of minimizing energy consumption without sacrificing user experience. This is achieved by training with large amounts of data collected from RAN operation, rather than “hand-crafted” rules. This approach therefore implicitly considers both the hardware activation time and energy saving feature interactions. Also, the RL-based mechanism learns from new experiences, allowing the RL-based mechanism to adapt to a variety of network conditions (e.g., traffic patterns, radiofrequency channel conditions) and changing hardware and software developments in the RAN with minimal or no manual intervention.

[0047] Figure 1 illustrates one example embodiment of a RAN node 100. As illustrated, the RAN node 100 includes a Central Unit (CU) 102, a Distributed Unit (DU) 104, and a radio 106. The CU 102 includes CU application software 108, an Operating System (OS) 110, and compute hardware 112. The CU application software 108 includes software that, when executed by the compute hardware 112, causes the CU 102 to perform the desired CU operations of the RAN node 100. For example, these CU operations include operations performed by one or more higher layers of the RAN node 100 such as, e.g., Radio Resource Control (RRC) layer operations, Packet Data Convergence Protocol (PDCP) layer operations, or the like. The compute hardware 112 includes any type(s) or combination of compute hardware components such as, e.g., one or more CPUs, one or more DSPs, one or more GPUs, one or more hardware accelerators, or the like.

[0048] The DU 104 includes DU application software 114, an OS 116, and compute hardware 118. The DU application software 114 includes software that, when executed by the compute hardware 118, causes the DU 104 to perform the desired DU operations of the RAN node 100. For example, these DU operations include operations performed by one or more lower layer of the RAN node 100 such as, e.g., Radio Link Control (RLC) layer operations, Medium Access Control (MAC) layer operations, or the like. The compute hardware 118 includes any type(s) or combination of compute hardware components such as, e.g., one or more CPUs, one or more DSPs, one or more GPUs, one or more hardware accelerators, or the like.

[0049] The radio 106 includes radio application software 120, an OS 122, and radio hardware 124. The radio application software 120 includes software that, when executed by the radio hardware 124, causes the radio 106 to perform the desired radio operations of the RAN node 100. For example, these radio operations include operations performed by one or more lower layers of the RAN node 100 such as, e.g., Physical (PHY) layer operations, or the like. The radio hardware 124 includes any type(s) or combination of compute hardware components such as, e.g., one or more CPUs, one or more DSPs, one or more GPUs, one or more hardware accelerators, or the like.

[0050] Importantly, the compute hardware 112 of the CU 102, the compute hardware 118 of the DU 104, and / or the radio hardware 124 of the radio 106 have energy saving features (or include hardware component(s) that have energy saving features) that are controlled in accordance with embodiments of the present disclosure. In the preferred embodiments disclosed herein, the control procedure disclosed herein (e.g., the RL-based control mechanism or the heuristics-based control mechanism) is used to control the energy savings feature(s) of only one compute hardware (i.e., the compute hardware 112 of the CU 102, the compute hardware 118 of the DU 104, or the radiohardware 124 of the radio 106). For example, the RL-based control mechanism or the heuristicsbased control mechanism disclosed herein may be deployed to control the energy efficiency features of the compute hardware 118 of the DU 104, but use state information of both the CU application software 108 and the DU application software 114 as inputs. However, even though controlling the energy saving feature(s) of only one compute hardware, the control mechanism disclosed herein considers both a hardware state of that hardware component and the application software state of any or all of the CU application software 108, the DU application software 114, and the radio application software 120. It should also be noted that separate control mechanisms (e.g., separate instances of the RL-based control mechanism or the heuristics-based control mechanism disclosed herein) may be used to separately control the energy savings features of any two or more of the compute hardware 112 of the CU 102, the compute hardware 118 of the DU 104, and the radio hardware 124 of the radio 106.

[0051] It should also be noted that the example of the RAN node 100 shown in Figure 1 is only an example. The RAN node 100 may more generally be any RAN node having one or more application software components and respective compute hardware.

[0052] In regard to controlling the energy saving features of compute hardware (e.g., the compute hardware 112 or the compute hardware 118 or the radio hardware 124), a first set of embodiments relate to a RL-based control mechanism, and a second set of embodiments relate to a heuristics-based control mechanism, each of which is described in detail below.RL-Based Control of Energy savings features

[0053] In a first set of embodiments, the energy efficiency versus performance (e.g., user experience) of the RAN node 100 is modelled as a MDP. The general structure of an MDP is shown in Figure 2. The MDP model of Figure 2 includes an agent 200 that chooses an action (Ar) based on a state (St) of an environment 202 at a particular time t using a policy 204. The environment 202 then transitions to state St+iand provides a reward Rt+ithat describes a value of the action towards a particular objective. The objective is described by a user-defined reward function. A RL function 206 operates to modify or train the policy 204, using state, action and reward data from interacting with the environment 202, to maximize the cumulative rewards of actions.

[0054] Figure 3 illustrates a MDP model for applying RL to optimize control of one or more energy saving features available on compute hardware (e.g., the compute hardware 112 or the compute hardware 118 or the radio hardware 124) of the RAN node 100 in accordance with an embodiment of the present disclosure. As illustrated, an energy efficiency agent 300 repeatedlychooses actions to control the energy saving features of the compute hardware of the RAN node 100 and receives: (a) updated states from the RAN application software (e.g., the CU application software 108, the DU application software 114, and the radio application software 120), (b) updated states from the compute hardware of the RAN node 100, and optionally (c) updates states from an environment of the RAN node 100. In Figure 3, the RAN application software, the RAN compute hardware, and the environment are represented by a block denoted as “RAN SW, HW, and optionally environment 302). In addition, particularly for training or updating of a policy 304 used by the energy efficiency agent 300 to select the actions based on the combined input state (i.e., the combination of the state(s) of the RAN application software, the state of the RAN compute hardware, and optionally the state of the environment). The incoming states indicates what is happening in the system (RAN application software and RAN compute hardware and optionally environment), which gives the energy efficiency agent 300 the basis to choose its next action based on the policy 304. A reward computed for a particular action indicates a quality of the action by considering metrics indicative of the resulting power consumption and user experience outcomes. Using RL function 306, the energy efficiency agent 300 updates the policy 304 to optimize user experience and energy efficiency across various combined states of the RAN application software, the RAN compute hardware, and optionally the environment.

[0055] It should be noted that the RL-based control mechanism, or procedure, disclosed herein can be applied to any hardware running any component of RAN application software. However, it may be most beneficial in the lower layers of user plane processing where the processing demands are highest and with the most stringent real-time requirements. It should also be noted that the RL-based control mechanism can be trained for different levels of software constructs, such as for an entire application, for individual processes, or for individual threads. For example, a thread that is dedicated to physical layer processing may have dedicated cores, which are isolated from a core that is dedicated to running a scheduling thread, and the RL-based control mechanism can learn to generalize for both threads, or optimize each independently. It should also be noted that the RAN software, RAN hardware, and optionally environment 302 can be completely simulated, can be a controlled lab RAN node deployment, or a live RAN node deployment.

[0056] The state of the compute hardware of the RAN node 100 may include one or more metrics that provide any measure of the state of the compute hardware such as, e.g., core utilization, core frequency, power consumption, residence time in a particular state, or the like. The state of the compute hardware may further include one or more metrics indicative of related measurements such as, e.g., a temperature reading of a temperature of the compute hardware, informationindicative of a power consumption of one or more fans used to cool the compute hardware, or the like.

[0057] The state of the RAN application software may include one or more metrics that provide any measurement of the RAN application software domain such as, e.g., (a) traffic load in a PDCP layer of the RAN node, (b) MAC throughput, (c) MAC layer and / or Radio Link Control (RLC) layer latency, (d) PRB utilization, (e) RLC buffer size, (f) one or more inputs about a state of one or more UEs scheduled for a transmit time interval, (g) an average or combined state of multiple UEs scheduled for a transmit time interval, (h) a predicted value(s) for any one or more of (a)-(g), or the like. The state of the RAN application software may additionally or alternatively include one or more metrics indicative of any configuration settings of the RAN node 100 or a cell(s) operated by the RAN node 100 such as, e.g., cell configuration, carrier frequency, carrier aggregation, MIMO configuration, a number of cells operated by the RAN node 100, a cell bandwidth of each of one or more cells operated by the RAN node 100, or the like. The state of the environment may include one or more metrics that provide environmental or situational knowledge such as, e.g., a location of the RAN node 100, a location of the DU 104 of the RAN node 100, a traffic type being handled by the RAN node 100, one or more weather conditions (e.g., temperature, humidity, precipitation, or the like) for the geographic location at which the RAN node 100 is locate, or the like.

[0058] The actions output by the energy efficiency agent 300 may each include any one or more of the following: one or more outputs for controlling one or more energy saving features of the compute hardware. These energy saving features can include any feature of the compute hardware that directly or indirectly impacts the energy consumption or performance of the compute hardware. Examples include core sleep states (e.g., C-state on Intel® processors), core clock frequency, uncore frequency, or the like. Some further examples includes less direct actions such as, e.g., setting limits for maximum and minimum settings are also possible. For example, setting an upper limit on core frequency, but allowing a subsystem (in OS or HW) to freely change the frequency below the threshold.

[0059] The reward function is used by the RL function 306 to compute rewards for actions taken by the energy efficiency agent 300. The policy 304 is updated based on the computed rewards. In some embodiments, the reward function is a function of one or more metrics related to energy efficiency or energy consumption of the RAN compute hardware and one or more metrics related to user experience (e.g., user experience of users UEs served by the RAN node 100). In other words, the reward function will impact the policy 304 and thus dictate how the energy efficiency agent 300 will choose actions to optimize the desired objective. Any reward function ispossible if it has an energy saving component and a user experience or RAN performance component. In some embodiments, the objective of the energy efficiency agent 300, which is expressed via the reward function, is configurable, e.g., by an MNO. As such, different MNOs may configure the objective or reward function in such a way as to meet the needs or desired tradeoff of that MNO. For example, in one embodiment, the reward function includes weighting of energy saving vs. user experience which is flexible and can be customized by the MNO. Some specific, non-limiting examples of the reward function are as follows:• weighted sum or polynomial of user experience and power consumption o Example: wi x latency + W2 x power, where wi and W2 are positive integer values• discontinuous function of user experience o Example:■ negative reward only when latency is above a threshold, zero otherwise.■ negative reward for power consumption, still continuous• state adjusted reward o Example: The reward is scaled based on the traffic such that during high traffic (high power and latency) the reward is of similar magnitude to when there is low traffic (low power and latency).

[0060] The energy efficiency agent 300, in general, complies to reinforcement learning and MDP formalized problems. Example implementations of an agent which may be used for the energy efficiency agent 300 include an agent that utilizes a policy gradient method (e.g., proximal policy optimization), an agent that utilizes a value estimation method (e.g., deep Q-network), or an agent that utilizes a model-based method (e.g., imagination-augmented agent).

[0061] Both during training of the policy 304 and during inference using the trained policy 304, the energy efficiency agent 300 decides a new action at a “decision-making frequency.” The decision-making frequency offers a tradeoff between accuracy and computation cost. Setting the energy efficiency agent 300 to take a new action at a high frequency (e.g., every slot) gives the highest accuracy and reactivity to changing conditions but is computationally expensive and requires highly reactive / low-latency input measurements. Conversely, acting over longer periods (such as once per second) is less accurate but reduces the computation cost and is generally easier to correlate actions input measurements.

[0062] Figure 4 is a flow chart that illustrates the operation of the energy efficiency agent 300 for online training of the policy 304 in accordance with an embodiment of the present disclosure. In this context, the energy efficiency agent 300, or at least the part(s) of the energy efficiency agent 300 responsible for performing the steps of Figure 4, is referred to herein as a “training entity”.Optional steps are represented by dashed lines / boxes. As illustrated, the energy efficiency agent 300 obtains, or collects, a first set of inputs related to (e.g., indicative of) the state of the RAN application software (step 400), a second set of inputs related to (e.g., indicative of) the state of the RAN compute hardware (402), and optionally a third set of inputs related to (e.g., in indicative of) the environment (step 404). The combination of the state of the RAN application software, the state of the RAN compute hardware, and optionally the state of the environment is referred to here as a “combined state.” When the energy efficiency agent 300 is collecting these inputs as training samples to be used by the RL function 306 for training the policy 304, a reward is calculated based on the obtained input sets (i.e., based on the combined state) (step 406). In some embodiments, a new training sample is created with the previous combined state (i.e., the combined state from the previous iteration of the training procedure for time t-1), the previous action taken at time t-1, the current combined state (i.e., the combined state for time t), and the computed reward at time t. The energy efficiency agent 300 updates the policy 304 based on the computed reward (e.g., based on the recorded training sample). Note that, in an alternative embodiment, multiple training samples may be recorded before the policy 304 is updated (e.g., the update step 408 may be performed every N iterations of the process, where N is equal to or greater than 1). Note that the update procedure itself is not described here as numerous variations of the update procedure for RL are well-known in the art. It should also be noted that the energy efficiency agent 300 may be implemented at a single node (e.g., the RAN node 100 or a computing node separate from but communicatively coupled to the RAN node 100) or implemented in a distributed manner across two or more nodes (e.g., steps 400-406 may be implemented at a first node (e.g., the RAN node 100 or a first computing node that is separate from but communicatively coupled to the RAN node 100) and step 408 may be implemented at a second computing node that receives the training sample(s) from the RAN node 100 or first computing node, updates the policy 304 based on the training sample(s), and provides an updated policy 304 to the RAN node 100 possibly via the first computing node).

[0063] The energy efficiency agent 300 applies the policy 304 to the combined state (step 410). By applying the policy 304 to the combined state, the energy efficiency agent 300 obtains a probability distribution for multiple candidate actions (i.e., for multiple different sets of outputs for controlling the hardware energy savings features) achieving the desired objective. The energy efficiency agent 300 then selects one of the candidate actions as an action to be taken (step 412). In one embodiment, if the energy efficiency agent 300 is exploring different actions (step 412, YES), the energy efficiency agent 300 selects the action to be taken from the list of candidate actions using epsilon greedy action selection (step 412B), otherwise (step 412A, NO), the energyefficiency agent 300 selects the action to be taken from the list of candidate actions using greedy action selection (i.e., best known action or action having the highest probability of achieving the desired outcome as indicated by the probability distribution) (step 412C). Finally, the energy efficiency agent 300 controls the energy saving feature(s) of the RAN compute hardware (i.e., sets the energy saving state of the RAN compute hardware) according to the selected action (step 414). The energy efficiency agent 300 then waits for the next decision time (step 416) and then returns to step 400 start another iteration of the training procedure.

[0064] Figure 5 is a flow chart that illustrates the operation of the energy efficiency agent 300 for offline training of the policy 304 in accordance with an embodiment of the present disclosure. In this context, the energy efficiency agent 300, or at least the part(s) of the energy efficiency agent 300 responsible for performing the steps of Figure 5, is referred to herein as a “training entity”. Optional steps are represented by dashed lines / boxes. Note that, in this example, the RL function 306, or policy update entity, is a separate entity implemented on a node that is separate from that on which the energy efficiency agent 300 is implemented.

[0065] As illustrated, the energy efficiency agent 300 obtains, or collects, a first set of inputs related to (e.g., indicative of) the state of the RAN application software (step 500), a second set of inputs related to (e.g., indicative of) the state of the RAN compute hardware (502), and optionally a third set of inputs related to (e.g., in indicative of) the environment (step 504). The combination of the state of the RAN application software, the state of the RAN compute hardware, and optionally the state of the environment is referred to here as a “combined state.” When the energy efficiency agent 300 is collecting these inputs as training samples to be used by the RL function 306 for training the policy 304, a reward is calculated based on the obtained input sets (i.e., based on the combined state) (step 506). In some embodiments, a new training sample is created with the previous combined state (i.e., the combined state from the previous iteration of the training procedure for time t-1), the previous action taken at time t-1, the current combined state (i.e., the combined state for time t), and the computed reward at time t. The energy efficiency agent 300 stores this new training sample (step 508).

[0066] The energy efficiency agent 300 sends the training sample, or alternatively sends (e.g., periodically) a batch of training samples to the RL function 306 (i.e., the policy update entity) (step 510). For example, in one embodiment, the energy efficiency agent 300 determines whether to send the training sample or a batch of training samples including the training sample and a number of previously collected training samples, to the RL function 306 (step 510A). If so (step 510A, YES), the energy efficiency agent 300 sends the training sample, or the batch of training samples, to the RL function 306 (step 510B).

[0067] The energy efficiency agent 300 performs an exploratory action whereby it selects a new action (i.e., a new set of outputs) for controlling the HW energy savings features (step 512). The energy efficiency agent 300 controls the energy saving feature(s) of the RAN compute hardware (i.e., sets the energy saving state of the RAN compute hardware) according to the selected action (step 514). The energy efficiency agent 300 then waits for the next decision time (step 514) and then returns to step 500 start another iteration of the training procedure.

[0068] Figure 6 a flow chart that illustrates the operation of the energy efficiency agent 300 for decision making (i.e., inference) in accordance with an embodiment of the present disclosure. In this context, the energy efficiency agent 300, or at least the part of the energy efficiency agent 300 utilized during inference, is also referred to herein as an evaluation entity or an evaluation function. Optional steps are represented by dashed lines / boxes. This process is beneficial where the energy efficiency agent 300 is used for decision making (inference) and not training. This may be advantageous if the energy efficiency agent 300 is required to run in the fewest possible cycles, or if training data is unnecessary. In this case, the energy efficiency agent 300 chooses the action it believes has the highest value given the current combined state, based on previous training.

[0069] As illustrated in Figure 6, the energy efficiency agent 300 obtains, or collects, a first set of inputs related to (e.g., indicative of) the state of the RAN application software (step 600), a second set of inputs related to (e.g., indicative of) the state of the RAN compute hardware (step 602), and optionally a third set of inputs related to (e.g., in indicative of) the environment (step 604). The combination of the state of the RAN application software, the state of the RAN compute hardware, and optionally the state of the environment is referred to here as a “combined state.” When the energy efficiency agent 300 is collecting these inputs for inference only, the energy efficiency agent 300 applies the policy 304 (which may be trained via, e.g., the procedure of Figure 4 or Figure 5) to the combined state (step 606). By applying the policy 304 to the combined state, the energy efficiency agent 300 obtains a probability distribution for multiple candidate actions (i.e., for multiple different sets of outputs for controlling the hardware energy savings features) achieving the desired objective. The energy efficiency agent 300 then selects one of the candidate actions as an action to be taken (step 608). Here, the selection is a greedy action selection (i.e., best known action or action having the highest probability of achieving the desired outcome as indicated in the probability distribution). Finally, the energy efficiency agent 300 controls the energy saving feature(s) of the RAN compute hardware (i.e., sets the energy saving state of the RAN compute hardware) according to the selected action (step 610). The energy efficiency agent 300 then waits for the next decision time (step 612) and then returns to step 600 start another iteration of the inference procedure.

[0070] The energy efficiency agent 300, and in particular the policy 304 and the RL function 306 that trains the policy 304, may be implemented at a single node such as, e.g., the RAN node 100 or a computing node that is separate from the RAN node 100. Alternatively, the policy 304 (and part of the energy efficiency agent 300 that makes decisions based on the policy 304) and the RL function 306 may be implemented on separate nodes. For example, the policy 304 (and part of the energy efficiency agent 300 that makes decisions based on the policy 304) may be implemented on a first node (e.g., the RAN node 100 or a first computing node that is separate from and communicatively coupled to the RAN node 100) and the RL function 306 may be implemented on a second node (e.g., a second computing node that is separate from the RAN node 100).

[0071] In other words, the RAN node 100 is highly computationally constrained and additional processing for an RL algorithm may not fit into the processing budget. Therefore, it is easier to consider the policy 304 and the RL function 306 that trains the policy 304 separately, as the policy 304 requires much less processing, as it only maps states to actions. In regard to the policy 304, the choice of where the policy 304 is implemented depends on the decision-making frequency. A high decision-making frequency (e.g., select new action every 10 milliseconds (ms) or less) requires the policy 304 to run inside the RAN application software that is using the compute hardware being controlled (e.g., in the DU application software 114 if controlling the compute hardware 118 of the DU 104). Conversely, a decision-making frequency having a period of more than 10 ms would allow the policy 304 to be implemented off the compute hardware being controlled, e.g., at a low-latency controller such as the near-real time RAN Intelligent Controller defined by Open RAN (O-RAN), although it does not exclude the implementation of the policy 304 in the RAN application software.

[0072] In regard to the RL function 306 that trains the policy 304, the RL function 306 is computationally expensive, requiring a database for training samples and processing to run an optimization algorithm (e.g., for a neural network RL algorithm, gradient descent with back propagation). Where the RL function 306 is implemented depends on how frequently the policy 304 is to be updated. For frequent policy updates, allowing the policy 304 to adapt quickly, the RL function 306 can be implemented close to the RAN application software (e.g., in an adjacent RAN network function running simultaneously in the cloud, or the non-real time RAN Intelligent Controller defined by O-RAN). For less frequent policy updates, the RL function 306 can run completely “offline”, meaning outside of the RAN context, using data collected and exported from the RAN application software.

[0073] Figure 7 illustrates one example implementation of the energy efficiency agent 300 and the RAN node 100 in the O-RAN architecture. In this example, the RAN node 100 includes a Non-RT-RIC 700, a Near-RT-RIC 702, an O-DU 704, and RAN hardware 706. In this example, the policy 304 and the RL function 306 are split between the Non-RT-RIC 700 and the Near-RT- RIC 702, based on their requirements. An Agent Proxy 708 in the O-DU 704 only facilitates setting of energy efficiency features and measuring state in the RAN Hardware 706. Note that the RAN hardware 706 is the RAN compute hardware being controlled in this example implementation. Also note that a training sample database 710 that stores the training samples collected by the energy efficiency agent 300 (more specifically collected by the agent proxy 708 in this example) that are used by the RL function 306 to train the policy 304.

[0074] The above description of the RL-based control of the energy saving features of the RAN compute hardware covers a range of possible implementations and variations. One example implantation is as follows:• The RAN node includes a DU that includes cloud RAN DU software and COTS hardware consisting of a multi-core processor.• The compute hardware being controlled by the energy efficiency agent 300 are the cores of the multi-core processor used to run a baseband application software.• RAN application software state includes the PDCP traffic incoming for baseband processing.• RAN hardware state includes the percent utilization of each core, the percent of time each core spends in each C-state, and the frequency of each core.• Actions include controlling C-states of cores, and the maximum core frequency.• The reward function uses a weighted sum of the throughput, CPU power consumption, and latency.• The energy efficiency agent 300 is implemented with the DQN algorithm (i.e., the RL function 306 uses a Deep-Q Network algorithm), although other types of RL algorithms may be used.• The decision-making frequency is once per second.• The policy 304 is implemented with a deep neural network (e.g., as is described by DQN) and runs inside the baseband pod / process.• Training of the policy 304 is done offline using data collected from a lab environment with an exploring agent.Heuristics-based Control of Energy saving Features

[0075] In other embodiments of the present disclosure, one or more energy saving features of the RAN compute hardware are controlled via heuristics (e.g., based on one or more rules which may be predefined or configurable). The heuristic-based control mechanism is a rule-based procedure when one or more predefined or configurable rules are applied to the combined state (i.e., the combination of the state of the compute hardware being controlled, the state(s) of the RAN application software(s), and optionally the state of the environment) to determine an output(s) for controlling the energy saving feature(s) of the compute hardware. In some embodiments, the one or more rules and / or the inputs to be used for the one or more rules area are configurable, e.g., by the MNO to achieve a desired objective(s) (e.g., network cost, RAN performance, user experience, or the like).

[0076] For example, the MNO may provide a weekly guideline (e.g., with 30-60 minute granularity) for the desired setting of the energy saving feature(s) of the compute hardware of the RAN node, where this weekly guideline is applied directly as instructed by the MNO. The weekly guideline may define the desired setting of the energy saving feature(s) of the compute hardware of the RAN node for each time period (e.g., each 30-60 minute time period). In this case, the MNO may use historical data and or expected data (e.g., historical or expected traffic patterns) to determine the optimal setting of the energy saving feature(s) of the compute hardware of the RAN node for each time period.

[0077] As another example, the MNO may provide a weekly traffic pattern (e.g., with 30-60 minute granularity) as a configurable input. This traffic pattern may provide the amount of traffic or expected amount of traffic (e.g., based on historical traffic data) in one or more cells operated by the RAN node 100 during each of multiple time periods (each 30 or 60 minute time period) during the week. The control mechanism may then select a setting for the energy saving feature(s) of the compute hardware of the RAN node for each time period based on an expected traffic load for the RAN node for that time period as defined by the weekly traffic pattern provided by the MNO. Alternatively, the MNO may also provide the desired or recommended setting for the energy saving feature(s) of the compute hardware of the RAN for each time period defined in the weekly traffic pattern or for each of a number of traffic levels, where the control mechanism would determine the expected traffic load for a particular time period using the provided weekly traffic pattern and then determine the desired setting of the energy saving feature(s) for that traffic load (also provided by the MNO).

[0078] As another example, the control mechanism controls the energy saving feature(s) of the compute hardware of the RAN node within a limit (e.g., every 1 second, but not going beyond 20%change compared to an MNO recommended setting). For this implementation, the control mechanism may use information from the RAN application software domain (e.g., incoming traffic load, latency) and / or hardware state (e.g., CPU / core utilization) to make more informed decisions.

[0079] Note that, in some embodiments, the control mechanism may raise an alarm to recommend when the MNO configuration should change. This alarm may be raised with a predefined or configured alarm threshold(s) is reached, where the alarm threshold(s) may be for any network performance related metric or a user experience related metric (e.g., when latency exceeds a predefined or configured alarm threshold, an alarm is raised).

[0080] Still further, the control mechanism may use rules (e.g., configured by the MNO) for controlling the energy saving feature(s) of the compute hardware of the RAN node in a manner similar to the RL-based control mechanism described above. For example, the rules may provide a calibrated mapping of traffic level to a particular setting of the energy saving feature(s) (e.g., a rule that states that, for 1 Gigabits per second downlink on three cells, set eight cores active at 70% max frequency). As another example, the rules may enable monitoring of the capacity of the compute hardware and traffic demand with a feedback loop (e.g., current active cores are 90% utilized and traffic is expected to decrease, therefore start to enable more energy saving but never bring system above 95% utilization). As another example, the rules may enable monitoring of user experience metric(s) and feedback to energy efficiency feature control (e.g., when the RLC latency starts to exceed 1000 ms, increase the processing capacity (i.e., decrease the energy saving)).

[0081] Figure 8 is a flow chart that illustrates the operation of a control entity (e.g., a control function implemented, e.g., at the RAN node 100) for controlling one or more energy saving features of compute hardware of the RAN node, in accordance with one example embodiment of the present disclosure. Optional steps are represented by dashed lines / boxes. As illustrated, the control entity optionally obtains, or collects, a first set of inputs related to (e.g., indicative of) the state of the RAN application software (step 800), a second set of inputs related to (e.g., indicative of) the state of the RAN compute hardware (802), and a third set of inputs related to (e.g., in indicative of) the environment (step 804). The combination of the state of the RAN application software, the state of the RAN compute hardware, and optionally the state of the environment is referred to here as a “combined state.” The control entity applies heuristic-based rules (e.g., to the combined state) to determine a set of outputs for controlling one or more energy saving features of the compute hardware of the RAN node (step 806). In some embodiments, the heuristic-based rules are configurable, e.g., by the MNO. For example, as discussed above, the MNO may provide a weekly guideline that defines a desired setting of the energy saving feature(s) of the compute hardware, where the heuristic-based rules then are time based rules, or some equivalent mechanismsuch as a look-up table or data structure, that enables the control entity to determine the set of outputs (i.e., the desired setting) for the one or more energy saving features based on the current time period. As another example, the MNO may provide a weekly traffic pattern, where the weekly traffic pattern (e.g., in lieu of or in addition to the state of the RAN application software obtained in step 800) is used along with one or more predefined or configured rules to determine the set of outputs (i.e., the setting of the energy saving feature(s) of the compute hardware) based on the expected traffic load of the RAN node as indicated by the configured traffic pattern. As another example, the MNO may configure one or more rules that define which setting of the energy saving feature(s) of the compute hardware are selected as a function of one or more inputs related to the RAN application software state, one or more inputs related to the compute hardware state, and / or one or more inputs related to the environment.

[0082] Finally, the control entity controls, or sets, the energy saving feature(s) of the RAN compute hardware (i.e., sets the energy saving state of the RAN compute hardware) according to the determined set of outputs from step 806 (step 808). The energy efficiency agent 300 then waits for the next decision time (step 810) and then returns to step 800 start another iteration of the inference procedure.

[0083] Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as Read Only Memory (ROM), Random Access Memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according to one or more embodiments of the present disclosure.

[0084] While processes in the figures may show a particular order of operations performed by certain embodiments of the present disclosure, it should be understood that such order is exemplary (e.g., alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, etc.).

[0085] Those skilled in the art will recognize improvements and modifications to the embodiments of the present disclosure. All such improvements and modifications are considered within the scope of the concepts disclosed herein.

Claims

Claims1. A method performed by an evaluation entity for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, the method comprising: obtaining (600; 800) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtaining (602; 802) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; and determining (606-608; 806) a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs.

2. The method of claim 1, further comprising providing (610; 808) the set of outputs to the hardware component of the RAN node for control of the one or more energy savings features of the hardware component of the RAN node.

3. The method of claim 1 or 2, wherein the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following:(a) Packet Data Convergence Protocol, PDCP, traffic load;(b) Medium Access Control, MAC, throughput;(c) MAC layer and / or Radio Link Control, RLC, layer latency;(d) Physical Resource Block, PRB, utilization;(e) RLC buffer size;(f) one or more inputs about a state of one or more UEs scheduled for a transmit time interval;(g) average or combined state of multiple UEs scheduled for a transmit time interval;(h) a predicted value for any one or more of (a)-(g).

4. The method of any of claims 1 to 3, wherein the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following:one or more configurations of the RAN node; one or more configurations of a cell operated by the RAN node; a carrier frequency of a cell operated by the RAN node; an indication of whether carrier aggregation is being utilized; a Multiple Input Multiple Output, MIMO, configuration of the RAN node; a number of cells operated by the RAN node; a cell bandwidth of each of one or more cells operated by the RAN node.

5. The method of any of claims 1 to 4, wherein the second set of inputs related to the hardware state of the hardware component of the RAN node comprise any one or more of the following: core utilization of one or more multi-core processors comprised on the hardware component; core frequency of one or more cores of the one or more multi-core processors comprised in the hardware component; power consumption of the hardware component; power consumption of the one or more cores of the one or more multi-core processors comprised in the hardware component; a residence time in a particular state of the hardware component; power consumption of one or more fans used for cooling the one or more multi-core processors comprised in the hardware component.

6. The method of any of claims 1 to 5, wherein the set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node comprise any one or more of the following: an output for controlling a core sleep state of one or more cores of one or more multi-core processors comprised in the hardware component; an output for controlling a clock frequency of one or more cores of the one or more multicore processors comprised in the hardware component; an output for controlling an uncore clock frequency of the one or more multi-core processors comprised in the hardware component.

7. The method of any of claims 1 to 6, wherein determining (806) the set of outputs for control of the one or more energy saving features of the hardware component comprises:applying (806) the first set of inputs and the second set of inputs to a heuristic model that outputs the set of outputs for control of the one or more energy savings features of the hardware component based on the first and second sets of inputs.

8. The method of claim 7, wherein the heuristic model comprises one or more rules that are configurable by a network operation of the wireless communication system.

9. The method of claim 8, wherein the one or more rules comprise one or more rules that define a direct relationship between an amount of traffic or expected amount of traffic in one or more cells operated by the RAN node and which of the one or more energy savings features are to be enabled.

10. The method of claim 8 or 9, wherein the one or more rules comprise any one or more of the following: one or more rules that define a mapping between different traffic levels in a cell operated by the RAN node and different sets of outputs that provide different setting of the one or more energy savings features; one or more rules that selects the set of outputs based on a capacity of the hardware component and expected traffic demand in the cell operated by the RAN node; one or more rules that select the set of outputs based on a user experience metric.

11. The method of any of claims 1 to 6, wherein determining (606-608) the set of outputs for control of the one or more energy saving features of the hardware component comprises: applying (606) a reinforcement learning, RL, policy to a combined state formed by the first set of inputs and the second set of inputs to thereby obtain a probability distribution for a plurality of different sets of possible outputs for the combined state achieving a desired result; selecting (608), based on the probability distribution, one of the plurality of different sets of possible outputs having a highest probability of achieving the desired result as the set of outputs for control of the one or more energy savings features of the hardware component.

12. The method of claim 11, wherein the desired result is an optimal tradeoff between user experience of users of a plurality of user equipments, UE, served by the RAN node and power consumption of the RAN node.

13. The method of claim 8 or 11, wherein the RL policy is trained to achieve an objective that is configurable by a network operator of the wireless communication system.

14. The method of claim 8 or 11, wherein the RL policy is trained based on a reward function that is configurable by a network operator of the wireless communication system.

15. The method of claim 14, wherein the reward function is a function of one or more metrics indicative of user experience of users associated to user equipments, UEs, served by the RAN node and one or more metrics indicative of power consumption of the hardware component.

16. The method of claim 14, wherein the reward function is a weighted sum of one or more metrics indicative of user experience of users associated to user equipments, UEs, served by the RAN node and one or more metrics indicative of power consumption of the hardware component.

17. The method of claim 14, wherein the reward function is a discontinuous function of one or more metrics indicative of user experience of users associated to user equipments, UEs, served by the RAN node.

18. The method of claim 14, wherein the reward function is a state-adjusted reward function.

19. The method of any of claims 1 to 18, further comprising: obtaining (604; 804) a third set of inputs related to an environment of the hardware component of the RAN node and / or an environment of the RAN node; wherein determining (606-608; 806) the set of outputs for control of the one or more energy saving features of the hardware component is further based on the third set of inputs.

20. The method of claim 19, wherein the third set of inputs comprises any one or more of the following: a location of the hardware component; a location of the RAN node; a current time of day; current weather information for geographic area in which the hardware component and / or RAN node is located.

21. The method of any of claims 1 to 20, wherein the steps of obtaining the first set if inputs, obtaining the second set of inputs, and determining the set of outputs is repeated at a predefined or configured frequency.

22. The method of any of claims 1 to 21, wherein the evaluation entity is implemented at the RAN node.

23. The method of any of claims 1 to 21, wherein the evaluation entity is implemented at computing node that is separate from the RAN node.

24. A node for implementing an evaluation entity for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, the node adapted to: obtain (600; 800) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtain (602; 802) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; and determine (606-608; 806) a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs.

25. The node of claim 24, wherein the node is the RAN node.

26. The node of claim 24, wherein the node is a computing node that is separate from the RAN node.

27. The node of any of claims 24 to 26, wherein the node is further adapted to perform the method of any of claims 2 to 20.

28. A node for implementing an evaluation entity for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components andone or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, the node comprising processing circuitry configured to cause the node to: obtain (600; 800) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtain (602; 802) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; and determine (606-608; 806) a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs.

29. The node of claim 28, wherein the node is the RAN node.

30. The node of claim 28, wherein the node is a computing node that is separate from the RAN node.

31. The node of any of claims 28 to 30, wherein the processing circuitry is further configured to cause the node to perform the method of any of claims 2 to 20.

32. A computer program comprising instructions which, when executed on at least one processor, cause the processor to carry out the method according to any of claims 1 to 23.

33. A carrier containing the computer program of claim 32, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, or a computer readable storage medium.

34. A non-transitory computer-readable medium comprising instructions executable by processing circuitry of a node for implementing an evaluation entity for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, whereby the node is operable to: obtain (600; 800) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node;obtain (602; 802) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; and determine (606-608; 806) a set of outputs for control of the one or more energy saving features of the hardware component, based on the first set of inputs and the second set of inputs.

35. A method performed by a training entity for training a Reinforcement Learning, RL, policy for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, the method comprising: obtaining (400; 500) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtaining (402; 502) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; calculating (406; 506), based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node; selecting (412; 512) a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node; and setting (414; 514) the one or more energy savings features of the hardware component based on the selected set of outputs.

36. The method of claim 35, further comprising: updating (408) the RL policy at a previous combined state and the previous selected set of outputs based on the calculated reward and the current combined state; and applying (410) the RL policy to the combined state to thereby obtain, for each of a plurality of different sets of possible outputs for the combined state, a probability of that set of possible outputs providing a desired result;wherein selecting (412) the new set of outputs comprises selecting (412) one of the plurality of different sets of possible outputs as the new set of outputs for controlling the one or more energy savings features of the hardware component.

37. The method of claim 35, further comprising sending (510) the reward, the current combined state, and the previous selected set of outputs to a policy update entity.

38. The method of any of claims 35 to 37, further comprising performing a plurality of iterations of the steps to thereby train the RL policy.

39. The method of claim 35 or 38, wherein the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following:(a) Packet Data Convergence Protocol, PDCP, traffic load;(b) Medium Access Control, MAC, throughput;(c) MAC layer and / or Radio Link Control, RLC, layer latency;(d) Physical Resource Block, PRB, utilization;(e) RLC buffer size;(f) one or more inputs about a state of one or more UEs scheduled for a transmit time interval;(g) average or combined state of multiple UEs scheduled for a transmit time interval;(h) a predicted value for any one or more of (a)-(g).

40. The method of any of claims 35 to 39, wherein the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node comprise any one or more of the following: one or more configurations of the RAN node; one or more configurations of a cell operated by the RAN node; a carrier frequency of a cell operated by the RAN node; an indication of whether carrier aggregation is being utilized; a Multiple Input Multiple Output, MIMO, configuration of the RAN node; a number of cells operated by the RAN node; a cell bandwidth of each of one or more cells operated by the RAN node.

41. The method of any of claims 35 to 40, wherein the second set of inputs related to the hardware state of the hardware component of the RAN node comprise any one or more of the following: core utilization of one or more multi-core processors comprised on the hardware component; core frequency of one or more cores of the one or more multi-core processors comprised in the hardware component; power consumption of the hardware component; power consumption of the one or more cores of the one or more multi-core processors comprised in the hardware component; a residence time in a particular state of the hardware component; power consumption of one or more fans used for cooling the one or more multi-core processors comprised in the hardware component.

42. The method of any of claims 35 to 41, wherein the previous selected set of outputs and the new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node comprise each comprise any one or more of the following: an output for controlling a core sleep state of one or more cores of one or more multi-core processors comprised in the hardware component; an output for controlling a clock frequency of one or more cores of the one or more multicore processors comprised in the hardware component; an output for controlling an uncore clock frequency of the one or more multi-core processors comprised in the hardware component.

43. The method of any of claims 35 or 42, wherein the reward function is configurable by a network operator of the wireless communication system.

44. The method of any of claims 35 or 43, wherein the reward function is a function of one or more metrics indicative of user experience of users associated to user equipments, UEs, served by the RAN node and one or more metrics indicative of power consumption of the hardware component.

45. The method of any of claims 35 or 43, wherein the reward function is a weighted sum of one or more metrics indicative of user experience of users associated to user equipments, UEs,served by the RAN node and one or more metrics indicative of power consumption of the hardware component.

46. The method of any of claims 35 or 43, wherein the reward function is a discontinuous function of one or more metrics indicative of user experience of users associated to user equipments, UEs, served by the RAN node.

47. The method of any of claims 35 or 43, wherein the reward function is a state-adjusted reward function.

48. The method of claim 35, further comprising: obtaining (404) a third set of inputs related to an environment of the hardware component of the RAN node and / or an environment of the RAN node; wherein the current combined state further comprises the third set of inputs.

49. The method of claim 48, wherein the third set of inputs comprises any one or more of the following: a location of the hardware component; a location of the RAN node; a current time of day; current weather information for geographic area in which the hardware component and / or RAN node is located.

50. The method of any of claims 35 to 49, wherein the training entity is implemented at the RAN node.

51. The method of any of claims 35 to 49, wherein the training entity is implemented at computing node that is separate from the RAN node.

52. A node for implementing a training entity for training a Reinforcement Learning, RL, policy for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software componentsoperating on the one or more hardware components, respectively, of the RAN node, the node adapted to: obtain (400; 500) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtain (402; 502) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; calculate (406; 506), based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node; select (412; 512) a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node; and set (414; 514) the one or more energy savings features of the hardware component based on the selected set of outputs.

53. The node of claim 52, wherein the node for implementing the training entity is the RAN node.

54. The node of claim 52, wherein the node for implementing the training entity is a computing node that is separate from the RAN node.

55. The node of any of claims 52 to 54, further adapted to perform the method of any of claims 36 to 49.

56. A node for implementing a training entity for training a Reinforcement Learning, RL, policy for controlling one or more energy savings features of a hardware component of a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, the node comprising processing circuitry configured to cause the node to:obtain (400; 500) a first set of inputs related to a RAN application state of the one or more RAN application software components of the RAN node; obtain (402; 502) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; calculate (406; 506), based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node; select (412; 512) a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node; and set (414; 514) the one or more energy savings features of the hardware component based on the selected set of outputs.

57. The node of claim 56, wherein the node for implementing the training entity is the RAN node.

58. The node of claim 56, wherein the node for implementing the training entity is a computing node that is separate from the RAN node.

59. The node of any of claims 56 to 58, wherein the processing circuitry is further configured to cause the node to perform the method of any of claims 36 to 49.

60. A computer program comprising instructions which, when executed on at least one processor, cause the processor to carry out the method according to any of claims 35 to 51.

61. A carrier containing the computer program of claim 60, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, or a computer readable storage medium.

62. A non-transitory computer-readable medium comprising instructions executable by processing circuitry of a node for implementing a training entity for training a Reinforcement Learning, RL, policy for controlling one or more energy savings features of a hardware componentof a Radio Access Network, RAN, node of a wireless communication system where the RAN node comprises one or more hardware components and one or more RAN application software components operating on the one or more hardware components, respectively, of the RAN node, whereby the node is operable to: obtain (400; 500) a first set of inputs related to a RAN application state of the one or moreRAN application software components of the RAN node; obtain (402; 502) a second set of inputs related to a hardware state of a hardware component of the RAN node, the hardware component being one of the one or more hardware components of the RAN node; calculate (406; 506), based on a reward function, a reward associated to a current combined state and a previous selected set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node, the current combined state comprising the first set of inputs related to the RAN application state of the one or more RAN application software components of the RAN node and the second set of inputs related to the hardware state of the hardware component of the RAN node; select (412; 512) a new set of outputs for controlling the one or more energy savings features of the hardware component of the RAN node; and set (414; 514) the one or more energy savings features of the hardware component based on the selected set of outputs.

Citation Information

Patent Citations

  • Reinforcement learning for multi-access traffic management

    US20220014963A1

  • Radio access network intelligent application manager

    WO2023091664A1