Scalable mobility-aware energy efficiency management for cell-free networks
A deep reinforcement learning-based framework with two neural networks optimizes user handover and access point activation in cell-free massive MIMO networks, addressing inefficiencies by enhancing mobility awareness and energy efficiency while maintaining stable connections and reducing computational complexity.
Patent Information
- Application Number
- PCT/IB2024/051383
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-14
- Publication Date
- 2025-08-21
AI Technical Summary
Existing wireless communication systems, particularly user-centric cell-free massive MIMO networks, lack scalable and mobility-aware solutions for optimizing user handover and access point activation decisions, leading to inefficient energy consumption and frequent handover events.
Implementing two deep neural networks trained via deep reinforcement learning to determine user handover and access point activation decisions, considering mobility and network dynamics, using reward functions for signal-to-noise ratio and energy efficiency to optimize connection and activation strategies.
The solution provides a scalable, mobility-aware framework that reduces frequent handovers, optimizes energy efficiency, and ensures fair prioritization of users, maintaining stable connections and reducing computational complexity.
Smart Images

Figure IB2024051383_21082025_PF_FP_ABST
Abstract
Description
SCALABLE MOBILITY-AWARE ENERGY EFFICIENCY MANAGEMENT FOR CELL- FREE NETWORKS Technical Field
[0001] The present disclosure relates to a method for enabling a scalable, mobility- aware solution for improving the energy efficiency of a cell-free network in a wireless communication system. Background
[0002] A user-centric cell-free massive Multiple Input Multiple Output (UC-mMIMO) network as depicted in Figure 1B comprises a dense deployment of access points (APs) (104-1-104-N) that jointly serve the users, e.g., User Equipment (UE) 106. However, as the traffic load of the network changes, the service of some APs may not be needed, hence allowing them to enter a power-saving sleep mode, or simply get deactivated. In UC-mMIMO networks, each user is served by a set of neighboring APs, which eliminates the cells on the access channel, and hence solves the problem of poor performance of cell-edge users. Namely, each user is surrounded by APs that cooperate and jointly serve the user, which results in enhanced desired signal strength, interference mitigation, macro diversity, no cell-edges, and a smaller serving distance as compared to a cell-centric architecture as shown in Figure 1A where each AP 104 belongs to one or more cells 102.
[0003] As users move, the set of APs serving each user is updated through handoff (HO) operations, where the HO operations represent the change in the connection decisions of the users. It is evident that HOs can affect the activation and deactivation decisions of the APs. Hence, it is natural to jointly take the two decisions represented by the connection decisions of the users and the activation decisions of the APs.
[0004] Despite the importance of this topic, the activation decisions of the APs in both conventional and UC-mMIMO networks have been based on static snapshot of the network neglecting the impact of mobility. As a result, efficient solutions to jointly determine the connection decisions of the users and the activation decisions of the APs have not been proposed. Most importantly, for the UC-mMIMO scheme, the mobility aspect is more pronounced, because each user is served by a set of APs instead of a single base station (BS) as in conventional networks.Summary
[0005] Various embodiments disclosed herein provide for methods for enabling a scalable, mobility-aware solution for improving the energy efficiency of a cell-free network in a wireless communication system. Deep reinforcement learning (DRL) is used to train two deep neural network (DNNs) to serve as policies for users’ handover (HO) decisions and access points (APs) activation decisions in a user-centric cell-free massive Multiple Input Multiple Output (UC-mMIMO). The first DNN and second DNN can be trained at the gnB Central Unit (gnB-CU). The first DNN can be implemented at either the gNB-CU or one or more User Equipments (UEs) to determine the connection preferences while the second DNN is implemented at the gNB-CU to determine the activation decisions of the APs.
[0006] In an embodiment, a method can be implemented in a gNB-CU for enabling scalable mobility-aware energy efficient management of a cell-free network. The method can include receiving, as an output of a first deep neural network, DNN, connection preferences for a user equipment device, UE, wherein the output of the first DNN is based on a large scale fading between the UE and a set of Access Points, APs, of the cell-free network; a number of UEs served by each AP of the set of APs; previous connection decisions of the UE; and a first parameter related to the likelihood of channels between each AP and the UE having a threshold quality for a predefined time period, The method can also include providing, as an input to a second DNN, connection preferences for a plurality of UEs, wherein the connection preferences for the plurality of UEs are summed in a voting system, and an output of the second DNN results in activation decisions for each AP of the set of APs. The method can also include providing the activation decisions for each AP of the set of APs (104) to each AP of the set of APs from a previous decision or time slot.
[0007] In an embodiment, the method includes training the first DNN using a first reward function based on an achievable rate based on signal-to-noise ratio at the UE, with a penalty for time-loss due to performing a handover from a first AP to a second AP.
[0008] In an embodiment, the method includes training the second DNN using a second reward function associated with energy efficiency that is based on a function ofa sum of achievable data rates divided by power consumption of active APs and wherein the second reward function is based on activation and deactivation of APs.
[0009] In an embodiment, the first DNN is implemented in a first reinforcement learning environment, and the second DNN is implemented in a second reinforcement learning environment.
[0010] In an embodiment, the first reinforcement learning environment is implemented at the UE, and the second reinforcement learning environment is implemented at the gNB-CU.
[0011] In an embodiment, the method further includes providing the first DNN to the UE to implement in the first reinforcement learning environment.
[0012] In an embodiment, the receiving the connection preferences comprises receiving the connection preferences from the UE.
[0013] In an embodiment, the first reinforcement learning environment and the second reinforcement learning environment are implemented at the gNB-CU.
[0014] In an embodiment, the method includes providing connection decisions for the plurality of UEs to each UE of the plurality of UEs, wherein the connection decisions are based on the connection preferences for the plurality of UEs and the activation decisions for each AP of the set of APs.
[0015] In an embodiment, the first parameter related to the likelihood of channels between each AP and the plurality of UEs having the threshold quality for the predefined time period is based on a second parameter related to a movement direction of the plurality of UEs and a third parameter related to a history of large scale fading between the plurality of UEs and the set of APs.
[0016] In an embodiment, a gNB-CU can be configured for enabling scalable mobility-aware energy efficient management of a cell-free network where the gNB-CU comprises processing circuitry configured to receive, as an output of a first DNN connection preferences for a user equipment device, UE, wherein the output of the first DNN is based on large scale fading between the UE and a set of APs of the cell-free network; a number of UEs served by each AP of the set of APs; previous connection decisions of the UE; and a first parameter related to likelihood of channels between each AP and the UE having a threshold quality for a predefined time period. The processing circuitry can also provide, as an input to a second DNN, connection preferences for a plurality of UEs, wherein the connection preferences for the pluralityof UEs are summed in a voting system, and an output of the second DNN results in activation decisions for each AP of the set of APs. The processing circuitry can also provide the activation decisions for each AP of the set of APs to each AP of the set of APs.
[0017] In an embodiment, a method can be implemented in a UE for enabling scalable mobility-aware energy efficient management of a cell-free network. The method can include, based on network information received from a gNB-CU, generating, as an output of a first DNN, connection preferences for the UE, wherein the output of the first DNN is based on large scale fading between the UE and a set of APs of the cell- free network; a number of UEs served by each AP of the set of APs; previous connection decisions of the UE; and a first parameter related to likelihood of channels between each AP and the UE having a threshold quality for a predefined time period.
[0018] The method can include receiving, from the gNB-CU, a connection decision associated with handovers to one or more APs, wherein the connection decision is based on the connection preferences for the UE and activation decisions for each AP of the set of APs.
[0019] In an embodiment, a UE can be configured for enabling scalable mobility- aware energy efficient management of a cell-free network. The UE can include a radio interface, and processing circuitry configured to, based on network information received from a gNB-CU, generate, as an output of a first DNN, connection preferences for the UE, wherein the output of the first DNN is based on large scale fading between the UE and a set of APs of the cell-free network; a number of UEs served by each AP of the set of APs; previous connection decisions of the UE; and a first parameter related to likelihood of channels between each AP and the UE having a threshold quality for a predefined time period.
[0020] Some of the advantages of the methods and techniques disclosed herein are that the system is mobility aware. Unlike solutions used for activating access points (APs), the solution disclosed herein considers the mobility of the users. In a considered scenario, the users are assumed to be moving and hence handoffs (HOs) should be executed to update the serving set of each user. By using information from the network, a predictive HO scheme can be provided that gathers HOs in a single time step. By this, the serving set of each user can be kept the same for a longer timecompared to conventional schemes, and hence APs can be kept in activation / deactivation modes for a longer time.
[0021] Another advantage is that the system has a fast response time. Once two deep neural networks (DNNs) are trained using deep reinforcement learning (DRL), these trained DNNs can be used as the HO policy and AP activation policy, respectively. The use of DNNs provides fast feedback control where the connection decisions of the users, and the activation decision of the APs can be immediately obtained once the needed information is supplied (observation vectors) to the DNNs. This is a huge advantage over conventional mathematical optimization algorithms that employ iterative optimization approaches.
[0022] Usability and scalability: the same trained DNNs can still be used when the number of users changes in the network. The use of two separate DNN’s results in two major benefits: the first is that the solution disclosed herein is scalable in the sense that the complexity does not become higher as the number of users increase, and the second is that the same trained DNN can be used even when the number of users changes.
[0023] Another advantage is that the system is able to provide user prioritization. By using fairness weights when adding the votes of the users, high priority can be given to users that have been experiencing low data rates, and hence high priority is set for activating the APs that can help these users to get a good performance. This approach results in a solution that is more fair compared to a scheme that prioritizes users with good channels and neglecting those with weak channels. Brief Description of the Drawings
[0024] The accompanying drawing figures incorporated in and forming a part of this specification illustrate several aspects of the disclosure, and together with the description serve to explain the principles of the disclosure.
[0025] Figure 1A illustrates an example of a cell-centric network scheme according to an embodiment of the present disclosure;
[0026] Figure 1B illustrates an example of a user-centric cell-free network scheme according to an embodiment of the present disclosure;
[0027] Figure 2 illustrates a hierarchy of proposed reinforcement learning environments according to some embodiments of the present disclosure;
[0028] Figure 3 illustrates the operation of two deep neural networks (DNNs) according to some embodiments of the present disclosure;
[0029] Figure 4 illustrates an exemplary algorithm for implementing the two DNNs according to some embodiments of the present disclosure;
[0030] Figure 5 illustrates a user-centric massive Multiple Input Multiple Output (UC- mMIMO) network according to some embodiments of the present disclosure;
[0031] Figure 6 illustrates a message sequence chart for enabling scalable mobility- aware energy efficient management of a cell-free network according to some embodiments of the present disclosure;
[0032] Figure 7 illustrates one example of a cellular communications system according to some embodiments of the present disclosure;
[0033] Figure 8 is a schematic block diagram of a gNB-Central Unit (CU) according to some embodiments of the present disclosure;
[0034] Figure 9 is a schematic block diagram that illustrates a virtualized embodiment of the gNB-CU of Figure 8 according to some embodiments of the present disclosure;
[0035] Figure 10 is a schematic block diagram of the gNB-CU of Figure 8 according to some other embodiments of the present disclosure;
[0036] Figure 11 is a schematic block diagram of a User Equipment device (UE) according to some embodiments of the present disclosure; and
[0037] Figure 12 is a schematic block diagram of the UE of Figure 11 according to some other embodiments of the present disclosure. Detailed Description
[0038] The embodiments set forth below represent information to enable those skilled in the art to practice the embodiments and illustrate the best mode of practicing the embodiments. Upon reading the following description in light of the accompanying drawing figures, those skilled in the art will understand the concepts of the disclosure and will recognize applications of these concepts not particularly addressed herein. It should be understood that these concepts and applications fall within the scope of the disclosure.
[0039] Wireless Communication Device: One type of communication device is a wireless communication device, which may be any type of wireless device that hasaccess to (i.e., is served by) a wireless network (e.g., a cellular network). Some examples of a wireless communication device include, but are not limited to: a User Equipment device (UE) in a 3GPP network, a Machine Type Communication (MTC) device, and an Internet of Things (IoT) device. Such wireless communication devices may be integrated into, a mobile phone, smart phone, sensor device, meter, vehicle, household appliance, medical appliance, media player, camera, or any type of consumer electronic, for instance, but not limited to, a television, radio, lighting arrangement, tablet computer, laptop, or PC. The wireless communication device may be a portable, hand-held, computer-comprised, or vehicle-mounted mobile device, enabled to communicate voice and / or data via a wireless connection.
[0040] Note that the description given herein focuses on a 3GPP cellular communications system and, as such, 3GPP terminology or terminology similar to 3GPP terminology is oftentimes used. However, the concepts disclosed herein are not limited to a 3GPP system.
[0041] Note that, in the description herein, reference may be made to the term “cell”; however, particularly with respect to 5G NR concepts, beams may be used instead of cells and, as such, it is important to note that the concepts described herein are equally applicable to both cells and beams.
[0042] There are challenges with existing implementations, whereby current solutions that maximize the energy efficiency of wireless networks through optimizing the activation decisions of the access points (APs) do not consider the mobility aspects of the users. Instead, these solutions are developed to work on a snapshot of the network while neglecting the effect of the temporal changes in the network due to the movement of the users. Moreover, frameworks that consider both the connection decisions for the users (especially the handoff (HO) decisions) and the activation of the APs do not exist. Another important aspect is that most mathematical optimization frameworks do not consider the scalability aspect of the network. For example, most of these solutions have a growing computational complexity as the number of users increases.
[0043] Even in cases where solutions aim to enhance energy efficiency by using trained neural networks, the employed neural network may still expand in proportion to the number of users within the network. Moreover, most of these solutions aredeveloped for conventional networks instead of the user-centric cell-free massive MIMO (UC-mMIMO) network scheme disclosed herein.
[0044] Provided this, a framework is needed to allow for determining both the connection decisions of the users as they move and the activation decisions of the APs so that energy efficiency is maximized. Most importantly, the solution developed should be scalable with respect to the number of users found in the network.
[0045] The present disclosure provides a solution to these challenges, where the present disclosure provides a scalable solution that optimizes both the HO decisions of the users and the activation decisions of the APs in the user-centric cell-free massive MIMO (UC-mMIMO) network scheme. Deep reinforcement learning (DRL) can be used to train two deep neural networks (DNNs) to serve as policies for users’ HO decisions and APs’ activation decisions.
[0046] The first DNN, denoted as DNN A, uses information about the large-scale fading (LSF) between each user and the APs, the number of users served by each AP, the previous connection decisions of the user, and a metric that characterizes the likelihood of each channel between the user and the APs to be in a good state in the future. Using this information, the DNN that is trained using DRL, provides the connection preferences of the user to the APs. The reward function used in the DRL framework is the achievable rate based on the signal-to-noise ratio (SNR) with a time overhead for executing HOs. The time overhead is defined as a non-linear function that defines the cost for initiating HOs. By tuning this cost, HO decisions can be executed during the same time step, which provides a stable and longer connection for the user with the APs and decreases frequent HO decisions. This also provides an advantage for the activation decisions of the AP, where APs can be deactivated for a longer time compared to the case when the users execute frequent HOs at sparse time steps.
[0047] The second DNN, denoted as DNN B, collects the connection preferences of each user and sums them up as a voting system. The summation of these connection preferences (treated now as votes) is done through a predefined weighted sum function, where the weights can be any metric that gives priority to chosen users. These weights are updated each time step to change the priority. One typical metric is proportional fairness, where the users’ weights are defined through an exponentially weighted window that balances achievable rates of the users. Hence, users experiencing high data rates get low weights over time, while users getting low data rates get highweights over time. In this manner, the voting system can be tuned more toward users with a low data rate, making their votes more important.
[0048] The reward function used in the DRL framework for training DNN B is the energy efficiency (EE), which is a fraction function that has the sum achievable rates of the users in the numerator and the power consumption of the APs in the denominator. The power consumption model considers the transmit power and the power consumption of the radio frequency circuit of the APs. An additional penalty can be included in the reward function to penalize the frequent activation / deactivation of the APs, which allows us to put the APs into a deactivated mode for a longer time than normal setting.
[0049] By using two separate DNNs that take their observations from modeled reinforcement learning (RL) environments, a scalable framework is provided that is independent of the number of users found in the network.
[0050] Some of the advantages of the methods and techniques disclosed herein are that the system is mobility aware. Unlike solutions used for activating APs, the solution disclosed herein considers the mobility of the users. In a considered scenario, the users are assumed to be moving and hence HOs should be executed to update the serving set of each user. By using information from the network, a predictive HO scheme can be provided that gathers HOs in a single time step. By this, the serving set of each user can be kept the same for a longer time compared to conventional schemes, and hence APs can be kept in activation / deactivation modes for longer time.
[0051] Another advantage is that the system has a fast response time. Once two DNNs are trained using deep reinforcement learning (DRL), these trained DNNs can be used as the HO policy and AP activation policy, respectively. The use of DNNs provides fast feedback control where the connection decisions of the users, and the activation decision of the APs can be immediately obtained once the needed information is supplied (observation vectors) to the DNNs. This is a huge advantage over conventional mathematical optimization algorithms that employ iterative optimization approaches.
[0052] Usability and scalability: the same trained DNNs can still be used when the number of users changes in the network. The use of two separate DNN’s results in two major benefits: the first is that the solution disclosed herein is scalable in the sense that the complexity does not become higher as the number of users increases, and thesecond is that the same trained DNN can be used even when the number of users changes.
[0053] Another advantage is that the system is able to provide user prioritization. By using fairness weights when adding the votes of the users, high priority can be given to users that have been experiencing low data rates, and hence high priority is set for activating the APs that can help these users to get a good performance. This approach results in a solution that is more fair compared to a scheme that prioritizes users with good channels and neglecting those with weak channels.
[0054] In an embodiment, a DNN framework is developed that manages the connection decisions of the moving users and the activation decisions of the APs. As the solution manages the changes in the connections decisions of the users with respect to time, it inherently manages the handoff (HO) decisions. There are a couple of difficulties for developing a solution that directly determines these control decisions through a single DNN. These difficulties can be summarized as follows.
[0055] Training Difficulty: The dimension of a DNN that needs to jointly determine the connection decisions of the users and the activation decisions of the APs is verylarge. Namely, it will be ^ + ^^, where ^ is the number of APs, and ^ is the number ofusers. As a result, training such a DNN using deep reinforcement learning (DRL) will be very hard and achieving good performance may not be feasible with the level of compute processing available in today’s basestations. Furthermore, the size of the DNN will be very large, which will take up a large disk space. In summary, such a solution is not scalable.
[0056] Reusability: Any direct solution will be dependent on the number of users found within the network. Hence, when the number of users changes, a new DNN is reconstructed with different input and output dimensions, and then it is re-trained. Hence, such a solution is not reusable.
[0057] To overcome these difficulties, the present disclosure provides for a voting system that is composed of two DNNs. The first DNN is named DNN A, and it is used to determine the connection preferences of each user independently. While the second DNN is named DNN B and is used to determine the APs that will be active and the APs that will be deactivated. After these decisions are obtained, a scheme can be designed to determine the connection decisions (or HO connections) of the users. This allows theDNN architecture to be solution independent of the number of users in the network in terms of both dimension and complexity, which leads to a scalable solution.
[0058] Figure 2 illustrates a hierarchy of proposed reinforcement learning environments according to some embodiments of the present disclosure.
[0059] A hierarchical structure of reinforcement learning (RL) environments composed of two types, which is denoted as the 1) Main RL environment 202 and the 2) Subset RL environment 204-1 – 204-N as shown in Figure 2.
[0060] The main RL environment 202 hosts all the network information needed to execute the control decisions of interest to this study. Most importantly, these parameters compose the observation vector that will serve as the input for DNN B. It also includes the network information that will be communicated for the subset RL environments. To allow for inter-communication between the main and the subsets RL environments, the main RL environment and the subset RL environments are equipped with inter-RL communication functions that compose their application programming interface (API). The API implemented by the main RL is composed of the following functions: • ^^(^)= getNetInfo(user_id): this function provides the information needed for the RL environment for user identified by the ID: user_id. The information returned by this function includes the large-scale fading (LSF) statistics between the user and the APs, the number of users served by each AP, the movement direction indicator (discussed later), and the history of LSF state indicator (discussed later). This information will be used by the subset RL environment as the observation vector that will be supplied to DNN A. • updateVotes(^^(^)): this function will take the connection preferences of the users from the subset RL environments and treat them as the users’ vote.
[0061] The subset RL environment will be generated independently for each user, so each user will have its own environment. It will take its input from the main RL environment using the provided API. The subset RL environment interacts with DNN A by supplying the observation and reward, and then obtaining the action. The API implemented by the subset RL environment for intercommunication with the main RL environment include:• setNetInfo(^^(^)): this function sets the network info needed by the user to construct the observation vector at time step ^.
[0062] Figure 2 depicts this hierarchy with the main communication streams between these environments. Also included is the reset() and step() functions, which are standard functions used by the OpenAI's Gym / Gymnasium library which is a standard API to implement RL environments.
[0063] DNN A interacts with the subset RL environment for each user, here, indexed by ^. When interacting with subset RL environment ^, the output of DNN A is a vector (^) (^) ^^^ = ^^^^ … ^(^)^^^ ∈ ℝ^×^ from a continuous action space, where ^ is the number ofrepresents the connection preferences of user ^ to the APs found inthe network. The higher the value of ^(^)^^, the higher the preference of user ^ to connect to AP ^, hence the action ^^(^)can be treated as a vote of the user to connect to the APs.
[0064] Using ^^(^)obtained from both DNN A and the APs activation decisions ^^(^)∈ ,denoted as ^(^)^^^×^ obtained from DNN B, the connection decisions (^) (^)^ = ^^^^ … ^^^^ ∈ℝ^×^, are determined by connecting user ^ to the ^^^^activeconnection preferences. Mathematically, each connection decision between user ^ and AP ^ can be defined as: (^)^(^) 1, if ^ ∈ arg maxK*^ ⊙ ^^(^); ^ -^^ = ^ ^ ^^^ Eqn. 1where ⊙of thelargest 6 values in vector 5. The first condition in Eqn.1 means that user ^ will be served by AP ^ at time step ^ if this AP is active and it was among the top ^^^^of the user’s vote. Otherwise, the user will not be served by this AP.
[0065] The input for DNN A is the observation vector 7(^) 8^×^^ ∈ ℝ defined as:7(^) = S : log <(^) (^) (^) (^) ( ) ( )^ … log <^^ = S : ℰ^ … ^A^ ^A^^ ^ ^ * ^ - * -^ ^> > >ℰ^ >^= @ :^^^^ … ^^^ ^=2whereofusers served by the terms B(^),(C) (^),(D)^^and B^^are two metrics that describe the likelihood of the LSF between userin a good state, which is definednext. The term ^(^^^A^)is the connection decision between user ^ and AP ^ at the previous (^ − 1).
[0066] Thefor the LSF statistics is used to limit the dynamic range of the LSF. The operator S{⋅} applies a min-max normalization, then shifts and scales the input to make it between −1 and 1, hence, any entry in the observation vector 7^(^)lies in the range [−1, 1]. This operation can be applied on the elements of a vector 6 =[… 6K … ]^ through 6L^MNOPK = 2 * RSATU^{R}TMV{R}ATU^{R} − 0.5-, which leads to a scaled version6L^MNOP = Y… 6L^MNOP ^K … Z . Itoperator to be applied independently foreach different piece of information as shown in Eqn. 2. Investigations show that this leads to an easier configuration for the hyperparameters of the DRL algorithm used to train the DNN, which helps in making the training of the DNN smooth and fast.
[0067] The reason behind using :B(^),(C) (^),(D)^^ : ^ ∈ ℬ= and :B^^ : ^ ∈ ℬ=, where ℬrepresents the APs, in the that are more likelyto have a good channel toward the user. By this, predictive connection preferences can be facilitated for the users to reduce unnecessary HOs. Both the two parameters :B(^),(C)^^ : ^ ∈ ℬ= and :B(^),(D)^^ : ^ ∈ ℬ= use different network information to predict theand by using two different metrics it allows for a robust predictive HO.
[0068] The first parameter 0 ≤ B(^),(C)^^ ≤ 1 uses the movement direction of the userto give higher priority for APsis moving toward, and it is defined at time step ^ as: B(^),(C)^^ = ^^L(^_`)a^b , Eqn. 3where cof the line connecting this user to AP ^. This angle can be calculated using the locations of the user and the APs. In 5G networks, this information can be accessed because the locations of the APs are known, while the location of the user can be accurately tracked using either global positioning system (GPS), received signal strength indicator (RSSI), or other network identifiers.
[0069] The second parameter 0 ≤ B(^),(D)^^ ≤ 1 uses the history of the state of the LSFbetween user ^ and AP ^ toit is that this LSF is in a good state. A LSF is defined to be in a good state if it is larger than some chosen threshold <defOLe^NP.The intuition behind B(^),(D)^^is that if the LSF was in a good state in the last few time steps, then there is probability for it to be in a good state in the next time step. (^),(D)Namely, B at ^ is efined as: ∑k klm (mlo)mno ij p:q_` rqstuvwtjxy== g ∑k klmmno ij , if ^ > 0 Eqn. 40, otherwisewhere < for the LSF to be classified in achannel classification history, and p{⋅}the indicator function where p{}} = 1, if condition } is valid, and equals 0, otherwise.The term {|^A~^ : ^ ≤ ^} found in the numerator of Eqn. 4 prioritizes recent history.
[0070] The intuition from B(^),(C)^^and B(^),(D)^^in Eqn. 3 and Eqn. 4, respectively, is to provide information about the of the LSF in the future for the DRLalgorithm, which eventually allows for predictive connections preferences and hence predictive HOs.
[0071] The subset RL environment is defined for each user separately. Using both the connection decisions ^^(^)and the observation vector 7^(^)of user ^, a HOs penalized reward function can beas:^*7(^) (^) (^) U^P (^) (^)^ , ^^ - = ^ ^^ *7^ , ^^ -, Eqn. 5where ^from the actual achievable rate derived below. The term ^(^)is the fractional time remaining for data transmission after the weighted time penalty for the HOs is deducted; this penalty parameter is defined as: ^(k)^(^) = ^^^A^^^,sjs^x^ 6where ^transmission phases, ^^is the number of communication cycles contained in a single time step, and ^(^)D^,d^dMNis the weighted overhead from executing the HOs at the beginning of time step t. ^(^)D^,d^dMNis defined through a nonlinear function ^(⋅) using: ^(^) (^) (^) (^)D^,d^dMN = ^^^ ^ = min^p^^ > 0^^^ + ^ ^D^, ^^^^^ Eqn. 7where the term ^^is a basic overhead for initiating HOs; it represents the weighted time spent doing radio resource control (RRC) reconfiguration, handshakes, reporting, etc. The term ^D^is the weighted overhead resulting from each HO, and ^(^)is the number of HOs per time step ^. The minimum operator found in Eqn.7 simply indicates that theduration between any two decision steps (^ − 1) and ^ should be chosen to be morethan the maximum duration of any HOs decision. The intuition from the nonlinear function in Eqn.7 is to both allow for a flexible time overhead that is a function of the number of HOs and to use a penalty for initiating HOs.
[0072] The achievable rate ^^U^P(⋅)used in Eqn.5 provides an estimate of both the interference experienced by the user from the non-served APs and the power allocated to the user based on the number of users served by each AP. This approximation enables the implementation of the HO solution on an individual basis for each user. This achievable rate is defined as: ^^ (̀k) (̀k)o,^,`^ ,^ ,^ ^^U^P^ *7(^)^ , ^(^)^ - = ^^^ ∑^^^¢^vws log ^1 +^ ^,^(̀k),^(̀k)^a^^ (̀k),^(̀k)- Eqn.9whereb£ ¥, 7(^), ^(^) = ¦b§b[¥ − ¥ ] ∑ ^(^)© (^) (^)^,¤,^^ ^ ^ ^ ^ OLd ¨ ^∈ℬ ^^ ª^^« ¨ Eqn.10where^̀ ( ) (k) ^° [^ AK]± ² *q -«(^) = vws _`^ for ^ ∈ ³ 13and theª(^) ±(y)^^ =´* (k) (k)
[0073] B is a vector of continuous output ^¶ (^) ∈ ℝ^×^ that ismapped to the activation / deactivation decisions for the APs. The activation decisions arerepresented by the vector ^^(^) = ^^^(^)^ … ^^(^) ^^×^^ ^ ∈ ^ where ^^(^)^ = ¶(^)^ > 0=. Intheory, during the training of^^(^)^ ; however, using a threshold 0 is the most logical as it is the decisive boundary betweenthe negative and positive values of ^¶ (^).
[0074] The input for DNN B is determined through the observation vector: 7(̅^) = Y S^¸(^)^ S^^^(^A^)^Z^Eqn. 14where ¸¸(^) = ∑ (^) (^)^∈³ ¹^ S{^^ } Eqn. 15
[0075] foruser ^ at time earlier, the operator S{⋅} performs min-maxnormalization vector, then shifts the input to make it between −1 and 1.
[0076] The intuition from ¸(^)is to collect the votes of the users while applyingweights :¹(^)^ : ^ ∈ ³= that enables prioritization of individual user contributions. Thereare to define these weights; one way is through the usage of proportional fairness as: (^) 1, if ^ = 0¹^ = º ^if t > ^^(^)^ = Eqn. 16with:() 0, if ^ = 0^^ ^^ = ^ª(¼½)^(^A^) ^ − ª(¼½)^^^(^A^) Eqn. 17where ^^datarate of user ^ averaged over previous decision steps with ^^(^) (¼½)^ = 0, and 0 ≤ ª ≤ 1 isthe forgetting factor.
[0077] As for the reward used to train DNN B, the energy efficiency (EE) metric is used which is measured in bits / joules, and it is defined in the downlink at time step ^ as: (EE(^) = ¿ ∑ ̀k)`∈³ » Eqn. 18where Áof the system, ^^(^)is the downlink achievable rate of the user measured in bits / s / Hz which can be defined in detail in the system model, and Â(^, d^dMN)is the total power consumption of the APs on the downlink. This power consumption model is measured in Watts or Joules.
[0078] In a dense network, where the communication occurs at short distances, the transmit power becomes comparable to the circuit power consumption. Hence, thepower consumption of the circuit of the APs needs to be taken into consideration in any study that optimizes the EE.
[0079] We use a generic downlink power consumption model written as: Â(^, d^dMN) = ∑ ^^(^) Â(ÃÄ) + Â(^, ÅÆANP) + Â(ÅÆAÃÅ) (ÅÆA½ÇÄ)^∈ℬ * ^ * ^ ^ ^ - + Â^ - Eqn. 19to the power (PA), load-dependent circuit operations, the transceiver chain,and basic The last three terms of the power consumptionrepresent the circuit power consumption, where Â(^, ÅÆANP)^ is load-dependent while bothÂ(ÅÆAÃÅ)^ and Â(ÅÆA½ÇÄ)^ are load-independent.of the PA includes the sum of the radiatedpower needed to send data for the users and the dissipated power, and it can be represented through: (à ±(y) Ä)^ = È(ÉÊ) Eqn. 20where 0=0.39, and ®(P) is the power budget of each AP. The formula in Eqn. 20 simply states thatthe total power consumed will be more than the power used for transmitting data and pilot signals because some dissipated power is due to an imperfect PA.
[0081] The term Â(ÅÆAÃÅ)^ accounts for the power consumption of the transceiver chains, and it is defined as: Â(ÅÆAÃÅ)^ = ¦Â(Ë)^ + Â(¤ÎÏ)^ Eqn. 21where ¦the circuit components attached to each antenna, and Â(¤ÎÏ)^ is the power consumed by the local oscillator.
[0082] The load dependent power consumed by the circuit can be approximated as: Â(^, ÅÆANP) ≃ Â(^, ÅÑ) + Â(^, ÅC) + Â(^, ½D) + Â(^, ÒÆ)^ ^ ^ ^ ^ Eqn. 22where ÂÂ(^^, ÅC)accounts for channel coding and decoding units, Â(^^, ½D)accounts for the load-dependent fronthaul, and Â(^^, ÒÆ)accounts for the transmit and receive beamforming at the AP.
[0083] The power consumption of the channel estimation can be approximated as: b´ (k)(^, Å ) ¿ ^Ó>ℰ_ >Â Ñ^ = ^^ Ô(ÊÉ) Eqn. 23where ^Õthelength of of the AP measured in flops / Watt (operations per Joule). The first fraction term in Eqn. 23 represents the number of coherence blocks per second where the AP performs a single pilot-based CSI estimation per block. In the uplink, the AP receives the pilot signal as an¦ × ^ matrix, and it needs to estimate the channe (^)± ls of each user ^ ∈ ℰ^ bymultiplying the matrix with the corresponding pilot sequence of each user.
[0084] The power consumption for the channel coding and decoding units is proportional to the number of bits and hence can be approximated as Â(^, ½D) (½Ã) ∑ (^)^ =  ^∈ℰ(k)_ ×:^^ = Eqn. 24where Â
[0085] The term Â(^^, ÒÆ)accounts for the transmit and receive beamforming at the AP. For conjugate beamforming, it can be approximated as: Â^, ÒÆ) Ø b (k)( ^ ´>ℰ_ > (^, Ò )^ = Á *1 − ^^- (ÊÉ) + Â ÆÙ^ Eqn. 25whereper data symbol, while the second term Â(^^, ÒÆÙ)is a beamforming dependent term that accounts for the computation of the transmit and receive beamformers. For maximum ratio transmission and receiving, Â(^^, ÒÆÙ)is defined as: ¬(k)(^, ÒÆ ) ¿ ´>ℰ_ >Â ÙAÚÛ^ = 26where
[0086] Figure 3 illustrates the operation of two deep neural networks (DNNs) according to some embodiments of the present disclosure.
[0087] The main RL environment 202 and the subset RL environment 204 can be separately used when training DNN B and DNN A, respectively. This allows for implementing the RL environments as a standard OpenAI's Gym / Gymnasium library RL environment, which includes implementing the reset() and the step() methods, and removing the intercommunication between the main RL 202 and the subset RL environments 204. In doing so, the standard implementation of DRL algorithms is usedwhich are tailored to interact with RL environments using these aforementioned methods.
[0088] Separate training of the main RL and the subset RL environments can be enabled by introducing small modifications during the training. For the main RL environment, the updateVotes*^^(^)- method is not used to obtain the votes of the users and instead the log of the can be directly used.
[0089] For the subset RL a single user network can be generated instead of using the setNetInfo() method (check Figure 1). Another simple modification is to obtain the connection decisions of the user from the connection preferences. In this regard, each user can choose the ^^^^connections with the largest preference cores resulting in the binary vector ^( ^s Ü^)^ = ^^Ü(^)^^ … ^Ü(^) ^×^^^^ ∈ ℝ which represents theAPs that the user prefers to connect to APs are active.Mathematically, this can be described as: ^1, if ^ ∈ arg ma (^)Ü(^) xK*^^ ; ^^^^-^^ = ^ Eqn. 270,where arg maxK*^(^)(^)^ ; ^^^ in vector ^^ with the largest values.approach, the Soft Actor-Critic (SAC) algorithm is used to implement DRL. The SAC algorithm is an off-policy learning technique, which means it reuses past experience through a replay buffer Ý to perform the training. This buffer saves past experiences from the RL environment, and it contains the following entries: [state, action, reward, next-state, episode termination indicator]. Then, to train the DNNs, a batch (set of samples) is extracted through random sampling without repetition. As opposed to on-policy learning which includes collecting new samples for every update to the policy, off-policy learning offers better utilization of data, and simpler exploration-exploitation strategies for actions.
[0091] Once done with training, the framework using the intercommunication between the two types of the RL environments, is then used which is described next.
[0092] Figure 4 illustrates an exemplary algorithm for implementing the two DNNs according to some embodiments of the present disclosure.
[0093] In Algorithm 1, the operation of the scalable and reusable DNN framework and RL environment architecture is detailed. The input for the algorithm is the trainedDNNs, i.e., DNN A and DNN B, while the output is the APs’ activation decisions and the users’ connection decisions.
[0094] In Steps 3 and 4, the RL environments are created that will hold the information of the network. The main RL environment is modeled to provide the information and evolution of the main parameters of the network including all the users,while each subset RL environment ^ ∈ ³ is modeled to hold those for its user ^. Steps 5and 6 initialize the parameters of the RL environments to prepare them to interact with the DNNs.
[0095] After that, the episode starts by providing the info needed by each subset RL environment to operate (Steps 10-12). In Step 14, the votes of the users are collected and provided to the main RL environment. In Step 15, the votes are processed and added. In Step 16, the observation of the main RL environment is constructed. In Step 17, the action for the main RL environment is obtained using the trained DNN B. In Step 18, the activation decisions of the APs are determined from the action of DNN B.
[0096] The connection decisions of each user are, then, determined as shown in Step 19 using the connection decisions which were obtained from the trained DNN A and the activation decisions of the APs. Then, in Steps 20 and 21, these actions are executed by the subset and main RL environments, and new rewards are obtained. Due to the complex RL architecture, the observations are not directly obtained after executing the actions. This is different from standard RL environments. Finally, in Step 22, the fairness weights of the users are updated, and now the next time step can proceed as shown in Step 23.
[0097] Figure 5 illustrates a user-centric massive Multiple Input Multiple Output (UC- mMIMO) network according to some embodiments of the present disclosure.
[0098] We consider a network comprising ^ APs represented through the set ℬ and connected to a gNB-CU 502 through a wired fronthaul solution, as shown in Figure 5. Each AP is equipped with ¦ antennas and serves the users in the set ³ according to theUC-mMIMO scheme. Specifically, at any time step ^, each user ^ ∈ ³ is served by a setof APs represented by the set Þ^(^)using the coherent transmission mode. The serving set for each user is chosenthe neighboring APs, irrespective of the cells. Hence, in a UC-mMIMO scheme the cells have no relevance on the access channel. Consequently, the users will experience comparable performance, and the concept of cell-edge user is eliminated.
[0099] Due to the importance of mobility considerations, it is assumed that communication is divided into time steps, where at each step the user moves a specific distance determined by the velocity of the user ß^and the fixed time duration Δ¶between any ^ and ^ + 1. As the user moves, Þ(^)^ needs to be updated by adding andremoving APs through HOs. Hence, HOs the beginning of each decision cycle ^if the connection decisions of the user i.e., Þ(^) (^A^)^ ≠ Þ^ . The users to beserved by AP ^ are represented through the set ℰ :ℰ(^): ^ ∈ ℬ= can be^directly obtained from the sets :Þ(^) (^)^ : ^ ∈ ³=,∈ ℰ^ ⇔ ^.
[0100] At each time step ^, deactivate theAPs, aiming to optimize energy proper activation decision must consider the connection preferences (i.e., HOs) for the users. An objective is to develop a DNN framework that obtains both the users' connection decisions and the APs' activation decisions. This framework should be mobility aware and allows to minimize HO decisions and APs' transitions between active and non-active modes. In this regard, the predictive behavior is invaluable in optimizing these critical decisions.
[0101] The mobility of the users causes temporal variations in the wireless channels. Consequently, when a user moves, utilizing a block fading channel model with a specific coherence time is inaccurate; instead, a time-varying channel model is necessary. This is especially important for high-speed users.
[0102] We consider a communication cycle that contains ^^channel uses comprising an uplink pilot training phase of length ^ãand a downlink data transmission phase of length ^P. Within a communication cycle, the small-scale fading part of the channel is considered to be a wide-sense stationary (WSS) process, while the large-scale fading (LSF) is considered constant because it changes at a slower rate compared to the duration ^^, even for a high mobility scenario.
[0103] The channel realization between AP b and user ^ at channel use instant ¥ =0, … , ^^ − 1 is modeled as ℎ^^[¥] ≜ æ<^^ç^^[¥] ∈ ℂ´×^, where ç^^[¥] ∼ Þê(0 ,́ ) isthe small-scale fading,and path loss.
[0104] The channel is temporally correlated, which means for any channel uses ¥ and ¥′, ç^^[¥] and ç^^[¥′] are correlated. The temporal correlation of the channel at ¥and ¥′; for ¥, ¥¯ = 0, … , ^^ − 1, can be characterized using Jakes' model through thetemporal correlation coefficient §^[¥ − ¥′] defined as§^[¥ − ¥¯](¥ − ¥′)^C`îL^ Eqn. 28where ì^periodof eachC` = ^ on themobility speed ß^of the user,carrier frequency ^^, and the speed of light c.
[0105] Based on Eqn. 28, for each communication cycle of length ^^, the small-scale fading at instant ¥ can be written as a function of an initial state ç^^[0] and an innovation component ¸^^[¥] as: ç^^[¥] = §^[¥]ç^^[0] + §̅^[¥]¸^^[¥] Eqn. 29where ¸^^[¥] ∼ Þê(0 ,́ ) is the innovation component at instant ¥ and it isindependent of the small-scale fading, §^[¥] is the temporal correlation coefficient ofuser ^ between channel realizations at instants 0 and ¥, with 0 ≤ §^[¥] ≤ 1, and§̅^[¥] = æ1 − |§^[¥]|b.
[0106] To construct a proper signal model, the channel at the instant it was estimated, which is denoted as instant ¥OLd, is related to the other channel use instants.
[0107] For the pilot training sequence, the delta-function is used, where the set ³Kisdefined which corresponds to the users using the pilot signal ¹[¥ − ò]. Consequently,user ^ ∈ ³K uses the pilot signal ó^[¥] = ¹[¥ − ò], with óD^[¥]ó^[¥] = 1 and 1 ≤ ò ≤ ^ã,where ^ã is the length of the pilot transmission¹[¥] = 0 for ¥ ≠ 0 and¹[¥] = 1 for ¥ = 0. At instant ò during the uplink pilot training phase, i.e., ò ≤ ^ã, thesignal received at AP ^ can be written as: 5^[ò] = ∑^ö∈³S æ®(ô)ℎ^^¯[ò] + õ^[ò] Eqn. 30where ®
[0108] Channel estimation occurs after all the training signals are received, i.e., at¥OLd = ^ã + 1. Based on Eqn. 29, ç^^[¥] at time instant ¥ ≤ ^ã can be related toç^^[¥OLd] as:Eqn. 31in the channel ℎ^^[ò]inEqn. 30, and assuming that AP ^ is estimating the channel for user ^ ∈ ³K, the signal inEqn. 30 received at AP ^ at instant ò during the uplink pilot training phase can be rewritten as:5^[ò] = æ®(ô)§^[¥OLd − ò]ℎ^^[¥OLd] + æ®(ô)<^^§^̅[¥OLd − ò]¸^^[ò] +32the°[ ]æ (²)ℎú^^[¥OLd] = ` ^vwsAK ± q_`∑`ö∈³ 5 [ò], for ^ ∈ ³ Eqn.33S ±(²)q_`ö a^ ^ ^ K
[0111] errorû = ℎ^^ OLd − ^^ OLd ^^ OLddistributed as û^^ ∼ Þê(0, Θ^^), where the covariance Θ^^ ≜ <^^́ − «^^́ , with «^^being defined as: °^̀[^ A ] (²) ^« vws K ± q_`^^ =∑ ^`ö∈³ , for ^ ∈ ³K Eqn.34S ±(²)q_`ö a^ý
[0112] beamforming (maximum ratio transmission) toserve the users, where AP ^ uses the channels estimated at time instant ¥OLdtoconstruct the beamformers used to transmit data to the users. Based on this, the signal6^[¥] ∈ Þ´×^ sent by AP ^ at time instant ¥ ≥ ¥OLd to users ℰ^ can be written as:6^[¥] = ∑^∈ℰ_ æª^^ℎú∗^^ [¥OLd]^^[¥] Eqn.35where ℎú∗¥^ [¥] complex data symbol for the user with ×{|^^[¥]|b} = 1, and ª^^ isstatistical normalizing term for the transmit power allocated by AP ^ to user ^. Thisterm allows the AP to satisfy, on-average, the available power budget ®(P), i.e.,×{‖6 b (^[¥]‖ } ≤ ®P).
[0113] Channel aging also affects the data transmission phase. Using Eqn.29 for the channel evolution inside a communication cycle, the small-scale fading at time instant¥ ≥ ¥OLd can also be represented asç^^I¥J = §^I¥ − ¥OLdJç^^I¥OL^J + §̅^[¥ − ¥OLd]¸^^[¥] Eqn.36
[0114] By employing Eqn.36, the^ during the downlink datatransmission phase, i.e., at ¥ ≥ ¥OLd, can be written as:5^I¥J = ^ ℎÃ^^ I¥J6^I¥J + õ^[¥]^∈ℬ [¥]37where ^^^¯[¥] = ∑^ö∈Þ`ö æª^¯^¯ℎ^¯^ [¥] ^^ [¥OLd]. The first term in Eqn. 37 is the desiredsignal (DS), the uncertainty (BU), the third term is thechannel aging (CA), the fourth term is the multiuser interference (MI); finally, the fifth term represents the noise.
[0115] Using the signal model presented in Eqn. 37 and considering the uncorrelated interference signal as a Gaussian noise, the rate performance can be characterized through a lower bound for the ergodic achievable rate as: ^^ ×^^^ b^^ 1 > :DS^I¥J=>^interference can be, respectively, written in closed-form as b− ^^^ ^^^ ^^^ ¨ 3940£^^^^,^^¯ I¥J =b³K³K^^on decision of ^ can be formally defined as: ^^^^ ^ = 1=pF^ ∈ Þ H = ^^^^^ ^^^^^ ^ ^ ^ ^^^ Eqn. 42
[0117] thepower user ^^^ ±^y^ª^^ = ´ ∑ µ^k^` _`
[0118] B may be implemented on the central unit (CU, while thetrained DNN A can be implemented either at the user or at the CU. To implement DNN A at the user, the single trained DNN A is copied to the user equipment. Whether DNN A is implemented on the user’s side or at the CU, it is trained only once, hence both implementations are feasible.
[0119] Figure 6 illustrates a message sequence chart for enabling scalable mobility- aware energy efficient management of a cell-free network according to some embodiments of the present disclosure.
[0120] It is to be appreciated that in the message sequence chart of Figure 6, the lines and boxes that are dashed are optional, whereas the solid lines and boxes denote steps that are non-optional with respect to one or more aspects of the present disclosure.
[0121] The method can begin at step 602, where the method includes training the first DNN. In an embodiment, the first DNN (or DNN A, as used herein) can be trained using a first reward function based on an achievable rate based on signal to noise ratio at the UE 106, with a penalty for time-loss due to performing a handover from a first AP to a second AP.
[0122] At step 604, the method includes training the second DNN (or DNN B as used herein) using a second reward function associated with energy efficiency that is based on a function of a sum of achievable data rates divided by power consumption of activeAPs 104 and wherein the second reward function is based on activation and deactivation of APs. In an embodiment, steps 602 and 604 can be performed once, or at regular intervals to update the DNN models.
[0123] At step 606, the method includes optionally, providing the first DNN to the UE 106 so that the UE 106 may optionally implement the first DNN. Alternatively, the first DNN can be implemented at the gNB-CU 502 at step 610.
[0124] When implementing the first DNN, at either 608 or 610, the implementing can include receiving, as an output of a first DNN, connection preferences for the UE 106 that are based on large scale fading between the UE (106) and a set of Access Points, APs, (104) of the cell-free net, work, a number of UEs served by each AP of the set of APs, previous connection decisions of the UE 106, and a first parameter related to likelihood of channels between each AP and the UE (106) having a threshold quality for a predefined time period.
[0125] In an embodiment, the first parameter related to the likelihood of channels between each AP and the plurality of UEs having the threshold quality for the predefined time period is based on a second parameter related to a direction of movement of the plurality of UEs and a third parameter related to a history of large scale fading between the plurality of UEs and the set of APs.
[0126] If the first DNN is implemented at the UE 106, the UE 106 can provide the connection preferences to the gNB-CU 502 at step 612.
[0127] At step 614, the method includes implementing the second DNN, which includes providing, as an input to a second DNN, connection preferences for a plurality of UEs, wherein the connection preferences for the plurality of UEs are summed in a voting system, and an output of the second DNN results in activation decisions for each AP of the set of APs. The activation decisions can then be provided to the APs 104 at step 618.
[0128] In an embodiment, the first DNN is implemented in a first reinforcement learning environment 204, and the second DNN is implemented in a second reinforcement learning environment 202.
[0129] In the embodiment where the gNB-CU 502 implements the first DNN at step 610, the connection decisions can optionally be provided to the UE 106 at step 616.
[0130] Figure 7 illustrates one example of a cellular communications system 700 in which embodiments of the present disclosure may be implemented. In theembodiments described herein, the cellular communications system 700 is a 5G system (5GS) including a Next Generation RAN (NG-RAN) and a 5G Core (5GC) or an Evolved Packet System (EPS) including an Evolved Universal Terrestrial RAN (E-UTRAN) and an Evolved Packet Core (EPC). In this example, the RAN includes base stations 702-1 and 702-2, which in the 5GS include NR base stations (gNBs) and optionally next generation eNBs (ng-eNBs) (e.g., LTE RAN nodes connected to the 5GC) and in the EPS include eNBs, controlling corresponding (macro) cells 704-1 and 704-2.
[0131] In the embodiments disclosed herein, the base stations 702 could be gNB-CUs such as gNB-CU 502 in a cell free network, while the low power nodes 706 depicted in Figure 7 could be the APs 104 described herein.
[0132] The base stations 702 and the low power nodes 706 provide service to wireless communication devices 712-1 through 712-5 in the corresponding cells 704 and 708. The wireless communication devices 712-1 through 712-5 are generally referred to herein collectively as wireless communication devices 712 and individually as wireless communication device 712. In the following description, the wireless communication devices 712 are oftentimes UEs, but the present disclosure is not limited thereto.
[0133] Figure 8 is a schematic block diagram of a gNB-CU 502 according to some embodiments of the present disclosure. Optional features are represented by dashed boxes. The gNB-CU 502 may be, for example, a base station 702 or 706 or a network node that implements all or part of the functionality of the base station 702 or gNB described herein. As illustrated, the gNB-CU 502 includes a control system 802 that includes one or more processors 804 (e.g., Central Processing Units (CPUs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), and / or the like), memory 806, and a network interface 808. The one or more processors 804 are also referred to herein as processing circuitry.
[0134] The one or more processors 804 operate to provide one or more functions of a gNB-CU 502 as described herein. In some embodiments, the function(s) are implemented in software that is stored, e.g., in the memory 806 and executed by the one or more processors 804.
[0135] Figure 9 is a schematic block diagram that illustrates a virtualized embodiment of the gNB-CU 502 according to some embodiments of the present disclosure. This discussion is equally applicable to other types of network nodes.Further, other types of network nodes may have similar virtualized architectures. Again, optional features are represented by dashed boxes.
[0136] As used herein, a “virtualized” gNB-CU 502 is an implementation of the gNB- CU 502 in which at least a portion of the functionality of the gNB-CU 502 is implemented as a virtual component(s) (e.g., via a virtual machine(s) executing on a physical processing node(s) in a network(s)). As illustrated, in this example, the gNB- CU 502 may include the control system 802 as described above. The gNB-CU 502 includes one or more processing nodes 900 coupled to or included as part of a network(s) 902. If present, the control system 802 is connected to the processing node(s) 900 via the network 902. Each processing node 900 includes one or more processors 904 (e.g., CPUs, ASICs, FPGAs, and / or the like), memory 906, and a network interface 908.
[0137] In this example, functions 910 of the gNB-CU 502 described herein are implemented at the one or more processing nodes 900 or distributed across the one or more processing nodes 900 and the control system 802 in any desired manner. In some particular embodiments, some or all of the functions 910 of the gNB-CU 502 described herein are implemented as virtual components executed by one or more virtual machines implemented in a virtual environment(s) hosted by the processing node(s) 900. As will be appreciated by one of ordinary skill in the art, additional signaling or communication between the processing node(s) 900 and the control system 802 is used in order to carry out at least some of the desired functions 910.
[0138] In some embodiments, a computer program including instructions which, when executed by at least one processor, causes the at least one processor to carry out the functionality of gNB-CU 502 or a node (e.g., a processing node 900) implementing one or more of the functions 910 of the gNB-CU 502 in a virtual environment according to any of the embodiments described herein is provided. In some embodiments, a carrier comprising the aforementioned computer program product is provided. The carrier is one of an electronic signal, an optical signal, a radio signal, or a computer readable storage medium (e.g., a non-transitory computer readable medium such as memory).
[0139] Figure 10 is a schematic block diagram of the gNB-CU 502 according to some other embodiments of the present disclosure. The gNB-CU 502 includes one or more modules including the first reinforcement learning environment 204 and the secondreinforcement learning environment 202, each of which is implemented in software. The first reinforcement learning environment 204 and the second reinforcement learning environment 202 provide the functionality of the gNB-CU 502 described herein. This discussion is equally applicable to the processing node 900 of Figure 9 where the first reinforcement learning environment 204 and the second reinforcement learning environment 202 may be implemented at one of the processing nodes 900 or distributed across multiple processing nodes 900 and / or distributed across the processing node(s) 900 and the control system 802.
[0140] Figure 11 is a schematic block diagram of a UE 106 according to some embodiments of the present disclosure. As illustrated, the UE 106 includes one or more processors 1102 (e.g., CPUs, ASICs, FPGAs, and / or the like), memory 1104, and one or more transceivers 1106 each including one or more transmitters 1108 and one or more receivers 1110 coupled to one or more antennas 1112. The transceiver(s) 1106 includes radio-front end circuitry connected to the antenna(s) 1112 that is configured to condition signals communicated between the antenna(s) 1112 and the processor(s) 1102, as will be appreciated by on of ordinary skill in the art. The processors 1102 are also referred to herein as processing circuitry. The transceivers 1106 are also referred to herein as radio circuitry. In some embodiments, the functionality of the UE 106 described above may be fully or partially implemented in software that is, e.g., stored in the memory 1104 and executed by the processor(s) 1102. Note that the UE 106 may include additional components not illustrated in Figure 11 such as, e.g., one or more user interface components (e.g., an input / output interface including a display, buttons, a touch screen, a microphone, a speaker(s), and / or the like and / or any other components for allowing input of information into the UE 106 and / or allowing output of information from the UE 106), a power supply (e.g., a battery and associated power circuitry), etc.
[0141] In some embodiments, a computer program including instructions which, when executed by at least one processor, causes the at least one processor to carry out the functionality of the UE 106 according to any of the embodiments described herein is provided. In some embodiments, a carrier comprising the aforementioned computer program product is provided. The carrier is one of an electronic signal, an optical signal, a radio signal, or a computer readable storage medium (e.g., a non-transitory computer readable medium such as memory).
[0142] Figure 12 is a schematic block diagram of the UE 106 according to some other embodiments of the present disclosure. The UE 106 includes a first reinforcement learning environment 204, which is implemented in software. The first reinforcement learning environment 204 provides the functionality of the UE 106 described herein.
[0143] Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as Read Only Memory (ROM), Random Access Memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according to one or more embodiments of the present disclosure.
[0144] While processes in the figures may show a particular order of operations performed by certain embodiments of the present disclosure, it should be understood that such order is exemplary (e.g., alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, etc.).
[0145] Those skilled in the art will recognize improvements and modifications to the embodiments of the present disclosure. All such improvements and modifications are considered within the scope of the concepts disclosed herein.
Claims
Claims 1. A method implemented in a gNB Central Unit, gNB-CU, (502) for enabling scalable mobility-aware energy efficient management of a cell-free network, the method comprising: receiving (610, 612), as an output of a first deep neural network, DNN, connection preferences for a user equipment device, UE, (106) wherein the output of the first DNN is based on: large scale fading between the UE (106) and a set of Access Points, APs, (104) of the cell-free network; a number of UEs served by each AP of the set of APs (104); previous connection decisions of the UE (106); and a first parameter related to likelihood of channels between each AP and the UE (106) having a threshold quality for a predefined time period; providing (614), as an input to a second DNN, connection preferences for a plurality of UEs, wherein the connection preferences for the plurality of UEs are summed in a voting system, and an output of the second DNN results in activation decisions for each AP of the set of APs (104); and providing (618) the activation decisions for each AP of the set of APs (104) to each AP of the set of APs (104).
2. The method of claim 1, further comprising: training (602) the first DNN using a first reward function based on an achievable rate based on signal to noise ratio at the UE (106), with a penalty for time-loss due to performing a handover from a first AP to a second AP.
3. The method of any of claim 1 to 2, further comprising: training (604) the second DNN using a second reward function associated with energy efficiency that is based on a function of a sum of achievable data rates divided by power consumption of active APs (104) and wherein the second reward function is based on activation and deactivation of APs (104).
4. The method of any of claims 1 to 3, wherein the first DNN is implemented in a first reinforcement learning environment (204), and the second DNN is implemented in a second reinforcement learning environment (202).
5. The method of claim 4, wherein the first reinforcement learning environment (204) is implemented at the UE (106), and the second reinforcement learning environment (202) is implemented at the gNB-CU (502).
6. The method of claim 5, further comprising: providing (606) the first DNN to the UE (106) to implement in the first reinforcement learning environment (204).
7. The method of any of claims 5 to 6, wherein the receiving (610, 612) the connection preferences comprises receiving the connection preferences from the UE (106).
8. The method of claim 4, wherein the first reinforcement learning environment (204) and the second reinforcement learning environment (202) are implemented at the gNB-CU (502).
9. The method of claim 8, further comprising: providing (616) connection decisions for the plurality of UEs to each UE (106) of the plurality of UEs, wherein the connection decisions are based on the connection preferences for the plurality of UEs and the activation decisions for each AP of the set of APs (104).
10. The method of any of claims 1 to 9, wherein the first parameter related to the likelihood of channels between each AP and the plurality of UEs having the threshold quality for the predefined time period is based on a second parameter related to a direction of movement of the plurality of UEs and a third parameter related to a history of large scale fading between the plurality of UEs and the set of APs (104).
11. A gNB Central Unit, gNB-CU, (502) configured for enabling scalable mobility- aware energy efficient management of a cell-free network, the gNB-CU (502) comprising processing circuitry configured to: receive (610, 612), as an output of a first deep neural network, DNN, connection preferences for a user equipment device, UE, (106) wherein the output of the first DNN is based on: large scale fading between the UE (106) and a set of Access Points, APs, (104) of the cell-free network; a number of UEs served by each AP of the set of APs (104); previous connection decisions of the UE (106); and a first parameter related to likelihood of channels between each AP and the UE (106) having a threshold quality for a predefined time period; provide (614), as an input to a second DNN, connection preferences for a plurality of UEs, wherein the connection preferences for the plurality of UEs are summed in a voting system, and an output of the second DNN results in activation decisions for each AP of the set of APs (104); and provide (618) the activation decisions for each AP of the set of APs (104) to each AP of the set of APs (104).
12. The gNB-CU (502) of claim 11, wherein the processing circuitry is further configured to: train (602) the first DNN using a first reward function based on an achievable rate based on signal to noise ratio at the UE (106), with a penalty for time-loss due to performing a handover from a first AP to a second AP.
13. The gNB-CU (502) of any of claim 11 to 12, wherein the processing circuitry is further configured to: train (604) the second DNN using a second reward function associated with energy efficiency that is based on a function of a sum of achievable data rates divided by power consumption of active APs (104) and wherein the second reward function is based on activation and deactivation of APs (104).
14. The gNB-CU (502) of any claims 11 to 13, wherein the first DNN is implemented in a first reinforcement learning environment (204), and the second DNN is implemented in a second reinforcement learning environment (202).
15. The gNB-CU (502) of claim 14, wherein the first reinforcement learning environment (204) is implemented at the UE (106) and the second reinforcement learning environment (202) is implemented at the gNB-CU (502).
16. The gNB-CU (502) of claim 15, wherein the processing circuitry is further configured to: provide (606) the first DNN to the UE (106) to implement in the first reinforcement learning environment (204).
17. The gNB-CU (502) of any of claims 15 to 16, wherein receiving (610, 612) the connection preferences comprises receiving the connection preferences from the UE (106).
18. The gNB-CU (502) of claim 14, wherein the first reinforcement learning environment (204) and the second reinforcement learning environment (202) are implemented at the gNB-CU (502).
19. The gNB-CU of claim 18, wherein the processing circuitry is further configured to: provide (616) connection decisions for the plurality of UEs to each UE (106) of the plurality of UEs, wherein the connection decisions are based on the connection preferences for the plurality of UEs and the activation decisions for each AP of the set of APs (104).
20. The gNB-CU (502) of any of claims 11 to 19, wherein the first parameter related to the likelihood of channels between each AP and the plurality of UEs having the threshold quality for the predefined time period is based on a second parameter related to a direction of movement of the plurality of UEs and a third parameter related to a history of large scale fading between the plurality of UEs and the set of APs (104).
21. A method implemented in a User Equipment device, UE, (106) for enabling scalable mobility-aware energy efficient management of a cell-free network, the method comprising: based on network information received from a gNB Central Unit, gNB-CU, (502) generating (608), as an output of a first deep neural network, DNN, connection preferences for the UE (106), wherein the output of the first DNN is based on: large scale fading between the UE (106) and a set of Access Points, APs, (104) of the cell-free network; a number of UEs served by each AP of the set of APs (104); previous connection decisions of the UE (106); and a first parameter related to likelihood of channels between each AP and the UE (106) having a threshold quality for a predefined time period; and providing (612) the connection preferences to the gNB-CU (502).
22. The method of claim 21, further comprising: receiving (616), from the gNB-CU (502), a connection decision associated with handovers to one or more APs (104), wherein the connection decision is based on the connection preferences for the UE (106) and activation decisions for each AP of the set of APs (104).
23. A User Equipment device, UE, (106) configured for enabling scalable mobility- aware energy efficient management of a cell-free network, the UE (106) comprising a radio interface and processing circuitry configured to: based on network information received from a gNB Central Unit, gNB-CU, (502), generate (608), as an output of a first deep neural network, DNN, connection preferences for the UE (106), wherein the output of the first DNN is based on: large scale fading between the UE (106) and a set of Access Points, APs, (104) of the cell-free network; a number of UEs served by each AP of the set of APs (104); previous connection decisions of the UE (106); and a first parameter related to likelihood of channels between each AP and the UE (106) having a threshold quality for a predefined time period; and provide (612) the connection preferences to the gNB-CU (502).
24. The UE (106) of claim 23, wherein the processing circuitry is further configured to: receive (616), from the gNB-CU (502), a connection decision associated with handovers to one or more APs (104), wherein the connection decision is based on the connection preferences of the UE (106) and activation decisions for each AP of the set of APs (104).