Scalable, mobility-aware energy efficiency management for cell-less networks

CN122603548APending Publication Date: 2026-08-18TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480084918.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-14
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]尽管该主题很重要,但常规网络和UC-mMIMO网络两者中AP的激活决策已经基于网络的静态快照,而忽略了移动性的影响

Benefits of technology

[0020] Some advantages of the methods and techniques disclosed herein are that the system is mobility-aware. Unlike solutions for activating access points (APs), the solution disclosed herein considers user mobility. In the considered scenario, it is assumed that users are moving, and therefore a handover (HO) should be performed to update the service set for each user. By using information from the network, a predictive HO scheme for collecting HOs within a single time step can be provided. In this way, the service set for each user can remain unchanged for a longer period compared to conventional schemes, thus allowing the AP to maintain its active/deactivated mode for a longer period.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122603548A_ABST
    Figure CN122603548A_ABST
Patent Text Reader

Abstract

Various embodiments disclosed herein provide methods for implementing scalable, mobility-aware solutions to improve the energy efficiency of cell-free networks in wireless communication systems. Deep reinforcement learning (DRL) is used to train two deep neural networks (DNNs) to act as policies for user handover (HO) decisions and access point (AP) activation decisions in user-centric cell-free massive multiple-input multiple-output (UC-mMIMO). The first DNN and the second DNN can be trained at a gNB central unit (gNB-CU). The first DNN can be implemented at the gNB-CU or one or more user equipment (UEs) to determine connection preferences, while the second DNN is implemented at the gNB-CU to determine activation decisions for APs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods for implementing scalable, mobility-aware solutions to improve the energy efficiency of cellless networks in wireless communication systems. Background Technology

[0002] Figure 1B The user-centric, cell-free, massive MIMO (UC-mMIMO) network illustrated comprises a dense deployment of access points (APs) (104-1 to 104-N) jointly serving a user (e.g., User Equipment (UE) 106). However, as network traffic load changes, the service of some APs may no longer be needed, thus allowing them to enter a power-saving sleep mode or simply be deactivated. In a UC-mMIMO network, each user is served by a set of neighboring APs, which eliminates cells on the access channel, thus addressing the poor performance issues of cell-edge users. That is, each user is surrounded by multiple APs that cooperate and jointly serve that user, with... Figure 1A Compared to the cell-centric architecture shown (where each AP 104 belongs to one or more cells 102), this results in enhanced desired signal strength, interference mitigation, macro diversity, no cell edge, and shorter service distance.

[0003] As a user moves, the set of APs serving each user is updated via handover (HO) operations, where an HO operation represents a change in the user's connection decision. Clearly, HO can influence the activation and deactivation decisions of APs. Therefore, it is natural to jointly employ these two decisions, represented by the user's connection decision and the AP's activation decision.

[0004] While this topic is important, AP activation decisions in both conventional and UC-mMIMO networks have been based on static snapshots of the network, neglecting the impact of mobility. Therefore, no effective solution has yet been proposed to jointly determine user connectivity decisions and AP activation decisions. Crucially, mobility is even more prominent in UC-mMIMO schemes because each user is served by a set of APs, rather than by a single base station (BS) as in conventional networks. Summary of the Invention

[0005] The various embodiments disclosed herein provide methods for implementing scalable, mobility-aware solutions to improve the energy efficiency of cellless networks in wireless communication systems. Deep reinforcement learning (DRL) is used to train two deep neural networks (DNNs) to act as strategies for user handover (HO) decisions and access point (AP) activation decisions in user-centric, cellless massive multiple-input multiple-output (UC-mMIMO). The first and second DNNs can be trained at the gNB central unit (gnB-CU). The first DNN can be implemented at the gNB-CU or one or more user equipments (UEs) to determine connectivity preferences, while the second DNN is implemented at the gNB-CU to determine AP activation decisions.

[0006] In an embodiment, the method can be implemented in a gNB-CU for scalable, mobility-aware, and energy-efficient management of a cellless network. The method may include: receiving connection preferences of a user equipment (UE) as the output of a first deep neural network (DNN), wherein the output of the first DNN is based on: large-scale fading between the UE and a set of access points (APs) in the cellless network; the number of UEs served by each AP in the AP set; the UE's previous connection decisions; and a first parameter related to the probability that the channel between each AP and the UE has threshold quality within a predefined time period. The method may further include providing connection preferences of a plurality of UEs as input to a second DNN, wherein the connection preferences of the plurality of UEs are summed in a voting system, and the output of the second DNN results in an activation decision for each AP in the AP set. The method may further include providing activation decisions for each AP in the AP (104) set to each AP in the AP set based on previous decisions or time slots.

[0007] In one embodiment, the method includes training a first DNN using a first reward function based on an achievable rate according to the signal-to-noise ratio at the UE, wherein time loss due to performing a switch from a first AP to a second AP is penalized.

[0008] In one embodiment, the method includes training a second DNN using a second reward function associated with energy efficiency, the second reward function being based on a function of sum of achievable data rates divided by the power consumption of the active AP, and wherein the second reward function is based on the activation and deactivation of the AP.

[0009] In one embodiment, the first deep neural network is implemented in a first reinforcement learning environment, and the second DNN is implemented in a second reinforcement learning environment.

[0010] In this embodiment, the first reinforcement learning environment is implemented at the UE, and the second reinforcement learning environment is implemented at the gNB-CU.

[0011] In one embodiment, the method further includes providing a first DNN to the UE for implementation in a first reinforcement learning environment.

[0012] In this embodiment, receiving connection preferences includes receiving connection preferences from the UE.

[0013] In this embodiment, the first reinforcement learning environment and the second reinforcement learning environment are implemented at the gNB-CU.

[0014] In one embodiment, the method includes providing each of the plurality of UEs with a connection decision, wherein the connection decision is based on the connection preferences of the plurality of UEs and the activation decision of each AP in the AP set.

[0015] In this embodiment, the first parameter related to the probability that the channel between each AP and the plurality of UEs has a threshold quality within a predefined time period is based on a second parameter related to the movement direction of the plurality of UEs and a third parameter related to the history of large-scale fading between the plurality of UEs and the AP set.

[0016] In an embodiment, the gNB-CU can be configured to achieve scalable, mobility-aware, and energy-efficient management of a cellless network. The gNB-CU includes processing circuitry configured to receive connection preferences from user equipment (UE) devices as the output of a first DNN. The output of the first DNN is based on: large-scale fading between the UE and a set of access points (APs) in the cellless network; the number of UEs served by each AP in the AP set; the UE's previous connection decisions; and a first parameter related to the probability that the channel between each AP and the UE has threshold quality within a predefined time period. The processing circuitry can also provide connection preferences from multiple UEs as input to a second DNN, wherein the connection preferences of the multiple UEs are summed in a voting system, and the output of the second DNN results in an activation decision for each AP in the AP set. The processing circuitry can also provide each AP in the AP set with its own activation decision.

[0017] In an embodiment, the method can be implemented in the UE for scalable, mobility-aware, and energy-efficient management of a cellless network. The method may include: generating the UE's connection preferences as the output of a first DNN based on network information received from the gNB-CU, wherein the output of the first DNN is based on: large-scale fading between the UE and a set of APs in the cellless network; the number of UEs served by each AP in the AP set; the UE's previous connection decisions; and a first parameter related to the probability that the channel between each AP and the UE has threshold quality within a predefined time period.

[0018] The method may include receiving from the gNB-CU a connection decision associated with a handover to one or more APs, wherein the connection decision is based on the UE’s connection preferences and the activation decision of each AP in the set of APs.

[0019] In an embodiment, the UE can be configured to achieve scalable, mobility-aware, and energy-efficient management of a cellless network. The UE may include a radio interface and processing circuitry configured to: generate the UE's connection preferences as the output of a first DNN based on network information received from the gNB-CU, wherein the output of the first DNN is based on: large-scale fading between the UE and a set of APs in the cellless network; the number of UEs served by each AP in the AP set; the UE's previous connection decisions; and a first parameter related to the probability that the channel between each AP and the UE has threshold quality within a predefined time period.

[0020] Some advantages of the methods and techniques disclosed herein are that the system is mobility-aware. Unlike solutions for activating access points (APs), the solution disclosed herein considers user mobility. In the considered scenario, it is assumed that users are moving, and therefore a handover (HO) should be performed to update the service set for each user. By using information from the network, a predictive HO scheme for collecting HOs within a single time step can be provided. In this way, the service set for each user can remain unchanged for a longer period compared to conventional schemes, thus allowing the AP to maintain its active / deactivated mode for a longer period.

[0021] Another advantage is the system's fast response time. Once two deep neural networks (DNNs) are trained using deep reinforcement learning (DRL), these trained DNNs can be used as the HO policy and the AP activation policy, respectively. The use of DNNs provides fast feedback control, where the user's connection decisions and AP activation decisions can be obtained immediately once the required information (observation vectors) is provided to the DNN. This is a significant advantage compared to conventional mathematical optimization algorithms that employ iterative optimization methods.

[0022] Availability and Scalability: The same trained DNN can still be used when the number of users in the network changes. Using two separate DNNs offers two main benefits: the first benefit is that the solution presented in this paper is scalable in the sense that the complexity does not increase with the number of users, and the second benefit is that the same trained DNN can be used even when the number of users changes.

[0023] Another advantage is that the system can provide user prioritization. By using fair weights when adding user votes, users already experiencing low data rates can be given high priority, and this high priority is then used to activate APs that can help these users achieve good performance. This approach results in a fairer solution compared to schemes that prioritize users with good channels while ignoring those with weak channels. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate several aspects of this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0025] Figure 1A An example of a cell-centric network scheme according to embodiments of the present disclosure is shown;

[0026] Figure 1B An example of a user-centric cellless network scheme according to embodiments of the present disclosure is shown;

[0027] Figure 2 The hierarchical structure of a reinforcement learning environment proposed according to some embodiments of the present disclosure is shown;

[0028] Figure 3 The operation of two deep neural networks (DNNs) according to some embodiments of the present disclosure is illustrated;

[0029] Figure 4 Exemplary algorithms for implementing the two DNNs according to some embodiments of the present disclosure are shown;

[0030] Figure 5 A user-centric massively multi-input multiple-output (UC-mMIMO) network according to some embodiments of the present disclosure is illustrated;

[0031] Figure 6 A message sequence diagram for implementing scalable, mobility-aware, and energy-efficient management of cellless networks is shown according to some embodiments of the present disclosure.

[0032] Figure 7 An example of a cellular communication system according to some embodiments of the present disclosure is shown;

[0033] Figure 8 A schematic block diagram of a gNB central unit (CU) according to some embodiments of the present disclosure is shown;

[0034] Figure 9 Some embodiments according to this disclosure are shown. Figure 8 A schematic block diagram of a virtualization embodiment of gNB-CU;

[0035] Figure 10 Some other embodiments according to this disclosure are shown. Figure 8 A schematic block diagram of the gNB-CU;

[0036] Figure 11 These are schematic block diagrams of user equipment (UE) devices according to some embodiments of the present disclosure; and

[0037] Figure 12 It is based on some other embodiments of this disclosure Figure 11 A schematic block diagram of the UE. Detailed Implementation

[0038] The embodiments described below provide information for those skilled in the art to practice the embodiments and illustrate the best mode for practicing the embodiments. After reading the following description in conjunction with the accompanying drawings, those skilled in the art will understand the concepts of this disclosure and will recognize the applications of these concepts not specifically given herein. It should be understood that these concepts and applications fall within the scope of this disclosure.

[0039] Wireless communication device: One type of communication device is a wireless communication device, which can be any type of wireless device that can access a wireless network (e.g., a cellular network) (i.e., served by it). Some examples of wireless communication devices include, but are not limited to: User Equipment (UE) devices, Machine-Type Communication (MTC) devices, and Internet of Things (IoT) devices in 3GPP networks. Such wireless communication devices can be integrated into mobile phones, smartphones, sensor devices, meters, vehicles, home appliances, medical devices, media players, cameras, or any type of consumer electronic device (e.g., but not limited to: televisions, radios, lighting fixtures, tablet computers, laptop computers, or PCs). Communication devices can be portable, handheld, computer-integrated, or in-vehicle mobile devices capable of transmitting voice and / or data via a wireless connection.

[0040] Note that the descriptions presented herein focus on 3GPP cellular communication systems, and therefore frequently use 3GPP terminology or similar terms. However, the concepts disclosed herein are not limited to 3GPP systems.

[0041] Note that the term “cell” may be referenced in the description herein; however, in particular with respect to the 5G NR concept, the term “beam” may be used instead of “cell.” Therefore, it is important to note that the concepts described herein apply equally to both “cell” and “beam.”

[0042] Existing implementations present several challenges. Current solutions for maximizing the energy efficiency of wireless networks by optimizing access point (AP) activation decisions do not consider user mobility. Instead, these solutions are developed to address network snapshots, ignoring the impact of time variations within the network due to user movement. Furthermore, a framework that considers both user connection decisions (specifically, handover (HO) decisions) and AP activation is lacking. Another significant aspect is that most mathematical optimization frameworks do not account for network scalability. For example, most of these solutions exhibit computational complexity that increases with the number of users.

[0043] Even in solutions designed to improve energy efficiency by using trained neural networks, the neural networks employed may still scale proportionally to the number of users within the network. Furthermore, most of these solutions are developed for conventional networks, rather than the user-centric, cell-free, massive MIMO (UC-mMIMO) network schemes disclosed herein.

[0044] Therefore, a framework is needed to allow for determining both user connectivity decisions and access point activation decisions during user mobility, thereby maximizing energy efficiency. Most importantly, the developed solution should be able to scale relative to the number of users discovered in the network.

[0045] This disclosure provides solutions to these challenges, specifically a scalable solution for optimizing both user HO (Hosting Decision) and AP (Activation Decision) decisions in a user-centric, cell-free, massive MIMO (UC-mMIMO) network scheme. Deep reinforcement learning (DRL) can be used to train two deep neural networks (DNNs) to act as strategies for user HO and AP activation decisions.

[0046] The first DNN (denoted as DNN A) uses information about the large-scale fading (LSF) between each user and the AP, the number of users served by each AP, the user's previous connection decisions, and a metric characterizing the likelihood that each channel between the user and the AP will be in good condition in the future. Using this information, the DNN trained with DRL provides the user's connection preference for the AP. The reward function used in the DRL framework is based on the achievable rate of signal-to-noise ratio (SNR), taking into account the time overhead of performing HO (Hopping-Off) decisions. The time overhead is defined as a non-linear function that defines the cost of initiating HO. By adjusting this cost, HO decisions can be performed within the same time steps, providing users with more stable and longer connections to the access point and reducing frequent HO decisions. This also provides an advantage for the AP's activation decisions, where the AP can be deactivated for a longer period compared to when users perform frequent HO decisions in sparse time steps.

[0047] The second DNN (denoted as DNN B) collects each user's connection preferences and sums them as a voting system. This summation of connection preferences (now considered votes) is performed using a predefined weighted summation function, where the weights can be any metric that assigns priority to selected users. These weights are updated at each time step to change the priority. A typical metric is proportional fairness, where user weights are defined by an exponentially weighted window balancing the users' achievable data rates. Thus, users experiencing high data rates receive lower weights over time, while users experiencing low data rates receive higher weights. In this way, the voting system can give more weight to users with low data rates, making their votes more significant.

[0048] In the DRL framework, the reward function used to train the DNN B is Energy Efficiency (EE), which is a fractional function with the total user-achievable rate as the numerator and the AP's power consumption as the denominator. This power consumption model takes into account the transmit power and power consumption of the AP's RF circuitry. Additional penalties can be included in the reward function to penalize frequent AP activation / deactivation, allowing us to keep the AP in deactivated mode for longer periods than normal.

[0049] By using two separate DNNs that acquire their observations from the modeled reinforcement learning (RL) environment, a scalable framework independent of the number of users discovered in the network is provided.

[0050] Some advantages of the methods and techniques disclosed in this paper are that the system is mobility-aware. Unlike solutions for activating the AP, the solution disclosed in this paper takes into account user mobility. In the scenario under consideration, it is assumed that users are moving, and therefore, a HO (Hosting Operation) should be performed to update the service set for each user. By using information from the network, a predictive HO scheme for collecting HOs within a single time step can be provided. In this way, the service set for each user can remain unchanged for a longer period of time compared to conventional schemes, and therefore the AP can maintain its activated / deactivated mode for a longer period of time.

[0051] Another advantage is the system's fast response time. Once two DNNs are trained using Deep Reinforcement Learning (DRL), these trained DNNs can be used as the HO policy and the AP activation policy, respectively. The use of DNNs provides fast feedback control, where the user's connection decisions and AP activation decisions can be obtained immediately once the required information (observation vectors) is provided to the DNN. This is a significant advantage compared to conventional mathematical optimization algorithms that employ iterative optimization methods.

[0052] Availability and Scalability: The same trained DNN can still be used when the number of users in the network changes. Using two separate DNNs offers two main benefits: the first benefit is that the solution presented in this paper is scalable in the sense that the complexity does not increase with the number of users, and the second benefit is that the same trained DNN can be used even when the number of users changes.

[0053] Another advantage is that the system can provide user prioritization. By using fair weights when adding user votes, users already experiencing low data rates can be given high priority, and this high priority is then used to activate APs that can help these users achieve good performance. This approach results in a fairer solution compared to schemes that prioritize users with good channels while ignoring those with weak channels.

[0054] In this embodiment, a DNN framework was developed for managing connection decisions for mobile users and activation decisions for access points (APs). Because this solution manages changes in user connection decisions over time, it inherently manages handover (HO) decisions. Developing a solution that directly determines these control decisions through a single DNN presents several challenges. These challenges can be summarized as follows.

[0055] Training difficulty: DNNs that need to jointly determine user connection decisions and access point activation decisions have very high dimensionality. That is, it will be... ,in, It is the number of APs, and The number of users is a significant factor. Therefore, training such a DNN using deep reinforcement learning (DRL) would be extremely difficult and likely to fail to achieve good performance given the computational capabilities available in today's base stations. Furthermore, the DNN would be very large, consuming substantial disk space. In short, this solution lacks scalability.

[0056] Reusability: Any direct solution will depend on the number of users discovered in the network. Therefore, when the number of users changes, a new DNN with different input and output dimensions is reconstructed and then retrained. Thus, this solution is not reusable.

[0057] To overcome these difficulties, this disclosure provides a voting system consisting of two DNNs. The first DNN, named DNN A, is used to independently determine each user's connection preferences. The second DNN, named DNN B, is used to determine the APs that will be active and those that will be deactivated. After obtaining these decisions, a scheme can be designed to determine the user's connection decisions (or HO connections). This allows the DNN architecture to be a solution that is independent of the number of users in the network in terms of both dimensionality and complexity, resulting in a scalable solution.

[0058] Figure 2 The hierarchical structure of a proposed reinforcement learning environment according to some embodiments of the present disclosure is shown.

[0059] The hierarchical structure of reinforcement learning (RL) environments consists of two types, represented as 1) the main RL environment 202, and 2) subset RL environments 204-1 to 204-N, as follows: Figure 2 As shown.

[0060] The main RL environment 202 hosts all the network information required to execute the control decisions of interest in this study. Most importantly, these parameters constitute the observation vectors that will serve as inputs to the DNN B. It also includes the network information to be passed to the subset RL environment. To allow communication between the main and subset RL environments, both environments are equipped with inter-RL communication functions that constitute their application programming interfaces (APIs). The API implemented by the main RL consists of the following functions:

[0061] · This function provides the necessary information for a subset of the RL environment, identified by the user ID:user_id. The information returned by this function includes large-scale fading (LSF) statistics between the user and the AP, the number of users served by each AP, movement direction indicators (discussed later), and the history of LSF status indicators (discussed later). This information will be used by the subset of the RL environment as the observation vector to be provided to DNN A.

[0062] · This function will retrieve the user's connection preferences from a subset of the RL environment and treat them as the user's vote.

[0063] A subset RL environment will be generated independently for each user, so each user will have their own environment. It will obtain its input from the main RL environment using the provided APIs. The subset RL environment interacts with the DNN A by providing observations and rewards and then obtaining actions. The APIs implemented by the subset RL environment for mutual communication with the main RL environment include:

[0064] · This function sets the user's time step. The network information required to construct the observation vector.

[0065] Figure 2 This describes the hierarchy and the main communication flows between these environments. It also includes the `reset()` and `step()` functions, which are standard functions used by OpenAI's Gym / Gymnasium library, a standard API for implementing RL environments.

[0066] DNN A and a subset of RL environments for each user (here derived from...) (Index) interaction. When interacting with a subset of the RL environment. During interaction, the output of DNN A is a vector from the continuous action space. ,in, This is the number of access points (APs). This output represents the number of users. Connection preferences with APs discovered in the network. The higher the value, the better for users. For connection to AP The higher the preference, the more likely the action will be. This can be seen as a vote by users on connecting to the AP.

[0067] Used from DNN A and AP activation decisions obtained from DNN B By using users Connect to the one with the highest connection preference. Each active AP determines the connection decision (represented as...). In mathematics, the user and AP Each connection decision between them can be defined as:

[0068] Equation 1

[0069] in, It is element-wise multiplication, and Select vector The largest The index of each value. The first condition in Equation 1 means: if AP is active and it ranks among the top in user votes. In each AP, then at time step At the user This AP Provide services. Otherwise, the user will not be served by the AP.

[0070] The input to DNN A is the observation vector. It is defined as follows:

[0071] Equation 2

[0072] Among them, in time step hour, User and AP LSF between It is by AP The number of users served, and and Items are used to describe users and AP Two measures of the probability that the LSF is in a good state are given below. User and AP In the previous decision step Connection decisions at the location.

[0073] Log operations on LSF statistics are used to limit the dynamic range of LSF. Operators Apply min-max normalization, then shift and scale the input to make it between -1 and 1; therefore, observe the vector. Any entry within the scope This operation can be applied to vectors. to The element, which leads to a scaled version It is logical to apply this operator independently for each different piece of information, as shown in Equation 2. Surveys indicate that this makes configuring the hyperparameters of the DRL algorithm used to train DNNs easier, which helps to make DNN training smoother and faster.

[0074] Using the observation vector and (in, The reason for prioritizing APs (Access Points) is to prioritize APs that are more likely to have a good channel with the user. This encourages predictive connection preferences from users, reducing unnecessary connection requests (HOs). These two parameters... and Both methods use different network information to predict the quality of future channels and allow for robust predictive HO by using two different metrics.

[0075] First parameter Using the user's movement direction, the AP that the user is moving towards is given higher priority, and it is in time step The location is defined as:

[0076] Equation 3

[0077] in, User The direction of movement and the connection of the user to the AP The angle between the directions of the lines. This angle can be calculated using the locations of the user and the access point (AP). In a 5G network, since the location of the AP is known, this information can be accessed, and the user's location can be accurately tracked using the Global Positioning System (GPS), Received Signal Strength Indicator (RSSI), or other network identifiers.

[0078] Second parameter User and AP The probability that an LSF is in a good state is predicted by analyzing the history of its states over time. If the LSF is greater than a certain selected threshold... If so, LSF is defined as being in a good state. The intuitive implication is that if the LSF has been in a good state in the most recent few time steps, it is more likely to be in a good state in the next time step. Specifically, time steps... place Defined as:

[0079] Equation 4

[0080] in, It is the threshold at which LSF is classified as a good condition. It is a discount factor for the channel classification history, and It is an indicator function, where if the condition If it is effective, then Otherwise, it equals 0. The numerator of equation 4 contains... The most recent history should be given priority.

[0081] In Equation 3 In Equation 4 The intuitive meanings are that they provide the DRL algorithm with information about the possible future states of the LSF, which ultimately allows for predictive connection preferences and thus predictive HO.

[0082] Define a subset of RL environments separately for each user. This is done by using users... Connection decision and observation vector The HO penalty-reward function can be defined as:

[0083] Equation 5

[0084] in, This is the achievable rate based on the signal-to-noise ratio (SNR) version derived from the actual achievable rate below. (Item) This is the fractional time remaining for data transmission after deducting the weighted time penalty for HO; this penalty parameter is defined as:

[0085] Equation 6

[0086] in, It is the length of the communication cycle, including the pilot training and data transmission phases. It is the number of communication cycles contained in a single time step. In time step The weighted overhead of the HO is executed at the beginning. It uses the following equation through a nonlinear function Define:

[0087] Equation 7

[0088] Equation 8

[0089] Among them, item This represents the basic overhead of initiating a HO; it indicates the weighted time spent on Radio Resource Control (RRC) reconfiguration, handshake, reporting, etc. (Item) It is the weighted cost generated by each HO, and It is each time step The number of HOs. The minimal operator found in Equation 7 indicates only any two decision steps. and The duration between decisions should be chosen to be greater than the maximum duration of any HO decision. The intuitive meaning of the nonlinear function in Equation 7 is that it allows for flexible time overhead (which is a function of the number of HOs) and allows for penalties for initiating HOs.

[0090] Realizable rate used in Equation 5 It provides estimates of both the interference experienced by a user from non-serving APs and the power allocated to the user based on the number of users served by each AP. This approximation enables the implementation of a HO solution individually for each user. The feasibility is defined as:

[0091] Equation 9

[0092] in,

[0093] Equation 10

[0094] Equation 11

[0095] Equation 12

[0096] The variance of the estimated channel can be written as:

[0097] Equation 13

[0098] And users transmit power normalization Selected as:

[0099]

[0100] The output of DNN B is a continuously output vector. It maps to the activation / deactivation decision of AP. The activation decision is determined by a vector. It means that, among them, Theoretically, during the training of a DNN B, any threshold can be chosen to determine... However, using a threshold of 0 is the most logical approach because it is... The decisive boundary between negative and positive values.

[0101] The input to a DNN B is determined by the observation vector:

[0102] Equation 14

[0103] in, It is a weighted sum of user votes (connection preferences) and is calculated as follows:

[0104] Equation 15

[0105] in, and User At time step Priority weights and join preferences at each location. As mentioned earlier, operators... Perform min-max normalization on the input vector, thereby shifting the input so that it is between Between 1 and 2.

[0106] The intuitive meaning is to collect user votes while applying weights. This prioritizes the contributions of individual users. There are many ways to define these weights; one way is by using proportional fairness, as follows:

[0107] Equation 16

[0108] in:

[0109] Equation 17

[0110] in, It is the user's achievable spectral efficiency. User The average long-term data rate within the previous decision steps (where, ),and It is a forgetting factor.

[0111] For the reward used to train the DNN B, an energy efficiency (EE) metric (measured in bits per joule) is used, and at time step... Being in the downlink is defined as:

[0112] Equation 18

[0113] in, It is the system bandwidth. This is the user's downlink achievable rate (measured in bits per second per Hz), which can be defined in detail in the system model, and This is the total power consumption of the AP on the downlink. This power consumption model is measured in watts or joules.

[0114] In dense networks where communication occurs over short distances, transmit power becomes comparable to circuit power consumption. Therefore, the power consumption of the AP's circuitry must be considered in any study optimizing the electrical efficiency (EE).

[0115] We use a general downlink power consumption model, which is written as:

[0116] Equation 19

[0117] in, It is AP At time step Activation decision at the location, among which Instructing the activity AP and This indicates an inactive access point (AP). , , and It is AP The average power consumed during operation is due to the power amplifier (PA), load-related circuitry, transceiver chain, and basic operations. The last three terms of power consumption represent the circuit's power consumption, where... It's load-related, and and The two are unrelated to load.

[0118] The average power consumption of a PA includes the sum of the radiated power required to transmit user data and the dissipated power, and it can be expressed by the following equation:

[0119] Equation 20

[0120] in, This is the average efficiency of the PA, with a typical value of ,and This is the power budget for each AP. The formula in Equation 20 simply states that the total power consumed will be greater than the power used to transmit data and pilot signals, because some power dissipation is due to imperfect PAs.

[0121] item This represents the power consumption of the transceiver chain, and it is defined as:

[0122] Equation 21

[0123] in, It refers to the number of antennas in the AP. It is the power required to operate the circuitry components attached to each antenna, and This is the power consumed by the local oscillator.

[0124] The load-related power consumed by the circuit can be approximated as:

[0125] Equation 22

[0126] in, This represents the channel estimation process. Represents the channel coding and decoding unit. This indicates a load-dependent fronthaul, and This indicates the transmit and receive beamforming at the AP.

[0127] The power consumption for channel estimation can be approximated as:

[0128] Equation 23

[0129] in, It is the length of the regular coherent block of the channel. It is the length of the uplink pilot training phase, and This refers to the AP's computational efficiency, measured in flops / Watt (operations per joule). The first fractional term in Equation 23 represents the number of coherent blocks per second, where the AP performs a single pilot-based CSI estimate for each block. In the uplink, the AP receives pilot signals as... The matrix is ​​required to estimate the value of each user's pilot sequence by multiplying it by the corresponding pilot sequence for each user. The channel.

[0130] The power consumption of the channel coding and decoding units is proportional to the number of bits, and therefore can be approximated as:

[0131] Equation 24

[0132] in, It is the fronthaul service power (measured in watts per (Gbit / s)).

[0133] item This represents the transmit and receive beamforming at the AP. For conjugate beamforming, it can be approximated as:

[0134] Equation 25

[0135] The first term describes the power consumed by one matrix multiplication for each data symbol, while the second term... This is a beamforming-related term, representing the calculations for the transmit and receive beamformers. To achieve the maximum transmit-to-receive ratio, Defined as:

[0136] Equation 26

[0137] This power cost is due to the normalization of the channel to construct the beamformer.

[0138] Figure 3 The operation of two deep neural networks (DNNs) according to some embodiments of the present disclosure is illustrated.

[0139] When training DNN B and DNN A separately, a main RL environment 202 and a subset RL environment 204 can be used respectively. This allows the RL environment to be implemented as a standard OpenAI Gym / Gymnasium library RL environment, which includes implementations of the reset() and step() methods, and eliminates mutual communication between the main RL 202 and the subset RL environment 204. For this purpose, standard implementations of DRL algorithms are used, which are customized to interact with the RL environment using these aforementioned methods.

[0140] By introducing subtle modifications during training, separate training of the main RL environment and the subset RL environment can be achieved. For the main RL environment, updateVotes is not used. Instead of using methods to obtain user votes, you can directly use the logs containing LSF statistics.

[0141] For subset RL environments, a single user network can be generated instead of using the `setNetInfo()` method (see Figure 1). Another simple modification is to obtain the user's connection decision from connection preferences. In this respect, each user can choose the network with the highest preference score. A series of connections are used to obtain a binary vector. This represents the AP that the user prefers to connect to (assuming all APs are active). Mathematically, this can be described as:

[0142] Equation 27

[0143] in, Select vector The one with the maximum value Indexes.

[0144] To illustrate this method, the Soft Actor-Critic (SAC) algorithm is used to implement DRL. The SAC algorithm is an offline policy learning technique, meaning it learns through a replay buffer. Training is performed using past experience. This buffer stores past experience from the RL environment and contains the following entries: [state, action, reward, next state, simulation termination indicator]. Then, to train the DNN, batches (sample sets) are extracted through random sampling without repetition. In contrast to on-policy learning, which involves collecting new samples with each policy update, off-policy learning offers better utilization of the data and provides a simpler explore-and-exploit policy for actions.

[0145] Once training is complete, the framework for communication between these two types of RL environments, as described below, is used.

[0146] Figure 4 Exemplary algorithms for implementing the two DNNs according to some embodiments of this disclosure are shown.

[0147] Algorithm 1 details the operation of the scalable and reusable DNN framework and RL environment architecture. The algorithm takes trained DNNs (i.e., DNN A and DNN B) as input and outputs the activation decisions of the AP and the user's connection decisions.

[0148] In steps 3 and 4, RL environments that will store network information are created. The main RL environment is modeled to provide information on the main parameters and evolution of the network, including all users, while each subset of RL environments... Modeled to save its users This information and its evolution. Steps 5 and 6 initialize the parameters of the RL environment, preparing it to interact with the DNN.

[0149] Next, the simulation process (episode) begins by providing the information required for operation in each subset of the RL environment (steps 10-12). In step 14, user votes are collected and provided to the main RL environment. In step 15, the votes are processed and summed. In step 16, observations of the main RL environment are constructed. In step 17, actions of the main RL environment are obtained using a trained DNN B. In step 18, activation decisions for the AP are determined based on the actions of the DNN B.

[0150] Then, as shown in step 19, the activation decision of AP and the connection decision obtained from the trained DNN A are used to determine the connection decision for each user. Then, in steps 20 and 21, the subset RL environment and the main RL environment execute these actions and obtain new rewards. Due to the complex RL architecture, observations are not directly obtained after executing the actions. This differs from a standard RL environment. Finally, in step 22, the fairness weights for the users are updated, and the next time step can now proceed, as shown in step 23.

[0151] Figure 5 A user-centric massively multi-input multiple-output (UC-mMIMO) network according to some embodiments of this disclosure is illustrated.

[0152] We consider including through sets Indicated Each AP is connected to the gNB-CU 502 network via a wired fronthaul solution, such as... Figure 5 As shown. Each AP is equipped with There are 1 antenna, and according to the UC-mMIMO scheme, it is a set of 1 antennas. Services are provided to users within the platform. Specifically, at any given time... At this location, each user From set The represented set of access points (APs) provides services using coherent transmission modes. Each user's service set is selected from neighboring APs, independent of the cell. Therefore, in the UC-mMIMO scheme, the cell is independent of the access channel. Consequently, users will experience comparable performance, and the concept of cell-edge users is eliminated.

[0153] Due to the importance of mobility considerations, it is assumed that communication is divided into several time steps, where, at each time step, user movement is determined by the user... Speed ​​and any and Fixed duration between A specific distance is defined. As the user moves, updates are needed by adding and removing access points (APs) via homepages (HOs). Therefore, if the user's connection decision changes, i.e. Then at the beginning of each decision cycle HO occurs at this time. It needs to be handled by AP. Users of the service are accessed through a collection These sets are represented. You can directly from the set Obtain, that is .

[0154] At each time step Furthermore, there are options for activating and deactivating APs, aimed at optimizing energy consumption within the network. Correct activation decisions must consider user connectivity preferences (i.e., homeostasis, HO). The goal is to develop a DNN framework that obtains user connectivity decisions and AP activation decisions. This framework should consider mobility awareness and allow minimizing HO decisions and AP transitions between active and inactive modes. In this regard, predictive algorithms are invaluable in optimizing these critical decisions.

[0155] User mobility causes time variations in the wireless channel. Therefore, block fading channel models with specific coherence times are inaccurate when users are moving; instead, time-varying channel models are needed. This is especially important for high-speed users.

[0156] We consider including Channel usage (including length) Uplink pilot training phase and length The communication cycle (the downlink data transmission phase) is defined as follows: Within this communication cycle, the small-scale fading portion of the channel is considered a generally stationary (WSS) process, while the large-scale fading (LSF) is considered constant because it is related to the duration... In contrast, it changes at a slower rate, even in highly mobile scenarios.

[0157] At the time of channel use AP b and users The channel implementation between them is modeled as ,in, It was a small-scale decline, and It is a large-scale fading that takes into account shadows and path loss.

[0158] This channel is time-dependent, which means that for any channel used and , and They are all related. Regarding , and The time correlation of the channels can be determined using Jakes' model through the time correlation coefficient. To characterize this, the time correlation coefficient is defined as...

[0159] Equation 28

[0160] in, It is a zeroth-order Bessel function of the first type. It is the sampling period used for each channel. It is the largest Doppler shift, which depends on the user's mobility speed. carrier frequency and the speed of light .

[0161] Based on Equation 28, for length Each communication cycle, time The small-scale decay at a certain point can be written as the initial state. and innovation weight The function is shown below:

[0162] Equation 29

[0163] in, It is a moment The innovative component is significant, and it is unrelated to small-scale declines. User At time 0 and The time correlation coefficient between the channel implementations at the location, where, and .

[0164] To construct a correct signal model, the time at which it is estimated (which is represented as time) is used. The channel at point ) is correlated with other channels at the time of use.

[0165] For pilot training sequences, an increment (delta) function is used, where the increment is defined relative to the pilot signal. The corresponding set of users Therefore, users Using pilot signals ,in, and ,in, This is the length of the pilot transmission phase. Here, for , And for , During the uplink pilot training phase place (i.e.) ), in AP The signal received at point can be written as:

[0166] Equation 30

[0167] in, It is the uplink transmit power of each user, and .

[0168] Channel estimation is performed after all training signals have been received (i.e., after...). (At) this point. Based on Equation 29, at time (…). Place, Can be with

[0169] Related, as shown below:

[0170] Equation 31

[0171] Through the channel in equation (30) In the small-scale fading component, equation (31) is used, and it is assumed that AP Estimating users The channel, during the uplink pilot training phase. At AP The signal received in equation (30) can be rewritten as:

[0172] Equation 32

[0173] By using the linear minimum mean square error (MMSE), AP Estimated users The channels are as follows:

[0174] Equation 33

[0175] When using MMSE to estimate the channel, the channel estimation error With the estimated channel Unrelated, and distributed as Wherein, the covariance is , Defined as:

[0176] Equation 34

[0177] The AP utilizes conjugate beamforming (maximum ratio transmission) to provide services to users, whereby the AP... Use at time The estimated channel is used to construct a beamformer for transmitting data to users. Based on this, at time... By AP To users The signal sent It can be written as:

[0178] Equation 35

[0179] in, Indicates at time The conjugate of the estimated channel, It is the user's complex data symbol, where, ,and It is by AP Assigned to user The statistical normalization term for the transmit power. This term allows the APs to be averaged to meet the available power budget. ,Right now .

[0180] Channel aging also affects the data transmission phase. By applying Equation 29 to the channel evolution within the communication cycle, at time... Small-scale fading at a certain point can also be represented as

[0181] Equation 36

[0182] By employing Equation 36, during the downlink data transmission phase (i.e., in...) (at the user's location) The received signal can be written as:

[0183] Equation 37

[0184] in, The first term in Equation 37 is the desired signal (DS), the second term is the beamformer uncertainty (BU), the third term is the channel aging (CA), the fourth term is the multi-user interference (MI), and finally, the fifth term represents noise.

[0185] By using the signal model represented in Equation 37 and treating uncorrelated interference signals as Gaussian noise, the rate performance can be characterized by iterating through the lower bound of the achievable rate, as follows:

[0186] Equation 38

[0187] The power terms for the desired signal, beamformer uncertainty + channel aging, and interference can be written in closed-form as follows:

[0188] Equation 39

[0189] Equation 40

[0190] Equation 41

[0191] here, Depends on AP Activation and users The connection decision, and can be formally defined as:

[0192] Equation 42

[0193] In addition, regarding the statistical power allocation for user usage, the power allocated by the serving AP is... To users The power control factor of the transmitted signal is

[0194]

[0195] The trained DNN B can be implemented on the central unit (CU), while the trained DNN A can be implemented at the user or at the CU. To implement DNN A at the user, a single trained DNN A is copied to the user device. Regardless of whether DNN A is implemented on the user side or at the CU, it is trained only once, so both implementations are feasible.

[0196] Figure 6 A message sequence diagram is shown for implementing scalable, mobility-aware, and energy-efficient management of cellless networks according to some embodiments of the present disclosure.

[0197] It should be understood that, Figure 6 In the message sequence diagram, dashed lines and dashed boxes are optional, while solid lines and solid boxes represent steps that are non-optional for one or more aspects of this disclosure.

[0198] The method may begin at step 602, wherein the method includes training a first DNN. In an embodiment, the first DNN (or DNN A, as used herein) may be trained using a first reward function based on an achievable rate according to the signal-to-noise ratio at UE 106, wherein time loss due to performing a switch from the first AP to the second AP is penalized.

[0199] At step 604, the method includes training a second DNN (or DNN B, as used herein) using a second reward function associated with energy efficiency, the second reward function being based on a function of the sum of achievable data rates divided by the power consumption of the active AP 104, and wherein the second reward function is based on the activation and deactivation of the AP. In embodiments, steps 602 and 604 may be performed once or at fixed intervals to update the DNN model.

[0200] In step 606, the method includes optionally providing a first DNN to the UE 106, such that the UE 106 can optionally implement the first DNN. Alternatively, the first DNN can be implemented at gNB-CU 502 in step 610.

[0201] When implementing the first DNN, at 608 or 610, the implementation may include receiving the connection preference of UE 106 as the output of the first DNN, the output of which is based on: large-scale fading between UE (106) and a set of access points (APs) (104) of the cellless network; the number of UEs served by each AP in the set of APs; the previous connection decisions of UE 106; and a first parameter related to the probability that the channel between each AP and UE (106) has a threshold quality within a predefined time period.

[0202] In the embodiment, the first parameter related to the probability that the channel between each AP and multiple UEs has a threshold quality within a predefined time period is based on the second parameter related to the movement direction of the multiple UEs and the third parameter related to the history of large-scale fading between the multiple UEs and the AP set.

[0203] If the first DNN is implemented at UE 106, then at step 612, UE 106 can provide connection preferences to gNB-CU 502.

[0204] In step 614, the method includes implementing a second DNN, which includes providing connection preferences of a plurality of UEs as input to the second DNN, wherein the connection preferences of the plurality of UEs are summed in a voting system, and the output of the second DNN results in an activation decision for each AP in the AP set. Then, at step 618, an activation decision may be provided to AP 104.

[0205] In one embodiment, the first deep neural network is implemented in the first reinforcement learning environment 204, and the second DNN is implemented in the second reinforcement learning environment 202.

[0206] In an embodiment where the first DNN is implemented in gNB-CU 502 at step 610, a connection decision may optionally be provided to UE 106 at step 616.

[0207] Figure 7An example of a cellular communication system 700 capable of implementing embodiments of the present disclosure is shown. In the embodiments described herein, the cellular communication system 700 is a 5G system (5GS) including a next-generation RAN (NG-RAN) and a 5G core (5GC), or including an evolved universal terrestrial RAN (E-UTRAN) and an evolved packet core (EPC). In this example, the RAN includes: base stations 702-1 and 702-2, which include NR base stations (gNBs) in the 5GS; and optionally, a next-generation eNB (ng-eNB) (e.g., an LTE RAN node connected to the 5GC), which includes an eNB in ​​the EPS, thereby controlling the corresponding (macro)cells 704-1 and 704-2.

[0208] In the embodiments disclosed herein, base station 702 may be a gNB-CU (such as gNB-CU502) in a cellless network, while Figure 7 The low-power node 706 shown can be the AP 104 described herein.

[0209] Base station 702 and low-power node 706 provide services to wireless communication devices 712-1 to 712-5 in corresponding cells 704 and 708. Wireless communication devices 712-1 to 712-5 are generally referred to herein as wireless communication device 712, and are referred to individually as wireless communication device 712. In the following description, wireless communication device 712 is generally referred to as UE, but this disclosure is not limited thereto.

[0210] Figure 8 This is a schematic block diagram of a gNB-CU 502 according to some embodiments of the present disclosure. Optional features are indicated by dashed boxes. The gNB-CU 502 may be, for example, a base station 702 or 706, or a network node implementing all or part of the functions of the base station 702 or gNB described herein. As shown, the gNB-CU 502 includes a control system 802, which includes one or more processors 804 (e.g., a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), a memory 806, and a network interface 808. The one or more processors 804 are also referred to herein as processing circuitry.

[0211] The one or more processors 804 operate to provide one or more functions of the gNB-CU 502 described herein. In some embodiments, the functions are implemented in software, for example, stored in memory 806 and executed by the one or more processors 804.

[0212] Figure 9This is a schematic block diagram illustrating a virtualization embodiment of the gNB-CU 502 according to some embodiments of this disclosure. This discussion is equally applicable to other types of network nodes. Furthermore, other types of network nodes may have similar virtualization architectures. Similarly, optional features are indicated by dashed boxes.

[0213] As used herein, a “virtualized” gNB-CU 502 is an implementation of gNB-CU 502 in which at least a portion of the functionality of gNB-CU 502 is implemented as a virtual component (e.g., via a virtual machine executing on a physical processing node in the network). As illustrated, in this example, gNB-CU 502 may include a control system 802 as described above. gNB-CU 502 includes one or more processing nodes 900 coupled to or included as part of network 902. If a control system 802 is present, it is connected to the processing nodes 900 via network 902. Each processing node 900 includes one or more processors 904 (e.g., CPU, ASIC, FPGA, etc.), memory 906, and a network interface 908.

[0214] In this example, the functionality 910 of the gNB-CU 502 described herein is implemented at one or more processing nodes 900, or distributed in any desired manner across one or more processing nodes 900 and the control system 802. In some specific embodiments, some or all of the functionality 910 of the gNB-CU 502 described herein are implemented as virtual components executed by one or more virtual machines implemented in a virtual environment hosted by the processing node 900. As will be appreciated by those skilled in the art, additional signaling or communication between the processing node 900 and the control system 802 is used to perform at least some of the desired functionality 910.

[0215] In some embodiments, a computer program including instructions is provided that, when executed by at least one processor, causes the at least one processor to perform the functions of one or more nodes (e.g., processing node 900) implementing the functions of gNB-CU 502 or a virtual environment according to any embodiment described herein. In some embodiments, a carrier including the computer program product described above is provided. The carrier is one of an electronic signal, an optical signal, a radio signal, or a computer-readable storage medium (e.g., a non-transitory computer-readable medium such as a memory).

[0216] Figure 10This is a schematic block diagram of a gNB-CU 502 according to some other embodiments of the present disclosure. The gNB-CU 502 includes one or more modules, each implemented in software, including a first reinforcement learning environment 204 and a second reinforcement learning environment 202. The first reinforcement learning environment 204 and the second reinforcement learning environment 202 provide the functionality of the gNB-CU 502 described herein. This discussion also applies to... Figure 9 The processing node 900, wherein the first reinforcement learning environment 204 and the second reinforcement learning environment 202 may be implemented at one of the processing nodes 900 or distributed among multiple processing nodes 900 and / or distributed between the processing node 900 and the control system 802.

[0217] Figure 11 This is a schematic block diagram of a UE 106 according to some embodiments of the present disclosure. As shown, the UE 106 includes one or more processors 1102 (e.g., CPU, ASIC, FPGA, etc.), a memory 1104, and one or more transceivers 1106. Each transceiver 1106 includes one or more transmitters 1108 and one or more receivers 1110 coupled to one or more antennas 1112. The transceiver 1106 includes radio front-end circuitry connected to the antennas 1112, configured to modulate signals transmitted between the antennas 1112 and the processor 1102, as will be understood by those skilled in the art. The processor 1102 is also referred to herein as processing circuitry. The transceiver 1106 is also referred to herein as radio circuitry. In some embodiments, the functionality of the UE 106 described above may be implemented wholly or partially by software, for example, stored in the memory 1104 and executed by the processor 1102. Note that the UE 106 may include... Figure 11 Additional components not shown, such as one or more user interface components (e.g., input / output interfaces including displays, buttons, touchscreens, microphones, speakers, etc., and / or any other components that allow information to be input to and / or output from the UE 106), power supplies (e.g., batteries and associated power circuitry), etc.

[0218] In some embodiments, a computer program including instructions is provided that, when executed by at least one processor, causes the at least one processor to perform the functions of UE 106 according to any embodiment described herein. In some embodiments, a carrier including the computer program product described above is provided. The carrier is one of an electronic signal, an optical signal, a radio signal, or a computer-readable storage medium (e.g., a non-transitory computer-readable medium such as a memory).

[0219] Figure 12This is a schematic block diagram of a UE 106 according to some other embodiments of the present disclosure. UE 106 includes a first reinforcement learning environment 204 implemented in software. The first reinforcement learning environment 204 provides the functionality of the UE 106 described herein.

[0220] Any suitable steps, methods, features, functions, or benefits disclosed herein can be performed by one or more functional units or modules of one or more virtual devices. Each virtual device may include multiple such functional units. These functional units may be implemented by processing circuitry, which may include one or more microprocessors or microcontrollers and other digital hardware (which may include digital signal processors (DSPs), application-specific digital logic, etc.). The processing circuitry may be configured to execute program code stored in memory, which may include one or more types of memory, such as read-only memory (ROM), random access memory (RAM), cache memory, flash memory devices, optical storage devices, etc. The program code stored in the memory includes program instructions for executing one or more telecommunications and / or data communication protocols and instructions for executing one or more techniques described herein. In some implementations, the processing circuitry may be used to cause the various functional units to perform corresponding functions according to one or more embodiments of this disclosure.

[0221] While the processes in the accompanying drawings illustrate a particular sequence of operations performed in certain embodiments of this disclosure, it should be understood that such sequence is exemplary (e.g., alternative embodiments may perform operations in a different order, combine certain operations, overlap certain operations, etc.).

[0222] Those skilled in the art will recognize improvements and modifications to the embodiments of this disclosure. All such improvements and modifications are considered to fall within the scope of the concept disclosed herein.

Claims

1. A method implemented in the gNB central unit gNB-CU (502) for achieving scalable, mobility-aware, and energy-efficient management of cellless networks, the method comprising: The connection preferences of the user equipment device (UE) (106) (610, 612) are received as the output of a first deep neural network (DNN), wherein the output of the first DNN is based on: Massive fading between the UE (106) and the set of access points (APs) (104) of the cellless network; The number of UEs served by each AP in the set of APs (104); The previous connection decision of the UE (106); and A first parameter relating to the probability that the channel between each AP and the UE (106) has a threshold quality within a predefined time period; Provide (614) multiple UE connection preferences as input to a second DNN, wherein the connection preferences of the multiple UEs are summed in a voting system, and the output of the second DNN results in the activation decision for each AP in the AP (104) set; and Provide (618) an activation decision for each AP in the set of APs (104).

2. The method according to claim 1, further comprising: The first DNN is trained (602) using a first reward function, which is based on the achievable rate of the signal-to-noise ratio at the UE (106), wherein time loss due to performing a switch from the first AP to the second AP is penalized.

3. The method according to any one of claims 1 to 2, further comprising: The second DNN is trained (604) using a second reward function associated with energy efficiency, the energy efficiency being a function of the sum of achievable data rates divided by the power consumption of the active AP (104), and wherein the second reward function is based on the activation and deactivation of the AP (104).

4. The method according to any one of claims 1 to 3, wherein, The first DNN is implemented in the first reinforcement learning environment (204), and the second DNN is implemented in the second reinforcement learning environment (202).

5. The method according to claim 4, wherein, The first reinforcement learning environment (204) is implemented at the UE (106), and the second reinforcement learning environment (202) is implemented at the gNB-CU (502).

6. The method according to claim 5, further comprising: The first DNN is provided (606) to the UE (106) for implementation in the first reinforcement learning environment (204).

7. The method according to any one of claims 5 to 6, wherein, Receiving (610, 612) connection preferences includes receiving the connection preferences from the UE (106).

8. The method according to claim 4, wherein, The first reinforcement learning environment (204) and the second reinforcement learning environment (202) are implemented at the gNB-CU (502).

9. The method according to claim 8, further comprising: Provide (616) a connection decision for each of the plurality of UEs (106), wherein the connection decision is based on the connection preferences of the plurality of UEs and the activation decision of each AP in the set of APs (104).

10. The method according to any one of claims 1 to 9, wherein, The first parameter, which is related to the probability that the channel between each AP and the plurality of UEs has a threshold quality within the predefined time period, is based on the second parameter, which is related to the movement direction of the plurality of UEs, and the third parameter, which is related to the history of large-scale fading between the plurality of UEs and the AP (104) set.

11. A gNB central unit (gNB-CU) (502) configured to implement scalable, mobility-aware, and energy-efficient management of a cellless network, the gNB-CU (502) including processing circuitry configured to: The connection preferences of the user equipment device (UE) (106) (610, 612) are received as the output of the first deep neural network (DNN), wherein, The output of the first DNN is based on: Massive fading between the UE (106) and the set of access points (APs) (104) of the cellless network; The number of UEs served by each AP in the set of APs (104); The previous connection decision of the UE (106); as well as A first parameter relating to the probability that the channel between each AP and the UE (106) has a threshold quality within a predefined time period; Provide (614) multiple UE connection preferences as input to the second DNN, wherein the multiple UE connection preferences are summed in a voting system, and the output of the second DNN results in the activation decision of each AP in the AP (104) set; as well as Provide (618) an activation decision for each AP in the set of APs (104).

12. The gNB-CU (502) according to claim 11, wherein, The processing circuit is further configured to: The first DNN is trained (602) using a first reward function, which is based on the achievable rate of the signal-to-noise ratio at the UE (106), wherein time loss due to performing a switch from the first AP to the second AP is penalized.

13. gNB-CU (502) according to any one of claims 11 to 12, wherein, The processing circuit is further configured to: The second DNN is trained (604) using a second reward function associated with energy efficiency, the energy efficiency being a function of the sum of achievable data rates divided by the power consumption of the active AP (104), and wherein the second reward function is based on the activation and deactivation of the AP (104).

14. gNB-CU (502) according to any one of claims 11 to 13, wherein, The first DNN is implemented in the first reinforcement learning environment (204), and the second DNN is implemented in the second reinforcement learning environment (202).

15. The gNB-CU (502) according to claim 14, wherein, The first reinforcement learning environment (204) is implemented at the UE (106), and the second reinforcement learning environment (202) is implemented at the gNB-CU (502).

16. The gNB-CU (502) according to claim 15, wherein, The processing circuit is further configured to: The first DNN is provided (606) to the UE (106) for implementation in the first reinforcement learning environment (204).

17. The gNB-CU (502) according to any one of claims 15 to 16, wherein receiving (610, 612) connection preference includes receiving the connection preference from the UE (106).

18. The gNB-CU (502) according to claim 14, wherein, The first reinforcement learning environment (204) and the second reinforcement learning environment (202) are implemented at the gNB-CU (502).

19. The gNB-CU according to claim 18, wherein, The processing circuit is further configured to: Provide (616) a connection decision for each of the plurality of UEs (106), wherein the connection decision is based on the connection preferences of the plurality of UEs and the activation decision of each AP in the set of APs (104).

20. gNB-CU (502) according to any one of claims 11 to 19, wherein, The first parameter, which is related to the probability that the channel between each AP and the plurality of UEs has a threshold quality within the predefined time period, is based on the second parameter, which is related to the movement direction of the plurality of UEs, and the third parameter, which is related to the history of large-scale fading between the plurality of UEs and the AP (104) set.

21. A method implemented in a user equipment (UE) (106) for implementing scalable, mobility-aware, energy-efficient cell-free network management, the method comprising: Based on the network information received from the gNB central unit gNB-CU (502), the connection preference of the UE (106) is generated (608) as the output of the first deep neural network DNN, wherein the output of the first DNN is based on: Massive fading between the UE (106) and the set of access points (APs) (104) of the cellless network; The number of UEs served by each AP in the set of APs (104); The previous connection decision of the UE (106); and A first parameter related to the probability that the channel between each AP and the UE (106) has a threshold quality within a predefined time period; and Provide the connection preference (612) to the gNB-CU (502).

22. The method of claim 21, further comprising: The gNB-CU (502) receives (616) a connection decision associated with a handover to one or more APs (104), wherein the connection decision is based on the connection preference of the UE (106) and the activation decision of each AP in the set of APs (104).

23. A user equipment (UE) apparatus (106) configured to implement scalable, mobility-aware, and energy-efficient management of a cellless network, the UE (106) including a radio interface and processing circuitry configured to: Based on the network information received from the gNB central unit gNB-CU (502), the connection preference of the UE (106) is generated (608) as the output of the first deep neural network DNN, wherein, The output of the first DNN is based on: Massive fading between the UE (106) and the set of access points (APs) (104) of the cellless network; The number of UEs served by each AP in the set of APs (104); The previous connection decision of the UE (106); as well as A first parameter related to the probability that the channel between each AP and the UE (106) has a threshold quality within a predefined time period; and Provide the connection preference (612) to the gNB-CU (502).

24. The UE (106) according to claim 23, wherein, The processing circuit is further configured to: The gNB-CU (502) receives (616) a connection decision associated with a handover to one or more APs (104), wherein the connection decision is based on the connection preference of the UE (106) and the activation decision of each AP in the set of APs (104).