Joint distributed learning of signaling and policies for radio resource allocation
Patent Information
- Application Number
- EP2022822092
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-09-17
AI Technical Summary
Current radio resource allocation methods in wireless networks face challenges in scalability and performance due to the need for either local decision-making without considering neighbor AP states or centralized approaches with high communication overhead and latency, and existing AI-based solutions suffer from stability and convergence issues.
A decentralized radio resource allocation system using a network device with a policy processor and training processor that exchanges messages to determine radio resource allocation configurations and updates models based on local and global rewards, leveraging deep neural networks and reinforcement learning to adapt to dynamic environments and consider neighbor AP decisions.
This approach enables efficient decentralized radio resource allocation with improved stability and performance by reducing communication overhead and latency, allowing APs to learn optimal resource allocation strategies while considering dynamic environments and neighbor behaviors.
Smart Images

Figure 1.1
Abstract
Description
[0001] JOINT DISTRIBUTED LEARNING OF SIGNALING AND POLICIES FOR RADIO RESOURCE ALLOCATION
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to allocating radio resources in a network, for example, a wireless network. The disclosure presents a joint distributed learning approach of signaling and policies for the radio resource allocation. The disclosure describes a network device and a method for the radio resource allocation and for the distributed learning.
[0004] BACKGROUND
[0005] Next generation wireless networks are expected to have a large number of terminal devices connected to network devices, for instance, access points (APs). In such a setting, radio resource allocation, which is a key aspect for wireless networks, will become a very challenging task. In practice, each AP needs to collect state information for its connected terminal devices, and then solve an optimization problem to decide which terminal device to schedule and / or how to allocate radio resources to the terminal devices, e.g., transmission power, subcarrier frequency, and beam selection. Interference among terminal devices that belong to nearby (neighbor) APs can have a significant impact on the overall performance, and should thus be considered when taking these decisions.
[0006] According to an approach that is used similarly in existing wireless networks, each AP can allocate radio resources by relying only on local information from its connected terminal devices. This is a scalable approach, since it does not require any communication among APs of the network. However, it does not consider states of neighbor APs, and can lead to poor performance. Alternatively, the problem could be handled centrally, for example, by a master single entity, which would need to collect the bulky information from all the APs of the network, then solve an even more challenging large-scale optimization problem, and then send back to each AP the respective radio resource allocation decision. This is not a practical approach, as it requires a huge communication overhead, and may also lead to a significant increase in latency. SUMMARY
[0007] This disclosure anticipates that communications will rely on distributed resource management, wherein APs exchange only the necessary information with a small amount of transmitted data. While this is a very attractive approach, it is also quite challenging and most proposed ideas are either impractical or suffer from stability and convergence issues that lead to poor system performance.
[0008] Distributed radio resource allocation has become more attractive with the recent success of artificial intelligence (Al) and specifically of deep neural networks (DNNs), which have already been used with success in very complex optimization problems. The need to have a decisionmaking framework for distributed resource management that adapts to the dynamic environment, naturally leads to solutions that combine the representation power of DNNs with reinforcement learning (RL). RL is an area of machine learning where agents (e.g., APs) learn an optimal behavior based on their trial-and-error interaction with the environment, in order to eventually maximize a long-term objective. The original RL methods suffer from the curse of dimensionality, which limits their application to models with a small number of states. This limitation can be tackled with the help of DNNs, which can be used to approximate functions of the RL algorithms.
[0009] Moreover, there is another challenge in RL-based distributed approaches for radio resource allocation that needs to be addressed. Besides the dynamics of the environment, each AP must also consider the behavior of other agents, for example, an AP must learn to predict the decisions of other APs and take strategic decisions in response. Convergence can be a problem in such a setting and thus, it is often assumed that there is centralized training, which however is impractical for the resource allocation problem. Another approach is to have independent distributed training on each AP, but this usually has stability issues and leads to unfair solutions, which is a poor performance indicator for wireless networks.
[0010] In view of the above, an objective of this disclosure is to provide a solution for an efficient decentralized radio resource allocation in a wireless network. Another objective is to provide a configuration of a network device that enables the network device to participate in the decentralized radio resource allocation. These and other objectives are achieved by the solutions described in the independent claims. Advantageous implementations are further defined in the dependent claims.
[0011] A first aspect of this disclosure provides a network device for radio resource allocation, the network device comprising a policy processor and a training processor, wherein the policy processor is configured to, in a current time slot: determine an initial local state of the network device, the initial local state indicating a current state of one or more terminal devices connected to the network device; generate a first message according to a first message generation model and based on the initial local state; cause the network device to send the first message to one or more neighboring network devices, and to obtain one or more first messages from the one or more neighboring network devices; and determine a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message and the one or more first messages of the one or more neighboring network devices; and allocate radio resource according to the radio resource allocation configuration to the one or more terminal devices connected to the network device; wherein the training processor is configured to, in the current time slot: generate a second message indicating a local reward obtained by allocating the radio resources according to the radio resource allocation configuration; cause the network device to send the second message to the one or more neighboring network devices, and to obtain one or more second messages from the one or more neighboring network devices, wherein each second message indicates a respective local reward of one of the network devices; calculate a global reward based on all the local rewards; generate a third message, indicating gradients for respectively updating the allocation policy model and the first message generation model, based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources to the one or more terminal devices connected to the network device, and global rewards, of one or more consecutive time slots before the current time slot; cause the network device to send the third message to one or more neighboring network devices, and to obtain one or more third messages from the one or more neighboring network devices; and update the first message generation model and the allocation policy model based on the third messages.
[0012] The network device of the first aspect is able to participate in decentralized radio resource allocation in a wireless network together with other network devices. The network devices of the wireless network may thus implement decentralized radio resource allocation in a wireless network. By exchanging the first messages, the second messages, and the third messages, and by calculating and taking into account the global and local rewards, an efficient joint distributed learning approach of signaling (using and training the first message generation model) and learning approach for policies for the radio resource allocation (using and training the allocation policy model) is achieved.
[0013] In an implementation form of the first aspect, the policy processor is configured to determine an initial local state, generate a first message, cause the network device to send the first message, compute a radio resource allocation configuration, and allocate radio resources repeatedly every time slot of a plurality of consecutive time slots; and the training processor is configured to generate a second message, cause the network device to send the second message, and calculate a global reward repeatedly every time slot of the plurality of consecutive time slots.
[0014] This may also be done at other network devices of the wireless network, and in this way the distributed learning is implemented.
[0015] In an implementation form of the first aspect, wherein the training processor is configured to generate a third message, cause the network device to send the third message, and update the first message generation model and the allocation policy model based on the third messages repeatedly every few time slots of the plurality of consecutive time slots.
[0016] This allows taking into account sufficient information about rewards obtained in the network, which is required to determine the gradients in the third messages that are exchanged with the one or more neighboring network devices, and thus leads to an improved training and radio resource allocation.
[0017] In an implementation form of the first aspect, the training processor is further configured to compute the local reward based on the changed local state as a result of allocating the radio resources according to the radio resource allocation configuration to the one or more terminal devices connected to the network device.
[0018] In an implementation form of the first aspect, the state of the one or more terminal devices comprises a queue length of each terminal device connected to the network device. In an implementation form of the first aspect, the training processor is configured to compute the global reward as a sum of all the local rewards.
[0019] In an implementation form of the first aspect, the training processor is configured to use a gradient algorithm to update the first message generation model and the allocation policy model based on the third messages each indicating gradients for respectively updating the allocation policy model and the first message generation model.
[0020] In an implementation form of the first aspect, the policy processor is configured to determine the initial local state based on an observation of one or more performance parameters of the one or more terminal devices connected to the network device.
[0021] In an implementation form of the first aspect, the policy processor is configured to input into the allocation policy model the initial local state and at least one of the generated first message and the one or more first messages of the one or more neighboring network devices, to receive the radio resource allocation configuration as an output of the allocation policy model.
[0022] In an implementation form of the first aspect, the policy processor is configured to aggregate the generated first message and the one or more first messages of the one or more neighboring network devices, and to compute the radio resource allocation configuration based on the aggregation result used as input into the allocation policy model.
[0023] In an implementation form of the first aspect, the radio resources allocated to the one or more terminal devices comprise at least one of transmission power, scheduling grant, and precoding and beamforming options.
[0024] In an implementation form of the first aspect, the policy processor is configured to: calculate directly the radio resource allocation configuration; or calculate a probability distribution over different radio resource allocation configurations and randomly select the radio resource allocation configuration based on the probability distribution.
[0025] In an implementation form of the first aspect, each first message comprises a vector of real numbers, wherein a dimension of the vector depends on a bandwidth that is allocated among the network device and the one or more neighboring network devices for exchanging the first messages.
[0026] In an implementation form of the first aspect, the policy processor is configured to generate the first message according to an initial first message generation model, if the time slot is an initial time slot.
[0027] In an implementation form of the first aspect, the policy processor and / or the training processor are each configured to run at least one neural network to perform their respective operations.
[0028] In an implementation form of the first aspect, at least one of the first message generation model, the allocation policy model and an aggregator configured to aggregate the generated first message and the one or more first messages of the one or more neighboring network devices, is implemented by a neural network.
[0029] A second aspect of this disclosure provides a method for radio resource allocation by a network device, the method comprising, in a current time slot: determining an initial local state of the network device, the initial local state indicating a current state of one or more terminal devices connected to the network device; generating a first message according to a first message generation model and based on the initial local state; causing the network device to send the first message to one or more neighboring network devices, and to obtain one or more first messages from the one or more neighboring network devices; and determining a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message and the one or more first messages of the one or more neighboring network devices; allocating radio resources according to the radio resource allocation configuration to the one or more terminal devices generating a second message indicating a local reward obtained by allocating the radio resources according to the radio resource allocation configuration; causing the network device to send the second message to the one or more neighboring network devices, and to obtain one or more second messages from the one or more neighboring network devices, wherein each second message indicates a respective local reward of one of the network devices; calculating a global reward based on all the local rewards; generating a third message, indicating gradients for respectively updating the allocation policy model and the first message generation model, based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources to the one or more terminal devices connected to the network device, and global rewards, of one or more consecutive time slots before the current time slot; causing the network device to send the third message to one or more neighboring network devices, and to obtain one or more third messages from the one or more neighboring network devices; and updating the first message generation model and the allocation policy model based on the third messages.
[0030] The method of the second aspect may be performed by the network device of the first aspect. The method of the second aspect may have implementation forms that correspond to the implementation forms of the network device of the first aspect. The method of the second aspect and its implementation forms achieve the same advantages as described above for the network device of the first aspect.
[0031] A third aspect of this disclosure provides a computer program comprising instructions which, when the program is executed by at least one processor, cause the at least one processor to perform the method of the second aspect or any implementation form thereof.
[0032] A fourth aspect of this disclosure provides a non-transitory storage medium storing executable program code which, when executed by a processor, causes the method according to the second aspect or any of its implementation forms to be performed.
[0033] In summary of the above aspects and implementation forms, this disclosure proposes a system for efficient decentralized radio resource allocation in a wireless network. The system may comprise a policy processor and a training processor at each participating network device (e.g., each AP or equivalently base station (BS)), as described for the network device of the first aspect.
[0034] At each time slot, the policy processor of each AP: may observe its local state (i.e., the state of the terminal devices associated to the AP, such as their channel states, delay experiences, and traffic information); may generate the first message and send it to the neighbor APs; may receive similar first messages from other APs; may output a radio resource allocation configuration and decision based on its current allocation policy model, according to its local state and the received first messages from its neighbor APs; and may observe a new local state as the outcome of the radio resource allocation decision (action, which allocates radio resources).
[0035] The training module of each AP: may exchange a second message with the other APs so that a global reward (or equivalently cost) can be calculated; may store in a replay buffer a tuple that includes action, local state transitions, and global reward; and, after a specified period of time where the APs have used their current policy models, may run a policy gradient algorithm, in order to update the policy module, wherein it may update the way to generate the first messages and it may update its own radio resource allocation policy model, by exchanging the third messages indicating the relevant gradients.
[0036] It has to be noted that all devices, elements, units and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof.
[0037] BRIEF DESCRIPTION OF DRAWINGS
[0038] The above described aspects and implementation forms will be explained in the following description of specific embodiments in relation to the enclosed drawings, in which
[0039] FIG. 1 shows a network device for radio resource allocation according to this disclosure.
[0040] FIG. 2 shows blocks of a network device according to this disclosure for decentralized resource allocation.
[0041] FIG. 3 shows a flowchart of a network device according to this disclosure for decentralized resource allocation.
[0042] FIG. 4 shows an exemplary implementation of a policy processor.
[0043] FIG. 5 shows an exemplary implementation of a policy processor without aggregators. FIG. 6 shows an exemplary implementation of a policy processor with different kinds of encoders.
[0044] FIG. 7 illustrates an exemplary policy module update when a policy gradient algorithm is used.
[0045] FIG. 8 illustrates an exemplary policy module update when policy Modules are updated via minimizing training loss functions.
[0046] FIG. 9 shows a global reward over training for the solution of this disclosure and a comparison against baselines.
[0047] FIG. 10 a method for radio resource allocation according to this disclosure.
[0048] DETAILED DESCRIPTION OF EMBODIMENTS
[0049] FIG. 1 shows a network device 100 according to this disclosure. The network device 100 is configured to perform radio resource allocation, and may participate in a joint distributed learning approach for the radio resource allocation with other network devices in a wireless network. FIG. 1 shows a neighboring network device 120, for instance, in the vicinity of the network device 100, which may also participate in the joint distributed learning approach for the radio resource allocation. There may be multiple such neighboring network devices 120 that participate. FIG. 1 also shows terminal devices 110, which are connected to the network device 100. Other terminal devices (not shown) may be connected to one or more neighboring network devices 120, respectively. The devices 100, 110, 120, which are shown in FIG. 1, may be of or in the same wireless network.
[0050] The network device 100 comprises a policy processor 101 and a training processor 102. Also the neighboring network devices 120 may each comprise such a policy processor 101 and training processor 102, wherein these processors 101, 102 may work likewise in all network devices 100, 120. The policy processor 101 may also be referred to as policy module, and the training processor 102 may also be referred to as training module in this disclosure. Both the policy processor 101 and the training processor 102 may be implemented by processing circuitry of the network device 100. Some actions of the policy processor 101 and the training processor 102 may be performed repeatedly at every time slot of a plurality of consecutive time slots.
[0051] In a current time slot, the policy processor 101 is configured to, determine an initial local state of the network device 100. The initial local state indicates a current state of the one or more terminal devices 110 connected to the network device 100. For instance, the state of the one or more terminal devices 110 may comprise a queue length of each terminal device 110 connected to the network device 100, or other one or more performance parameters of the terminal devices 110. The state of the terminal devices 110 may comprise channel states, delay experiences, and traffic information at these terminal devices 110. For example, the initial local state may be determined based on one or more performance parameters of the one or more terminal devices 110 connected to the network device 100.
[0052] The policy processor 101 is further configured to generate a first message 103 according to a first message generation model and based on the initial local state. The policy processor 101 is further configured to cause the network device 100 to send the first message 103 to one or more neighboring network devices 120, and to obtain one or more first messages 103 from the one or more neighboring network devices 120.
[0053] The policy processor 101 is further configured to determine a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message 103 and the one or more first messages 103 of the one or more neighboring network devices 120. The policy processor 101 may also aggregate the generated first message 103 and the one or more first messages 103 received from the one or more neighboring network devices 120, and then to determine the radio resource allocation configuration based on the aggregation result, which may be used as input into the allocation policy model.
[0054] The policy processor 101 is further configured to allocate radio resources 104 according to the radio resource allocation configuration to the one or more terminal devices 110 connected to the network device 100. The radio resources 104 allocated to the one or more terminal devices 110 may comprise at least one of transmission power, scheduling grant, and precoding and beamforming options.
[0055] In the current time slot, the training processor 102 is configured to generate a second message 105, which indicates a local reward. The training processor 102 is further configured to cause the network device 100 to send the second message 105 to the one or more neighboring network devices 120, and to obtain one or more second messages 105 from the one or more neighboring network devices 120. The local reward indicated by the second message generated by the network device 100 is obtained by allocating the radio resources 104 according to the radio resource allocation configuration. Each second message 105 obtained from the neighboring network devices 120 indicates a respective local reward of these neighboring network devices 120. The training processor 102 is further configured to calculate a global reward based on all the local rewards.
[0056] The training processor 102 is also configured to generate a third message 107, which indicates gradients for respectively updating the allocation policy model and the first message generation model. The third message 107 is generated by the training processor 102 based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources 104 to the one or more terminal devices 110 connected to the network device 100, and global rewards, of one or more consecutive time slots before the current time slot. The training processor 102 is further configured to cause the network device 100 to send the third message 107 to the one or more neighboring network devices 120, and to obtain one or more third messages 107 from the one or more neighboring network devices 120. Each third message 107 obtained from a respective neighboring network device 120 also indicates gradients for respectively updating the allocation policy model and the first message generation model, as determined by the respective neighboring network device 120.
[0057] The training processor 102 is then configured to update 106 the first message generation model and the allocation policy model - generally to update the policy processor 101 - based on the third messages 107 (the generated one and the obtained ones), in particular, based on the gradients indicated in the third messages 107. The sending and receiving of the third message(s) 107 to and from the one or more neighboring network devices 120, as well as the updating 106 of the first message generation model and the allocation policy model, may be repeated every few time slots of the plurality of consecutive time slots, while the other actions of the processors 101, 102 (i.e., determining the initial local state, sending and receiving the first message(s) 103, allocating the radio resources 104 according to a computed radio resource allocation, sending and receiving the second message(s) 105, and calculating the global reward) may be performed repeatedly every time slot of the plurality of consecutive time slots.
[0058] The network device 100 may comprise processing circuitry (not shown) configured to perform, conduct or initiate the various operations of network device 100 described in this disclosure, particularly to implement the policy processor 101 and the training processor 102, respectively. The processing circuitry may comprise hardware and / or the processing circuitry may be controlled by software. The hardware may comprise analog circuitry or digital circuitry, or both analog and digital circuitry. The digital circuitry may comprise components such as applicationspecific integrated circuits (ASICs), field-programmable arrays (FPGAs), digital signal processors (DSPs), or multi-purpose processors. The network device 100 may further comprise memory circuitry, which stores one or more instruction(s) that can be executed by the processor or by the processing circuitry, in particular under control of the software. For instance, the memory circuitry may comprise a non-transitory storage medium storing executable software code which, when executed by the processor or the processing circuitry, causes the various operations of the network device 100 to be performed. In one embodiment, the processing circuitry comprises one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code which, when executed by the one or more processors, causes the network device 100, particularly the policy processor 101 and the training processor 102, to perform, conduct or initiate the operations or methods described herein.
[0059] FIG. 2 shows a network device 100 - specifically an AP 100, as example for the network device 100 - according to this disclosure, which builds on the network device 100 shown in FIG. 1. In particular, FIG. 2 shows exemplary blocks of the AP 100. Also neighboring network devices 120 - specifically neighbor APs 120 - are shown. The blocks of the AP 100 are described in the following. An exemplary flowchart of proposed functionalities of the blocks is shown in FIG. 3.
[0060] The policy module (policy processor 101) receives the local observations of the AP 100 from its connected terminal devices 110, e.g., channel states and queue lengths, as inputs at the beginning of each time slot. Based on this local information, its function is mainly twofold: it firstly computes and broadcasts to the neighboring APs 120 a first message 103 (also referred to as intermediate message in FIG. 2), and secondly, based on the received first messages 103 from the neighbor APs 120 and the local state, takes an action, for this time slot.
[0061] The action can be the allocation of any radio resources 104, such as transmission power, device scheduling, precoding and / or beamforming if multiple antennas are used, allocation of resource blocks to terminal devices 110, etc. Taking the resource allocation action can be implemented by the policy processor 101 by: either outputting a probability distribution over different radio resource allocation configurations, if they are discrete (e.g., a probability distribution over which device to schedule or which power level to use) and then sample the configuration to use from this distribution; or specifying the action directly (e.g. which beamforming to use). For example, the policy processor 101 may directly calculate the radio resource allocation configuration; or calculate a probability distribution over different radio resource allocation configurations and randomly select the radio resource allocation configuration based on the probability distribution.
[0062] In all cases, the probability TT^itf lSt), that the specific radio resource allocation action taken at a time slot t had to be selected, is recorded for further use in the training. This radio resource allocation 104 is then implemented by the corresponding Physical and MAC layer modules of the AP 100. The first message 103 may be a vector of real numbers, whose dimension depends on the bandwidth that is allocated for signaling among the APs 100, 120 over the control interface.
[0063] FIG. 3 shows an exemplary flowchart for the network device 100 of FIG. 2. At 301, the network device 100 observes channel and traffic states of associated terminal devices 110. At 302, it sends the first message 103 to the neighbor APs 120. At 303, it obtains the first messages 103 from the neighbor APs 120, and uses the first messages 103, and the local observations 201 to determine a radio resource allocation configuration, for instance, to schedule a terminal device 110. At 304, the network device 100 may then exchange sums of queue length and / or delays of the terminal devices 110 (local cost or local reward) with the other APs 120. It then computes, at 305, a global cost or global reward based on the local rewards, for instance, as the sum of the local rewards. At 307, the network device determines, whether it is time to update the policy model. If no, then it stores at 306 the observations 201 in a buffer and proceeds to 301. If yes, at 310 the network device 100 exchanges the third messages 107 indicating the gradients (and possibly first messages 103 from forward passes on trajectories used fortraining) with the other APs 120. It then updates the allocation policy model at 309. To this end, it may run an iteration of a policy gradient-based algorithm that uses the indicated gradients of the third messages 107, which are based on the computed global reward. At 308 it clears the buffer and proceeds to 301. In practice, the policy processor 101 can be parametrized as a modular function, with each module being, for example, a neural network. As shown in FIG. 4, two basic modules for each AP i may be: a. A state encoder 403, which may be a function that takes as an input the local observation 201 and outputs the first message mlt, which may be a vector of specified dimensions (e.g., according to the capacity of the control interface as mentioned above). b. A policy function 402, which may output the policy to be followed (radio resource allocation configuration) after receiving inputs from the state encoder 403, the local observations 201 and the first messages 103 of the neighbor APs 120.
[0064] In order to process the first messages 103 from the neighbor APs 120, the AP 100 can use an aggregator 401, which may be a block that takes input from its own state encoder 403, but also from the ones of the neighboring APs 120.
[0065] An alternative implementation of the policy module (policy processor 101) is shown in FIG. 5, and may include using the first messages 103 from the neighbor APs 120 directly without using the aggregator 401.
[0066] Another alternative implementation of the policy module (policy processor 101) is shown in FIG. 6, and may include having the local observation 201 of each AP 100 to pass from a different state encoder 403 (self-state encoder), before it is fed into the policy module. This architecture can be combined with the use of aggregators 401 (as in FIG. 4) for the outputs of the state encoders 403.
[0067] The parameters 0lof the policy module of AP i (for example, the weights and biases of the neural networks that constitute the blocks of the architectures shown in FIG. 4 - FIG. 6) may be periodically updated with the help of the training module (training processor 102), as will be detailed next.
[0068] The function of the training processor 102 of the AP 100 (e.g., generally any AP i) is firstly to exchange a specific second message 105 of the local reward (or cost) with the other APs 120 so that a global reward (or cost) can be calculated. Secondly, at each time slot, to save the transition to be used for updating the (parameters of the) policy module 101 in a suitable form for future processing. Thirdly, periodically updating the parameters of the policy module 101 ofthe AP 100.
[0069] The first two parts of the function of the training processor 102 may be accomplished using a method for distributed collaborative multi-agent learning with minimal communication. Specifically, at the end of each time slot, the AP 100 may: a. Calculate the local reward, which may consist of the minus of the sum of the queue lengths Xf of the terminal devices 110 associated to this AP 100 - which are denoted by K(i),rt=~ keKfi)xt ~and exchange, as a second message 105, only this local reward with the respective one or more neighbor APs 120, in order to compute the global reward as the sum of all the local rewards, i.e. Then, store the global reward as the reward for this state transition. b. Store the state transitions and actions of the AP 100, as well as the global rewards (or equivalently cost), accrued by the radio resource allocation policy, as calculated by the previous step. In more detail, it stores in a replay buffer the tuple s , utl, Rt, slt+1), consisting of the state sltat the beginning of that time slot t, the resource allocation decision utltaken (i.e., the action of allocating the radio resources 104 according to the radio resource allocation configuration), the global reward Rtcomputed from the exchange of the second messages 105 described in the previous point and the next state St+i, which is a result of the resource allocation decision utland the interference by the resource allocation decisions of the neighbor APs 120. The state slthere includes the queue length x and may also include the channel states g ’1of its K connected devices,
[0070] In order to update the parameters of the policy processor 101, this disclosure specifies as further exchange of information the third messages 107 between the APs 100, 120. Specifically, for the last step of the training update, the training processor 101 at the AP 100 (e.g., generally any AP i) works as follows (in the following, A(i) denotes the neighbor APs 120 of AP i): a. After a specified period of time where the APs 100, 120 have used their current policy, e.g. M episodes of length T time slots, the training processor of each AP 100, 120 computes for M trajectories the reward-to-go R™ fromeach slot t until the end of trajectory m, based on the previously calculated global reward Rt. The term trajectory denotes a sequence of observations, actions and rewards obtained when taking actions according to a given policy, which can be sampled from the replay buffer. b. In order to reduce the variance of the estimation, it may optionally train a baseline function b f the local state slt. This baseline function can be used to compute the advantage function which quantifies how much is a certain action utla good or a bad decision at a given state slt. c. It then updates the policy parameters of the policy nl. Is^; 0l), for example, with Monte- Carlo estimation for the rewards-to-go by using the computed advantage function and appropriate signaling. Note here that, due to the first messages 103, the policy of an AP 100 depends on the local state of neighbor APs 120 in the network (depends always on the local state of its neighbor APs 120, and may also depend on the n-hop neighbors if n levels of aggregation are used) as well as the parameters of the state encoders and perhaps aggregators of the neighbors.
[0071] A policy gradient algorithm may be used to estimate the gradient in the direction of which to update the allocation policy model and the parameters for the blocks responsible for the first message generation, i.e., the first message generation model. In this context, the AP i may use the following training loss iog(1r‘(u ^.9 e «1)» (1) and may update the parameters of the policy processor 101 by a gradient descent on the above loss and the corresponding losses of its neighbor APs 120. This may be done by exchanging the third messages 107 (which indicate the gradients) between the APs 100, 120, in order to then back-propagate the gradients, as seen in FIG. 7. The procedure for the alternative policy modules (shown in FIG. 5 and FIG. 6) is completely analogous. Accordingly, FIG. 7 illustrates the policy module update when a policy gradient algorithm is used, i.e. a single update towards the direction of the gradient of equation (1) is done. The arrows indicate the directions where the gradients are back-propagated and exchanged between the APs 100, 120 by indication in the third messages 107. Alternatively, as shown in FIG. 8, a more advanced policy gradient algorithm like Trust Region Policy Optimization (TRPO) or Proximal Policy Optimization (PPO) can be used. These methods involve that each AP now minimizes a training loss function Li For example, in the case of PPO, this loss function takes the form where Alt e is a parameter and isthe allocation policy model used in the previous M episodes. In this case, training is done in multiple rounds, each including a backward and a forward pass: In the backward pass, the parameters of the policy modules 101 of the APs 100, 120 are updated in the direction of the gradient of equation (2); for this, a back-propagation between modules, and signaling where neighboring APs 100, 120 exchange the third messages 107 with the relevant gradients, is needed. In the forward pass, the local observation 201 of each transition, in the trajectories used for training, is fed in the updated policy module 101 and the corresponding probabilities of the resource allocation decisions are calculated; for this, first messages 103 are exchanged between APs for each transition. Accordingly, FIG. 8 illustrates a policy module update, wherein the policy modules are updated via minimizing training loss functions (e.g. PPO, TRPO). The arrows indicate the directions for the backward and forward passes.
[0072] As before, the procedure for training in this way for the alternative policy modules (in FIG. 5 and FIG. 6) is completely analogous. This method leads to more signaling to execute each policy update, however it makes each policy update more effective and may enable the algorithm to learn faster. The above process implies that, if the global reward at each time slot t is known to all APs 100, 120, the factorized policy applied in the system will be updated in the direction of the stochastic gradient of the global average reward in a totally distributed manner. Thus, there is a theoretical guarantee that the radio resource allocation policies as well as the communication among APs (through the state encoder modules) are updated in a way that improves the performance and at the same time this is achieved by requiring much less communication overhead than the methods with centralized training.
[0073] A representative implementation of the above-proposed system of network devices 100, 120 is in a cellular network environment where the APs 100, 120 (or BSs) have to perform uplink scheduling for their connected terminal devices 110. Therefore, each AP 100, 120 has to decide which of its associated devices 110 to schedule at each time slot. Regarding the traffic model, it may be assumed that traffic A (in bits) arrives at a device k at each slot t according to a random process. In detail, for AP i and its K connected devices one has:
[0074] • State: where:
[0075] ■ x Amount of data bits of device k at time slot t, not yet delivered (queue length).
[0076] ■ g ’1Channel state between device k and AP i at time slot t.
[0077] • Action: The action utlG {0, 1, ... , K} is the selected device to schedule at each slot t. The choice k = 0 denotes the decision to schedule no device. Given the actions and the channel states of all APs at slot t, the number of transmitted bits sent by device i at slot t is calculated using the Shannon formula: where W is the bandwidth used for transmission, Tsis the duration of each time slot and P is the uplink transmission power normalized by the receiver noise power. The number of bits remaining in the queue of device k in the beginning of the next slot is given
[0078] • Reward: The local reward for each AP 100, 120 is defined as the minus of the sum of the queue lengths of its devices, i.e. rt(= — x .
[0079] The policy processor 101 used in this implementation may be a deep neural network (DNN) with K + 1 outputs with parameters 01. The training processor 102 may use a Generalized Advantage Estimation (GAE) method for advantage estimation and a policy gradient may be used to update the parameters of the policy modules 101 (see FIG. 7).
[0080] For the performance evaluation of the solutions of this disclosure, particularly the abovedescribed implementation, simulations were performed for a system of N = 4 APs 100, 120 in a topology of a square. Traffic follows a Poisson distribution with the same mean for all devices. There are K = 20 terminal devices 110 placed at random within the coverage area and each device is associated with the closest AP 100, 120. The allocation policy models are updated every epoch of 4 episodes and each episode consists of 1000 time slots.
[0081] In order to evaluate the performance of the proposed solution, it was compared against the standard scheduling algorithms of Proportional Fairness (PF) and Max-Weight. FIG. 9 presents the global reward as a function of the training epochs, for results averaged over 5 random realizations. We can see that after a short training time, the information exchange among the APs 100, 120 in the execution and training phases allows them to learn quickly a way to coordinate and outperform both baselines.
[0082] Fig. 10 shows a method 1000 for radio resource allocation by a network device 100. The method 1000 comprises, in a current time slot: a step 1001 of determining an initial local state of the network device 100, the initial local state indicating a current state of one or more terminal devices 110 connected to the network device 100; as step 1002 of generating a first message 103 according to a first message generation model and based on the initial local state; a step 1003 of causing the network device 100 to send the first message 103 to one or more neighboring network devices 120, and to obtain one or more first messages 103 from the one or more neighboring network devices 120; a step 1004 of determining a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message 103 and the one or more first messages 103 of the one or more neighboring network devices 120; and a step 1005 of allocating radio resources 104 according to the radio resource allocation configuration to the one or more terminal devices 110. The steps 1001-1005 may be performed by a policy processor 101 of the network device 100.
[0083] The method 1000 further comprises, in the current time slot: a step 1006 of generating a second message 105 indicating a local reward obtained by allocating the radio resources 104 according to the radio resource allocation configuration; a step 1007 of causing the network device 100 to send the second message 105 to the one or more neighboring network devices 120, and to obtain one or more second messages 105 from the one or more neighboring network devices 120, wherein each second message 105 indicates a respective local reward of one of the network devices 100, 120; a step 1008 of calculating a global reward based on all the local rewards; a step 1009 of generating a third message 107, indicating gradients for respectively updating the allocation policy model and the first message generation model, based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources 104 to the one or more terminal devices 110 connected to the network device 100, and global rewards, of one or more consecutive time slots before the current time slot; a step 1010 of causing the network device 100 to send the third message 107 to one or more neighboring network devices 120, and to obtain one or more third messages 107 from the one or more neighboring network devices 120; and a step 1011 of updating the first message generation model and the allocation policy model based on the third messages 107.
[0084] The advantages of the solution of this disclosure can be summarized as follows:
[0085] • The solution enables APs 100, 120 to learn a radio resource allocation policy and a signaling between APs 100, 120, which is tunable to the communication capabilities of the control interface: training adjusts to the possible dimension of the first messages 103.
[0086] • The first message generation model learned may be different for each AP 100, 120, adapting to the traffic and channel characteristics of its associated devices. • Radio resource allocation policy model and first message generation model are updated in a way that the total reward accrued is statistically guaranteed to improve.
[0087] • The solution is distributed, therefore can be executed at each network device 100, 120 (e.g., AP or BS) concurrently.
[0088] The present disclosure has been described in conjunction with various embodiments as examples as well as implementations. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed matter, from the studies of the drawings, this disclosure and the independent claims. In the claims as well as in the description the word “comprising” does not exclude other elements or steps and the indefinite article “a” or “an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.
Claims
CLAIMS1. A network device (100) for radio resource allocation, the network device (100) comprising a policy processor (101) and a training processor (102), wherein the policy processor (101) is configured to, in a current time slot: determine an initial local state of the network device (100), the initial local state indicating a current state of one or more terminal devices (110) connected to the network device (100); generate a first message (103) according to a first message generation model and based on the initial local state; cause the network device (100) to send the first message (103) to one or more neighboring network devices (120), and to obtain one or more first messages (103) from the one or more neighboring network devices (120); and determine a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message (103) and the one or more first messages (103) of the one or more neighboring network devices (120); and allocate radio resources (104) according to the radio resource allocation configuration to the one or more terminal devices (110) connected to the network device (100); wherein the training processor (102) is configured to, in the current time slot: generate a second message (105) indicating a local reward obtained by allocating the radio resources (104) according to the radio resource allocation configuration; cause the network device (100) to send the second message (105) to the one or more neighboring network devices (120), and to obtain one or more second messages (105) from the one or more neighboring network devices (120), wherein each second message (105) indicates a respective local reward of one of the network devices (100, 120); calculate a global reward based on all the local rewards; and generate a third message (107), indicating gradients for respectively updating the allocation policy model and the first message generation model, based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources(104) to the one or more terminal devices (110) connected to the network device (100), and global rewards, of one or more consecutive time slots before the current time slot; cause the network device (100) to send the third message (107) to one or more neighboring network devices (120), and to obtain one or more third messages (107) from the one or more neighboring network devices (120); and update (106) the first message generation model and the allocation policy model based on the third messages (107).
2. The network device (100) according to claim 1, wherein: the policy processor (101) is configured to determine an initial local state, generate a first message (103), cause the network device (100) to send the first message (103), compute a radio resource allocation configuration, and allocate radio resources (104) repeatedly every time slot of a plurality of consecutive time slots; and the training processor (102) is configured to generate a second message (105), cause the network device (100) to send the second message (105), and calculate a global reward repeatedly every time slot of the plurality of consecutive time slots.
3. The network device (100) according to claim 1 or 2, wherein the training processor (102) is configured to generate a third message (107), cause the network device (100) to send the third message (107), and update (106) the first message generation model and the allocation policy model based on the third messages (107) repeatedly every few time slots of the plurality of consecutive time slots.
4. The network device (100) according to one of the claims 1 to 3, wherein the training processor (102) is further configured to compute the local reward based on the changed local state as a result of allocating the radio resources (104) according to the radio resource allocation configuration to the one or more terminal devices (110) connected to the network device (100).
5. The network device (100) according to one of the claims 1 to 4, wherein the state of the one or more terminal devices (110) comprises a queue length of each terminal device (110) connected to the network device (100).
6. The network device (100) according to one of the claims 1 to 5, wherein the training processor (102) is configured to compute the global reward as a sum of all the local rewards.
7. The network device (100) according to one of the claims 1 to 6, wherein the training processor (102) is configured to use a gradient algorithm to update the first message generation model and the allocation policy model based on the third messages each indicating gradients for respectively updating the allocation policy model and the first message generation model.
8. The network device (100) according to one of the claims 1 to 7, wherein the policy processor (102) is configured to determine the initial local state based on an observation (201) of one or more performance parameters of the one or more terminal devices (110) connected to the network device (100).
9. The network device (100) according to one of the claims 1 to 8, wherein the policy processor (101) is configured to input into the allocation policy model the initial local state and at least one of the generated first message (103) and the one or more first messages (103) of the one or more neighboring network devices (120), to receive the radio resource allocation configuration as an output of the allocation policy model.
10. The network device (100) according to one of the claims 1 to 9, wherein the policy processor (101) is configured to aggregate the generated first message (103) and the one or more first messages (103) of the one or more neighboring network devices (120), and to compute the radio resource allocation configuration based on the aggregation result used as input into the allocation policy model.
11. The network device (100) according to one of the claims 1 to 10, wherein the radio resources (104) allocated to the one or more terminal devices (110) comprise at least one of transmission power, scheduling grant, and precoding and beamforming options.
12. The network device (100) according to one of the claims 1 to 11, wherein the policy processor (101) is configured to: calculate directly the radio resource allocation configuration; orcalculate a probability distribution over different radio resource allocation configurations and randomly select the radio resource allocation configuration based on the probability distribution.
13. The network device (100) according to one of the claims 1 to 12, wherein each first message (103) comprises a vector of real numbers, wherein a dimension of the vector depends on a bandwidth that is allocated among the network device (100) and the one or more neighboring network devices (120) for exchanging the first messages (103).
14. The network device (100) according to one of the claims 1 to 13, wherein the policy processor (101) is configured to generate the first message (103) according to an initial first message generation model, if the time slot is an initial time slot.
15. The network device (100) according to one of the claims 1 to 14, wherein the policy processor (101) and / or the training processor (102) are each configured to run at least one neural network to perform their respective operations.
16. The network device (100) according to one of the claims 1 to 15, wherein at least one of the first message generation model, the allocation policy model and an aggregator (401) configured to aggregate the generated first message (103) and the one or more first messages (103) of the one or more neighboring network devices (120), is implemented by a neural network.
17. A method (1000) for radio resource allocation by a network device (100), the method (1000) comprising, in a current time slot: determining (1001) an initial local state of the network device (100), the initial local state indicating a current state of one or more terminal devices (110) connected to the network device (100); generating (1002) a first message (103) according to a first message generation model and based on the initial local state; causing (1003) the network device (100) to send the first message (103) to one or more neighboring network devices (120), and to obtain one or more first messages (103) from the one or more neighboring network devices (120);determining (1004) a radio resource allocation configuration based on an allocation policy model, the initial local state, and at least one of the generated first message (103) and the one or more first messages (103) of the one or more neighboring network devices (120); allocating (1005) radio resources (104) according to the radio resource allocation configuration to the one or more terminal devices (110); generating (1006) a second message (105) indicating a local reward obtained by allocating the radio resources (104) according to the radio resource allocation configuration; causing (1007) the network device (100) to send the second message (105) to the one or more neighboring network devices (120), and to obtain one or more second messages (105) from the one or more neighboring network devices (120), wherein each second message (105) indicates a respective local reward of one of the network devices (100, 120); calculating (1008) a global reward based on all the local rewards; generating (1009) a third message (107), indicating gradients for respectively updating the allocation policy model and the first message generation model, based on a history of one or more radio resource allocation configurations, initial local states, changed local states obtained as a result of allocating radio resources (104) to the one or more terminal devices (110) connected to the network device (100), and global rewards, of one or more consecutive time slots before the current time slot; causing (1010) the network device (100) to send the third message (107) to one or more neighboring network devices (120), and to obtain one or more third messages (107) from the one or more neighboring network devices (120); and updating (1011) the first message generation model and the allocation policy model based on the third messages (107).
18. A computer program comprising instructions which, when the program is executed by at least one processor, cause the at least one processor to perform the method (1000) according to claim 17.