Method and apparatus for radio resource allocation
Patent Information
- Application Number
- PCT/SE2024/050207
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing resource allocation methods in wireless networks fail to maintain fairness and quality of service (QoS) requirements, leading to inefficiencies and sub-optimal performance, particularly in scenarios like 5G networks where diverse QoS demands exist.
A method and apparatus that utilize a multi-armed bandit (MAB) approach to allocate discrete radio resources, incorporating a lower-bound optimization problem and a probabilistic sampling strategy to ensure fairness and QoS, with a convex combination of estimated optimal resource allocation and an exploration term to achieve fast convergence.
The method ensures fair and efficient allocation of radio resources, achieving optimal throughput while meeting QoS requirements, with fast convergence and minimal sample complexity, outperforming existing methods in terms of fairness and sample efficiency.
Smart Images

Figure SE2024050207_02102025_PF_FP_ABST
Abstract
Description
[0001]METHOD AND APPARATUS FOR RADIO RESOURCE ALLOCATION TECHNICAL FIELDThe application relates to methods of allocating discrete radio resources of anaccess node of a telecommunication network to wireless devices. Further disclosedare corresponding network apparatuses, a corresponding computer program, and acorresponding computer program product. BACKGROUND Fairness in wireless networks is a critical aspect to ensure equitable access anddistribution of network resources among different users and devices. In a wirelessnetwork optimization task, fairness considerations are of importance to prevent certain users or devices from monopolizing the available resources, leading to imbalance in Quality of Service (QoS) / Quality of Experience (QoE). In particular,fairness must be considered by wireless network schedulers. Scheduling is the taskof assigning time slots / frequency channels among competing wireless devices tomaximize a given utility function such as the total throughput of all wireless devicesattached to a particular base station (e.g., maximize the achievable sum-rate oftransmitted packets over all devices attached to a given base station). In scenarioslike Time Division Multiple Access (TDMA), Frequency Division Multiple Access(FDMA), or Orthogonal Frequency-Division Multiple Access (OFDMA) technologies,fairness mechanisms are important to ensure opportunities for transmission to / fromwireless devices corresponding to required quality of service (QoS) requirements.The stochastic Multi-Armed Bandit (MAB) is a sequential decision-making problemwhere a learner must strategically allocate resources among multiple alternatives tomaximize cumulative rewards.In MABs, in each round, an agent or learner selects an action, and in turn, receives arandom reward. More precisely, in MAB the allocation proceeds in rounds ^ ≥ 1 asfollows: At round ^, the learner selects an arm (or action) a^ from a set of ^ discretearms receives a reward ^^^ = ^^^ + ^^ , where ^^ ∈ ^ is anindependent and identically distributed (i.i.d.) noise sample from a given probabilitydistribution, and θ = {^^}^∈[^] is the unknown parameter of the MABproblem. Examples of probability distributions considered in the bandit literature areGaussian, Bernoulli, etc. The choice of the probability distribution depends on the modeling assumptions of the considered application and is a degree of freedom in MABs. Generally, there are two main objectives that can be considered in MABs: Best Arm Identification (BAI) and Regret Minimization (RM). These two objectives are detailed in the following. In BAI with fixed confidence, the learner aims at devising an algorithm that estimatesthe best arm^⋆ ^with the least number of samples. The best arm is defined as thearm with the largest expected reward and is denoted as: In RM, the learner aims at minimizing the regret up to a given time horizon T ≥ 1. Analgorithm in RM is defined as a sequence of actions (^^)^^^.The regret up to time T is defined as the difference between the best possible actiona⋆^ and the action selected by the algorithm, or more formally as Where ^^[. ] Is the expectation operator and the expectation is taken over all therandom quantities. Resource allocation problems in wireless networks can be modelled as MABs,wherein each arm corresponds to a wireless device to be allocated discrete radioresources. RM and BAI can be understood as adding a QoS requirement to theresource allocation problem – a naïve distribution of resources using an MAB maysimply maximize throughput without maintaining QoS for all devices. With 5G and future standards, allocation of radio resources capable of maintaining QoS fordifferent devices with different QoS requirements is of great interest.Much of the work in bandit literature has focused on the RM or BAI problems withoutconsidering the fairness aspect. There are a few works studying bandits under fairness constraints mostly in the RM setting. However, these works consider different objectives from ours and do not come with any optimality guarantees, i.e.,they are sub-optimal in terms of regret. In particular, when applying MABs toresource scheduling in wireless networks, known algorithms fail to maintain fairnessat each step in the allocation process and / or require large amounts of training data / many rounds of allocation before converging toward an optimal allocation, making them unsuitable for implementation in a live network. SUMMARYIt is an object of the present disclosure to solutions for resource allocation in wirelessnetworks with fast convergence taking QoS requirements into account.According to a first aspect, there is a method of allocating discrete radio resources ofan access node of a telecommunication network to wireless devices. The method isperformed by a scheduler node of the telecommunication network. The methodcomprises, sequentially, for each radio resource solving a lower-bound optimizationproblem {^^}^ which is dependent on a respective mean reward and arespective value ^^, for each wireless device ^ to obtain an estimated optimal resource allocation. The respective parameters are updated, for each wireless device, by sampling from a probability function defined as a convex combination of the estimated optimal resource allocation and an exploration term. The methodcomprises initiating allocation of the radio resource to the wireless device having thelargest mean reward. According to an embodiment of the first aspect, the respective mean reward forwireless device ^ at iteration ^ is a sum over the reward for selecting wireless device^ at each iteration 1, 2, … , ^, divided by the number of previous iterations wherewireless device ^ was selected.According to an embodiment of the first aspect, the exploration term is set to 1 / ^,wherein ^ is the number of wireless devices.According to an embodiment of the first aspect, the convex combination at iteration ^is (1 − + ^^ / ^, where ^^^ is the estimated optimal resource allocation towireless device ^ at iteration ^ and ^^ = 1 / (2√^).According to an embodiment of the first aspect, the method is iterated at least until anumber of radio resources ^^ = min{^ > 0: ^ ≥ havebeen allocated wherein is an exploration threshold such that the predictedwireless device to allocate the discrete radio resource to by learned vector ^ isoptimal with a probability of at least 1 − ^.According to a second aspect, there is a method of allocating discrete radioresources of an access node of a telecommunication network to wireless devices.The method is performed by a scheduler node of the telecommunication network.The method comprises sequentially, for each radio resource, solving a lower-boundoptimization problem {^^}^ which is dependent on a respective meanreward and a respective value ^^, for each wireless device ^, to obtain anestimated optimal resource allocation. The respective parameters are updated, foreach wireless device, by sampling from a probability distribution at least dependent on: a wireless device having the largest mean reward, a number of previous allocations where a best wireless device was not selected, and a current solution tothe lower-bound optimization problem. The method further comprises initiatingallocation of the radio resource to the wireless device having the largest mean reward. According to an embodiment of the second aspect, the respective mean reward forwireless device ^ at iteration ^ is defined as a sum over the reward for selectingwireless device ^ at each iteration 1, 2, … , ^ divided by the number of previousiterations where wireless device ^ was selected.According to an embodiment of the second aspect, the probability distribution isdefined by: if each wireless device associated to a sub-optimal reward has been allocated a resource more than a threshold number of times, select the wireless device having the largest mean reward; else if the wireless device which has been sampled the least number of timeshas been sampled less than ^^(^) times, select the wireless device which has beensampled the least number of times; else select the wireless device which most deviates from the respective value ^^. According to an embodiment of the second aspect, the predetermined threshold forwireless device ^ is ^^(^)log (^), where is the solution to the lower boundmaximization problem at iteration ^ for wireless device ^.According to an embodiment of the second aspect, the method steps are iterated atleast until a number of radio resources ^ have been distributed such that the learnedvector ^ allocates wireless resources in fairly in expectation.According to a third aspect, there is a network apparatus for allocating discrete radioresources of an access node of a telecommunication network to wireless devices.The network apparatus is configured to, sequentially for each radio resource, solve alower-bound optimization problem {^^}^ which is dependent on arespective mean reward and a respective value for each wireless device ^, toobtain an estimated optimal resource allocation. The respective parameters areupdated, for each wireless device, by sampling from a probability function defined as a convex combination of the estimated optimal resource allocation and an exploration term. The network apparatus is configured to initiate allocation of the radio resource to the wireless device having the largest mean reward. According to an embodiment of the third aspect, the respective mean reward forwireless device ^ at iteration ^ is a sum over the reward for selecting wireless device^ at each iteration 1, 2, … , ^, divided by the number of previous iterations wherewireless device ^ was selected.According to an embodiment of the third aspect, the exploration term is set to 1 / ^,wherein ^ is the number of wireless devices.According to an embodiment of the third aspect, the convex combination at iteration^ is (1 − is the estimated optimal resource allocationto wireless device ^ at iteration ^ and ^^ = 1 / (2√^).According to an embodiment of the third aspect, the network apparatus is furtherconfigured to iterate at least until a number of radio resources ^^ = min{^ >0: ^ ≥ have been allocated wherein ^(^, ^) is an explorationthreshold such that the predicted wireless device to allocate the discrete radioresource to by learned vector ^ is optimal with a probability of at least 1 − ^.According to a fourth aspect, there is a network apparatus allocating discrete radioresources of an access node of a telecommunication network to wireless devices.The network apparatus is configured to, sequentially for each radio resource, solve alower-bound optimization problem ^, {^^}^ which is dependent on arespective mean reward and a respective value ^^, for each wireless device ^, toobtain an estimated optimal resource allocation. The respective parameters areupdated, for each wireless device, by sampling from a probability function at least dependent on: a wireless device having the largest mean reward, a number of previous iterations where the best arm was not selected, and a current solution to the lower-bound optimization problem. The network apparatus is configured to initiate allocation of the radio resource to the wireless device having the largest mean reward.According to an embodiment of the fourth aspect, the respective mean reward forwireless device ^ at iteration ^ is defined as a sum over the reward for selectingwireless device ^ at each iteration 1, 2, … , ^ divided by the number of previousiterations where wireless device ^ was selected.According to an embodiment of the fourth aspect, the probability distribution is defined by: if each wireless device associated to a sub-optimal reward has been allocated a resource more than a threshold number of times, select the wireless device having the largest mean reward; else if the wireless device which has been sampled the least number of timeshas been sampled less than ^^(^) times, select the wireless device which has beensampled the least number of times; else select the wireless device which most deviates from the respective value ^^. According to an embodiment of the fourth aspect, the predetermined threshold forwireless device ^ is ^^(^)log (^), where is the solution to the lower boundmaximization problem at iteration ^ for wireless device ^.According to an embodiment of the fourth aspect, the network apparatus isconfigured to iterate at least until a number of radio resources ^ have beendistributed such that the learned vector ^ allocates wireless resources in fairly inexpectation. According to a fifth aspect, there is a computer program comprising computer readable instruction which, on execution by the processor of a network apparatus, cause the network apparatus to perform a method according to any embodiment of the first or second aspects.According to a sixth aspect, there is a computer program product comprising a non-transient storage medium on which a computer program according to the fifth aspect is stored. BRIEF DESCRIPTION OF THE DRAWINGS Fig.1 is a flowchart of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices, accordingto embodiments. Fig. 2 is a flowchart of a method of allocating discrete radio resources of anaccess node of a telecommunication network to wireless devices, accordingto embodiments. Fig.3 is a box plot of an outcome of an embodiment of a method of allocating discrete radio resources of an access node of a telecommunication network towireless devices, according to embodiments.Fig.4 is a line graph of an outcome of an embodiment of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices, according to embodiments.Fig.5 is a line graph of an outcome of an embodiment of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices, according to embodiments.Fig.6 is a line graph of an outcome of an embodiment of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices, according to embodiments.Fig. 7 is an example of a network apparatus configured to performembodiments of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices according toembodiments. Fig. 8 is an example of a telecommunication network according toembodiments in which embodiments of a method according to the disclosuremay be implemented. Fig.9 is a table of an outcome of an embodiment of a method of allocating discrete radio resources of an access node of a telecommunication network to wireless devices according to embodiments. DETAILED DESCRIPTION OF THE DRAWINGS Wireless communication between wireless devices and network nodes requires allocation of radio resources for uplink and / or downlink transmissions of data packetsby the network. When a plurality of wireless devices are connected to a single radioaccess network node, such as a Radio Base Station (RBS), a NodeB, an eNodeB, agNodeB, etc, the wireless devices share radio resources in time and frequency. Theradio nodes may implement a channel access method such as Time DivisionMultiple Access (TDMA), Frequency Division Multiple Access (FDMA), or OrthogonalFrequency-Division Multiple Access (OFDMA), for managing access to the radioresources. However, the technical channel access methods must be combined witha method for selecting which wireless device should be assigned a discrete radioresource, i.e., a certain frequency range) during a time slot (e.g., 1 ms).Fig.1 is a flowchart of a method 100 of allocating discrete radio resources of anaccess node of a telecommunication network to wireless devices. In practice, eachiteration of the method 100 results in one time / frequency slot being allocated to onewireless device of the plurality of wireless devices.The method 100 is performed by a scheduler node of the telecommunicationnetwork. The scheduler node may be a distinct physical node in thetelecommunication network, or it may be a logical component of a node with multiple functions. For example, the scheduler node performing the method may becomprised in a network node in a radio access network.The method 100 is based on a version of a multi-armed bandit (MAB) adapted forbest-arm identification (BAI) in a wireless network where arms correspond towireless devices. The method is a probabilistic method which is ProbablyApproximately Correct (PAC) with confidence ^ ∈ (0,1), that is, the optimal wirelessdevice is selected with probability 1 − ^, where the parameter ^ may be configuredby the network operator. By way of example, the parameter may be set to 10^^ ≥^ ≥ 10^^. The precise value may be determined by weighing available computationalresources and stringency of the quality of service (QoS) requirements for the radioaccess node – larger values of ^ lead to faster convergence, but a higher probabilityof learning a non-optimal allocation.A BAI method comprises a sampling rule, a stopping rule, and a recommendationrule. The sampling rule is a time-varying discrete probability distribution assigning aprobability to each wireless device for allocation of a discrete wireless resource. Thesampling rule varies with time, i.e., the probability distribution depends on the currentestimate of the vector ^. The stopping rule is a rule determining the number ofiterations needed to obtain convergence of the BAI method, guaranteeing optimalallocation. The recommendation rule is the means by which a wireless device isselected when the stopping rule has been fulfilled. The result of iterating the method 100 is learning a parameter vector ^, where ^ represents an optimal allocation of discrete radio resource satisfying the QoS constraints. That is, in each iteration of the method 100 a more preciseapproximation of ^ is obtained.The method 100 may be implemented in a wireless communications network forallocation of radio resources. The fairness constraint imposes that, as the method isiterated to distribute radio resources, each wireless device is allocated a radioresource at least a given fraction of the total rounds. In particular, a vector ^ = where ^ is the number of wireless devices currently attached to thenetwork node and [^] = {1, is a set indexing the ^ wireless devices,determines a desired minimum allocation level for each wireless device. The vector ^may, for example, be determined using QoS requirements for each wireless device.The method 100 guarantees that the constraints set by the vector ^ are satisfied withprobability 1 − ^, so that throughput of the network may be maximized whileensuring that all devices satisfy a minimal QoS. The fairness constraints may be afixed vector of parameter values ^ = , or, alternatively, may comprise avector of functions ^ = {^^(^)}^∈[^] of the parameter ^, so that the fairnessconstraints may vary as a function of the parameter ^. This allows for dynamic control of the fairness.The method 100 is ^-fair ^-PAC if the expected fraction of times a wireless device ^is selected is at least ^^ for all wireless devices in the set of wireless devices, and theprobability of selecting the best wireless device at a stopping time when the methodhas converged is at least 1 − ^. It will be shown that the method 100 is ^-fair ^-PACfor a specific stopping criterion. That is, the method 100 is mathematically optimal forresource distribution under a fairness constraint as the number of rounds the methodis iterated increases. Once a number of rounds corresponding to the specificstopping criterion have been performed, the obtained parameter vector ^ is ^-fair ^-PAC optimal for allocation of discrete radio resources in a network.The method comprises performing the method steps sequentially for each discreteradio resource.The method comprises solving S101 a lower-bound optimization problem {^^}^ which is dependent on a respective mean reward and theprobability vector ^^to obtain a resource allocation. The respective mean reward for each wireless device is initialized during an initialization phase and then updated each time the method 100 is performed. In the initialization phase, each wireless device is allocated a radio resource once, and a counter for each wireless device isinitiated to ^^(^) = 1, ^ = 1, … , for a set of ^ wireless devices attached to the radioaccess node, so that the counter for wireless device ^ is set to 1 at time step ^. Therespective mean rewards are initialized to ^^^ (^) = (∑^∈[^]:^^^^ ^^^ ) / ^^(^), is the reward for selecting the wireless device ^ at time ^. The initialrewards may be set equal for all wireless devices, or the initial rewards may be setproportionally to the probability vector, or the initial rewards may be set randomly. Insubsequent iterations, the respective mean rewards may be calculated recursivelyfor efficiency. The counter ^^(^) counts the number of times a wireless device ^ hasbeen selected up to round ^, or, in other words, the number of times the wirelessdevice ^ has been allocated a radio resource.The lower bound allocation problem may be written as: is the variable to be optimized for wireless device ^, = {^^: ∑^ ^^ = 1} is the associated clipped probability simplex, ^∗ = max ^^ (^),^^and Δ = is the sub-optimality gap for wireless device ^. Solvinglower bound optimization problem comprises obtaining a solution ^∗(^) =(^^∗(^))^∈[^]for the optimal resource allocation at time ^. The allocation vector ^ may be understood as the asymptotic proportion of allocations in which a wirelessdevice ^ is selected.The method further comprises updating the respective parameters using theobtained solution to the lower bound allocation problem based on sampling from aprobability function defined as a convex combination of the estimated optimal resource allocation and an exploration term. The estimated allocation ^^∗iscalculated in the previous step. The exploration term comprises a weighting factor toprevent the scheduling node from selecting wireless devices solely because of theassociated reward, i.e., to force the scheduling node to explore allocation to alldevices to ensure an asymptotically correct estimate of ^, i.e., a fair distribution ofradio resources. In embodiments, the exploration term may be set to 1 / ^, where ^ isthe number of wireless devices. The convex combination may be a combination Preferably, ^ = ^^ = 1 / (2√^), so thatthe exploration term vanishes asymptotically in the number of iterations ^,or,equivalently, the number of distributed radio resources, while ensuring that eachwireless device is asymptotically allocated a radio resource infinitely often.Based on the sampling rule, a random received reward is sampled from theprobability distribution and the wireless device with the largest observed reward isselected to be allocated the wireless resource.The provably optimal result is obtained if a number of wireless resourcescorresponding to a stopping rule have been distributed. To obtain the provablyoptimal result, a stopping rule based on a Generalized Log-Likelihood Ratio Test and the Chernoff stopping rule may be used. The optimal stopping rule may be chosen as: an exploration threshold such that the algorithm is PAC. The function ^^^^isdefined in Theorem 7, page 11, of “Mixture martingales revisited with applications tosequential tests and confidence intervals”, by Kaufmann, E. and Koolen, W. M.,Journal of Machine Learning Research, vol. 22, pages 11140–11183, 2021. It can be shown that the sample complexity, the minimum number of resources to beallocated before convergence to the optimal allocation, is bounded below by aconstant from Eq. 1 which may be referred to as characteristictime. The characteristic time may be understood as a measure of the difficulty ofidentifying the best arm for a given fairness vector.In particular, it can be shown that any ^-fair ^-PAC algorithm satisfies that: Additionally, an asymptotically ^-fair ^-PAC algorithm instead satisfies:liminf ^ [τ ] / ^^^( ⁄ ) ∗^→^ ^ ^ 1 2.4^ ≥ 2^^ ∀^ ∈ Θ.In these expressions, Θ is the sample space for the parameter vector ^ and ^[. ] isthe expectation operator. The proof leverages classical change-of-measurearguments from Lai, T. L. and Robbins, H. (“Asymptotically efficient adaptiveallocation rules”, Advances in Applied Mathematics, vol. 6, pages 4–22, Elsevier,1985) to extend the proof by Garivier, A. and Kaufmann, E. (“Optimal best armidentification with fixed confidence”, Proceedings of Machine Learning Research,vol. 49, pages 1–30, 2016).In view of this lower bound, we can show that the method 100 satisfies the fairnessconstraints and is asymptotically optimal in expectation. That is, the method 100 isprovably optimal for resource allocation under fairness constraints. In particular, it can be shown that the method 100 is ^-fair ^-PAC (or asymptotically ^-fair ^-PAC in the case of variable fairness constraints). Furthermore, it can beshown that for the case where the vector ^ is constant, the method 100 satisfies thefairness constraints for every round (i.e., not just in the limit, but the allocation ofevery radio resource satisfies the fairness constraint). That is, the QoS requirementsare met even during the training phase, when learning an optimal allocation. It can also be shown that the method 100 achieves optimal sample complexityasymptotically. In particular, for all 0 < ^ < 1 / 2, the method 100 has a finiteexpected sample complexity^^[^^], and it satisfies:1. almost sure asymptotic optimality, in the sense that and asymptotic optimality in expectation: ^^^ That is, the method 100 achieves a provably optimal combination of minimal samplecomplexity while satisfying fairness constraints. Algorithm A illustrates a pseudocode implementation of the method 100. Algorithm A Input: fairness vector ^, confidence ^Sample each device once and initialize ^^^ , ^^(^) = 1 ∀^Set ^ ←K compute ^ (^) and set ^ play ^^ ∼ ^^ and observe reward ^^update and set ^ ← ^ +end while arg The method 100 further comprises initiating S103 allocation of the radio resource tothe selected wireless device having the largest mean reward. Initiating allocation may comprise initiating a transmission of an identifier of the selected device to the network node serving the selected device. Alternatively, initiating allocation may comprise initiating a transmission to the selected wireless device. Preferably, the method 100 may be iterated a number of times exceeding the number of times needed for convergence, i.e., to the stopping criterion, at which time subsequent allocations may be performed using the learned vector ^. In simulations,a number of iterations as low as 50 results in convergence, so an optimal resourceallocation may be learned in as little as 50 ms if the discrete resource to be allocatedis 1 ms time slots for transmission.With reference to Fig.9, there is a comparison between the performance of the method 100 (F-BAI), Track-and-Stop (TaS), and a uniform-fair allocation of radioresources. The TaS algorithm is further explored in Garivier and Kaufmann (“Optimalbest arm identification with fixed confidence”, Proceedings of Machine LearningResearch, vol. 49, pages 1–30, 2016). The uniform-fair method is a naïve allocationscheme which selects device ^ in allocation round ^ with probability ^^ +^1 / ^. The comparison is in terms of fairness violations, where the fairness constraints represent the minimal fraction of rounds in which eachwireless device of ten wireless devices is scheduled for transmission. Theassociated reward to be maximized is the sum throughput across all ten wirelessdevices. The method 100 outperforms the TaS allocation method on fairness andoutperforms the uniform-fair allocation method on sample complexity.Fig. 2 is a flowchart of a second method 200 of allocation of discrete radio resourcesof an access node of a telecommunication network to wireless devices. The methodis performed by a scheduler node of the telecommunication network. Themethod 200 is a regret minimization method, aiming to allocate radio resources in away which minimizes regret while satisfying a fairness constraint. More precisely,each wireless device ^ has an associated value ^^, where an allocation is said to be^-fair consistent if every non-optimal wireless device is allocated a radio resource atan asymptotic rate of at least ^^^^^(^). The value of ^^may thus be considered asa lower bound on the regret of the method 200, i.e., the rate at which a non-optimalwireless device (a wireless device with a non-maximal associated reward) isselected. The method 200 thus allocates resources while minimizing the regret. Theregret may be bounded below by the following optimization problem: where the variables ^^determine the rate at which each sub-optimal wireless device must be selected. Preferably, the method 200 is performed for at least apredetermined number of times ^, where ^ may be referred to as a time horizon. Thetime horizon is such that the parameter vector ^ is learned and may be used forsubsequent allocations. In embodiments with at least a predetermined number ofiterations, the regret of the learned vector satisfies:lim inf ^^[^(^)] ^→^log(T)≥ ^∗(^, ^)where ^(^) = ^^[∑^ ^^∗^ − where the expectation is calculated over allpossible actions in the sample space, and ^(^) can be understood as the expecteddifference between the reward of best possible action^∗ ^and the reward of theaction ^^selected by the algorithm.The method 200 comprises, sequentially for each radio resource, solving S201 thelower-bound optimization problem ^∗^^^^^^, {^^}^, which is dependent on arespective mean reward ^ = and the regret vector ^ ={^^}^∈, [^]where the vector indices run over the total number of wireless devices ^ attached tothe network node. Solving the lower-bound optimization problem results in anestimated optimal resource allocation. During an initialization phase, each wireless device is allocated a wireless resourceonce, and counters for the number of times each wireless device of the ^ wirelessdevices has been allocated a resource are set to ^^(^) = 1 for each wirelessdevice ^.In all subsequent iterations ^, an estimator for the optimal resource allocation iscalculated as is the reward obtained at time ^ for selecting device ^. Solving the lowerbound estimation problem ^∗^^^^^^, {^^}^ comprises using the current estimate tosolve the lower bound estimation problem, thereby obtaining a solution vectorcurrent best wireless device estimate is ^⋆= The respective mean rewards may be updated for the next allocation by sampling from a probability distribution dependent on at least the wireless device having the largest mean reward, the number of previous allocations where the best arm was notselected, and the current solution ^(^) to the lower bound optimization problem. Inparticular, the respective mean rewards may be updated after selecting a wireless device to allocate the radio resource to. The wireless device may be selected by, ifeach sub-optimal arm ^ has been samples at least ^^(^)log (^) times, selectingthe current estimate of the best wireless device. However, otherwise the counter ^(^)for the number of times the best wireless device was not selected may be increasedby one and1. ifmin^^(^) , select the wireless device which has been selected^ the least number of times; or else 2. select the wireless device ^ which violates the fairness constraint the most,that is, the wireless device ^ which minimizes (^).The parameter ^ is an exploration parameter which ensures that each device isexplored and the exploration of the devices does not overly favor a device with ahigh associated reward. It can be shown that the method converges to an optimal resource allocation within a time horizon ^.The method 200 further comprises initiating S203 allocation of the radio resource tothe selected wireless device. Initiating allocation may comprise initiating a transmission of an identifier of the selected device to the network node serving the selected device. Alternatively, initiating allocation may comprise initiating a transmission to the selected wireless device. Algorithm B is a pseudo-code implementation of the method 200. Algorithm B Input: fairness vector ^, time horizon ^, explorationparameter ^Sample each device once and initialize ^(^) = 1, ^^(^) =1 ∀^for ^ = ^, … , ^ do: thenplay^(^) = ^(^ − 1)else ^(^) = ^(^ − 1) + 1 end if end if update and set ^ ← ^ + 1end for.Fig. 3 is a box plot of an outcome of the method 100 (F-BAI) compared to twobaseline state-of-the-art methods, Track-and-Stop (TaS) and a uniform fair method,for resource distribution. Note that TaS does not account for fairness constraints. The sample complexity of F-BAI can be seen to be considerably lower (the meanbelow 1000 for F-BAI and the mean around 1400 for Uniform fair) and with a smallerspread. TaS achieves a lower sample complexity, but does not take fairness intoaccount.Fig. 4 is a line graph of an outcome of the method 100. The line graph represents amaximum fairness violation of the method 100 compared to a fairness violation of theTaS method and the Uniform fair method as an average percentage, where theaverage is taken over all iterations. The maximum fairness violation is defined as amaximum difference between the proportion of allocated resources to device ^ andthe desired proportion of allocated resources ^^. The TaS method can be seen to have considerably higher fairness violations of around 11%, while the method 100 achieves an average maximum around 1.5%. Uniform fair performs slightly better, but has considerably higher sample complexity.Fig. 5 is a line graph of an outcome in terms of regret of the method 200 (Fair-RM)according to the disclosure compared to a baseline algorithm Optimal Sampling forStructured Bandits (OSSB) from section 5 of “Minimal Exploration in StructuredStochastic Bandits” by Combes, R., Magureanu, S., and Proutiere, A., 31stConference on Neural information Processing Systems (NIPS 2017), ArXiV:1711.00400v1. The baseline OSSB algorithm does not account for the fairnessconstraints.Fig. 6 is a line graph of an outcome in terms of fairness for the method 200 (Fair-RM)and OSSB. As can be seen, the fairness residual is essentially zero already after afew hundred resource allocations, while the fairness residual of OSSB starts andremains very high. The fairness violation for Fig. 6 is calculated and plotted for eachiteration.Fig. 7 is an example of a network apparatus 700 for allocating discrete radioresources of an access node of a telecommunications network to wireless devices.The network apparatus 700 may comprise a non-transient memory 701 andprocessing circuitry 702, the processing circuitry adapted to perform methods according to embodiments described herein. The memory may comprise a computer program 703, the computer program comprising computer readable instructionswhich, when executed by the processing circuitry, cause the network apparatus toperform an embodiment of the method 100 or 200. Alternatively, the memory maycomprise a computer program product 704, comprising the computer program 703. The network apparatus may be comprised in a radio access node of the telecommunications network.Fig. 8 depicts a telecommunications network 800 in which methods 100, 200according to the disclosure may be performed. The telecommunications networkcomprises at least a radio access node 801 and wireless devices 802, 803. Theradio access node 801 nodes provide access to the telecommunication network towireless devices 802, 803, where access may be improved by employingmethods 100, 200 for scheduling / allocating wireless resources to the wirelessdevices.The telecommunication network may implement a communication protocol accordingto a standard defined by the 3rdGeneration partnership Program (3GPP), the Institute of Electrical and electronics Engineers (IEEE), or any other standardizingbody. The communication protocol may include a broadband cellular networkprotocol such as a version of 4G Long Term Evolution (LTE), 5G New Radio (NR), orany future standard, or a wireless network protocol such as Wi-Fi. Alternatively, a combination of multiple communication protocols may be implemented in the telecommunication network. For example, some wireless devices may be connectedto the telecommunication network using a 5G NR protocol, some may be connectedusing a Wi-Fi protocol, and some may be connected using a device-to-devicecommunication such as Bluetooth. The methods 100, 200 may be implemented inany part of the network which requires sharing of wireless resources.The network apparatus implementing the method 100, 200 may be comprised in thesame node which allocates the radio resources such as a radio access node, or themethod may be performed in a separate node and the selected wireless device maybe communicated to the node performing the allocation.
Claims
CLAIMS 1. A method (100) of allocating discrete radio resources of an access node ofa telecommunication network to wireless devices, the method performed by ascheduler node of the telecommunication network, the method comprising,sequentially, for each radio resource:solving (S101) a lower-bound optimization problem{^^}^ which isdependent on a respective mean rewardand a respective valuefor eachwireless device ^, to obtain an estimated optimal resource allocation, wherein therespective parameters are updated, for each wireless device, by sampling from aprobability function defined as a convex combination of the estimated optimal resource allocation and an exploration term; and initiating (S103) allocation of the radio resource to the wireless device havingthe largest mean reward.
2. The method (100) according to claim 1, wherein the respective meanreward for wireless device ^ at iteration ^ is a sum over the reward for selectingwireless device ^ at each iteration 1, 2, … , ^, divided by the number of previousiterations where wireless device ^ was selected.
3. The method (100) according to claims 1 or 2, wherein the exploration termis set to 1 / ^, wherein ^ is the number of wireless devices.
4. The method (100) according to claim 3, wherein the convex combination atiteration ^ is (1 −^^^ is the estimated optimal resourceallocation to wireless device ^ at iteration ^ and ^^ = 1 / (2√^).
5. The method (100) according to any one of claims 1-4, wherein the methodis iterated at least until a number of radio resources ^^ = min{^ > 0: ^ ≥have been allocated wherein ^(^, ^) is an explorationthreshold such that the predicted wireless device to allocate the discrete radioresource to by learned vector ^ is optimal with a probability of at least 1 − ^.
6. A method (200) of allocating discrete radio resources of an access node ofa telecommunication network to wireless devices, the method performed by ascheduler node of the telecommunication network, the method comprising, sequentially, for each radio resource:solving (S201) a lower-bound optimization problem {^^}^ which isdependent on a respective mean reward and a respective value ^^, for eachwireless device ^, to obtain an estimated optimal resource allocation, wherein therespective parameters are updated, for each wireless device, by sampling from aprobability distribution at least dependent on: a wireless device having the largestmean reward, a number of previous allocations where a best wireless device was notselected, and a current solution to the lower-bound optimization problem, and initiating (S203) allocation of the radio resource to the wireless device havingthe largest mean reward.
7. The method (200) according to claim 6, wherein the respective meanreward for wireless device ^ at iteration ^ is defined as a sum over the reward forselecting wireless device ^ at each iteration 1, 2, … , ^ divided by the number ofprevious iterations where wireless device ^ was selected.
8. The method (200) according to claim 6 or 7, wherein the probabilitydistribution is defined by:if each wireless device associated to a sub-optimal reward has been allocateda resource more than a threshold number of times, select the wireless device havingthe largest mean reward; else if the wireless device which has been sampled the least number of timeshas been sampled less than ^^(^) times, select the wireless device which has beensampled the least number of times; else select the wireless device which most deviates from the respective value^^.
9. The method (200) according to claim 8, wherein the predeterminedthreshold for wireless device ^ is ^^(^)log (^), whereis the solution to thelower bound maximization problem at iteration ^ for wireless device ^.
10. The method (200) according to any one of claims 6-9, wherein the methodsteps are iterated at least until a number of radio resources ^ have been distributedsuch that the learned vector ^ allocates wireless resources in fairly in expectation.
11. A network apparatus (700) for allocating discrete radio resources of anaccess node of a telecommunication network to wireless devices, the networkapparatus configured to, sequentially for each radio resource:solve a lower-bound optimization problem {^^}^ which isdependent on a respective mean reward and a respective value ^^, for eachwireless device ^, to obtain an estimated optimal resource allocation, wherein therespective parameters are updated, for each wireless device, by sampling from aprobability function defined as a convex combination of the estimated optimalresource allocation and an exploration term; andinitiate allocation of the radio resource to the wireless device having the largest mean reward.
12. The network apparatus (700) according to claim 11, wherein therespective mean reward for wireless device ^ at iteration ^ is a sum over the rewardfor selecting wireless device ^ at each iteration 1, 2, … , ^, divided by the number ofprevious iterations where wireless device ^ was selected.
13. The network apparatus (700) according to claims 11 or 12, wherein theexploration term is set to 1 / ^, wherein ^ is the number of wireless devices.
14. The network apparatus (700) according to claim 13, wherein the convexcombination at iteration ^ is (1 − ^^^ is the estimatedoptimal resource allocation to wireless device ^ at iteration ^ and ^^ = 1 / (2√^).
15. The network apparatus (700) according to any one of claims 11-14, furtherconfigured to iterate at least until a number of radio resources ^^ = min{^ >0: ^ ≥have been allocated wherein ^(^, ^) is an explorationthreshold such that the predicted wireless device to allocate the discrete radioresource to by learned vector ^ is optimal with a probability of at least 1 − ^.
16. A network apparatus (700) allocating discrete radio resources of anaccess node of a telecommunication network to wireless devices, the networkapparatus configured to, sequentially for each radio resource: solve a lower-bound optimization problem{^^}^ which isdependent on a respective mean rewardand a respective value ^^, for eachwireless device ^, to obtain an estimated optimal resource allocation, wherein therespective parameters are updated, for each wireless device, by sampling from aprobability function at least dependent on: a wireless device having the largest meanreward, a number of previous iterations where the best arm was not selected, and acurrent solution to the lower-bound optimization problem; and initiate allocation of the radio resource to the wireless device having the largest mean reward.
17. The network apparatus (700) according to claim 16, wherein therespective mean reward for wireless device ^ at iteration ^ is defined as a sum overthe reward for selecting wireless device ^ at each iteration 1, 2, … , ^ divided by thenumber of previous iterations where wireless device ^ was selected.
18. The network apparatus (700) according to claim 16 or 17, wherein theprobability distribution is defined by: if each wireless device associated to a sub-optimal reward has been allocated a resource more than a threshold number of times, select the wireless device having the largest mean reward; else if the wireless device which has been sampled the least number of timeshas been sampled less than ^^(^) times, select the wireless device which has beensampled the least number of times; else select the wireless device which most deviates from the respective value ^^.
19. The network apparatus (700) according to claim 18, wherein thepredetermined threshold for wireless device ^ is ^^(^)log (^),is thesolution to the lower bound maximization problem at iteration ^ for wireless device ^.
20. The network apparatus (700) according to any one of claims 16-19,wherein the network apparatus is configured to iterate at least until a number of radioresources ^ have been distributed such that the learned vector ^ allocates wirelessresources in fairly in expectation.
21. A computer program (703) comprising computer readable instructionwhich, on execution by the processor of a network apparatus (700), cause thenetwork apparatus to perform a method according to any one of claims 1-5 or 6-10.
22. A computer program product (704) comprising a non-transient storagemedium on which a computer program (703) according to claim 21 is stored.