Devices and methods for multi-objective radio resource allocation in wireless networks

WO2026175520A1PCT designated stage Publication Date: 2026-08-27HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/054833
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2026-08-27

Smart Images

  • Figure EP2025054833_27082026_PF_FP_ABST
    Figure EP2025054833_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A network device (120a), e.g. base station (120a), is disclosed for distributed Reinforcement Learning, RL, radio resource allocation with a plurality of further network devices (120b-n), e.g. base stations (120b-n), based on a radio resource allocation policy. The network device (120a) is configured to obtain a plurality of rewards, wherein the plurality of rewards are indicative of a plurality of performance objectives of the network device (120a), and to obtain an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device (120a). Moreover, the network device (120a) is configured to obtain from each of the plurality of further network devices (120b-n) a plurality of local scalarized advantages, wherein the local scalarized advantages from each further network device (120b-n) are scalarized based on an individual preference weight vector of the further network device (120b- n) and wherein the individual preference weight vector of the further network device (120b-n) is indicative of an importance measure of the plurality of performance objectives for the further network device (120b-n). The network device (120a) is further configured to adjust a meta radio resource allocation policy based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Devices and methods for multi-objective radio resource allocation in wireless networks

[0002] TECHNICAL FIELD

[0003] The present invention relates to wireless communication technology. More specifically, the present invention relates to devices and methods for cooperative multi-agent multi-objective radio resource allocation in wireless networks, in particular mobile networks.

[0004] BACKGROUND

[0005] In wireless networks, one of the most important tasks of a base station (BS) is to allocate radio resources to the wireless terminal devices connected to the BS so that these terminal devices can communicate efficiently with the rest of the world. For this reason, radio resource allocation strategies (also referred to as radio resource allocation policies) are designed in such a way that allows BSs to make allocation decisions based on the available network related information. Typically, each BS aims to optimize a combination of target objectives (also referred to as metrics), such as throughput, latency, reliability, energy efficiency, network coverage and the like. This combination may be determined based on the preference (weight) associated with each objective, which is usually captured in a preference vector. The size of the preference vector corresponds to the number of objectives. Each BS may have its individual preference vector that can be updated over time according to the applications it serves. It is desirable to have a radio resource allocation policy that can dynamically adapt to changing individual preference vectors, such as individual preference vectors associated with different communication services supported by a BS.

[0006] On top of that, another practical constraint is the adoption of decentralized BS policies, where each BS that has an individual preference vector can take actions by relying mainly on local information from its connected devices. However, a key problem that arises in such distributed policies is that BSs may act selfishly, resulting in high interference which effectively degrades the overall network performance. Hence, to manage the interference, it is expected that network operators will rely on decentralized resource policies but also allow BSs to exchange a small amount of necessary information.

[0007] Overall, the technical problem to be solved is how to efficiently train decentralized BS radio resource allocation policies with low communication overhead between neighboring BSs, which can be quickly adapted to any combination of individual preference vectors.

[0008] Resource allocation in traditional wireless networks with many BSs is based on solutions that use only local information, namely each BS uses state information of its associated devices.Although these policies are distributed and easy to implement, they come with certain disadvantages. Specifically, these policies can suffer when there is high interference among the BSs so that sophisticated interference management algorithms are required with, at least, some communication overhead for the interference management.

[0009] This challenge naturally leads to solutions that use Reinforcement Learning (RL), a branch of machine learning where agents learn good policies through interaction with the true environment, without making assumptions on the wireless network and its behavior. Specifically, the RL direction that is related is Multi Agent Multi-Objective (MAMO) RL. In a standard MAMO RL framework, agents learn a distributed policy optimized for a common preference vector i.e., BSs are constrained to optimize their policies for the same preference vector. This presents a limitation for wireless networks since each BS may have its individual preference vector that depends on the applications it serves.

[0010] SUMMARY

[0011] It is an object of the invention to provide improved devices and methods for cooperative multiagent multi-objective radio resource allocation in a wireless network, in particular a mobile network, such as a 3GPP mobile network.

[0012] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.

[0013] According to a first aspect a network device (herein also referred to as an agent) is disclosed for distributed Reinforcement Learning, RL, radio resource allocation with a plurality of further network devices, i.e. further agents based on a radio resource allocation policy. In an implementation form the network device according to the first aspect may be a base station and the plurality of further networks a plurality of further base stations for distributed RL radio source allocation to a plurality of terminal devices, e.g. user equipments UEs, connected to the base station and a plurality of further terminal devices, e.g. UEs, connected to the plurality of further base stations based on a radio resource allocation policy. In an RL inference phase the network device is configured to obtain a plurality of rewards indicative of a plurality of performance objectives of the network device, e.g. base station and to obtain an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device, e.g. base station. Moreover, the network device, e.g. base station is configured to obtain from each of the plurality of further base stations a plurality of local scalarizedadvantages (i.e. advantage values), wherein the local scalarized advantages from each further network device, e.g. base station are scalarized based on an individual preference weight vector of the further network device, e.g. base station, wherein the individual preference weight vector of the further network device, e.g. base station is indicative of an importance measure of the plurality of performance objectives for the further network device, e.g. base station. The network device, e.g. base station according to the first aspect is further configured to adjust a meta radio resource allocation policy learned during a training phase based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy. Thus, the network device, e.g. base station according to the first aspect allows for an efficient and fast adaptation of the meta radio resource allocation policy without requiring a re-training.

[0014] In a further possible implementation form, the network device, e.g. base station according to the first aspect is configured to determine a plurality of local advantages and to scalarize the plurality of local advantages based on the individual preference weight vector of the network device, e.g. base station for obtaining a plurality of local scalarized advantage values, wherein the network device, e.g. base station is configured to provide the plurality of local scalarized advantage values to the plurality of further network devices, e.g. base stations.

[0015] In a further possible implementation form, in a RL training phase, the network device, e.g. base station according to the first aspect is configured to determine an adapted radio resource allocation policy based on the meta policy and the plurality of local scalarized advantages.

[0016] In a further possible implementation form, in the RL training phase, the network device, e.g. base station according to the first aspect is configured to determine, i.e. learn the meta radio resource allocation policy based on a plurality of the adapted radio resource allocation policies and the plurality of local scalarized advantages.

[0017] In a further possible implementation form, the network device, e.g. base station according to the first aspect is configured to determine a global scalarized advantage value based on the plurality of local scalarized advantage values and to adjust the meta radio resource allocation policy based on the global scalarized advantage value for obtaining the radio resource allocation policy.

[0018] In a further possible implementation form, the network device, e.g. base station according to the first aspect is configured to determine the global scalarized advantage value as a sum of the plurality of local scalarized advantage values.According to a second aspect a method is provided for operating a network device, e.g. base station for distributed Reinforcement Learning, RL, radio resource allocation with a plurality of further network devices, e.g. base stations, for instance, to a plurality of terminal devices, e.g. UEs, connected to the network device, e.g. base station and a plurality of further terminal devices, e.g. UEs, connected to the plurality of further network devices, e.g. base stations, based on a radio resource allocation policy. In a RL inference phase the method comprises the steps of:

[0019] obtaining a plurality of rewards, wherein the plurality of rewards are indicative of a plurality of performance objectives of the network device, e.g. base station;

[0020] obtaining a single individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device, e.g. base station;

[0021] obtaining from each of the plurality of further network devices, e.g. base stations a plurality of local scalarized advantages, wherein the local scalarized advantages from each further network device, e.g. base station are scalarized based on an individual preference weight vector of the further network device, e.g. base station, wherein the individual preference weight vector of the further network device, e.g. base station is indicative of an importance measure of the plurality of performance objectives for the further network device, e.g. base station; and adjusting a meta radio resource allocation policy learned during a training phase based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.

[0022] In a further possible implementation form, the method according to the second aspect further comprises determining a plurality of local advantage values and scalarizing the plurality of local advantage values based on the individual preference weight vector of the network device, e.g. base station for obtaining a plurality of local scalarized advantage values, wherein the method according to the second aspect further comprises providing the plurality of local scalarized advantage values to the plurality of further network devices, e.g. base stations.

[0023] In a further possible implementation form, in a RL training phase, the method according to the second aspect comprises determining an adapted radio resource allocation policy based on the meta policy and the plurality of local scalarized advantages.

[0024] In a further possible implementation form, in the RL training phase, the method according to the second aspect comprises determining, i.e. learning the meta radio resource allocation policy based on a plurality of the adapted radio resource allocation policies and the plurality of local scalarized advantages.In a further possible implementation form, the method according to the second aspect comprises determining a global scalarized advantage based on the plurality of local scalarized advantages and adjusting the meta radio resource allocation policy based on the global scalarized advantage for obtaining the radio resource allocation policy.

[0025] In a further possible implementation form, the step of determining the global scalarized advantage comprises determining the global scalarized advantage as a sum of the plurality of local scalarized advantages.

[0026] The method according to the second aspect can be performed by the network device, e.g. base station according to the first aspect. Thus, further features of the method according to the second aspect result directly from the functionality of the network device, e.g. base station according to the first aspect and its different implementation forms described above and below.

[0027] According to a third aspect a computer program or a computer program product is provided, comprising a computer-readable storage medium carrying program code which causes a computer or a processor to perform the method according to the second aspect when the program code is executed by the computer or the processor.

[0028] Different aspects of the invention can be implemented in software and / or hardware.

[0029] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.

[0030] BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In the following embodiments of the invention are described in more detail with reference to the attached figures and drawings, in which:

[0032] Fig. 1 is a schematic diagram illustrating a network device in the form of a base station according to an embodiment interacting with a plurality of further network devices in the form of a plurality of further base stations for decentralized radio resource allocation;

[0033] Fig. 2 is a schematic diagram illustrating different operation stages of the network device of figure 1;

[0034] Fig. 3 is a schematic diagram illustrating in more detail the interactions of a network device in the form of a base station according to an embodiment with a further network device in the form of a further base station during different operation stages for decentralized radio resource allocation; andFig. 4 is a flow diagram illustrating a method according to an embodiment for operating a network device for decentralized radio resource allocation.

[0035] In the following identical reference signs refer to identical or at least functionally equivalent features.

[0036] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the invention or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the invention may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0038] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.

[0039] Figure 1 is a schematic diagram illustrating a communication system 100, including a network device 120a in the form of a base station 120a according to an embodiment and a plurality of further network devices 120b-n for decentralized radio resource allocation. As will be described in more detail below, the network device 120a, e.g. base station 120a is configured for distributed Reinforcement Learning, RL, radio resource allocation with the plurality of further network devices 120b-n, e.g. base stations 120b-n shown in figure 1, to a plurality of terminaldevices connected to the base station 120a and a plurality of further terminal devices connected to the plurality of further base stations 120b-n based on a radio resource allocation policy. As will be appreciated, the exchange of information between the network device 120a (referred to as BS i in figure 1) and the further network device 120b (referred to as BS j in figure 1), which will be described in more detail below, is analogous to the information exchange between the network device 120a and the further network device 120n (referred to as BS k in figure 1) and the information exchange between the further network device 120b and the further network device 120n.

[0040] More specifically, in a RL inference phase the network device 120a, e.g. base station 120a is configured per time slot to obtain a plurality of rewards indicative of a plurality of performance objectives of the network device 120a, e.g. base station 120a and to obtain an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device 120a, e.g. base station 120a. Moreover, in the RL inference phase the network device 120a, e.g. base station 120a is configured per time slot to obtain from each of the plurality of further network devices 120b-n, e.g. base stations 120b-n shown in figure 1 a plurality of local scalarized advantage values (herein referred to as local scalarized advantages). The local scalarized advantages from each further network device 120b-n, e.g. base station 120b-n are scalarized based on an individual preference weight vector of the respective further network device 120b-n, e.g. base station 120b-n, wherein the individual preference weight vector of the respective further network device 120b-n, e.g. base station 120b-n is indicative of an importance measure of the plurality of performance objectives for the further network device 120b-n, e.g. base station 120b-n. As will be described in more detail below, in the RL inference phase the network device 120a, e.g. base station 120a is further configured to adjust a meta radio resource allocation policy learned, for instance, during a RL training phase based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.

[0041] Thus, by considering the individual preference vectors of the plurality of network devices 120a-n, e.g. base stations 120a-n a distributed policy may be generated that may be represented by a single Deep Neural Network (DNN) at each network device 120a-n, e.g. base station 120a-n. As will be described in more detail below, the trained models may then be quickly adapted to any combination of individual preference vectors.

[0042] In an embodiment, the system comprising the plurality of BSs 120a-n operates in slotted time and at every time slot t, for a BS i, the following RL related quantities may be defined.An individual preference vector wit∈ ℝL, expressing the interest of the BS i towards the L different objectives.

[0043] A set of neighboring, i.e. further BSs 퓜i, |퓜i| = M (including i).

[0044] A preference matrix W ∈ ℝL×Mthat contains the individual preference vectors of BS i and the further, i.e. its neighboring BSs, i.e. W = [ivn

[0045]

[0046] The state representing the state of all devices connected to BS i.

[0047] The vector rewards rit∈ ℝL.

[0048] The resource allocation action ait.

[0049] The trajectory

[0050]

[0051] a sequence (s°, a°, r°, s,

[0052]

[0053] starting at state

[0054]

[0055] at time t = 0 and ending at state s at time t

[0056]

[0057] = H, i.e. T; = (s? a? r^, s),...,s^Y

[0058] A meta policy DNN with parameters θi.

[0059] An adapted policy DNN with parameters

[0060]

[0061] wherein

[0062]

[0063] is the adaptation

[0064]

[0065] of to the preference matrix I.

[0066] A critic DNN with parameters φi.

[0067] As illustrated in figure 1, each base station 120a-n may comprise a transceiver 130a, b configured to communicate with the environment 110, the further base stations and / or the terminal devices connected to the base station 120a-n and a controller 140a,b configured to implemented the functionality described above. For the sake of simplicity only the transceivers 130a, b and the controllers 140a, b of the base stations 120a, b are illustrated in figure 1. As will be appreciated, however, the other base stations, such as the base station 120n illustrated in figure 1, may comprise the same or similar transceivers and / or controllers as the base stations 120a, b. In an embodiment, the controller 140a, b of each base station 120a-n may comprise the following modules for implementing the functionality described above. A Generator module configure to generate an individual preference vector

[0068]

[0069] in the inference phase or a batch {iv }"=1of N individual preference vectors during training. An Actor DNN configured to receivestate sitand output the resource allocation action aitbased on the meta policy DNN θior the adapted policy θi,W. An Experience buffer configured to store the generated trajectories. A Critic DNN configured to compute the value vi(sit; φi) of a state

[0070]

[0071] based on φi. An Adapter module configured to adapt the meta policy DNN θito the preference matrix W by fine-tuning θi. A Trainer module configured to train the meta policy DNN θiand the critic DNN φi. In figure 1, the communication exchange indicated by the dashed arrows is happening during training, while the communication exchange indicated by the bold arrows is happening during inference. The Trainer module is used during training only, while the other modules are used during training and inference.

[0072] Under further reference to figure 2, the operation of the network device 120a, e.g. base station 120a (or equivalently of the other network devices 120b-n) may be broken down into two stages or phases, namely a training stage 210 and an inference stage 220.

[0073] As illustrated in figure 2, the training stage 210 comprises two substages, namely an Initialization substage 211 and a Meta-training substage 213.

[0074] In the Initialization substage 211 the plurality of BSs 120a-n agree on a common horizon length H, a common number D of trajectories and a common preference batch size N of preference vectors. BS i sends an initialization message mito its neighbors that contains its suggestion for these parameters mi= {Hi, di, Ni}. Then each common parameter is defined as the minimum among these suggestions.

[0075] The Meta-training substage 213 includes the process of learning an optimal meta-policy 0*. The Meta-training is based on values of advantage functions, which in RL are used to evaluate the difference between the value of a specific action and the average value of actions in a given state. This approach enhances the efficiency and stability of the training process. In an embodiment, each training iteration of the meta-training may be composed of two phases, namely Phase 1 and Phase 2.

[0076] Phase 1: Update the critic DNN φiand construct N adapted policy DNNs in the following way:

[0077] 1. Interact with the environment 110 using the meta policy DNN θito collect D trajectories

[0078]

[0079] .

[0080] 2. Update the critic DNN φibased on the generated trajectories 풟i.

[0081] 3. Create a preference batch composed of N individual preference vectors {win}n=1N.4. Compute, for each individual preference vector win, the D × H local scalarized advantages SADVi,nmetabased on the trajectories 풟i.

[0082] 5. Send the N × D × H local scalarized advantages N_SADVimeta= {SADVi,nmeta}n=1Nto neighboring BSs and receive theirs.

[0083] 6. Compute the N × D × H global scalarized advantages {SADVnmeta}n=1Nwhere global scalarized advantages are defined as a summation of local ones:

[0084] SADVnmeta= ∑j∈퓜SADVj,nmeta

[0085] 7. Construct the adapted policy DNNs {θi,W}n=1Nwhere θi,Wis the adaptation of θito the preference matrix Wn= [w1n, ..., wMn]. θi,Wis constructed by fine-tuning θithrough a gradient step based on SADVnmetaand 풟i.

[0086] Phase 2: Update the meta-policy DNN. This phase starts after the end of phase 1 and comprises the following steps:

[0087] 1. For each adapted policy θi,W, interact with the environment 110 to collect D trajectories 풟in.

[0088] 2. Compute the N × D × H local scalarized advantages N_SADViada= {SADVi,nada}n=1N, where SADVi,nadaare the local scalarized advantages computed based on winand the trajectories 풟in.

[0089] 3. Send the N × D × H local scalarized advantages N_SADViadato neighboring BSs and receive theirs.

[0090] 4. Compute the N × D × H global scalarized advantages {SADVnada}n=1N, where the global scalarized advantages are defined as a summation of the local ones, SADVnada= ∑j∈퓜SADVj,nada.

[0091] 5. Update the meta policy DNN through a gradient step based on the global scalarized advantages {SADVnada}n=1Nand the generated trajectories {풟in}n=1N.

[0092] The inference stage 220 starts after the end of the training stage 210 and comprises two substages, namely an Adaptation (fine-tuning) substage 221 and a Deployment substage 223.

[0093] The Adaptation (fine-tuning) substage 221 is similar to phase 1 of the meta-training except that the critic is not updated and it is done for a specific individual preference vector

[0094]

[0095] (not for abatch of N individual preference vectors). More precisely, the Adaptation (fine-tuning) substage 221 comprises the following steps:

[0096] 1. Interact with the environment 110 using the meta policy DNN 0* to collect / ) trajectories

[0097]

[0098] .

[0099] 2. Compute the D x H local scalarized advantages SADVaietabased on the generated trajectories 2)£and the individual preference vector W;.

[0100] 3. Send the D x H local scalarized advantages SADV"ietato neighboring BSs and receive theirs.

[0101] 4. Compute the D x H global scalarized advantages SADVmeta=

[0102]

[0103] SADVymeta.

[0104] The global scalarized advantages are defined as a summation of local ones.

[0105] 5. Construct the adapted policy DNN 0*w, by fine-tuning 0* through a gradient step based on SADVmetaand Z>£.

[0106] The model deployment substage 223 starts after the adaptation substage 221 has finished. During the model deployment substage 223 BS i uses its adapted policy DNN 0*wwith no further communication with its neighbors. If any BS decides to change its individual preference vector then its sends a message to its neighbors to restart the process of adaptation.

[0107] In the following a further embodiment is described under reference to figure 3 in the context of a cellular, i.e. mobile network for a downlink power allocation scenario. In this example M BSs and K terminal devices, e.g. UEs, are considered, wherein each terminal device is associated to its closest BS. As will be appreciated, the exchange of information between the network device 120a (referred to as BS i in figure 3) and the further network device 120b (referred to as BS j in figure 3), which will be described in more detail below, is analogous to the information exchange between the network device 120a and the further network device 120n and the information exchange between the further network device 120b and the further network device 120n. Each BS 120a-n wants to select a transmit power from [0, pmax] at each time slot, where Pmax is the maximum transmission power. In this example two objectives are considered, namely throughput and power consumption. For a BS i, letw; =

[0108]

[0109] be the individual preference vector,

[0110]

[0111] where is the weight associated to throughput and wi 2is the weight associated to power consumption.

[0112] In the initialization phase 211, the BSs 120a-n agree on the common parameters i.e., a common horizon length

[0113]

[0114] H =, a common number of trajectories D = mm{Di, Dj, Dk}, and a common batch size N = mm{Ni, Nj, Nk}.In the meta-training phase 213 and during each training iteration, the BSs 120a-n update their meta policy DNNs based on global scalarized advantages defined as a summation of local ones. A possible way to compute the local scalarized advantage for a given state-action pair (

[0115]

[0116] s^af) e Tt =...,sf) and a given individual preference vector w is A

[0117]

[0118] DV(s-,a-,ivp) = wp • rf - V; ($[;< / >;)), where y e [0,1) is the discount factor.

[0119] In the adaptation (fine-tuning) substage 221, the BSs 120a-n exchange the local scalarized advantages in order to adapt their meta-policy DNNs to the preference matrix.

[0120] In the model deployment phase 223, the BSs 120a-n use their adapted policy DNNs with no further communication. If any BS decides to change its individual preference vector then it may send a message to the other BS in order to restart the process of adaptation.

[0121] Figure 4 is a flow diagram illustrating a method 400 for operating the network device 120a, e.g. base station 120a for distributed RL radio resource allocation with the plurality of further network devices 120b-n, e.g. base stations 120b-n based on a radio resource allocation policy. The method 400 comprises a step 401 of obtaining a plurality of rewards, wherein the plurality of rewards are indicative of a plurality of performance objectives of the network device 120a, e.g. base station 120a, and a step 403 of obtaining an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device 120a, e.g. base station 120a. Moreover, the method comprises a step 405 of obtaining from each of the plurality of further network devices 120b-n, e.g. further base stations 120b-n a plurality of local scalarized advantages, wherein the local scalarized advantages from each further network device 120b-n, e.g. base station 120b-n are scalarized based on an individual preference weight vector of the further network device 120b-n, e.g. base station b-n and wherein the individual preference weight vector of the further network device 120b-n, e.g. base station 120b-n is indicative of an importance measure of the plurality of performance objectives for the further network device 120b-n, e.g. base station 120b-n. As will be appreciated, the steps 401, 403 and 405 may be performed in an order different to the order illustrated in figure 4 and / or at least partially in parallel. The method 400 further comprises a step 407 of adjusting a meta radio resource allocation policy based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.As will be appreciated, embodiments herein provide a scalable approach in that each network device 120a-n, e.g. BS 120a-n has a single DNN that can quickly adapt to any preference matrix i.e., any combination of individual preference vectors. The training is done once, no need to have one DNN per preference matrix, resulting in reduced hardware requirements at each network device 120a-n, e.g. BS 120a-n.

[0122] The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the invention (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step).

[0123] In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely exemplary. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0124] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.

[0125] In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

Claims

CLAIMS1. A network device (120a) for distributed Reinforcement Learning, RL, radio resource allocation with a plurality of further network devices (120b-n) based on a radio resource allocation policy, wherein the network device (120a) is configured to:obtain a plurality of rewards, wherein the plurality of rewards are indicative of a plurality of performance objectives of the network device (120a);obtain an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device (120a);obtain from each of the plurality of further network devices (120b-n) a plurality of local scalarized advantages, wherein the local scalarized advantages from each further network device (120b-n) are scalarized based on an individual preference weight vector of the further network device (120b-n), wherein the individual preference weight vector of the further network device (120b-n) is indicative of an importance measure of the plurality of performance objectives for the further network device (120b-n); andadjust a meta radio resource allocation policy based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.

2. The network device (120a) of claim 1, wherein the network device (120a) is configured to determine a plurality of local advantage values and to scalarize the plurality of local advantage values based on the individual preference weight vector of the network device (120a) for obtaining a plurality of local scalarized advantage values and wherein the network device (120a) is configured to provide the plurality of local scalarized advantage values to the plurality of further network devices (120b-n).

3. The network device (120a) of claim 1 or 2, wherein in a RL training phase the network device (120a) is configured to determine an adapted radio resource allocation policy based on the meta policy and the plurality of local scalarized advantages.

4. The network device (120a) of claim 3, wherein in the RL training phase the network device (120a) is configured to determine the meta radio resource allocation policy based on a plurality of the adapted radio resource allocation policies and the plurality of local scalarized advantages.

5. The network device (120a) of any one of the preceding claims, wherein the network device (120a) is configured to determine a global scalarized advantage based on the pluralityof local scalarized advantages and to adjust the meta radio resource allocation policy based on the global scalarized advantage for obtaining the radio resource allocation policy.

6. The network device (120a) of claim 5, wherein the network device (120a) is configured to determine the global scalarized advantage as a sum of the plurality of local scalarized advantages.

7. A method (400) of operating a network device (120a) for distributed Reinforcement Learning, RL, radio resource allocation with a plurality of further network devices (120b-n) based on a radio resource allocation policy, wherein the method (400) comprises: obtaining (401) a plurality of rewards, wherein the plurality of rewards are indicative of a plurality of performance objectives of the network device (120a);obtaining (403) an individual preference weight vector for the plurality of rewards, wherein the individual preference weight vector is indicative of an importance measure of the plurality of performance objectives for the network device (120a);obtaining (405) from each of the plurality of further network devices (120b-n) a plurality of local scalarized advantages, wherein the local scalarized advantages from each further network device (120b-n) are scalarized based on an individual preference weight vector of the further network device (120b-n), wherein the individual preference weight vector of the further network device (120b-n) is indicative of an importance measure of the plurality of performance objectives for the further network device (120b-n); andadjusting (407) a meta radio resource allocation policy based on the plurality of local scalarized advantages for obtaining the radio resource allocation policy.

8. The method (400) of claim 7, wherein the method (400) further comprises determining a plurality of local advantage values and scalarizing the plurality of local advantage values based on the individual preference weight vector of the network device (120a) for obtaining a plurality of local scalarized advantage values and wherein the method (400) further comprises providing the plurality of local scalarized advantage values to the plurality of further network device (120b-n).

9. The method (400) of claim 7 or 8, wherein in a RL training phase the method (400) comprises determining an adapted radio resource allocation policy based on the meta policy and the plurality of local scalarized advantages.

10. The method (400) of claim 9, wherein in the RL training phase the method (400) comprises determining the meta radio resource allocation policy based on a plurality of the adapted radio resource allocation policies and the plurality of local scalarized advantages.

11. The method (400) of any one of claims 7 to 10, wherein the method (400) comprises determining a global scalarized advantage based on the plurality of local scalarized advantages and adjusting the meta radio resource allocation policy based on the global scalarized advantage for obtaining the radio resource allocation policy.

12. The method (400) of claim 11, wherein determining the global scalarized advantage comprises determining the global scalarized advantage as a sum of the plurality of local scalarized advantages.

13. A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (400) of any one of claims 7 to 12, when the program code is executed by the computer or the processor.