Decentralised resource allocation policies

A neural network-based controller synchronizes training phases and defines a common macro time slot to optimize decentralized resource allocation across BSs with asynchronous TTIs, reducing interference and improving network performance.

WO2026098783A1PCT designated stage Publication Date: 2026-05-15HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-11-08
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In wireless networks with asynchronous Transmission Time Intervals (TTIs), base stations (BSs) acting selfishly lead to high interference, degrading overall network performance, as existing decentralized resource allocation policies fail to coordinate effectively.

Method used

A neural network-based controller at each BS implements a decentralized resource allocation policy, synchronizing training phases and defining a common macro time slot through communication with neighbors, using local and global macro rewards to optimize bandwidth allocation across BSs with asynchronous TTIs.

Benefits of technology

This approach reduces interference and enhances global network performance by coordinating BS decisions with low communication overhead, enabling efficient resource allocation even in asynchronous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024081645_15052026_PF_FP_ABST
    Figure EP2024081645_15052026_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, an apparatus for a controller of a first node of a wireless telecommunication system is provided, wherein the controller comprises a neural network configured to implement a decentralised resource allocation policy for the first node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DECENTRALISED RESOURCE ALLOCATION POLICIES

[0002] TECHNICAL FIELD

[0003] The present disclosure relates, in general, to decentralised resource allocation policies, and more specifically, although not exclusively, to multi-agent bandwidth resource allocation in wireless environments with asynchronous transmission time intervals.

[0004] BACKGROUND

[0005] In existing wireless networks, a base station or node (BS) is connected to a large number of devices, such as user equipment (UE) and one of its most important tasks is to allocate bandwidth resources to those devices so that the latter can communicate efficiently with the rest of the world. To this end, the bandwidth is divided into several resource blocks (RBs) based on the chosen numerology. Then, these resource blocks are assigned to the connected devices via a resource allocation strategy (policy).

[0006] The choice of the numerology for a BS depends on the services or the use cases that will be supported, as shown in release 15 and beyond of the 3rd Generation Partnership Project (3GPP) for example. Moreover, due to practical constrains, it is expected that operators will adopt decentralized BS policies, where each BS can take actions (e.g., bandwidth allocation) by relying mainly on local information from its connected devices. However, a key problem that arises in such distributed policies is that BSs may act selfishly, resulting in high interference which effectively degrades the overall network performance. Hence, to manage the interference, it is expected that network operators will rely on decentralized resource policies, but also allow BSs to exchange a small amount of necessary information. Such policies are designed to perform well when BSs have synchronous Transmission Time Intervals (TTIs), i.e., BSs use the same numerology and take synchronous actions. However, especially in wireless networks, BSs are expected to choose different numerologies when they serve applications with different requirements, which leads to asynchronous TTIs. In this situation, the operator needs an automatic way to coordinate BS decisions and learn good decentralized policies that can perform well for wireless networks with asynchronous TTIs.

[0007] SUMMARY

[0008] An objective of the present disclosure is to provide a mechanism to efficiently train and operate a decentralized BS bandwidth resource allocation policy with low communication overhead between neighbouring nodes, when nodes have asynchronous TTIs.

[0009] The foregoing and other objectives are achieved by the features of the independent claims.

[0010] Further implementation forms are apparent from the dependent claims, the description and the Figures.

[0011] A first aspect of the present disclosure provides an apparatus for a controller of a first node of a wireless telecommunication system, wherein the controller comprises a neural network configured to implement a decentralised resource allocation policy for the first node, wherein the controller is configured to receive, from a second node of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second node, compute, on the basis of the information representing the TTI of the second node and information representing a TTI of the first node, a duration for a common time slot for the first node and the second node, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second node and the TTI of the first node, receive, from the second node, information representing a macro reward for the second node, wherein the macro reward for the second node comprises a measure associated with performance by the second node, from the beginning of the common time slot and during the common time slot, of at least one action, compute, using the macro reward for the second node and a macro reward for the first node, a global macro reward, wherein the macro reward for the first node comprises a measure associated with performance by the first node, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second nodes to create a single reward associated to the first node and second node for the common time slot, and update a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

[0012] A controller at a first node can therefore be used for coordination at (re)initialization to synchronize the start of a training phase and to define a common macro time slot. Each node exchanges, with its neighbours, its local macro rewards and (re initialization messages.

[0013] Accordingly, bandwidth allocation when nodes have synchronous / asynchronous transmission time intervals can be optimised, and a stable distributed policy can be implemented since communication with neighbours helps to manage the interference between the nodes which improves global performance.

[0014] In an implementation of the first aspect, the controller can receive, from the second node, information representing a starting time selected by the second node for a training phase of the neural network, and select, on the basis of the starting time selected by the second node and a starting time selected by the first node, a common starting time for a training phase of the neural network. The controller can transmit, to the second node, the TTI of the first node and information representing the starting time selected by the first node for a training phase of the neural network. The controller can transmit, to the second node, the macro reward for the first node. The controller can determine an initial state of a set of devices served by the first node at the beginning of the common time slot, determine a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first node of the at least one action, and update the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state.

[0015] In an example, the controller select, using the neural network, the at least one action to be performed by the first node on the basis of the initial state of the set of devices served by the first node.

[0016] A second aspect of the present disclosure provides a first base station in a wireless telecommunication system, wherein the first base station comprises a controller comprising a neural network configured to implement a decentralised resource allocation policy for the first base station, wherein the controller is configured to receive, from a second base station of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second base station, compute, on the basis of the information representing the TTI of the second base station and information representing a TTI of the first base station, a duration for a common time slot for the first base station and the second base station, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second base station and the TTI of the first base station, receive, from the second base station, information representing a macro reward for the second base station, wherein the macro reward for the second base station comprises a measure associated with performance by the second base station, from the beginning of common time slot and during the common time slot, of at least one action, compute, using the macro reward for the second base station and a macro reward for the first base station, a global macro reward, wherein the macro reward for the first base station comprises a measure associated with performance by the first base station, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second base stations to create a single reward associated to the first base station and second base station for the common time slot, and update a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward. In an implementation of the second aspect, the controller receive, from the second base station, information representing a starting time selected by the second base station for a training phase of the neural network, and select, on the basis of the starting time selected by the second base station and a starting time selected by the first base station, a common starting time for a training phase of the neural network. The controller can transmit, to the second base station, the TTI of the first base station and information representing the starting time selected by the first base station for a training phase of the neural network. The controller can transmit, to the second base station, the macro reward for the first base station. The controller can determine an initial state of a set of devices served by the first base station at the beginning of the common time slot, determine a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first base station of the at least one action, and update the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state. The controller can select, using the neural network, the at least one action to be performed by the first base station on the basis of the initial state of the set of devices served by the first base station.

[0017] A third aspect of the present disclosure provides a method for configuring a decentralised resource allocation policy in a wireless telecommunication system, the method comprising, at a first base station of the wireless telecommunication system receiving, from a second base station of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second base station, computing, on the basis of the information representing the TTI of the second base station and information representing a TTI of the first base station, a duration for a common time slot for the first base station and the second base station, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second base station and the TTI of the first base station, receiving, from the second base station, information representing a macro reward for the second base station, wherein the macro reward for the second base station comprises a measure associated with performance by the second base station, from the beginning of the common time slot and during the common time slot, of at least one action, computing, using the macro reward for the second base station and a macro reward for the first base station, a global macro reward, wherein the macro reward for the first base station comprises a measure associated with performance by the first base station, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second base stations to create a single reward associated to the first base station and second base station for the common time slot, and updating a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

[0018] In an implementation of the third aspect, the method can further comprise receiving, from the second base station, information representing a starting time selected by the second base station for a training phase of the neural network, and selecting, on the basis of the starting time selected by the second base station and a starting time selected by the first base station, a common starting time for a training phase of the neural network. The method can further comprise transmitting, to the second base station, the TTI of the first base station and information representing the starting time selected by the first base station for a training phase of the neural network. The method can further comprise transmitting, to the second base station, the macro reward for the first base station. The method can further comprise determining an initial state of a set of devices served by the first base station at the beginning of the common time slot, determining a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first base station of the at least one action, and updating the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state. The method can further comprise selecting, using the neural network, the at least one action to be performed by the first base station on the basis of the initial state of the set of devices served by the first base station.

[0019] These and other aspects of the invention will be apparent from the embodiment(s) described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order that the present disclosure may be more readily understood, embodiments will now be described, by way of example, with reference to the accompanying drawings, in which:

[0021] Figure 1 is a schematic representation of a BS (node), according to an example;

[0022] Figure 2 is a schematic representation of initialisation and training aspects for a pair of nodes, according to an example; Figure 3 is a schematic representation of initialisation and training aspects for a trio of nodes, according to an example; Figure 4 is a flowchart of a method for configuring a decentralised resource allocation policy in a wireless telecommunication system, according to an example; and

[0023] Figure 5 is a schematic representation of a machine according to an example.

[0024] DETAILED DESCRIPTION

[0025] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.

[0026] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.

[0027] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof. The term “and / or” is only an association relationship for describing associated objects and represents that three relationships may exist such that A and / or B may indicate that A exists alone, A and B exist at the same time, or B exists alone. The character “ / ” generally represents that the associated objects are in an “or” relationship.

[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.

[0029] The following contains specific information related to implementations of the present disclosure. The drawings and their accompanying detailed disclosure are merely directed to implementations. However, the present disclosure is not limited to these implementations. Other variations and implementations of the present disclosure will be obvious to those skilled in the art. The phrases “in one implementation,” or “in some implementations,” may each refer to one or more of the same or different implementations. The term “coupled” is defined as connected whether directly or indirectly through intervening components and is not necessarily limited to physical connections. The expression “at least one of A, B and C” or “at least one of the following: A, B and C” means “only A, or only B, or only C, or any combination of A, B and C.”

[0030] The terms “system” and “network” may be used interchangeably.

[0031] The terms “neural network” and “deep neural network” or “DNN” may be used interchangeably.

[0032] For the purposes of explanation and non-limitation, specific details such as functional entities, techniques, protocols, and standards are set forth for providing an understanding of the present disclosure. In other examples, detailed disclosure of well-known methods, technologies, systems, and architectures are omitted so as not to obscure the present disclosure with unnecessary details.

[0033] Persons skilled in the art will immediately recognize that any network function(s) or algorithm(s) disclosed may be implemented by hardware, software or a combination of software and hardware. Disclosed functions may correspond to modules which may be software, hardware, firmware, or any combination thereof.

[0034] A software implementation may include machine- and / or computer- readable and / or executable instructions stored on a machine- and / or computer-readable medium such as memory or other types of storage devices. One or more microprocessors or general-purpose computers with communication processing capability may be programmed with corresponding executable instructions and perform the disclosed network function(s) or algorithm(s).

[0035] The microprocessors or general-purpose computers may include Applications Specific Integrated Circuitry (ASIC), programmable logic arrays, and / or using one or more Digital Signal Processor (DSPs). Although some of the disclosed implementations are oriented to software installed and executing on computer hardware, alternative implementations implemented as firmware or as hardware or as a combination of hardware and software are well within the scope of the present disclosure. The computer readable medium includes but is not limited to Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory, Compact Disc Read-Only Memory (CD-ROM), magnetic cassettes, magnetic tape, magnetic disk storage, or any other equivalent medium capable of storing computer-readable instructions.

[0036] Resource allocation in traditional wireless networks with many BSs is based on solutions that use only local information, namely each BS uses state information of its associated devices. Although these policies are distributed and easy to implement, they come with certain disadvantages. Specifically, they ignore the fact that the environment’s behaviour may exhibit some sort of statistical pattern (e.g., channel gains may change somewhat periodically) which, if taken correctly into account, could improve the performance dramatically.

[0037] The latter reason naturally leads to solutions that use Reinforcement Learning (RL), a branch of machine learning where agents learn good policies, by interacting with the true environment, without making assumptions on the wireless network and its behaviour. However, standard RL does not generalize when agents take asynchronous actions. In order to deal with this challenge, the concept of macro action was proposed. Compared to standard RL agents that decide the action at the beginning of each time slot, the macro action concept allows agents to decide a deterministic sequence of actions for several consecutive time slots. However, the standard macro action algorithms for multi-agent settings assume that agents have the same time slot duration, i.e., the same TTI. Therefore, they cannot be directly applied for wireless scenario since BSs may have asynchronous TTIs, such as when they use different numerologies. According to an example, by considering wireless networks with asynchronous TTIs and by taking interference into account, a method and a system are provided that learn decentralized policies represented by a Deep Neural Network (DNN) at each BS.

[0038] That is, a system and method are provided for resource allocation in wireless environment with asynchronous TTIs. A controller at the BS / node is used to implement signalling between BSs for efficient learning in order to enable coordination / agreement at (re)initialization to synchronize the start of training and to define a common macro time slot. BS i sends an (re)initialization message mi= {Start_t0i, σito its neighbour nodes and receives theirs. There is then an exchange of local macro rewards at the end of each macro time slot during training.

[0039] For a BS i, we define the following quantities:

[0040] • N: neighboring BSs including i

[0041] • σi: index of the deployed numerology

[0042] • μσi: deployed numerology

[0043] • local time slot, its duration is equal to the TTI of the chosen numerology,

[0044]

[0045] • h macro time slot, its duration is denoted by H. It is composed of ntlocal time slots (tl,..., t"1), where n,- =

[0046]

[0047] • of local macro resource allocation action. It is composed of a sequence of n,- local actions

[0048]

[0049] ...,

[0050]

[0051] • Rhi: local macro reward. It is defined based on the received local rewards (rh,1i,..., rh,ni) during h

[0052] • Rh: group macro reward; the group macro reward is defined based on all local macro rewards {Rhm}m∈N

[0053] • Thi: local macro transition; Thi= (shi, ahi, Rh, sh+1i), where shiis the local state at the beginning of the macro time slot h, i.e., the local monitored states of its associated devices and sh+1iis the new local state resulting from taking local macro action ahiat local state shi.

[0054] • θi: parameters of the DNN policy

[0055] Figure 1 is a schematic representation of a BS (node), according to an example. The node 101 comprises:

[0056] • Experience buffer 103, which stores macro transitions.

[0057] • Reward module 105, which has two roles:

[0058] a) Store the nilocal rewards (rh,1i,..., rh,ni) received during h

[0059] b) Compute the local macro reward Rhi= fi(rh,1i,..., rh,ni) and the global macro reward Rh= g({Rhm}m∈N) at the end of h, where:

[0060] fi is a mapping function that maps n, rewards received during h to a single reward representing the reward of the macro time slot. The choice of f depends on the metric and the number of local time slots ntduring h, e.g., if the reward is associated to throughput then ftcan be defined

[0061]

[0062] • g is a common function for all BSs that computes the global reward. The choice of g depends on the metric, e.g., if the global reward is defined as the global throughput then

[0063]

[0064] •)=S yew, Rj1

[0065] • Controller 107, which has the following roles:

[0066] a) Select the local numerology

[0067]

[0068] b) Send the local numerology index σito neighboring nodes and receive theirs (σ-i)

[0069] c) Compute the duration of the macro time slot h and the number n,- of local time slot during h. d) Decide the local macro action a1-' = (a-1,1,...,

[0070]

[0071] via the DNN policy, at the beginning of the macro time slot based on the local state s1-'

[0072] e) Play the local macro action during h

[0073] f) Exchange local macro rewards with neighbouring BSs and then compute the group macro reward. g) Train the DNN policy θibased on the stored macro transitions, e.g. with actor-critic algorithms.

[0074] According to an example, there are three phases that can be implemented using controller 107 in order to implement a decentralised resource allocation policy for the node 101:

[0075] 1. Initialization begins with a handshake signalling process to establish the start of learning. In this phase, node 101 synchronizes the start of training with another node and defines the macro time slot h. For example, node 101 (BS i) sends an (re initialization message to its neighbour nodes and receives theirs. The initialization message is composed of a suggested time to start the training and the index of the local numerology. An example of an initialization message m;sent from BS i to its neighbors is:

[0076] mi= {Start_t0i, σi

[0077] • The training starts at the latest suggested starting time

[0078] Start

[0079]

[0080] • The macro time slot duration is defined as the largest duration of local time slots

[0081] H -.= maxfTTI

[0082]

[0083] 2. Learning of the decentralized policies for asynchronous TTIs: Learn a single DNN for each BS that can perform well for wireless networks with asynchronous TTIs.

[0084] I. At the beginning of each macro time slot h, the controller 107 decides the local macro action af =

[0085]

[0086] via its DNN policy θi(109) based on the local state shi II. During the macro time slot which is composed of ntconsecutive local time slots (t ’1,..., tk,n‘), the controller 107 plays the local macro action a'1= (a?1,1,..., c 'n‘): The macro action is composed of n,- local actions. These actions will be played sequentially during the macro time slot: the kthaction ak,k113 for fc G {1,...,n,} is played at the beginning of the kthlocal time slot tk,k. Moreover, at the end of tk,k, BS i receives a local reward r'1,k115 from the environment 117 resulting from playing action ak,k113.

[0087] III. At the end of the macro time slot, the controller 107 computes the local macro reward / ??1= f(rk'1,...,rk,ni'), then sends it to neighbour BSs and receives theirs (using, e.g., the transmission / reception module 111) in order to compute the global macro reward Rh= 5({ / ?mlmew )■ Then, the controller 107 observes the new local state and stores the macro transition

[0088]

[0089] in the experience buffer 103.

[0090] IV. Every M macro time slots, the controller updates θibased on sampled transitions from the experience buffer 103.

[0091] 3. Inference: This phase starts at the end of the learning process described above. At the beginning of each macro time slot and based on the local state sk, the controller 107 decides the local macro action a1-' = ak,1>...,

[0092]

[0093] via its DNN policy θi109. Then, during the macro time slot, the controller 107 plays (or executes) the local macro action ak. This step ends if any BS changes its numerology. In this situation, it sends a re-initialization message to its neighbours and the learning process is restarted.

[0094] Figure 2 is a schematic representation of initialisation and training aspects for a pair of nodes, according to an example. With reference to the example of figure 2, a representative implementation in a cellular network environment for an uplink scheduling scenario is described, where each BS i has to allocate to its Ntassociated UEs a bandwidth SB composed of k^i blocks at the beginning of each local time slot with a duration equal to TTI Zg,).

[0095] Each UE n that belongs to BS i causes interference to neighboring BSs TV). After taking a local action ak,k, BS i obtains a reward rk,kdefined as the summation of throughputs of its associated UEs. At the end of the macro time slot, each BS i

[0096] computes its local macro reward defined as Rhi= (1 / ni) Σ rh,ki, it sends it to its neighbors and receives theirs in order to

[0097]

[0098] compute a global reward:

[0099]

[0100] With reference to figure 2a, during initialization, nodes 201, 203 agree on the duration of the macro time slot and when to start the training. In this example H = max{TTI(μ0), TTI(μ1)} = TTI(μ0).

[0101] With reference to figure 2b, during training, for each macro time slot: • At the beginning of the macro time slot, node (BS) i 201 with a local time slot duration equal to TTI( / z0) decides a macro action composed of a single action a'1= {a1,1} since TTI( / z0) = H. Node (BS) j 203 with a local time slot duration equal to TTI q) decides a macro action composed of two

[0102]

[0103] • During each macro time slot, node i 201 plays its local macro action a'1and receives a single reward r1,1associated to the played action. Node j 203 plays its local macro action aj1and receives a reward for each played action i.e., node j 203 receives two rewards {r ’1,r ’2}.

[0104] • At the end of the macro time slot, the nodes 201, 203 exchange local macro rewards i.e., node i 201 sends Rhi= rh,1ito node j 203 and node j 203 sends Rhj=

[0105]

[0106] + r^2) to node i 201. Then each node computes the global macro reward Rh= Rhi+ Rhjobserves the new local macro state and stores the local macro transition in its buffer.

[0107] After each M macro time slots: each node updates its DNN policy based on sampled macro transitions from its experience buffer 103.

[0108] The proposed solution can be easily extended to any network topology, e.g., the previous configuration can be extended by adding a third node.

[0109] Figure 3 is a schematic representation of initialisation and training aspects for a trio of nodes, according to an example.

[0110] With reference to figure 3a, in the initialization, nodes 201, 203, 301 agree on the duration of the macro time slot and when to start the training. In this example H = max{TTI(μ0), TTI(μ1), TTI(μ2)} = TTI(μ0).

[0111] With reference to figure 3b, during training:

[0112] For each macro time slot:

[0113] • At the beginning of each macro time slot, node i 201 and node j 203 operate as in the previous setting, node k 301 with a local time slot duration equal to TTI(μ2), decides a local macro action composed of four actions

[0114]

[0115] since H = TTI(μ0) = 4TTI(μ2).

[0116] • During each macro time slot, node i 201 and node j 203 operate as in the previous setting. Node k 301 plays its local macro action akand receives a reward for each played action i.e., node k 301 receives four rewardsSrh,l h,2 h,3 l

[0117] t' / c r

[0118] • At the end of the macro time slot, the nodes exchange local macro rewards i.e., node i 201 sends Rhi= rh,1ito

[0119] node j 203 and node k 301, node j 203 sends Rhj= (1 / 2)(rh,1j+ rh,2j) to node i 201 and node k 301; and node k 301

[0120] sends Rhk= (1 / 4)(rh,1k+ rh,2k+ rh,3k+ rh,4k) to node i 201 and node j 203. Then, each node computes the global macro reward Rh= Rhi+ Rhj+ Rhk, observes the new local state and stores the local macro transition in its buffer.

[0121] After each M macro time slots: each node updates its DNN policy based on sampled macro transitions from its experience buffer. With reference to figures 2 and 3, a node 201, 203, or 301 can therefore comprise an apparatus for a controller 107. For example, referring to figure 2b, node 201 can comprise a first node of a wireless telecommunication system, and the controller 107 of the first node 201 can comprise a neural network 109 configured to implement a decentralised resource allocation policy for the first node 201. The controller 107 can receive, from a second node 203 of the wireless telecommunication system, information representing a TTI of the second node 203 and compute, on the basis of the information representing the TTI of the second node 203 and information representing a TTI of the first node 201, a duration H for a common time slot for the first node 201 and the second node 203. The common time slot can comprise a duration that is a maximum selected from the TTI of the second node and the TTI of the first node, as described above.

[0122]

[0123] Figure 4 is a flowchart of a method for configuring a decentralised resource allocation policy in a wireless telecommunication system, according to an example. In the example of figure 4, the method comprises, at a first base station 201 of the wireless telecommunication system receiving, in block 401, from a second base station 203 of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second base station 203.

[0124] In block 403, the first base station 201, via its controller 107, computes, on the basis of the information representing the TTI of the second base station and information representing a TTI of the first base station, a duration for a common time slot for the first base station and the second base station, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second base station and the TTI of the first base station.

[0125] In block 405, the first base station 201 receives, from the second base station, information representing a macro reward for the second base station, wherein the macro reward for the second base station comprises a measure associated with performance by the second base station, from the beginning of the common time slot and during the common time slot, of at least one action.

[0126] In block 407, the first base station 201 computes, using the macro reward for the second base station and a macro reward for the first base station, a global macro reward, wherein the macro reward for the first base station comprises a measure associated with performance by the first base station, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second base stations to create a single reward associated to the first base station and second base station for the common time slot.

[0127] In block 409, the first base station 201 updates a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward. Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.

[0128] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.

[0129] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, a platform comprising user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a NN 109, and / or a rewards module 105, and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.

[0130] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non-transitory computer readable storage medium encoded with instructions, executable by a processor.

[0131] Figure 5 is a schematic representation of a machine according to an example. The machine 500 can be, e.g., a system or apparatus, user equipment, or part thereof, or a controller 107 (or part thereof) for a base station or node. The machine 500 comprises a processor 503, and a memory 505 to store instructions 502, executable by the processor 503. The machine comprises a storage 509 that can be used to store data 511 representing DNN parameters, buffered information, rewards, actions, policies and so on as described above with reference to figures 1 to 4 for example.

[0132] The instructions 502, executable by the processor 503, can cause the machine 500, which can be provided as or as part of a first node of a wireless telecommunication system, to receive, from a second node of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second node, compute, on the basis of the information representing the TTI of the second node and information representing a TTI of the first node, a duration for a common time slot for the first node and the second node, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second node and the TTI of the first node, receive, from the second node, information representing a macro reward for the second node, wherein the macro reward for the second node comprises a measure associated with performance by the second node, from the beginning of the common time slot and during the common time slot, of at least one action, compute, using the macro reward for the second node and a macro reward for the first node, a global macro reward, wherein the macro reward for the first node comprises a measure associated with performance by the first node, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second nodes to create a single reward associated to the first node and second node for the common time slot, and update a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

[0133] Accordingly, the machine 500 can implement a method for generating a decentralised resource allocation policy for a base station.

[0134] Such machine-readable instructions may also be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.

[0135] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.

[0136] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.

[0137] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer-readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.

[0138] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.

Claims

CLAIMS1. Apparatus for a controller of a first node of a wireless telecommunication system, wherein the controller comprises a neural network configured to implement a decentralised resource allocation policy for the first node, wherein the controller is configured to:receive, from a second node of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second node;compute, on the basis of the information representing the TTI of the second node and information representing a TH of the first node, a duration for a common time slot for the first node and the second node, wherein the common time slot comprises a duration that is a maximum selected from the TTI of the second node and the TTI of the first node;receive, from the second node, information representing a macro reward for the second node, wherein the macro reward for the second node comprises a measure associated with performance by the second node, from the beginning of the common time slot and during the common time slot, of at least one action;compute, using the macro reward for the second node and a macro reward for the first node, a global macro reward, wherein the macro reward for the first node comprises a measure associated with performance by the first node, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second nodes to create a single reward associated to the first node and second node for the common time slot; andupdate a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

2. The apparatus as claimed in claim 1, wherein the controller is further configured to:receive, from the second node, information representing a starting time selected by the second node for a training phase of the neural network; andselect, on the basis of the starting time selected by the second node and a starting time selected by the first node, a common starting time for a training phase of the neural network.

3. The apparatus as claimed in claim 2, wherein the controller is further configured to:transmit, to the second node, the TTI of the first node and information representing the starting time selected by the first node for a training phase of the neural network.

4. The apparatus as claimed in any preceding claim, wherein the controller is further configured to:transmit, to the second node, the macro reward for the first node.

5. The apparatus as claimed in any preceding claim, wherein the controller is further configured to:determine an initial state of a set of devices served by the first node at the beginning of the common time slot;determine a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first node of the at least one action; andupdate the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state.

6. The apparatus as claimed in claim 5, wherein the controller is further configured to:select, using the neural network, the at least one action to be performed by the first node on the basis of the initial state of the set of devices served by the first node.

7. A first base station in a wireless telecommunication system, wherein the first base station comprises a controller comprising a neural network configured to implement a decentralised resource allocation policy for the first base station, wherein the controller is configured to:receive, from a second base station of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second base station;compute, on the basis of the information representing the TTI of the second base station and information representing a TTI of the first base station, a duration for a common time slot for the first base station and the second base station, wherein the common time slot comprises a duration that is a maximum selected from the TH of the second base station and the TTI of the first base station;receive, from the second base station, information representing a macro reward for the second base station, wherein the macro reward for the second base station comprises a measure associated with performance by the second base station, from the beginning of common time slot and during the common time slot, of at least one action;compute, using the macro reward for the second base station and a macro reward for the first base station, a global macro reward, wherein the macro reward for the first base station comprises a measure associated with performance by the first base station, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second base stations to create a single reward associated to the first base station and second base station for the common time slot; and update a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

8. The first base station as claimed in claim 7, wherein the controller is further configured to: receive, from the second base station, information representing a starting time selected by the second base station for a training phase of the neural network; andselect, on the basis of the starting time selected by the second base station and a starting time selected by the first base station, a common starting time for a training phase of the neural network.

9. The first base station as claimed in claim 8, wherein the controller is further configured to: transmit, to the second base station, the TTI of the first base station and information representing the starting time selected by the first base station for a training phase of the neural network.

10. The first base station as claimed in any of claims 7 to 9, wherein the controller is further configured to:transmit, to the second base station, the macro reward for the first base station.

11. The first base station as claimed in any of claims 7 to 10, wherein the controller is further configured to:determine an initial state of a set of devices served by the first base station at the beginning of the common time slot;determine a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first base station of the at least one action; andupdate the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state.

12. The first base station as claimed in claim 11, wherein the controller is further configured to:select, using the neural network, the at least one action to be performed by the first base station on the basis of the initial state of the set of devices served by the first base station.

13. A method for configuring a decentralised resource allocation policy in a wireless telecommunication system, the method comprising, at a first base station of the wireless telecommunication system:receiving, from a second base station of the wireless telecommunication system, information representing a Transmission Time Interval, TTI, of the second base station;computing, on the basis of the information representing the TTI of the second base station and information representing a TTI of the first base station, a duration for a common time slot for the first base station and the second base station, wherein the common time slot comprises a duration that is a maximum selected from the TH of the second base station and the TTI of the first base station;receiving, from the second base station, information representing a macro reward for the second base station, wherein the macro reward for the second base station comprises a measure associated with performance by the second base station, from the beginning of the common time slot and during the common time slot, of at least one action;computing, using the macro reward for the second base station and a macro reward for the first base station, a global macro reward, wherein the macro reward for the first base station comprises a measure associated with performance by the first base station, from the beginning of the common time slot and during the common time slot, of at least one action, and wherein the global macro reward is determined using a function that combines the macro rewards for the first and second base stations to create a single reward associated to the first base station and second base station for the common time slot; andupdating a set of parameters associated with the neural network to modify the decentralised resource allocation policy on the basis of the global macro reward.

14. The method as claimed in claim 13, further comprising:receiving, from the second base station, information representing a starting time selected by the second base station for a training phase of the neural network; andselecting, on the basis of the starting time selected by the second base station and a starting time selected by the first base station, a common starting time for a training phase of the neural network.

15. The method as claimed in claim 14, further comprising:transmitting, to the second base station, the TTI of the first base station and information representing the starting time selected by the first base station for a training phase of the neural network.

16. The method as claimed in any of claims 13 to 15, further comprising:transmitting, to the second base station, the macro reward for the first base station.

17. The method as claimed in any of claims 13 to 16, further comprising:determining an initial state of a set of devices served by the first base station at the beginning of the common time slot;determining a subsequent state of the set of devices served by the first node at the end of the common time slot following performance by the first base station of the at least one action; andupdating the set of parameters associated with the neural network based on a set of transitions, wherein a transition comprises a tuple comprising the initial state, the at least one action, the global macro reward and the subsequent state.

18. The method as claimed in claim 17, further comprising:selecting, using the neural network, the at least one action to be performed by the first base station on the basis of the initial state of the set of devices served by the first base station.