System and reinforcement learning method for cooperative multi-agent multiobjective radio resource allocation in wireless networks

A reinforcement learning method using a single DNN at each access point optimizes radio resource allocation across multiple access points, addressing interference and scalability issues to enhance network performance in wireless cellular communication networks.

WO2025247469A1PCT designated stage Publication Date: 2025-12-04HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/064499
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing wireless network management systems face challenges in optimizing radio resource allocation across multiple access points due to interference and scalability issues, leading to suboptimal network performance and inefficiencies in throughput, latency, reliability, and energy consumption.

Method used

Implementing a reinforcement learning method using a single deep neural network (DNN) at each access point to manage interference by generating and optimizing resource allocation policies based on local and group preferences, allowing for scalable and adaptable network performance across diverse interference scenarios.

Benefits of technology

The solution provides stable, distributed policies that minimize hardware requirements and improve global network performance by optimizing interference management across various interference group preferences, enhancing throughput, latency, reliability, and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024064499_04122025_PF_FP_ABST
    Figure EP2024064499_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments herein provide a method for an access point, AP, configured to operate in a wireless cellular communication network including one or more access points, APs, belonging to at least two interference groups. The AP includes a communication interface 206, 516, 616, 716. The method includes receiving a local preference vector for each of the APs of all other interference group members (I), through the communication interface. The method includes generating a local preference vector (II) for the AP based on AP processing needs, where i indicates the AP, AP i . The method includes generating an interference group preference vector (III), the interference group preference vector (III) being based on the local preference vectors (IV) and the local preference vector (II). The method includes determining a local resource allocation action (V) ) by applying an actor deep neural network, DNN, on a local state (VI) and the interference group preference vector.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]SYSTEM AND REINFORCEMENT LEARNING METHOD FOR COOPERATIVE MULTI-AGENT MULTI-OBJECTIVE RADIO RESOURCE ALLOCATION IN WIRELESS NETWORKSTECHNICAL FIELDThe disclosure relates to wireless network management and optimization and more particularly, the disclosure relates to anaccess point, AP, configured to operate in a wireless cellular communication network including one or more access points,APs, belonging to at least two interference groups. The disclosure relates to a method for an access point, AP, configured tooperate in a wireless cellular communication network including one or more access points, APs, belonging to at least twointerference groups.BACKGROUNDIn wireless network environments, an access point, AP, manages radio resources among multiple connected devices to provideefficient communication. The AP achieves optimized performance by allocating the radio resources to the multiple connecteddevices using device-specific state information, such as queue length, channel gain, etc. Network operators are tasked withoptimizing network performance metrics, such as throughput, latency, reliability, energy conservation, and coverage.Traditionally, optimized network performance metrics are achieved manually by adjusting the AP's radio resource allocationpolicies. However, such manual adjustments often result in suboptimal performance outcomes. There is an absence ofautomated solutions that enable operators to fine-tune the radio resource allocation policies.FIG. 1 illustrates an exemplary representation of multi-objective distributed radio resource allocation in a wireless cellularcommunication network 100 in accordance with a prior art. The wireless cellular communication network 100 includes accesspoints, APs 108A and 102B. The APs 108A and 102B include base stations 104A, and 104B, and two scheduled devices 102A,and 102B. In environments with a large number of operational APs, consistent communication among the operational APsbecomes impractical due to complexity and resource demands. Therefore, the APs 108A and 108B allocate the radio resourcesfrom the base stations 104A, and 104B by relying only on local information from corresponding scheduled devices 102A, and102B, a scalable approach that minimizes communication overhead among the APs 108A, 108B. However, the scalableapproach that includes distributed policies may result in situations where the APs 108A and 108B independently optimize fortheir efficiency a condition often termed as 'acting selfishly'. This condition of acting selfishly by the APs 108A, and 108B cancause significant interference 106, as depicted in FIG. 1 among the APs 108A, and 108B . The interference 106 severelydegrades achieving the network performance objectives including throughput, latency, reliability, energy efficiency, andcoverage. The visual representation of the interference 106 among the APs 108A, and 108B underscores the detrimental effectsof autonomous actions of the APs 108A and 108B on the network’s overall functionality.The approach may employ policies such as proportional fairness, and max weight. The proportional fairness is a scheduler thatallocates fair service rates among the devices associated to same AP, serving as a baseline for performance comparisons. Themax weight is a mechanism directed at queue stability, where unserved traffic from the scheduled devices 102A, and 102B,await transmission in queues. While these policies optimize specific objectives, the policies lack easy adjustability for othermetrics. When there is high interference among the APs (e.g., 108A, 108B), the policies become less effective. This issuecompels the adoption of algorithms that incorporate, at a minimum, some level of communication overhead to manage the highinterference effectively. Moreover, channel gains may exhibit statistical patterns, for example, periodic changes, that thesepolicies often overlook. By accurately accounting for the statistical patterns could significantly enhance performance.A machine learning methodology, Reinforcement Learning, RL, is provided where the APs 108A and 108B can directly learnoptimal policies through interaction with the actual environment, without preconceived assumptions about behavior of thewireless cellular communication network 100. This includes (a) multi-agent RL, MARL, which effectively managesinterference 106, with robust coordination and learning stability; and (b) multi-objective RL, MORL that explicitly models theprioritization of the multiple objectives.In the MARL approach, the APs 108A, and 108B interact with the environment and exchange information during training tocoordinate actions to learn stable, distributed policies for resource allocation that optimize a single objective or a particularcombination of the objectives over the long term. However, the MARL approach does not address the scalability issue acrossthe objectives, as the learning process must be repeated for each new objective or combination thereof.A challenge with the MORL is the scalability of the solution to different operator-defined priorities / preferences. For example,with two objectives, e.g., latency and throughput, and three services, each with different prioritization for the two objectives,denoting the objective vector as ^ ∈ ℝ^ and the vector of preferences / prioritizations across the objectives as ^ ∈ ℝ^, aimingto optimize functions ^^^(^), ^^^(^), ^ where ^^^(^) = ^^ ⋅ ^. Training three different neural networks for eachfunction ^^^(^) is inefficient, and unscalable, in terms of time and space, for numerous preferences / prioritizations.In the MORL approach, the agent interacts with the environment and tries to learn a single policy capable of performing wellacross any combination of the objectives, as explicitly encoded in the learning algorithm through a single policy neural network.It is obvious that without proper coordination, the MORL approach may lead to unstable learning.Therefore, there arises a need to address the aforementioned technical problem / drawbacks for providing a secure copy-pastein open environments.SUMMARYIt is an object of the disclosure to provide an access point, AP, configured to operate in a wireless cellular communicationnetwork including one or more access points, APs, belonging to at least two interference groups, and a method for an accesspoint, AP, configured to operate in a wireless cellular communication network including one or more access points, APs,belonging to at least two interference groups.This objective is achieved by the features of the independent claims. Further, implementation forms are apparent from thedependent claims, the description, and the figures.According to a first aspect, there is an access point, AP, configured to operate in a wireless cellular communication network.The wireless cellular communication network includes one or more access points, APs belonging to at least two interferencegroups. The AP includes a controller and a communication interface. The controller is configured to receive a local preferencevector for each of the APs of all other interference group ^^^ through the communication link, where ^ indicates a time slot,and ^ indicates one of the APs of all other interference group , AP^. The controller is configured to generate a local preferencevector ^^^ for the AP based on AP processing needs, where ^ indicates the AP, AP^. The controller is configured to generate aninterference group preference vector ^^ . The interference group preference vector ^^being based on the local preferencevectors ^ ^ ^ ^ ^^^ ^^^ and the local preference vector ^^.^ = f(^^, ^^^ ) and where ^^^ is the set of all local preference vectors of allother interference group members. The controller is configured to determine a local resource allocation action ^^^ by applyingan actor deep neural network, DNN, on a local state ^^ ^^ and the interference group preference vector ^ .In the wireless cellular communication network, each AP includes a single DNN that selects the best local actions that optimizethe interference group performance for a given interference group preference vector. The training of the DNN is done oncefor all the interference group preference vector space, without the need for multiple DNNs tailored to each interference grouppreference vector, thereby minimizing hardware requirements at each AP. The controller provides stable distributed policiesto manage the interference between the APs, thereby improving global performance. The controller provides distributedpolicies over the entire space of interference group preferences, leading to optimal policies for a diverse range of interferencegroup preference vectors. The controller is adaptable to a wide variety of resource allocation tasks such as power allocation,modulation, and coding schemes.Optionally, the controller is further configured to determine which interference group the AP belongs to.Optionally, the controller is further configured to transmit the local preference vector to all other interference group members.Optionally, the controller is further configured to determine the local state based on a local environment.Optionally, the controller is further configured to determine the local state by receiving the local state from the localenvironment.Optionally, the controller is further configured to generate the interference group preference vector ^^ based on the localpreference and the local preference for the one or more other interference group member by applying a function which isdeterministic and is the same across APs of the same interference group.Optionally, the controller is further configured to generate the interference group preference vector ^^ and determine the localresource allocation action ^^^ for a plurality of time slots. The controller is further configured to monitor a local reward ^^^ andan updated local state ^^^^^ for each generated interference group preference vector ^^ and determined local resource allocationaction ^^^. The controller is further configured to transmit the local reward and the updated local states to the at least one otherinterference group member. The controller is further configured to receive local rewards ^^^ and updated local states ^^^fromthe at least one other interference group member. The controller is further configured to store the local action, an interferencegroup state including states for the AP and the at least one other AP, an interference group reward including rewards for theAP and the at least one other AP and an updated interference group state including updated states for the AP and the at leastone other AP to an experience buffer. The controller is further configured to create a batch, ^^(^^)] of ^^ local preference vectors, where is the generated ^^^ preference vector. The controlleris further configured to receive a local batch composed of ^^ local preference vectors from each of the at least one otherinterference group member. The controller is further configured to create an interference group batch ^^ composed of ^^interference group preference vectors{^^(1), … , ^^(^^)} . The controller is further configured to optimize a Critic DNN througha gradient descent, which Critic DNN is configured to take as input the interference group state ^^ and the interference grouppreference vector ^^, and to output a value ^(^^, ^^ ; ^), which value is defined as the expected discounted sum of futureinterference group reward starting from the interference group state ^ for the interference group preference ^ wherein ^ is a discount factor defined as ^ ∈ [0,1) . The controller is furtherconfigured to feed the Critic DNN with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer and a batch ^^ of ^^interference group preferences. The controller is further configured to update the parameters of the Critic DNN based on thegradient of a loss function. The controller is further configured to feed ^ × ^^ advantage values {Adv^,^}^×^^ to the ActorDNN. The controller is further configured to update the Actor DNN parameters based on the ^ transitions {^^ ^^ , ^^ }^ from theexperience buffer, the batch ^^ of ^^ interference group preference vectors {^^(1), … , ^^(^^)} and the ^ × ^^ advantagevalues {Adv^,^}^×^^ from the Critic DNN.Optionally, the controller is further configured to update the parameters of the Critic DNN based on a gradient of the lossfunction: Optionally, the controller is further configured to update the parameters of the Actor DNN based on a gradient of the lossfunction: log^^ (^^^ , ^^ )^,^ ^ ^|^^ ; ^^ ^Adv .Optionally, the advantage values Adv^,^ are the extra value that can be obtained by choosing action ^^ ^^ at state ^^ for thepreference ^^ compared to ^^^^(^^, ^^).Optionally, the advantage value for a given ^ ∈ {1, … , ^} and ^ ∈ {1, … , ^^}, Adv^,^ is given as. Optionally, the Critic DNN is common to all APs of an interference group.Optionally, the controller is further configured to initialize the Critic DNN with the same parameters ^, as the other APs in theinterference group.Optionally, the controller is further configured to select from Experience buffer the same transitions as the other APs of theinterference group in order to update the parameters of the Actor DNN and the Critic DNN.Optionally, the controller is further configured to exchange data with each of the one or more other APs of the interferencegroup simultaneously.Optionally, the controller is further configured to exchange data with a subset of the one or more other APs of the interferencegroup.Optionally, the controller is further configured to exchange data with a randomly selected AP of the interference group.Optionally, all APs in an interference group cause interference to at least one of the other APs in the interference group, whereinan interference graph is defined as a graph ^ = (^, ^) where ^ is the set of APs, and ^ is the set of directed edges, where anedge ^^→^ suggests that AP ^ causes interference to AP ^.Optionally, the interference groups are defined as the strongly connected components of the interference graph.Optionally, the interference groups are defined as the final coalitions of a merge and split algorithm applied based on a utilityfunction.According to a second aspect, there is provided a method for an access point, AP, configured to operate in a wireless cellularcommunication network including one or more access points, APs, belonging to at least two interference groups. The APbelongs to one of the at least two interference groups. The AP includes a communication interface. The method includesreceiving a local preference vector for each of the APs of all other interference group members ^^^, through the communicationinterface, where ^ indicates a time slot, and ^ indicates one of the APs of all other interference group, AP^ . The method includesgenerating a local preference vector ^^^ for the AP based on AP processing needs, where ^ indicates the AP, AP^. The methodincludes generating an interference group preference vector ^^, the interference group preference vector ^^being based on thelocal preference vectors ^ ^ and the local preference vector ^ ^ ^ ^ ^^^ ^^, wherein ^ = f(^^, ^^^ ), and where ^^^ is the set of all localpreference vectors of all other interference group members. The method includes determining a local resource allocation action(^^^) by applying an actor deep neural network, DNN, on a local state (^^^) and the interference group preference vector.The method selects the best local actions by a single DNN of each AP that optimizes the interference group performance for agiven interference group preference vector in the wireless cellular communication network. The method trains the DNN oncefor all the interference group preference vector space, without the need for multiple DNNs tailored to each interference grouppreference vector, thereby minimizing hardware requirements at each AP. The method provides stable distributed policies tomanage the interference between the APs, thereby improving global performance. The method provides distributed policiesover the entire space of interference group preferences, leading to optimal policies for a diverse range of interference grouppreference vectors. The method is adaptable to a wide variety of resource allocation tasks such as power allocation, modulation,and coding schemes.Optionally, the method includes generating the interference group preference vector ^^ and determining the local resourceallocation action ^^ f ^^ or a plurality of time slots. The method includes monitoring a local reward ^^ and an updated local state^^^^^ for each generated interference group preference vector ^^ and determined local resource allocation action ^^^ . Themethod includes transmitting the local reward and the updated local states to the at least one other interference group member.The method includes receiving local rewards ^^ and updated local states ^^ ^^ from the at least one other interference groupmember. The method includes storing the local action, an interference group state including states for the AP and the at leastone other AP, an interference group reward including rewards for the AP and the at least one other AP and an updatedinterference group state including updated states for the AP and the at least one other AP to an experience buffer. The methodincludes creating a batch = [^^(1), … , ^^(^^)] of ^^ local preference vectors, is the generated^^preference vector. The method includes receiving a local batch composed of ^^ local preference vectors from each of the atleast one other interference group member. The method includes creating an interference group batch ^^ composed of ^^interference group preference vectors {^^(1), … , ^^(^^)} . The method includes optimizing a Critic DNN through a gradientdescent, which Critic DNN is configured to take as input the interference group state ^^ and the interference group preferencevector ^^, and to output a value ^(^^, ^^ ; ^), which value is defined as the expected discounted sum of future interferencegroup reward starting from the interference group state ^ for the interference group preference ^ wherein ^ is a discount factor defined as ^ ∈ [0,1). The method includes feedingCritic DNN with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer and a batch ^^ of ^^ interference group preferences.The method includes updating the parameters of the Critic DNN based on the gradient of a loss function. The method includesfeeding ^ × ^^ advantage values {Adv^,^}^×^^ to the Actor DNN. The method includes updating the Actor DNN parametersbased on the ^ transitions {^^^ , ^^^ }^ from the experience buffer, the batch ^^ of ^^ interference group preference vectors{^^(1), … , ^^(^^)} and the ^ × ^^ advantage values {Adv^,^}^×^^ from the Critic DNN.According to a third aspect, there is provided a computer program product comprising program instructions for performing theabove method, when executed by one or more processors in an Access point.The wireless cellular communication network implements a full communication method, a coordinated communicationmethod, and a gossip method. The full communication method includes APs, where each AP simultaneously exchangesinformation with all other APs. This method ensures that all APs are consistently informed, potentially enhancing overallnetwork synchronization and reducing individual AP misalignments. The coordinated communication method includes APs,where each AP communicates with a selected few neighbors based on a predefined exchange order, thereby reducingcommunication costs. Furthermore, the gossip method includes APs, where each AP communicates with a randomly chosenneighbor at each time slot, suitable for environments with unknown or dynamic network topologies. The gossip methodeliminates the need for a predefined communication order. However, some convergence time is required. If the convergencetime is less than the duration of the time slot, then the interference group preference vector is computed. Each AP starts fromits local preference vector and keeps updating its local preference after exchanging with neighbor APs. The APs keepexchanging until the local preference vectors converge to the interference group preference vector. The interference group stateand reward vectors are constructed by broadcasting the local new state and reward vectors of each AP to all other APs. Thegossip method may be implemented as a rumor spreading approach so that each AP delivers a rumor to all other APs.Therefore, in contradistinction to the existing solutions, the AP is configured to perform the reinforcement learning method forcooperative multi-agent multi-objective radio resource allocation in the wireless cellular communication networks as describedabove.These and other aspects of the disclosure will be apparent from the implementation(s) described below.BRIEF DESCRIPTION OF DRAWINGSImplementations of the disclosure will now be described, by way of example only, with reference to the accompanyingdrawings, in which:FIG. 1 illustrates an exemplary representation of multi-objective distributed radio resource allocation in a wireless cellularcommunication network in accordance with a prior art;FIG. 2 illustrates a block diagram of an access point, AP, configured to operate in a wireless cellular communication networkincluding one or more access points, APs belonging to at least two interference groups in accordance with an implementationof the disclosure;FIG. 3 is an exemplary block diagram that illustrates co-operative multi-objective, multi-agent radio resource allocation inaccordance with an implementation of the disclosure;FIG. 4 is an exemplary block diagram that illustrates a preference module of an access point in accordance with animplementation of the disclosure;FIG. 5 is an exemplary diagram that illustrates collecting data to an experience buffer in accordance with an implementationof the disclosure;FIG. 6 is an exemplary diagram that illustrates optimizing an actor deep neural network, DNN, and a critic DNN in accordancewith an implementation of the disclosure;FIG. 7 is an exemplary diagram that illustrates inference of an access point in accordance with an implementation of thedisclosure; andFIGS. 8A and 8B are flow diagrams that illustrate a method for an access point, AP, configured to operate in a wireless cellularcommunication network including one or more access points, APs, belonging to at least two interference groups in accordancewith an implementation of the disclosure.DETAILED DESCRIPTION OF THE DRAWINGSImplementations of the disclosure provide an access point, AP, configured to operate in a wireless cellular communicationnetwork including one or more access points, APs, belonging to at least two interference groups and a method for the AP,configured to operate in a wireless cellular communication network including one or more access points, APs, belonging to atleast two interference groups.To make solutions of the disclosure more comprehensible for a person skilled in the art, the following implementations of thedisclosure are described with reference to the accompanying drawings. Terms such as "a first", "a second", "a third", and "afourth" (if any) in the summary, claims, and foregoing accompanying drawings of the disclosure are used to distinguish betweensimilar objects and are not necessarily used to describe a specific sequence or order. It should be understood that the terms soused are interchangeable under appropriate circumstances, so that the implementations of the disclosure described herein are,for example, capable of being implemented in sequences other than the sequences illustrated or described herein. Furthermore,the terms "include" and "have" and any variations thereof, are intended to cover a non-exclusive inclusion. For example, aprocess, a method, a system, a product, or a device that includes a series of steps or units, is not necessarily limited to expresslylisted steps or units but may include other steps or units that are not expressly listed or that are inherent to such process, method,product, or device.Definitions: Local state ^^^, represents the state of all connected devices.The policy ^^: it receives local state ^^^ and interference group preference^^, and outputs the resource allocation action ^^^.From the viewpoint of access point, AP ^, ^ ^^^ denotes the collection of all states without the state of AP ^; similarly, Vectorrewards ^ ^^ ^^ and Local preference vector ^^^ other interference group members. The controller 204 is configured to determine a local resource allocation action ^^^byapplying an actor deep neural network, DNN, on a local state ^^ ^^ and the interference group preference vector ^ .Optionally, the controller 204 is further configured to determine which interference group the AP belongs to. The localpreference vector is transmitted to all other interference group.The interference group preference vector ^^ may be generated based on the local preference and the local preference for theone or more other interference group member by applying a function which is deterministic and is the same across APs of thesame interference group.In the wireless cellular communication network, each AP 202A includes a single DNN that selects the best local actions thatoptimize the interference group performance for a given interference group preference vector. The training of the DNN is doneonce for all the interference group preference vector space, without the need for multiple DNNs tailored to each interferencegroup preference vector, thereby minimizing hardware requirements at each AP. The controller 206 provides stable distributedpolicies to manage the interference between the APs, thereby improving global performance. The controller 204 providesdistributed policies over the entire space of interference group preferences, leading to optimal policies for a diverse range ofinterference group preference vectors. The controller 206 is adaptable to a wide variety of resource allocation tasks such aspower allocation, modulation, and coding schemes.The controller 204 provides stable distributed policies for managing the interference between the APs 202A, and 202B,improving global network performance. The distributed policies span over the entire space of group preferences, leading tooptimal policies for a diverse range of interference group preference vectors. The controller 204 is adaptable to a wide varietyof resource allocation tasks including power allocation, modulation, and coding schemes.Optionally, the controller 204 is further configured to generate the interference group preference vector ^^ and determine thelocal resource allocation action ^^^for a plurality of time slots. The controller 204 is further configured to monitor a local reward^^^ and an updated local state ^^^^^ for each generated interference group preference vector ^^ and determined local resourceallocation action ^^^. The controller 204 is further configured to transmit the local reward and the updated local states to the atleast one other interference group member. The controller 204 is further configured to receive local rewards ^^^ and updatedlocal states ^^^ from the at least one other interference group member. The controller 204 is further configured to store the localaction, an interference group state including states for the AP and the at least one other AP, an interference group rewardincluding rewards for the AP and the at least one other AP and an updated interference group state including updated states forthe AP and the at least one other AP to an experience buffer. The controller 204 is further configured to create a batch ^^ =[^^(1), … , ^^(^^)] of ^^ local preference vectors, where ^^(^) is the generated ^^^ preference vector. The controller 204is further configured to receive a local batch composed of ^^ local preference vectors from each of the at least one otherinterference group member. The controller 204 is further configured to create an interference group batch ^^ composed of ^^interference group preference vectors {^^(1), … , ^^(^^)} The controller 204 is further configured to optimize a Critic deepneural network, DNN through a gradient descent, which the Critic DNN is configured to take as input the interference groupstate ^^ and the interference group preference vector ^^, and to output a value ^(^^, ^^ ; ^), which value is defined as theexpected discounted sum of future interference group reward starting from the interference group state ^ for the interferencegroup preference ^ wherein ^ is a discount factor defined as ^ ∈ [0,1). The controller 206 is furtherconfigured to feed the Critic DNN with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer and a batch ^^ of ^^interference group preferences. The controller 206 is further configured to update the parameters of the Critic DNN based onthe gradient of a loss function. The controller 206 is further configured to feed ^ × ^^ advantage values {Adv^,^}^×^^ to theActor DNN. The controller 204 is further configured to update the Actor DNN parameters based on the ^ transitionsfrom the experience buffer, the batch ^^ of ^^ interference group preference vectors {^^(1), … , ^^(^^)} and the^ × ^^ advantage values {Adv^,^}^×^^ from the Critic DNN.In an exemplary embodiment for allocating distributed radio resources for ‘M’ APs with interference graph G in the wirelesscellular communication network, there exist K devices, and each device is associated with the closest AP. ^(^) may be the setof devices associated with AP ^. The controller 204 of the AP 202A may select which of its associated devices to be allocatedin a given time slot. For allocation, both the throughput, THP, for the ‘M’ APs and minimum average throughput T^HP^^^ forthe ‘M’ APs are considered. At time slot ^, prior to action execution, the AP 202A has access to channel states and averagethroughputs for each of its devices at time slot t-1. The throughput THP for the ‘M’ APs is determined by if ^^^ = ^where ^ is the transmission bandwidth, ^ is the transmission power, ^ is the noise power and ^^^^^,^is the channel gain between ^^^(the selected device of AP ^) and AP ^.The minimum average throughput T^HP^^^ for the ‘M’ APs is determined byT^HP^^^ = ∑^ ^ ^^^∑^^^min^∈^(^)^t^h^^p^^ ^FIG. 3is an exemplary block diagram that illustrates co-operative multi-objective, multi-agent radio resource allocation inaccordance with an implementation of the disclosure. In the exemplary block diagram, the co-operative multi-objective, multi-agent radio resource allocation includes access points 300A and 300B. The access point 300A includes a preference module302A, an actor DNN 304A, an experience buffer 306A, and a critic DNN 308A. The access point 300B includes a preferencemodule 302B, an actor DNN 304B, an experience buffer 306B, and a critic DNN 308B. In the exemplary block diagram, (i)the experience buffer 306A, and the critic DNN 308A corresponding to the access point 302A, and (ii) the experience buffer306B, and the critic DNN 308B corresponding to the access point 302B are depicted as dashed blocks or arrows indicating thetraining of the access points 300A and 300B. The preference module 302A, and the actor DNN 304A, corresponding to theaccess point 302A, and the preference module 302B, and the actor DNN 304B corresponding to the access point 302B aredepicted as bolded blocks or arrows depicting the training and the inference of the access points 300A and 300B.In the exemplary block diagram, training and inference between the access points 300A and 300B include generating a localpreference vector ^^^, for the AP 300A by the preference module 302A based on AP processing needs, where ^ indicates theAP, AP^ , and then sends the local preference vector ^^^ to all other APs 300B of the interference group through acommunication interface.In the exemplary block diagram, training of the access points 300A and 300B includes generating an interference grouppreference vector ^^ by the preference module 302A. The interference group preference vector ^^may be generated based onthe local preference vectors ^ and the local preference vector ^^ ^^^ ^^ . ^ = f(^^,^^ ) and where ^^^ is the set of all localpreference vectors of all other interference group members. During training and inference of the access points 300A and 300B,the interference group preference vector ^^ is returned by the preference module 302A based on all local preference vectorsDuring the training of the access points 300A and 300B, the preference module 302A returns a batch ^^ = [^^(1), … ,^^(^^)] of ^^ local preference vectors to all other APs. During inference of the access points 300A and 300B, the preferencemodule 302A returns the local preference vector ^^ ^^. Optionally, a generator returns the local preference vector ^^, or thebatch ^^ = [^^(1), … , ^^(^^)]of ^^ local preference vectors.In the exemplary block diagram, training and inference of the access points 300A and 300B includes determining a localresource allocation action ^^ by the actor DNN 304 based on a local state ^^ and the interference gro ^^ ^ up preference vector ^ .Each AP has its own actor DNN. The local state may be determined based on a local environment. The actor DNN 304 is activeduring training and on inference. The local state may be determined by receiving the local state from the local environment.During the training of the access points 300A and 300B, the experience buffer 306A receives local rewards ^^^ and updatedlocal states ^^^ from the at least one other interference group member. The experience buffer 306A stores transition tuples ofthe form {^^ , ^^^ , ^^ , ^^^^}, where ^^ = ∑^ ^^^ ^^ ^is the interference group reward vector. The experience buffer 306A storesthe local action, an interference group state including states for the AP and the at least one other AP, an interference groupreward including rewards for the AP and the at least one other AP, and an updated interference group state including updatedstates for the AP and the at least one other AP. The experience buffer 306A is active during training.The critic DNN 308A is common to all APs of an interference group. The critic DNN 308A is active during training alone.During training of the access points 300A and 300B, the critic DNN 308A is configured to take the interference group state ^^and the interference group preference vector ^^, as input to output a value ^(^^, ^^ ; ^), which value is defined as the expecteddiscounted sum of future interference group reward starting from the interference group state ^ for the interference grouppreference ^ where ^ is a discount factor defined as ^ ∈ [0,1)FIG. 4 is an exemplary block diagram that illustrates a preference module 400 of an access point in accordance with animplementation of the disclosure. The preference module 400 of the access point includes a generator 404, and a mappingfunction 406. The generator 404 generates a local preference vector ^^^, for the AP by the based on AP processing needs,where ^ indicates the AP, AP ^^, and then sends the local preference vector ^^ to all other APs of the interference group througha communication interface. The generator 404 generates a batch ^^ = [^^(1), … , ^^(^^)] of ^^ local preference vectorsand returns the batch of the local preference vectors to the mapping function 406.The generator 404 returns the local preference vector ^^^ , or the batch ^^ = [^^(1), … , ^^(^^)]of ^^ local preferencevectors to all other APs.The mapping function 406 generates an interference group preference vector ^^ based on the local preference vectors and thelocal preference vector ^^^. ^^= f(^^^^, ^^^ ) and where ^ ^^^ is the set of all local preference vectors of all other interferencegroup members. The mapping function 406 returns the interference group preference vector ^^ or a batch of the interferencegroup preference vector ^^ based on all local preference vectors ^^^ , … , ^^^ to other components of the AP 402.FIG. 5 is an exemplary diagram that illustrates collecting data to an experience buffer 506 in accordance with an implementationof the disclosure. In the exemplary diagram, collecting the data over T timeslots includes gathering transitions for optimizingan actor deep neural network, DNN, 504, and a critic DNN 508. At each timeslot, a preference module 502 generates a localpreference vector ^^^ , for the AP based on AP processing needs, where ^ indicates the AP, AP^, and then sends the localpreference vector ^^^ to all other APs 514 of the interference group through a communication interface 516. A generatorassociated with the preference module 502 may send the local preference vector ^^^ to all other APs 514 of the interferencegroup. The preference module 502 generates an interference group preference vector ^^. The interference group preferencevector ^^is generated based on the local preference vectors ^ ^^^ and the local preference vector ^^^. ^^= f(^^^^, ^^^ ) and where^^^^ is the set of all local preference vectors of all other interference group members. The interference group preference vector^^ maybe an average or weighted average of preference vectors across all other interference group members. The actor DNN504 monitors a local reward ^^ and an updated local state ^^^^for each generated interf ^^ ^ erence group preference vector ^ anda local state ^^^ and determined local resource allocation action ^^^. The local state is determined by receiving the local statefrom a local environment 520. The local reward ^^, the local state ^ ^^^^ ^^ , and the updated local state ^^ are transmitted to the atleast one other interference group member. The local reward ^^ ^ ^^^^ , the local state ^^ , the updated local state ^^ , and thedetermined local resource allocation action ^^^ is called as a local transition tuple. The local action, an interference group stateincluding states for the AP and the at least one other AP, an interference group reward including rewards for the AP and the atleast one other AP, and an updated interference group state including updated states for the AP and the at least one other APare stored to the experience buffer 506.FIG. 6 is an exemplary diagram that illustrates optimizing an actor deep neural network, DNN 604, and a critic DNN 608 inaccordance with an implementation of the disclosure. In the exemplary diagram, the actor DNN 604 is optimized by (i) creatinga batch ^^ = [^^(1), … , ^^(^^)]of ^^ local preference vectors. ^^(^) is the generated ^^^ preference vector by a preferencemodule 602, (ii) receive a local batch composed of ^^ local preference vectors by the AP from each of the at least one otherinterference group member, and (iii) creating an interference group batch ^^ composed of ^^ interference group preferencevectors {^^(1), … , ^^(^^)} by the AP. The AP includes a communication interface 616. The actor DNN 604 parameters areupdated based on the ^ transitions {^^^ , ^^^ }^ from an experience buffer 606, the batch ^^ of ^^ interference group preferencevectors {^^(1), … , ^^(^^)} and the ^ × ^^ advantage values {Adv^,^}^×^^ from the critic DNN 608.Optionally, the parameters of the Actor DNN 604 are updated based on a gradient of the loss function: log^^^(^^^|^^ , ^^; ^ )^AdvOptionally, the advantage values Adv^,^ are the extra value that can be obtained by choosing action ^^ at st ^^ ate ^^ for thepreference ^^ compared to ^^^^(^^, ^^).Optionally, the advantage value for a given ^ ∈ {1, … , ^} and m∈ {1, … , ^^}, Adv^,^ is given as. The critic DNN 608 is common to all APs of an interference group.The critic DNN 608 is optimized through a gradient descent. The critic DNN 608 is initialized with the same parameters ^, asthe other APs in the interference group. The critic DNN 608 is configured to take as input the interference group state ^^ andthe interference group preference vector ^^ , and to output a value ^(^^, ^^ ; ^), which value is defined as the expecteddiscounted sum of future interference group reward starting from the interference group state ^ for the interference grouppreference ^ where ^ is a discount factor defined as ^ ∈ [0,1). The critic DNN 608 is fed withtransitions {^^ , ^^ , ^^^^ } from the experience buffer 606 and a batch ^^ of ^^ interference group preferences. Theparameters of the critic DNN 608 are updated based on the gradient of a loss function. The parameters of the critic DNN 608are updated based on a gradient of the loss function:The critic DNN 608 has access to the interference group state and interference group reward and hence the critic DNN 608 canhelp the AP during the training by sending the advantage values, thereby resolving the partial observability of the AP. Byreceiving the advantages, the AP understands how its local action affects other APs of the same interference group so that theAP can adjust its policy to optimize the interference group performance.Optionally, the data is exchanged with each of the one or more other APs of the interference group simultaneously.Optionally, the data is exchanged with a subset of the one or more other APs of the interference group.Optionally, the data is exchanged with a randomly selected AP of the interference group.Optionally, all APs in the interference group cause interference to at least one of other APs 614 in the interference group. Aninterference graph is defined as a graph ^ = (^, ^) where ^ is the set of APs, and ^ is the set of directed edges, where an edge^^→^ suggests that AP ^ causes interference to AP ^.Optionally, the interference groups are defined as the strongly connected components of the interference graph.Optionally, the interference groups are defined as the final coalitions of a merge and split algorithm applied based on a utilityfunction.FIG. 7 is an exemplary diagram that illustrates inference of an access point, AP 700 in accordance with an implementation ofthe disclosure. In the exemplary diagram, at each time slot ^, a preference module 702 generates a local preference vector ^^^,for the AP 700 and then sends the local preference vector ^^^ to all other APs 714 of the interference group through acommunication interface 716. The preference module 702 receives information ^ ^^^ , from all other APs 714 of the interferencegroup. The preference module 702 generates an interference group preference vector ^^ . The local preferencevectors {^^ ^^ , … , ^^ } are mapped to the interference group.An actor deep neural network, DNN 704 determines a local resource allocation action ^^^ based on a local state ^^^ , and theinterference group preference vector ^^. The local state and a local reward ^^^ may not be exchanged. The actor DNN 704observes the local state ^^^ from local environment 720 to select the local preference vector ^^^. The local state may be forexample channel rates and traffic queues. The preference module 702, and the actor DNN 704 are depicted as bolded blocksor arrows depicting the training and the inference of the access point, AP 700. The preferences of all APs 714 are consideredacross objectives and managing interference across all APs based only on local state, thereby improving the scalability.An experience buffer 706, and a critic DNN 708 are depicted as dashed blocks or arrows indicating the training of the accesspoint 700.FIGS. 8A and 8B are flow diagrams that illustrate a method for an access point, AP, configured to operate in a wireless cellularcommunication network including one or more access points, APs, belonging to at least two interference groups in accordancewith an implementation of the disclosure. The AP belongs to one of the at least two interference groups. The AP includes acommunication interface. At a step 802, a local preference vector for each of the APs of all other interference group members^^^, is received through the communication interface, where ^ indicates a time slot, and ^ indicates one of the APs of all otherinterference group, AP ^^. At a step 804, generating a local preference vector ^^ is generated for the AP based on AP processingneeds, where ^ indicates the AP,AP^. At a step 806, an interference group preference vector ^^ , is generated based on the localpreference vectors ^ ^ and the local preference vector ^^ , where ^ ^ ^ ^^^ ^ ^ = f(^^ , ^^^ ), and where ^^^ is the set of all localpreference vectors of all other interference group members. At a step 808, a local resource allocation action (^^^) is determinedby applying an actor deep neural network, DNN, on a local state (^^^) and the interference group preference vector.This method selects the best local actions by a single DNN of each AP that optimizes the interference group performance fora given interference group preference vector in the wireless cellular communication network. This method trains the DNN oncefor all the interference group preference vector space, without the need for multiple DNNs tailored to each interference grouppreference vector, thereby minimizing hardware requirements at each AP. This method provides stable distributed policies tomanage the interference between the APs, thereby improving global performance. This method provides distributed policiesover the entire space of interference group preferences, leading to optimal policies for a diverse range of interference grouppreference vectors. This method is adaptable to a wide variety of resource allocation tasks such as power allocation, modulation,and coding schemes.Optionally, the method includes generating the interference group preference vector ^^ and determine the local resourceallocation action ^^^ for a plurality of time slots. The method includes monitoring a local reward ^^^ and an updated local state^^^^^ for each generated interference group preference vector ^^ and determined local resource allocation action ^^^ . Themethod includes transmitting the local reward and the updated local states to the at least one other interference group member.The method includes receiving local rewards ^^ ^^ and updated local states ^^ from the at least one other interference groupmember. The method includes storing the local action, an interference group state including states for the AP and the at leastone other AP, an interference group reward including rewards for the AP and the at least one other AP, and an updatedinterference group state including updated states for the AP and the at least one other AP to an experience buffer. The methodincludes creating a batch = [^^(1), … , ^^(^^)] of ^^ local preference vectors, is the generated^^preference vector. The method includes receiving a local batch composed of ^^ local preference vectors from each of the atleast one other interference group member. The method includes creating an interference group batch ^^ composed of ^^interference group preference vectors {^^(1), … , ^^(^^)} . The method includes optimizing a Critic DNN through a gradientdescent, which Critic DNN is configured to take as input the interference group state ^^ and the interference group preferencevector ^^, and to output a value ^(^^, ^^ ; ^), which value is defined as the expected discounted sum of future interferencegroup reward starting from the interference group state ^ for the interference group preference ^ wherein ^ is a discount factor defined as ^ ∈ [0,1). The method includes feedingCritic DNN with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer and a batch ^^ of ^^ interference group preferences.The method includes updating the parameters of the Critic DNN based on the gradient of a loss function. The method includesfeeding ^ × ^^ advantage values {Adv^,^}^×^^ to the Actor DNN. The method includes updating the Actor DNN parametersbased on the ^ transitions {^^^ , ^^^ }^ from the experience buffer, the batch ^^ of ^^ interference group preference vectors{^^(1), … , ^^(^^)} and the ^ × ^^ advantage values {Adv^,^}^×^^ from the Critic DNN.According to a third aspect, there is provided a computer program product comprising program instructions for performing theabove method, when executed by one or more processors in an Access point.It should be understood that the arrangement of components illustrated in the figures described is exemplary and that otherarrangements may be possible. It should also be understood that the various system components (and means) defined by theclaims, described below, and illustrated in the various block diagrams represent components in some systems configuredaccording to the subject matter disclosed herein. For example, one or more of these system components (and means) may berealized, in whole or in part, by at least some of the components illustrated in the arrangements illustrated in the describedfigures.In addition, while at least one of these components is implemented at least partially as an electronic hardware component, andtherefore constitutes a machine, the other components may be implemented in software that when included in an executionenvironment constitutes a machine, hardware, or a combination of software and hardware.Although the disclosure and its advantages have been described in detail, it should be understood that various changes,substitutions, and alterations can be made herein without departing from the spirit and scope of the disclosure as defined by theappended claims.

Claims

CLAIMS1. An access point, AP (202A, 300A, 500, 600, 700), configured to operate in a wireless cellular communication networkcomprising a plurality of access points, APs (202A, 202B, 300A, 300B, 500, 600, 700), belonging to at least two interferencegroups , wherein the AP (202A, 300A, 500, 600, 700) belongs to one of the at least two interference groups, wherein the AP(202A, 300A, 700, 500, 600) comprises a controller (204) and a communication interface (206, 516, 616, 716), and whereinthe controller (204) is configured to:receive, through the communication link, a local preference vector for each of the APs of all other interference group^^^, wherein ^ indicates a time slot, and ^ indicates one of the APs of all other interference group ,AP^;generate a local preference vector ^^^ for the AP (202A, 300A, 500, 600, 700) based on AP processing needs, wherein^ indicates the AP, AP^i,generate an interference group preference vector ^^, the interference group preference vector ^^ being based on thelocal preference vectors ^ ^ and the local preference vector ^^, wher ^ ^ ^ ^^^ ^ ein ^ = f(^^, ^^^ ), and where ^^^ is the set ofall local preference vectors of all other interference group members, anddetermine a local resource allocation action ^^^ by applying an actor deep neural network, DNN (304, 504, 604,704),on a local state ^^ and the interf ^^ erence group preference vector ^ .

2. The AP (202A, 300A, 500, 600, 700) according to claim 1, wherein the controller (204) is further configured to determinewhich interference group the AP (202A, 300A, 500, 600, 700) belongs to.

3. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein the controller (204) is further configuredto transmit the local preference vector to all other interference group members.

4. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein the controller (204) is further configuredto determine the local state based on a local environment (520, 620, 720).

5. The AP (202A, 300A, 500, 600, 700) according to claim 4, wherein the controller (204) is further configured to determinethe local state by receiving the local state from the local environment (520, 620, 720).

6. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein the controller (204) is further configuredto generate the interference group preference vector ^^ based on the local preference and the local preference for the one ormore other interference group member by applying a function which is deterministic and is the same across APs of the sameinterference group.

7. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein the controller (204) is further configuredto: generate the interference group preference vector ^^ and determine the local resource allocation action ^^^ for aplurality of time slots,monitor a local reward ^^^ and an updated local state ^^^^^ for each generated interference group preference vector^^ and determined local resource allocation action ^^^transmit the local reward and the updated local states to the at least one other interference group member,receive local rewards ^^ ^^ and updated local states ^^ from the at least one other interference group member,store the local action, an interference group state including states for the AP (202A, 300A, 500, 600, 700) and the atleast one other AP, an interference group reward including rewards for the AP (202A, 300A, 500, 600, 700) and the atleast one other AP and an updated interference group state including updated states for the AP (202A, 300A, 500, 600,700) and the at least one other AP to an experience buffer (306, 506, 606, 706),create a batch= [^^(1), … , ^^(^^)] of ^^ local preference vectors, whereis the generated ^^^preference vector, receive a local batch composed of ^^ local preference vectors from each of the at least one other interference groupmember, and thencreate an interference group batch ^^ composed of ^^ interference group preference vectors {^^(1), … , ^^(^^)} ,wherein the controller (204) is further configured tooptimize a Critic DNN (308, 508, 608, 708) through a gradient descent, which Critic DNN (308, 508, 608, 708) isconfigured to take as input the interference group state ^^ and the interference group preference vector ^^, and to outputa value ^(^^ , ^^ ; ^), which value is defined as the expected discounted sum of future interference group reward startingfrom the interference group state ^ for the interference group preference ^^(^, ^; ^) =(.|^^,^) ∑^ (^)^^^ , wherein ^ is a discount factor defined as ^ ∈ [0,1),feed the Critic DNN (308, 508, 608, 708) with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer (306, 506, 606,706) and a batch ^^ of ^^ interference group preferences,Update the parameters of the Critic DNN (308, 408, 508, 608) based on the gradient of a loss function and thenfeed ^ × ^^ advantage values {Adv^,^}^×^^ to the Actor DNN (304, 504, 604,704), and wherein the controller (204)is further configured to:update the Actor DNN (304, 404, 504, 604, 704) parameters based on:the ^ transitions {^^^ , ^^^ }^ from the experience buffer (306, 506, 606, 706),the batch ^^ of ^^ interference group preference vectors {^^(1), … , ^^(^^)} andthe ^ × ^^ advantage values {Adv^,^}^×^^ from the Critic DNN (308, 508, 608, 708).

8. The AP (202A, 300A, 500, 600, 700) according to claim 7, wherein the controller (204) is further configured to update theparameters of the critic DNN (308, 508, 608, 708) based on a gradient of the loss function:

9. The AP (202A, 300A, 500, 600, 700) according to claim 7, wherein the controller (204) is further configured to update theparameters of the Actor DNN (304, 504, 604, 704) based on a gradient of the loss function:log^^^(^^^|^^ , ^^; ^ )^10. The AP (202A, 300A, 500, 600, 700) according to claim 7, 8 or 9, wherein the advantage values Adv^,^ are the extravalue that can be obtained by choosing action ^^^ at state ^^ ^ ^^ ^for the preference ^ compared to ^ ^(^^, ^^).value for a given ^ ∈ {1, … , ^} and12. The AP (202A, 300A, 500, 600, 700) according to any of claims 7 to 11, wherein the Critic DNN (308, 508, 608, 708) iscommon to all APs of an interference group.

13. The AP (202A, 300A, 500, 600, 700) according to claim 12, wherein the controller (204) is further configured to initializethe Critic DNN (308, 508, 608, 708) with the same parameters ^, as the other APs in the interference group.

14. The AP (202A, 300A, 500, 600, 700) according to any of claims 7 to 13, wherein the controller (204) is further configuredto: select from Experience buffer (306, 506, 606, 706) the same transitions as the other APs of the interference group inorder to update the parameters of the Actor DNN (304, 504, 604, 704) and the Critic DNN (308, 508, 608, 708).

15. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein the controller (204) is further configuredto exchange data with each of the one or more other AP of the interference group simultaneously.

16. The AP (202A, 300A, 500, 600, 700) according to any of claims 1 to 14, wherein the controller (204) is further configuredto exchange data with a subset of the one or more other AP of the interference group.

17. The AP (202A, 300A, 500, 600, 700) according to any of claims 1 to 14, wherein the controller (204) is further configuredto exchange data with a randomly selected AP of the interference group.

18. The AP (202A, 300A, 500, 600, 700) according to any preceding claim, wherein all APs in the interference group causeinterference to at least one of the other APs in the interference group, wherein an interference graph is defined as a graph ^ =(^, ^) where ^ is the set of APs, and ^ is the set of directed edges, where an edge ^^→^ suggests that AP ^ causes interferenceto AP ^.

19. The AP (202A, 300A, 500, 600, 700) according to claim 18, wherein the interference groups are defined as the stronglyconnected components of the interference graph.

20. The AP (202A, 300A, 500, 600, 700) according to claim 17, wherein the interference groups are defined as the finalcoalitions of a merge and split algorithm applied based on a utility function.

21. A method for an access point, AP (202A, 300A, 500, 600, 700), configured to operate in a wireless cellular communicationnetwork comprising a plurality of access points, APs (202A, 202B, 300A, 300B), belonging to at least two interference groups,wherein the AP (202A, 300A, 500, 600, 700) belongs to one of the at least two interference groups (204A, 204B),wherein the AP (202A, 300A, 500, 600, 700) comprises a controller (204) and a communication interface (206, 516,616, 716), wherein the method comprises:receiving, through the communication interface (206, 516, 616, 716), a local preference vector for each of theAPs of all other interference group members ^^^, wherein ^ indicates a time slot, and ^ indicates one of theAPs of all other interference group,AP^;generating a local preference vector ^^^ for the AP (202A, 300A, 500, 600, 700) based on AP processing needs,wherein ^ indicates the AP (202A, 300A, 500, 600, 700), AP^;generating an interference group preference vector ^^, the interference group preference vector ^^being basedon the local preference vectors ^ ^^^ and the local preference vector ^^^, wherein ^^ = f(^^^, ^ ^^^ ), and where^^^^ is the set of all local preference vectors of all other interference group members, anddetermining a local resource allocation action (^^^) by applying an actor deep neural network, DNN (304, 504,604, 704), on a local state (^^^) and the interference group preference vector.

22. The method according to claim 21, wherein the method further comprises:generating the interference group preference vector (^^ ) and determine the local resource allocation action (^^^ ) for aplurality of time slots,monitoring a local reward ^^^ and an updated local state ^^^^^ for each generated interference group preference vector(^^ ) and determined resource allocation action (^^^), transmitting the local reward and the updated local states to the at least one other interference group member,receiving local rewards ^^ ^^^^ and updated local states ^^ from the at least one other interference group member,storing the local action, an interference group state including states for the AP (202A, 300A, 500, 600, 700) and the atleast one other AP, an interference group reward including rewards for the AP (202A, 300A, 500, 600, 700) and theat least one other AP and an updated interference group state including updated states for the AP (202A, 300A, 500,600, 700) and the at least one other AP to an experience buffer (306, 506, 606, 706),creating a batch ^^ = [^^(1), … , ^^(^^)] of ^^ local preference vectors, where ^^(^) is the generated ^^^ preference vector,receiving a local batch composed of ^^ local preference vectors from each of the at least one other interference groupmember, and thencreating an interference group batch ^^ composed of ^^ interference group preference vectors {^^(1), … , ^^(^^)} ,wherein the method further comprises:optimizing a Critic DNN (308, 508, 608, 708) through a gradient descent, which Critic DNN (308, 508, 608,708) is configured to take as input the interference group state ^^ and the interference group preference vector^^ , and to output a value ^(^^, ^^ ; ^), which value is defined as the expected discounted sum of futureinterference group reward starting from the interference group state ^ for the interference group preference ^^(^, ^; ^) = ^ ^ ^ ^^^,^(.|^^,^) ∑ (^) ^ wherein ^ is a discount factor defined as ^ ∈ [0,1)feeding the Critic DNN (308, 508, 608, 708) with ^ transitions {^^ , ^^ , ^^^^ }^ from the experience buffer(306, 506, 606, 706) and a batch ^^ of ^^ interference group preferences,updating the parameters of the Critic DNN (308, 508, 608, 708) based on the gradient of a loss function andthen feed ^ × ^ advantage values {Advto the Actor DNN (304, 504, 604, 704), and wherein thefurther comprises:updating the Actor DNN (304, 504, 604, 704) parameters based on:the ^ transitions {^^^ , ^^^ }^ from the experience buffer (306, 506, 606, 706),the batch ^^ of ^^ interference group preferences {^^(1), … , ^^(^^)} , andthe ^ × ^^ advantage values {Adv^,^}^×^^ from the Critic DNN (308, 508, 608, 708).

22. A computer program product comprising program instructions for performing the method according to claim 21 or 22,when executed by one or more processors in an Access point.

Citation Information

Patent Citations

  • Interference coordination method and device and storage medium

    CN114885336A

  • Method and apparatus for data scheduling in wireless communication system

    US20230217478A1