Access point for wireless communication network and method of operating access point for wireless communication network

By configuring the first and second interfaces and optional policy modules at the access point, and combining the controller's selection policy module to make resource allocation decisions, the problems of large interference between devices and communication overhead in the wireless communication network are solved, and flexible and efficient resource allocation is achieved.

CN120476659APending Publication Date: 2025-08-12HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090811.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the existing wireless communication network, the access point ignores the actions of other parts of the network during resource allocation, resulting in serious interference between devices, and the centralized processing method has a large communication overhead and high latency.

Method used

The access point is configured with first and second interfaces, respectively, for receiving information from the communication device and other access points, and compute resource allocation decisions through the optional policy module, the controller selects the most suitable policy module to make decisions, and exchanges information only if necessary.

Benefits of technology

It realizes flexible resource allocation in a dynamic network environment, reduces unnecessary data exchange, and improves network performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476659A_ABST
    Figure CN120476659A_ABST
Patent Text Reader

Abstract

The present disclosure relates to an access point (10) for a wireless communication network. The access point (10) comprises: a first interface (11) configured to receive first information from at least one communication device (20) in the wireless communication network; a second interface (12) configured to receive second information from at least one other access point (10 ') in the wireless communication network; and M selectable policy modules (14-1, 14-2), each of the selectable policy modules (14-1, 14-2) being configured to compute a resource allocation decision. The M selectable policy modules comprise: a first policy module (14-1) configured to calculate the resource allocation decision based on the received first information without requiring the second information; and at least one second policy module (14-2) configured to calculate the resource allocation decision based on the received first information and the received second information. The access point (10) further comprises a controller (13) configured to select a policy module from the M selectable policy modules (14-1, 14-2) to calculate the resource allocation decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to resource allocation in a wireless communication network. More particularly, the present disclosure relates to an access point for a wireless communication network and a method of operating an access point for a wireless communication network. Background Art

[0002] Radio resource allocation is a critical aspect of wireless networks, such as cellular networks. With the advent of 6G and higher wireless networks, it is expected that a large number of devices will connect to each access point (AP) in the network. In such an environment, radio resource allocation will become an increasingly challenging task.

[0003] The current approach is for each AP to collect state information about its connected devices and then solve an optimization problem to decide how to allocate radio resources, such as transmit power, subcarrier frequency, and beam selection. As a result, the AP is oblivious to the rest of the network (e.g., other APs and their actions). While this approach is scalable, since no additional communication between devices is required, it can result in strong interference from devices belonging to nearby APs. This interference can have a significant negative impact on overall network performance.

[0004] An alternative approach is to centrally handle AP actions through a master entity, which needs to (a) collect a large amount of information from all APs, (b) solve a more challenging large-scale optimization problem, and (c) finally send the corresponding radio resource allocation decision back to each AP. This approach is not practical due to the huge communication overhead and the potential for significant increase in latency.

[0005] Therefore, it would be advantageous if wireless resource allocation in the next generation communication network architecture could be more flexible and require fewer resources than the conventional approaches described above. Summary of the Invention

[0006] In view of the foregoing, the present disclosure is directed to providing an improved access point for a wireless communication network and an improved method of operating an access point for a wireless communication network, which overcome the above-mentioned limitations and disadvantages.

[0007] These and other objects are achieved by the solutions of the present disclosure as described in the independent claims. Advantageous implementations are further defined in the dependent claims.

[0008] A first aspect of the present disclosure provides an access point for a wireless communication network. The access point includes: a first interface configured to receive first information from at least one communication device in the wireless communication network; a second interface configured to receive second information from at least one other access point in the wireless communication network; and M selectable policy modules, each of the selectable policy modules configured to calculate a resource allocation decision. The M selectable policy modules include: a first policy module configured to calculate the resource allocation decision based on the received first information without requiring the second information; and at least one second policy module configured to calculate the resource allocation decision based on the received first information and the received second information. The access point also includes a controller configured to select a policy module from the M selectable policy modules to calculate the resource allocation decision.

[0009] This provides the advantage that the access point can perform resource allocation that adapts to the dynamic environment of the wireless network. For example, the access point can select the policy module that best suits the current network conditions, thereby avoiding unnecessary data exchange and thus reducing communication costs.

[0010] For example, the access point is configured to subsequently execute the calculated resource allocation decision.

[0011] The resource allocation decision may be a scheduling decision. The scheduling decision may include a decision as to which communication device in the network communicates with the access point at a certain (scheduled) time step or time slot.

[0012] The access point may be a base station or an agent of the wireless communication network.

[0013] The communication device may be any device in a network that is wirelessly connected to the access point, for example, user equipment (UE) such as a smartphone.

[0014] The wireless communication network may be a cellular network.

[0015] The first interface and / or the second interface may be respective wireless interfaces.

[0016] The first information may include information about the status of the at least one communication device, such as a buffer size and / or a channel used by the at least one communication device. For example, the first information may be local information from a communication device that directly communicates with the access point.

[0017] The second information may include information about the status of the at least one other access point and / or the communication device connected to the other access point. For example, the second information may include neighbor information from any number of other access points in the communication network.

[0018] The at least one other access point in the wireless communication network may be substantially identical to the access point, ie, the other access point may include the same features and calculate its own resource allocation decisions.

[0019] In an implementation form of the first aspect, the controller is configured to select the policy module from the M selectable policy modules based on the received first information. This achieves the advantage of being able to select a policy module suitable for the current environment of the wireless network.

[0020] In an implementation form of the first aspect, the controller is configured to broadcast a message including the selection of the policy module to one, a plurality of or all other access points in the wireless communication network. This achieves the advantage of ensuring that necessary information is shared between access points.

[0021] In an implementation form of the first aspect, the second information includes a response of the at least one other access point in the wireless communication network to the broadcast message.

[0022] In an implementation form of the first aspect, the second interface is further configured to receive a broadcast message regarding the selection of the policy module from at least one other access point in the wireless communication network.

[0023] For example, the access point may subsequently respond to the received broadcast message and send the received first information to the at least one other access point to help one of the at least one other access point make a better resource allocation decision.

[0024] In an implementation form of the first aspect, the access point is configured to send other information back to the at least one other access point in response to receiving the broadcast message.

[0025] Preferably, the access point is configured to send back the further information only if the other access point requires it due to its policy selection. Thus, distributed resource management can be achieved, wherein one or more access points only exchange necessary information, thereby reducing the amount of transmitted data to the minimum necessary level.

[0026] For example, the other information may be “second information” for the other access point. The other information may include first information previously received by the access point.

[0027] In an implementation form of the first aspect, the controller is a trainable controller that can be trained for the selection of the policy module.

[0028] For example, the controller of each access point in the communication network may be trained based on the specific traffic and channel characteristics of its associated communication devices.The controllers of each access point in the network may be trained independently of each other, for example, in parallel.

[0029] In an implementation form of the first aspect, after executing the calculated resource allocation decision, the first interface is configured to receive reward feedback from at least one communication device in the wireless communication network, and / or the second interface is configured to receive reward feedback or an accumulation of reward feedback from at least one other access point in the wireless communication network. This provides the advantage of providing feedback on resource allocation for use, for example, in training and / or adjusting components of an access point (e.g., a controller thereof).

[0030] The reward feedback may include feedback on the current performance of the communication device and / or access point (eg, communication rate or utilization of a communication channel).

[0031] For example, a communication channel between two or more access points may be used to exchange messages regarding selection of a policy module, second information (eg, in response to such messages), and reward feedback.

[0032] In an implementation form of the first aspect, the controller includes a trainable neural network. For example, the trainable neural network can be configured to select the policy module.

[0033] In an implementation form of the first aspect, the controller is configured to train the trainable neural network by: calculating a global reward based on the reward feedback from the at least one communication device and the one or more reward feedbacks from the at least one other access point, for example, using a Monte Carlo simulation; inputting the global reward into a loss function; and adjusting the trainable neural network based on a result of the loss function.

[0034] For example, the weights of the trainable neural network are adjusted based on the result of the loss function.The training of the neural network by the controller may be performed at fixed time intervals.

[0035] In an implementation form of the first aspect, each of the selectable policy modules comprises a set of rules, such as an algorithm, stored in a memory of the access point.

[0036] The optional policy modules, and in particular their corresponding sets of rules or algorithms, may be configured to be executed by a processor of the access point in order to calculate the corresponding resource allocation decision.

[0037] In an implementation form of the first aspect, the controller (eg, a neural network of the controller) is further configured to adjust the rule of at least one selectable strategy module among the M selectable strategy modules at a certain time interval.

[0038] The adjustment may be achieved through the above-mentioned training. For example, the neural network of the controller may be trained for the selection decision and / or adjustment of the policy module.

[0039] In an implementation form of the first aspect, each of the selectable strategy modules includes a corresponding other trainable neural network.

[0040] For example, the trainable neural network (of the controller) and / or the other trainable neural network (of the policy module) may be a fully connected neural network (FCNN), i.e., a neural network comprising fully connected layers.

[0041] In an implementation form of the first aspect, the other trainable neural networks of the optional strategy module are configured to be trained individually by calculating an individual loss for each other trainable neural network and adjusting the corresponding other trainable neural networks based on the individual loss.

[0042] The separate losses may be calculated using separate loss functions for each selectable policy module.

[0043] In an implementation form of the first aspect, the other trainable neural network of the optional policy module is configured to be jointly trained with the trainable neural network of the controller by calculating a joint loss of the trainable neural network of the controller and the other trainable neural network of the optional policy module, and adjusting the trainable neural network and the other trainable neural network based on the joint loss.

[0044] For example, the joint loss may be computed as the sum of individual loss functions for each selectable policy module, where each loss function is weighted according to the probability of the corresponding policy module being selected by the controller.

[0045] A second aspect of the present disclosure provides a system including at least two access points according to the first aspect of the present disclosure.

[0046] Each of the at least two access points of the system may be configured to receive first information from at least one corresponding communication device in the wireless communication network and to receive second information from corresponding other access points of the at least two access points.

[0047] The access points of the system can exchange different types of messages. For example, the controller of a first access point in the system can broadcast its policy module selection to one or more other access points in the system. On the receiving end, the one or more other access points can interpret the received message and send a predefined (pre-agreed) level of information back to the first access point (e.g., in the form of second information). This can be performed by all access points in the system. Furthermore, the controller actions of each access point can determine the policy module used by that access point, which in turn affects the amount of information exchanged between the access points of the system. Therefore, environmental conditions (e.g., arrival strength, channel gain) can determine when access points exchange such messages. Therefore, in an isolated environment, changes in environmental conditions (e.g., poor channels and / or large queues) can result in a significant increase in signaling between access points.

[0048] A third aspect of the present disclosure provides a method for operating an access point for a wireless communication network. The method comprises the following steps: receiving first information from at least one communication device in the wireless communication network; selecting a policy module from M selectable policy modules, each of the selectable policy modules being configured to calculate a resource allocation decision; broadcasting a message regarding the selection of the policy module to at least one other access point in the wireless communication network; and receiving second information from the at least one other access point based on the selected policy module; wherein the M selectable policy modules include: a first policy module configured to calculate the resource allocation decision based on the received first information without requiring the second information; and at least one second policy module configured to calculate the resource allocation decision based on the received first information and the received second information. The method further comprises the step of calculating the resource allocation decision using the selected policy module.

[0049] For example, the step of receiving the second information from the at least one other access point does not always occur. This step depends on the selection of the policy module. For example, the second information is forwarded by the at least one other access point only when required by the selected policy module. This provides the advantage of exchanging the second information (which may have a high bit count) only when required.

[0050] The method according to the third aspect of the present disclosure may be performed by the access point according to the first aspect of the present disclosure.

[0051] The above description on the access point according to the first aspect of the present disclosure and the system according to the second aspect of the present disclosure is also applicable to the method according to the third aspect of the present disclosure.

[0052] It should be noted that all devices, elements, units and devices described in this application can be implemented with software or hardware elements or any combination thereof. All steps performed by each entity described in this application and the functions described to be performed by each entity are intended to indicate that the corresponding entity is applicable to or is configured to perform the corresponding steps and functions. Although in the description of the following specific embodiments, the specific functions or steps to be performed by the external entity are not reflected in the description of the specific detailed elements of the entity that performs the specific steps or functions, it should be clear to those skilled in the art that these methods and functions can be implemented with corresponding software or hardware elements or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The specific embodiments described below, in conjunction with the accompanying drawings, illustrate the above aspects and implementation forms. In the drawings:

[0054] Figure 1 A schematic diagram illustrating an access point for a wireless communication network according to an embodiment of the present disclosure is shown;

[0055] Figure 2 A schematic diagram illustrating a system including an access point according to an embodiment of the present disclosure is shown;

[0056] Figure 3 A schematic diagram illustrating a system including an access point according to an embodiment of the present disclosure is shown;

[0057] Figure 4 shows the results of performance evaluation according to an embodiment of the present disclosure;

[0058] Figure 5 A flowchart illustrating a method of operating an access point for a wireless communication network according to an embodiment of the present disclosure is shown;

[0059] Figure 6 A flowchart of a method of operating an access point for a wireless communication network according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0060] Figure 1 A schematic diagram of an access point 10 for a wireless communication network according to an embodiment of the present disclosure is shown.

[0061] Access point 10 includes a first interface 11 configured to receive first information from at least one communication device 20 in a wireless communication network; a second interface 12 configured to receive second information from at least one other access point 10' in the wireless communication network; and M selectable policy modules 14-1 and 14-2, each of which is configured to calculate a resource allocation decision. The M selectable policy modules 14-1 and 14-2 include a first policy module 14-1 configured to calculate a resource allocation decision based on the received first information without requiring the second information; and at least one second policy module 14-2 configured to calculate a resource allocation decision based on the received first information and the received second information. Access point 10 also includes a controller 13 configured to select a policy module from the M selectable policy modules 14-1 and 14-2 for calculating the resource allocation decision.

[0062] For example, the access point 10 is configured to subsequently execute the calculated resource allocation decision.

[0063] For example, when M=2, two selectable policy modules (a first policy module 14 - 1 and a second policy module 14 - 2 ) may be selected to calculate the resource allocation decision. However, the number of selectable policy modules may be greater (M>2). Figure 1 The two policy modules 14 - 1 , 14 - 2 in FIG are shown only as an example.

[0064] Resource allocation decisions can be scheduling decisions. Scheduling decisions can include decisions about which communication devices in the network communicate with the access point within a certain (scheduled) time step or time slot. Resource allocation decisions can also include decisions about radio resource allocation (e.g., transmit power, subcarrier frequency, and / or beam selection).

[0065] The access point 10 may be a base station or a proxy of a wireless communication network.

[0066] The communication device 20 may be any device in the network that is wirelessly connected to the access point 10 , for example, a user equipment (UE) such as a smartphone.

[0067] The first interface 11 and / or the second interface 12 may be corresponding wireless interfaces.

[0068] Each of the M selectable policy modules 14-1, 14-2 may be optimized for specific conditions, wherein one or more second policy modules 14-2 may require data exchange (eg, observed status) from one or more other access points 10' of the communication network to calculate resource allocation.

[0069] For example, the M selectable policy modules may include other policy modules. The other policy modules may include other "second policy modules," i.e., policy modules configured to calculate resource allocation decisions based on received first information and second information. The second information used by each of these second policy modules may be information from a corresponding subset of other access points in the wireless communication network. For example, each of the M selectable policy modules may perform better under a given set of conditions.

[0070] For example, each optional policy module 14-1, 14-2 may include a set of rules, such as an algorithm, stored in a memory of the access point. The optional policy modules 14-1, 14-2, and in particular the corresponding rule sets or algorithms of these policy modules, may be configured to be executed by a processor of the access point 10 to calculate corresponding resource allocation decisions.

[0071] exist Figure 1 At least one other access point 10′, depicted by a dashed block, can be substantially identical to access point 10, i.e., can include the same features and can calculate its own resource allocation decisions using a selected policy module. Access points 10, 10′ in the wireless communication network can communicate via their respective second interfaces 12, e.g., exchange one or more second information items.

[0072] The controller 13 may be configured to select the policy module 14 - 1 , 14 - 2 that should be deployed for a given first information. For example, the controller 13 may include a processor configured to perform said selection.

[0073] The controller 13 may be configured to select a policy module from the M selectable policy modules based on the received first information, for example based on information from the communication device 20 transmitting its status directly to the access point 10 .

[0074] The controller may also be configured to broadcast a message including the selection of the policy module to one, multiple, or all other access points 10' in the wireless communication network.

[0075] The broadcast of the message including the policy module selection may trigger additional signaling because the selected policy module may require additional (second) information from at least one other access point 10'. For example, the other access point 10' may respond to receiving the broadcast by forwarding the second information to access point 10. If the broadcast indicates, for example, that the second policy module requires the second information, the other access point 10' may forward the second information.

[0076] For example, the (second) policy module computes the resource allocation decision after receiving the second information (in response to the broadcasted message including the policy module selection).

[0077] Similarly, the second interface 12 may be configured to receive broadcasted messages about selections of policy modules from other access points 10 ′ and may respond to the messages by sending further (second) information to the other access points 10 ′.

[0078] Controller 13 can be a trainable controller, capable of training the selection of policy modules. For example, controller 13 learns to select the optimal policy module from M selectable policy modules under a given set of conditions, i.e., the policy module that provides the best resource allocation decision. Therefore, in the specific case where the resource allocation decision is a scheduling decision, controller 13 can be configured to perform "meta-scheduling." That is, controller 13 can select which policy module 14-1 or 14-2 will make the scheduling decision, for example based on reinforcement learning (RL). Therefore, controller 13 can include a trainable neural network, such as a fully connected neural network (FCNN).

[0079] For example, the communication between access points 10 in the communication network is not fixed. The controller 13 of each access point can adjust the amount of information exchanged between access points 10 based on the current network conditions. In this way, the controller 13 can control the communication (information exchange) between access points 10 and 10'.

[0080] Figure 2 A schematic diagram of a system 30 including a plurality of access points 10 for a wireless communication network according to an embodiment of the present disclosure is shown. Figure 2 Each access point 10 of the system 30 shown in FIG. 1 may be Figure 1 The access point 10 is shown in FIG. The wireless communication network may be a cellular network.

[0081] Each of the access points 10 can communicate directly with a plurality of communication devices 20 in an environment (depicted by a circle surrounding each access point 10). This communication can be performed through the respective first interfaces 11 of the access points 10 and can include receiving first information from the devices 10 in the respective environments. The access points 10 can also communicate with each other through their respective second interfaces 12, for example, to exchange second information.

[0082] In this way, a system 30 for efficient decentralized radio resource allocation in a wireless network with controlled information exchange between access points 10 may be provided.

[0083] Figure 3 Another schematic diagram of a system 30 including a plurality of access points 10 according to an embodiment of the present disclosure is shown. Figure 3 Each access point 10 of the system 30 shown in FIG. 1 may be Figure 1 The access point 10 is shown in FIG.

[0084] Although Figure 2 and Figure 3 The schematic system architecture in FIG. 3 shows only two access points 10 of the system 30 , but the system 30 may include any number of access points 10 for a wireless communication network.

[0085] In the following, Figure 3 The two access points (APs) 10 in FIG. 3 are referred to as AP i and AP j. Each of AP i and AP j includes at least two policy modules 14 - 1 , 14 - 2 (also referred to as Policy 1 and Policy 2 ). Both AP i and AP j can connect to multiple communication devices 20 in their respective environments 31 .

[0086] An exemplary implementation of the policy modules 14 - 1 , 14 - 2 may be to differentiate the M policy modules based on the amount and / or type of information they require to make resource allocation decisions.

[0087] For example, if is a vector encoding the state of the communication device 20 (eg, a device in the environment 31) covered by the API, then the selectable policy module may include the following M=2 policies:

[0088] Strategy 1: Receiving (First information, also called local information) as input and computation (return) resource allocation blocks;

[0089] Strategy 2: Receiving (First Information) and (The second information is also called neighbor information, where ,and represents the set of APs that are one hop away from AP i) as input and a block that computes (returns) resource allocation.

[0090] The trainable controller 13 of the API may be configured to receive the first information based on the received (Input) Action to perform. Essentially, the action for this block can be an integer m (where m = 1, ..., M) indicating which policy block will be selected to make the resource allocation decision. Furthermore, the selected action of controller 13 of AP i can be communicated to neighboring APs, as this decision can trigger information exchange between APs.

[0091] In the following, “resource allocation decision” refers to the resource allocation decision made by one of the M policy modules; “action” refers to the controller selecting policy module m.

[0092] Signaling between APs 10 of system 30 may be performed as follows:

[0093] The controller of API can observe its local state , and take action based on this , indicating which policy module is selected to calculate resource allocation decisions (eg, a scheduled device). This signal may be transmitted to other APs of the system 30 For example, there is an underlying mapping (known to all APs 10 of the system 30) that specifies that if AP j≠i receives from AP i , then it knows how to respond. AP j may or may not send back some of its first (local) information to AP i, depending on the signal it receives. Therefore, there can be such a protocol between APs: if AP j receives ,and It implies that AP i requires AP j to provide a certain amount of information, then AP j will provide this information immediately. From the perspective of AP i, the data received by AP j is recorded as .

[0094] For example, all AP j can send corresponding actions to AP i, which may require AP i to provide some data, and then AP i will also send data to AP j. Figure 3 Steps 3a and 3b in the visualization are shown, where the exchange is shown from the perspective of API i, omitting the actions received by API i. .

[0095] Finally, in order to reduce the signaling between APs 10, it is possible not to send ;No May mean No change. This can be informally translated as a message from AP i to AP j saying: “As long as I don’t send any new messages, follow the latest message you received from me. act”.

[0096] Furthermore, after executing the calculated resource allocation decision, the API may receive reward feedback (eg, via its first interface) from at least one communication device 20 in its environment 31. Additionally or alternatively, the API may also receive reward feedback or an accumulation of reward feedback from at least one other AP 10 in the wireless communication network (eg, from AP j) through its second interface.

[0097] In the following, possible steps that may occur in a single time slot t experienced by API are shown (these steps are also shown in Figure 3 (Highlighted in ):

[0098] • Step 1: AP i observes the status of its connected devices 20, such as the buffer sizes and channels of the devices in its environment 31.

[0099] Step 2: This information is passed to the controller 13, which outputs the instruction policy module that should be executed (in Figure 3 In the example above, the action is either Strategy 1 or Strategy 2).

[0100] Step 3a, Step 3b: Policy selection signal to AP j , and Optional reception (if required).

[0101] Step 4: The controller 13 potentially receives information and forms the input to its selected strategy m. This ultimately results in a resource allocation decision .

[0102] Step 5: When resource allocation decisions are made for all APs 10 in the system 30 When this is achieved, the environment 31 returns the reward to API and proceeds to the new state Response to the switch. AP i schedules device 20, adapting the application to the device's uplink schedule, and all APs 10 in the network perform the same operation. Device 20 scheduled by AP j could potentially cause interference to the selected device 20, which would affect the observed reward of AP i. Furthermore, some new packets may arrive (at device 20) and channel conditions may change; this results in a new state observed by the AP.

[0103] Step 6: Finally, AP i sends its collected rewards to other APs 10; it also receives corresponding rewards from other APs 10. This is useful for training the controllers 13 of all APs, which will be disclosed below.

[0104] Training and optimization:

[0105] The following describes AP 10 (e.g. Figures 1 to 3 A possible training and optimization routine for the AP 10 shown in FIG. 1 and its components (particularly the controller 13 and the policy modules 14 - 1 , 14 - 2 ):

[0106] In a first training routine, controller 13 training may be performed. For example, each AP's controller 10 may include a trainable neural network, such as a deep neural network (DNN), which may be trained by the controller using a first training method. For example, the trainable neural network may be a fully connected neural network (FCNN).

[0107] For example, in the first training routine, the policy modules 14-1 and 14-2 are decomposed from the controller 13. Essentially, in this case, it is assumed that M policy modules are given, and only the controller 13 of the AP is optimized. Therefore, when performing this training, the controller policy can be fixed for K episodes, each of length L, and then at each time step τ, the controllers of the AP can exchange their collected rewards and update their controller policies based on the following method.

[0108] Here, the controller strategy may refer to the rules according to which the strategy module is selected by the controller 13. When training / optimizing the controller 13, these controller strategies of the controller 13 may be trained / optimized separately.

[0109] For example, the controller 13 of API stores a tuple with local observations, actions, and rewards ,in is the observed state; It is the action taken, is the time of the controller of AP i in the kth round The reward received. In this way, the controller of API can:

[0110] Send the rewards it collects to all APs in the network and / or

[0111] • Receive rewards from every other AP in the network.

[0112] After the controller 13 has accumulated experience and it is time to update its controller policy (K rounds have been completed), for example, each controller 13 in the system can perform a global reward and policy DNN update. This global reward and policy update can be carried by each controller 13 as follows:

[0113] Controller 13 uses Monte Carlo estimation to calculate the global future cumulative rewards (rewards-to-go) value ,in yes , where A represents the set of APs, Indicates the time of APa in the kth round Observed rewards.

[0114] About parameters The gradient of (the controller of API) is calculated as follows:

[0115] , which is then used to update the weights of the controller neural network as follows: (in is the learning rate).

[0116] gradient It can be the gradient of the loss function, the global reward value can be input into the loss function. The weights of the trainable neural network of the controller 13 can be adjusted based on the results of the loss function.

[0117] For example, due to the way the future cumulative reward values are calculated, it is possible to avoid swapping these values at each time step τ and instead swap all of these values in the form of a matrix (of size K × L). In this way, the controller 13 that receives these values knows which rewards occurred in which time slot and in which round at the end of K rounds, and therefore swaps these reward values only when gradient updates must occur.

[0118] As an alternative or supplement to the above solution, the controller neural network may also be trained to adjust the rules, such as algorithms, of at least one selectable strategy module among the M selectable strategy modules within certain time intervals.

[0119] In addition to the controller 13, the M selectable policy modules 14-1, 14-2 may also be trainable. For example, each of the selectable policy modules 14-1, 14-2 may include a corresponding other trainable neural network.

[0120] Therefore, in the second training routine, joint training of the trainable controller 13 and the M trainable policy modules 14-1, 14-2 may be performed. Thus, optimization of the components of the AP (controller and policy modules) is achieved in a joint manner.

[0121] For example, for each API, you can 、 and controller Here, for example, only the policy modules strategy 1 and strategy 2 are considered. Therefore, a possible way to calculate the joint loss is as follows:

[0122]

[0123] in, is the probability that the controller chooses policy 1. The above method can be easily extended to any number (M) of different policy modules (and sub-losses).

[0124] The optional policy module's neural network and / or the controller neural network can be tuned based on the joint loss.

[0125] Alternatively, the other trainable neural networks of the optional strategy modules 14 - 1 , 14 - 2 may also be configured to be trained individually, for example, by calculating an individual loss for each other trainable neural network and adjusting the corresponding other trainable neural network based on the individual loss.

[0126] In addition to the aforementioned methods and conventional solutions, there are other conceivable approaches to utilizing deep reinforcement learning (DRL) for radio resource allocation. The common idea behind these approaches is that each AP 10 learns its scheduling or radio resource allocation strategy based on observations from its interactions with the environment and from observations from other APs. However, the following potential alternatives all have some drawbacks:

[0127] Independent RL: Here, each AP 10 will run a single-agent RL algorithm (e.g., deep Q-learning) based on its own local observations. The advantage of this is that there is no information exchange overhead. The disadvantage is that the learned policy may lead to undesirable behavior, such as non-convergence or failure to cooperate.

[0128] Centralized training, decentralized execution: Here, the policy is distributed—each agent (access point 10)'s DNN takes actions based only on local observations (primary information), but its training is centralized by a central entity with access to the entire system state. The central entity's purpose is to stabilize the training process. After training is complete, the updated DNN parameters representing the new policy are sent to the AP. The main drawback of this approach is that it is not scalable due to the significant overhead required to transfer state and updated policies between the central entity and the AP.

[0129] Consensus-based algorithms: The common idea here is to remove the central entity to avoid communication overhead, and allow agents to communicate through a sparse control network with only a subset of neighboring agents (AP 10), with the goal of reaching consensus on the learned variables and ultimately on the adopted policy. This approach still requires relatively high communication overhead for exchanging state and DNN parameters between AP10, and it also usually requires strong assumptions on complete state knowledge at each AP in order to obtain performance guarantees.

[0130] Using the above-described methods and conventional schemes, the shortcomings and deficiencies of these alternative resource allocation methods can be overcome.

[0131] Example implementation:

[0132] In the following, it is shown that in a cellular network environment, Figure 3 A representative example of system 30 in FIG.

[0133] In this example, AP 10 (e.g., base station) performs uplink scheduling for its connected devices 20. Therefore, each AP 10 decides which of its associated devices 20 to schedule in each time slot. For the traffic model, assume that the traffic (in bits) arrives at device k in each time slot t according to a random process. Thus, specifically, for AP i and its N connected devices, assume the following:

[0134] The status of connected devices can be represented as: (against ),in: is the number of data bits that have not yet been delivered by device n in time slot t (queue length); is the channel state between device n and AP i at time slot t.

[0135] The scheduling decision can be expressed as .therefore, is the number of selected devices to be scheduled at each time slot t. Choosing n=0 means deciding not to schedule any device. Given the actions of all APs and the channel state at time slot t, the number of transmission bits sent by device i at time slot t is calculated using Shannon's formula:

[0136] ,

[0137] Where W is the bandwidth used for transmission, T S is the duration of each time slot, and P is the uplink transmission power normalized by the receiver noise power. The number of bits remaining in the queue of device k at the beginning of the next time slot is given by given.

[0138] The policy modules (Policy 1, Policy 2) and information exchange can be described as follows: Policy 1 (non-sharing) is a reinforcement learning (RL)-based uplink scheduler that uses only local observations (i.e., the first information) for actions and is trained using local rewards (without seeing rewards from neighbors, i.e., other APs 10). Policy 2 (sharing) is an RL-based uplink scheduler that takes actions based on the local AP state and the states of neighboring APs (i.e., based on the first information and the second information) and is trained using global rewards (i.e., rewards from all APs 10). Regarding data exchange between APs, an AP 10 does not request any information from other APs (i.e., its neighbors) when using Policy 1, or requests the local states of other APs when using Policy 2.

[0139] The controller actions can be described as follows: When the controller 13 receives its local state, it performs an action And send this action to all AP j of the network; this indicates which neighboring APs need to send information back to AP i.

[0140] The rewards can be described as follows: The local reward of the controller at API can be defined as , where the first term corresponds to the negative of the sum of the queues of devices connected to AP i, and the second term becomes C when an AP requests the local status of a neighboring AP.

[0141] The policy module used in the above example can be a DNN with multiple outputs that can match the number of schedulable devices (one additional output for the "no schedule" action). Controller 13 can also be a DNN with two outputs, one for each selectable policy module.

[0142] Figure 4 FIG4 shows the result of the performance evaluation according to the embodiment of the present disclosure. Specifically, the performance evaluation is performed on the system 30 configured according to the above exemplary implementation.

[0143] For this performance evaluation, a system with N = 4 APs was simulated in a square topology. Therefore, traffic from all devices follows a Poisson distribution with the same mean. K = 20 devices were randomly placed within the coverage area, each associated with the nearest AP. The policy module was updated once every epoch (4 rounds), with each round consisting of 1000 time slots.

[0144] To evaluate the performance of the proposed method43 (i.e., selecting between two policy modules (Policy 1 and Policy 2)), it is compared with the following scheduling algorithms:

[0145] Proportional Fair (PF) Scheduler41: The PF scheduler is the standard uplink scheduler in cellular systems. Its goal is to distribute traffic rates fairly across devices. The PF scheduler allocates radio resources based solely on local information. The PF scheduler is considered the gold standard for performance comparison. However, it generally performs poorly when there is high interaction between APs due to interference.

[0146] Unshared42: An RL-based uplink scheduler that always takes actions based only on local observations (i.e., first information) and is trained using local rewards (unseen rewards from neighbors).

[0147] Share 44: RL-based uplink scheduler that always takes actions based on the local AP state and neighboring AP states (i.e., based on the first information and the second information); and is trained using a global reward.

[0148] Therefore, the two policy modules (Policy 1 and Policy 2) are also used as separate baselines, since the "unshared" algorithm 42 corresponds to always using Policy 1, while the "shared" algorithm 44 corresponds to always using Policy 2. In this way, it is possible to evaluate whether the controller is able to select when one of the two policy modules is appropriate for scheduling decisions based on environmental conditions.

[0149] In principle, we should expect Strategy 2 (sharing 44) to perform better than Strategy 1 (not sharing 42) in terms of the network goal (i.e., the sum of queue lengths on all devices). However, Strategy 2 incurs additional communication costs, so it is advantageous for the controller to avoid Strategy 2 at least a certain percentage of the time, especially when conditions (the state of the queues) are favorable. In this case, the controller should choose Strategy 1 to avoid unnecessary communication costs.

[0150] exist Figure 4 In Figure 1, the curve represents the sum of the queue lengths in the network over a 1000-slot round. The percentages shown on the right represent the percentage of time the AP requests information from its neighbors (i.e., second information). The following conclusions can be drawn from this performance evaluation:

[0151] Regarding the baselines, policies PF 41 and "No Sharing" 42 performed poorly in the chosen scenario because they acted greedily based only on local state (i.e., based only on the first piece of information). These algorithms 41 and 42 failed to maintain a stable queue length, as evidenced by the steady increase we observed. On the other hand, the "Sharing" 44 baseline was able to maintain a stable queue length.

[0152] The proposed strategy 43 is able to achieve a stable queue based on the possibility of switching between "sharing" and "not sharing". However, the significant difference is that strategy 43 achieves this by exchanging information between APs only 11% of the time, thus achieving a significant communication gain between controllers compared to the "sharing" strategy, which exchanges such information 100% of the time.

[0153] In addition to the scheduling algorithms presented above, other approaches can be considered. For example, other distributed algorithms can be used—these algorithms solve optimization problems at each time slot in the system, typically aiming to maximize a weighted sum rate. However, these algorithms often suffer from slow decision-making, high information exchange, and, by short-sightedly optimizing per-time slot objectives, they often fail to achieve the desired long-term behavior of the system.

[0154] Figure 5 A flow chart illustrating a method 50 of operating an access point 10 for a wireless communication network according to an embodiment of the present disclosure is shown.

[0155] The method 50 includes the following steps: receiving 51 first information from at least one communication device 20 in a wireless communication network; selecting 52 a policy module from M selectable policy modules 14-1, 14-2, each of which is configured to calculate a resource allocation decision; broadcasting 53 a message regarding the policy module selection to at least one other access point 10' in the wireless communication network; and receiving 54 second information from at least one other access point in the wireless communication network based on the selected policy module. The M selectable policy modules include: a first policy module 14-1 configured to calculate a resource allocation decision based on the received first information without requiring the second information; and at least one second policy module 14-2 configured to calculate a resource allocation decision based on the received first information and the received second information. The method also includes calculating 55 the resource allocation decision using the selected policy module.

[0156] Figure 6 Another flow chart shows a method 60 of operating an access point 10 for a wireless communication network according to an embodiment of the present disclosure. The method 60 may be based on Figure 5 The method 50 shown in FIG. 5 is further described and can be expanded.

[0157] Method 60 includes the following steps: observing 61 the local state of a communication device 20 associated with an access point 10. First information may thereby be collected. In a subsequent step 62, the controller 13 of the access point 10 selects a policy module, for example, based on the local state. The access point then signals 63 the selected policy module to other access points 10 in the wireless communication network. Thus, the controllers 13 of the access point 10 may exchange messages indicating the policy module to be used. The access point may also signal 64 the required input (e.g., second information) for the selected policy module. The controllers 13 of the access point 10 may exchange this data only when needed. In a subsequent step 65, the selected policy module may calculate a resource allocation decision. This decision may be executed by the access point 10.

[0158] After performing resource allocation, the access point controller 13 may observe 66 the environment's response. The response may include a local reward and a local new state (eg, the local new state and local reward of the communication device 20) as well as rewards exchanged between the access points 10.

[0159] In a subsequent step 67, access point 10 may decide whether to update its controller 13. For example, if controller 13 is a trainable controller. If not, access point 10 may store its observations in a replay buffer (step 68) and return to the initial step 61 of observing the local state of device 20. If access point 10 decides to update the controller, the update may include adjusting the controller DNN weights using the global reward (step 69). After the update, the replay buffer may be cleared (step 70), and the access point may return to the initial step 61 of observing the local state of device 20.

[0160] Method 50 and Method 60 can be used Figures 1 to 3 This may be performed by any of the access points 10, 10' shown in FIG.

[0161] The above method has the following advantages:

[0162] Each access point 10 can adapt to the dynamic environment of the wireless network. Each access point 10 can select the policy module 14-1 or 14-2 that best suits the current network conditions. This can avoid unnecessary data exchange and reduce communication costs.

[0163] • The learned controller policy can be different for each access point 10 , so each access point 10 can adapt to the traffic and channel characteristics of its associated communication devices 20 .

[0164] • The policy modules 14-1, 14-2 can be updated in a manner that statistically guarantees an increase in the cumulative total reward.

[0165] • The method is distributed and can therefore be executed concurrently at each access point 10.

[0166] Specifically, by allowing each access point 10 of the system 30 to have a different policy module 14-1, 14-2, some policy modules may calculate resource allocations based only on local (first) information, while other policy modules calculate resource allocations based on both local and neighbor (first and second) information. This system architecture and its accompanying signaling allows for more flexible communication between access points 10.

[0167] The access points 10 can communicate according to multiple different policy modules 14-1, 14-2, and the controller can know which of these modules is best for certain conditions. In this way, both network-related performance indicators and communication between access points 10 can be optimized.

[0168] Therefore, distributed wireless resource allocation can benefit from artificial intelligence (AI), particularly deep neural networks (DNNs), which can be implemented in the controller 13 of each access point 10. By combining the representational power of DNNs with reinforcement learning (RL), a decision-making framework for distributed resource management that adapts to dynamic environments can be implemented. Reinforcement learning is a field of machine learning in which intelligent agents (e.g., different access points in a mobile network) learn optimal behavior strategies based on trial-and-error interactions with the environment to ultimately maximize long-term goals.

[0169] However, it's worth noting that directly applying "raw" RL methods may not yield the expected results for two reasons. First, simple RL algorithms (e.g., Q-learning) may suffer from the curse of dimensionality, limiting their applicability to problems with a small number of states. Second, in addition to the dynamics of the environment (which a single RL agent can navigate), each agent should also consider the actions of other agents taking similar actions, leading to the so-called multi-agent reinforcement learning (MARL). Intuitively, the agents (APs) should learn to predict the decisions of other APs and, based on this, make strategic responses to maximize a global performance metric. The former (i.e., the curse of dimensionality) can be circumvented using DNNs, which are known for their ability to approximate RL functions; however, DNNs may not be able to address the latter. When multiple agents act in a distributed scenario (e.g., each agent determines which devices to schedule), coordination and information exchange between them can lead to significant performance improvements. Therefore, according to the above-described system 30 and methods 50 and 60 , the access points 10 can communicate with each other so as to: (a) maximize long-term performance indicators; and (b) exchange information between agents only when necessary.

[0170] The present disclosure has been described with reference to various exemplary embodiments and implementations. However, other variations will be apparent to and will be realized by those skilled in the art in practicing the claimed subject matter, based on a study of the drawings, the present disclosure, and the independent claims. In the claims and the specification, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single element or other unit may fulfil the functions of several entities or items set out in the claims. The recitation of certain measures in mutually different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.

Claims

1. An access point (10) for a wireless communication network, the access point (10) comprising: A first interface (11) configured to receive first information from at least one communication device (20) in the wireless communication network; a second interface (12) configured to receive second information from at least one other access point (10') in the wireless communication network; M selectable policy modules (14-1, 14-2), each of the selectable policy modules (14-1, 14-2) being configured to calculate resource allocation decisions, wherein the M selectable policy modules include: - a first policy module (14-1) configured to calculate the resource allocation decision based on the received first information without requiring the second information; and - at least one second policy module (14-2) configured to calculate the resource allocation decision based on the received first information and the received second information; The controller (13) is configured to select a policy module from the M selectable policy modules (14-1, 14-2) to calculate the resource allocation decision.

2. The access point (10) according to claim 1, The controller (13) is configured to select the policy module from the M selectable policy modules (14-1, 14-2) based on the received first information.

3. The access point (10) according to claim 1 or 2, The controller (13) is configured to broadcast a message including the selection of the policy module to one, a plurality of or all other access points (10') in the wireless communication network.

4. Access point (10) according to any one of the preceding claims, The second information includes a response of the at least one other access point (10') in the wireless communication network to the broadcast message.

5. Access point (10) according to any one of the preceding claims, The second interface (12) is further configured to receive a broadcast message regarding the selection of a policy module from at least one other access point (10') in the wireless communication network.

6. The access point (10) according to claim 5, The access point (10) is configured to send other information back to the at least one other access point (10') in response to receiving the broadcast message.

7. Access point (10) according to any one of the preceding claims, The controller (13) is a trainable controller that can be trained for the selection of the policy module.

8. Access point (10) according to any one of the preceding claims, wherein after executing the calculated resource allocation decision, - the first interface (11) is configured to receive reward feedback from at least one communication device (20) in the wireless communication network, and / or The second interface (12) is configured to receive reward feedback or an accumulation of reward feedback from at least one other access point (10') in the wireless communication network.

9. Access point (10) according to any one of the preceding claims, The controller (13) comprises a trainable neural network.

10. Access point (10) according to claims 8 and 9, The controller (13) is configured to train the trainable neural network by: - calculating a global reward based on the reward feedback from the at least one communication device and the one or more reward feedbacks from the at least one other access point, for example, using a Monte Carlo simulation; - inputting the global reward into a loss function; and - adjusting the trainable neural network based on a result of the loss function.

11. Access point (10) according to any one of the preceding claims, Each of the selectable policy modules (14-1, 14-2) comprises a set of rules, such as an algorithm, stored in a memory of the access point.

12. The access point (10) according to claim 11, The controller (13), such as a neural network of the controller (13), is further configured to adjust the rule of at least one of the M selectable strategy modules (14-1, 14-2) at a certain time interval.

13. Access point (10) according to any one of the preceding claims, Each of the selectable strategy modules (14-1, 14-2) includes a corresponding other trainable neural network.

14. The access point (10) according to claim 13, The other trainable neural networks of the optional strategy modules (14-1, 14-2) are configured to be trained individually by calculating an individual loss for each other trainable neural network and adjusting the corresponding other trainable neural networks based on the individual losses.

15. Access point (10) according to any one of claims 9 or 10 and claim 13, The other trainable neural networks of the optional policy modules (14-1, 14-2) are configured to be jointly trained with the trainable neural network of the controller (13) by calculating a joint loss of the trainable neural network of the controller (13) and the other trainable neural networks of the optional policy modules (14-1, 14-2), and adjusting the trainable neural network and the other trainable neural networks based on the joint loss.

16. A system (30) comprising at least two access points (10) according to any one of the preceding claims.

17. A method (50) of operating an access point (10) for a wireless communication network, the method (50) comprising the steps of: receiving (51) first information from at least one communication device (20) in the wireless communication network; selecting (52) a policy module from M selectable policy modules (14-1, 14-2), each selectable policy module of the selectable policy modules (14-1, 14-2) being configured to compute a resource allocation decision; broadcasting (53) a message regarding the selection of the policy module to at least one other access point (10') in the wireless communication network; receiving (54) second information from the at least one other access point (10') in accordance with the selected policy module; The M selectable strategy modules include: - a first policy module (14-1) configured to calculate the resource allocation decision based on the received first information without requiring the second information; - at least one second policy module (14-2) configured to calculate the resource allocation decision based on the received first information and the received second information; The resource allocation decision is calculated (55) using the selected policy module.