Method and base station for efficient and adaptive resource allocation

The method employs a DNN policy for base stations to optimize resource allocation by determining states, receiving rewards, and updating hyperparameters, addressing dynamic network challenges and enhancing wireless network efficiency through coordinated learning.

WO2026046490A1PCT designated stage Publication Date: 2026-03-05HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/073798
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing wireless networks face challenges in efficiently and adaptively allocating radio resources due to dynamically changing conditions, leading to inaccurate and suboptimal resource allocation, high interference, and high communication overhead, which conventional methods like reinforcement learning fail to address effectively.

Method used

A method for a base station using a deep neural network (DNN) policy to determine initial and resulting states, receive local and neighboring rewards, and update hyperparameters through a hyper DNN policy to optimize resource allocation, reducing interference and enhancing network efficiency through collaborative learning among base stations.

Benefits of technology

The method enables efficient, adaptive, and reliable resource allocation by leveraging DNN policies for real-time feedback and continuous optimization, reducing interference and improving network performance through coordinated decision-making among base stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024073798_05032026_PF_FP_ABST
    Figure EP2024073798_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A method for a base station associated with one or more terminal devices comprising determining an initial state at a first time, determining a resource allocation for the one or more terminal devices associated with the base station based on a Deep Neural Network, DNN, policy, causing the resource allocation to be executed, determining a resulting state at a second time preceding the execution of the resource allocation, determining a local reward for the base station, receiving a neighbouring reward from a neighbouring base station, determining a group reward based on the local reward and the received neighbouring reward, receiving a previous neighbouring hyper parameter, updating a local hyper parameter based on the previous neighbouring hyper parameter and the group reward, wherein updating the local hyper parameter utilizes a hyper DNN, policy and training the DNN policy based on transitions of the base station and the updated local hyper parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD AND BASE STATION FOR EFFICIENT AND ADAPTIVE RESOURCE ALLOCATION

[0002] TECHNICAL FIELD

[0003] The present disclosure relates generally to the field of wireless communication networks and more specifically, to a method for a base station associated with a one or more terminal devices and a base station associated with the one or more terminal devices, such as for resource allocation based on collaborative lifelong multi-agent reinforcement learning (MARL).

[0004] BACKGROUND

[0005] In existing wireless networks, a base station is connected to a large number of devices to allocate radio resources, ensuring effective communication within the network. Moreover, radio resource allocation policies are used to make these allocation decisions based on the available network-related information. However, due to the dynamically changing conditions of wireless networks, these resource allocation policies may often fail, requiring operators to manually relocate resources and adjust the policies for future allocations. Additionally, the adoption of decentralized base station policies, where each base station can take multiple actions based on local information received from connected devices, results in high interference and further degrades overall network performance.

[0006] Conventionally, in wireless communication networks, resource allocation policies do not consider the changing behavior of the environment within the network, resulting in inaccurate, unreliable, and suboptimal resource allocation. Attempts to improve this, such as using reinforcement learning and similar methods, have often failed due to various reasons, including low adaptation rates, high device interference among base stations, and communication overhead. Thus, there exists a technical problem of how to train decentralized base station (BS) radio resource allocation policies with low communication overhead between neighboring base stations, ensuring automatic adaptation to changing environments.

[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional base stations and conventional methods for resource allocation based on collaborative lifelong multi-agent reinforcement learning (MARL).

[0008] SUMMARY

[0009] The present disclosure provides a method for a base station and the base station associated with one or more terminal devices. The present disclosure provides a solution to the existing problem of how to train decentralized base station (BS) radio resource allocation policies with low communication overhead between neighboring base stations, ensuring automatic adaption to changing environments. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the base station and the method for base station associated with one or more terminal devices, such as for resource allocation based on collaborative lifelong multi-agent reinforcement learning (MARL).

[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.

[0011] In one aspect, the present disclosure provides a method for a base station and the base station being associated with one or more terminal devices. The method comprises determining an initial state at a first time, determining a resource allocation for the one or more devices associated with the base station based on a deep neural network (DNN) policy, causing the resource allocation to be executed, determining a resulting state at a second time preceding the execution of the resource allocation. Furthermore, the method includes determining a local reward for the base station, receiving a neighboring reward from a neighboring base station, determining a group reward based on the local reward and the received neighboring reward, receiving a previous neighboring hyper parameter from the neighboring base station, updating a local hyper parameter based on the previous neighboring hyper parameter and the group reward. Moreover, updating the local hyperparameter utilizes a hyper deep neural network (DNN) policy and training the DNN policy based on transitions of the base station and the updated local hyper parameter, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the group reward.

[0012] Advantageously, the method for the base station associated with one or more terminal devices provides efficient and adaptive resource allocation through a deep neural network (DNN) policy. By determining an initial state at the first time and a resource allocation based on the DNN policy, the method ensures that resource allocation decisions are data-driven and optimized for current wireless network conditions. Moreover, causing the resource allocation to be executed and determining a resulting state at a second time allows for real-time feedback and adjustment of resource allocation strategies, enabling quick adaptation to change in the wireless network environment. Furthermore, by determining a local reward for the base station and receiving a neighboring reward from a neighboring base station, the method enhances the overall performance of the wireless network. The determination of the group reward based on local and neighboring rewards, along with the exchange of hyper parameters and the agreement on the reward statistic function with neighboring base stations, helps mitigate network interference and improve overall wireless network efficiency by ensuring that base stations work together rather than in isolation. Additionally, updating the local hyper parameter based on the previous neighbouring hyper parameter and the group reward using a hyper DNN policy allows continuous optimization of hyper parameters for optimal resource allocation policies. As a result, the method provides a comprehensive and adaptive approach to resource allocation in wireless networks, leveraging advanced machine learning techniques to optimize performance, reduce interference, and enhance coordination among base stations, thereby providing a more efficient, reliable, and scalable wireless network.

[0013] In another aspect, the present disclosure provides a base station associated with one or more terminal devices, the base station is configured to determine an initial state at a first time, determine a resource allocation for the one or more devices associated with the base station based on a deep neural network (DNN) policy, cause the resource allocation to be executed, determine a resulting state at a second time preceding the execution of the resource allocation, determine a local reward for the base station, receive a neighbouring reward from a neighbouring base station, determine a group reward based on the local reward and the received neighbouring reward, receive a previous neighbouring hyper parameter from the neighbouring base station, update a local hyper parameter based on the previous neighbouring hyper parameter and the group reward. Moreover, updating the local hyper parameter utilizes a hyper deep neural network (DNN) policy, and trains the DNN policy based on transitions of the base station and the updated local hyper parameter, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state, and the group reward.

[0014] The base station achieves all the advantages and technical effects of the method of the present disclosure.

[0015] It is to be appreciated that all the aforementioned implementation forms can be combined.

[0016] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.

[0017] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.

[0018] BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.

[0020] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:

[0021] FIG. 1 is a flowchart of a method for a base station associated with one or more terminal devices;

[0022] FIG. 2 is a block diagram of a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure;

[0023] FIG. 3 is a diagram that illustrates an architecture and a message exchange between a base station and a neighboring base station of the same interference group, in accordance with an embodiment of the present disclosure;

[0024] FIG. 4 is a diagram that illustrates an exemplary scenario of a base station operating in a wireless network, in accordance with an embodiment of the present disclosure; and

[0025] FIG. 5 is a diagram that illustrates another exemplary scenario of a base station operating in a wireless network, in accordance with an embodiment of the present disclosure.

[0026] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.

[0027] DETAILED DESCRIPTION OF EMBODIMENTS

[0028] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.

[0029] FIG. 1 is a flowchart of a method for a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure. With reference to FIG. 1 , there is shown a flowchart of a method 100 that includes steps 102 to 120. The base station is configured to execute the method 100.

[0030] There is provided the method 100 for the base station associated with the one or more terminal devices. The method 100 is used to provide a comprehensive and adaptive resource allocation in wireless networks by using the DNN policy. Moreover, by continuously optimizing the hyperparameters and coordinating with neighboring base stations, the method 100 is used to ensure an efficient, data-driven resource allocation with improved overall network performance, reduced interference, and a reliable wireless network. At step 102, the method 100 includes determining an initial state at a first time. In an implementation, a controller of the base station is configured to collect relevant data, such as the number of connected devices, network conditions, traffic demands, and the like to understand the current network conditions of the base station. Moreover, by aggregating and analyzing the collected data, the method 100 is used to determine the initial state at the first time. In an implementation, the initial state refers to a state of the base station that is used to determine the changing network conditions within the wireless network. As a result, the determination of the initial state at the first time is to allow effective and efficient utilization of resources.

[0031] At step 104, the method 100 includes determining a resource allocation for the one or more terminal devices associated with the base station based on a deep neural network (DNN) policy. In an implementation, the DNN policy refers to the decisions that are derived from a deep neural network. The DNN policy processes data, such as network conditions, device requirements, channel states, and the like to determine optimal actions, for example, resource allocation, scheduling, power control, and the like in order to ensure enhanced and improved network performance for optimized resource allocation. The DNN policy is used to process the initial state data and other relevant inputs to determine the optimal distribution of resources among the connected one or more terminal devices. As a result, the determination of the resource allocation for the one or more terminal devices associated with the base station based on the DNN policy is used to allow an accurate, reliable, and adaptive resource utilization.

[0032] At step 106, the method 100 includes causing the resource allocation to be executed. In an implementation, the base station is configured to cause the resource allocation to be executed to the one or more terminal devices based on the DNN policy that includes the determination of bandwidth, power, time slots, and the like. As a result, by causing the execution of the resource allocation, the base station is configured to dynamically adjust resource allocation based on current network conditions that enhance the overall efficiency, reduce device interference, and enhance the overall resource utilization within the wireless network.

[0033] At step 108, the method 100 includes determining a resulting state at a second time preceding the execution of the resource allocation. In an implementation, the method 100 is used to collect data, such as device throughput, latency, signal strength, network conditions, and the like in order to determine the resulting state at the second time preceding the execution of the resource allocation. Moreover, the resulting state refers to a state that reflects the changes and outcomes caused by the execution of the resource allocation. As a result, the determination of the resulting state at the second time preceding the execution of the resource allocation is used to dynamically monitor the resources allocated to the one or more terminal devices within the wireless network in order to identify the changing network conditions and take necessary required measures.

[0034] At step 110, the method 100 includes determining a local reward for the base station. Firstly, the method 100 includes the determination of the initial state at the first time. After that, the resource allocation for the one or more terminal devices associated with the base station based on the DNN policy is determined and then the resource allocation is executed. After that, the resulting state at the second time preceding the execution of the resource allocation is determined. Moreover, after the determination of the resulting state, the local reward for the base station is determined. In an implementation, the local reward for the base station is determined based on the summation of the rewards (i.e., rf G R) of the one or more terminal devices associated with the base station.

[0035] At step 112, the method 100 includes receiving a neighboring reward from a neighboring base station. In an example, a first base station receives the neighbouring reward from a second base station. Similarly, in another example, the second base station receives the neighbouring reward from the first base station. Moreover, the transmission of neighboring rewards from the neighbouring base station is used to analyze the overall condition of the wireless network in order to reduce interference between base stations, leading to an optimized network environment with reduced interference, improved resource utilization, and an efficient wireless network.

[0036] At step 116, the method 100 includes receiving a previous neighbouring hyper parameter from the neighbouring base station. In an implementation, the base station is configured to communicate with the neighbouring base stations and exchange previous neighbouring hyperparameters that are used to train the DNN policies and include learning rates, exploration rates, or other configuration settings that affect the performance of the base station. The method 100 is used to receive the hyperparameter from the neighbouring base stations to ensure that each of the base station has access to the required information to align the resource allocation strategies. As a result, by sharing the previous neighbouring hyperparameters, the base stations are configured to coordinate the actions of each of the base stations that are included within the wireless network, which leads to an efficient and effective network-wide resource allocation.

[0037] In accordance with an embodiment, the method 100 further comprises transmitting the local reward for the base station to the neighbouring base station and transmitting the previous local hyper parameter to the neighbouring base station. Moreover, the base station is configured to transmit the previous local hyperparameter used to train the DNN policy to allow the neighbouring base station to incorporate the received data into their own decision-making processes, ensuring a synchronized and cooperative resource allocation with an improved wireless network.

[0038] In accordance with an embodiment, the method 100 further comprises exchanging an identifier of the hyper DNN policy with the neighbouring base station. In an implementation, the identifier of the hyper DNN policy refers to a unique identifier that represents the hyper DNN policy of the associated base station. As a result, exchanging the identifier of the hyper DNN policy with the neighboring base station is used to provide a reliable and accurate coordination and alignment of the resource allocation to optimize the wireless network that supports various changing network conditions.

[0039] In accordance with an embodiment, the method 100 further comprises exchanging an indicator of a function for determining the reward statistics. The indicator of the function refers to an indicator that is used to identify the function or an algorithm, which is used to calculate the reward statistics. Moreover, such indicators are transmitted to the neighbouring base stations. The exchange of the indicator of the function for determining the reward statistics is used to allow the neighbouring base stations to evaluate the network condition and the performance of the neighbouring base stations for coordinated and optimized resource allocation decisions across the wireless network, leading to reduced interference and optimized overall performance of the wireless network.

[0040] In accordance with an embodiment, the method 100 further comprises exchanging an initial hyper parameter with the neighbouring base station. The initial hyper parameter refers to a hyper parameter that is used at the start of the training of the DNN policy. At step 118, the method 100 includes updating a local hyper parameter based on the previous neighbouring hyper parameter and the group reward. Moreover, updating the local hyper parameter utilizes a hyper DNN policy. In an implementation, the base station is configured to collect the previous neighbouring hyperparameter and the group reward, which is derived from the local reward and the neighbouring reward. Thereafter, the base station is configured to calculate the local hyper parameter based on the previous neighbouring hyper parameter and the group reward. Moreover, the local hyper parameter based on the previous neighbouring hyper parameter and the group reward is updated to adapt to the changing network conditions during the resource allocation. As a result, the method 100 is used to ensure that the resource allocation through the DNN policy is optimized locally as well as within the wireless network with reduced interference among the one or more terminal devices associated with the base station.

[0041] In accordance with an embodiment, the method 100 further comprises determining reward statistics for the group reward and updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the reward statistics. In an implementation, the base station is configured to calculate the reward statistics from the group reward, such as by calculating mean, variance, and other relevant statistical properties of the rewards accumulated over a defined period. Moreover, such reward statistics are used to provide a comprehensive analysis of the wireless network. The reward statistics are further utilized to update the local hyper parameter in order to provide an enhanced and improved resource allocation within the wireless network with reduced device interference among the one or more terminal devices.

[0042] In accordance with an embodiment, updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the k latest group rewards. Moreover, k is a natural number. In other words, the base station is configured to identify the latest k group rewards, such as by aggregating updated rewards. Thereafter, based on the identified latest k rewards, the base station is configured to update the local hyper parameter, which is determined using the hyper DNN policy. As a result, by updating the local hyper parameter based on the group reward based on the k latest group rewards is used to ensure that the base station and the associated one or more terminal devices are in compliance with the changing conditions of the network's environment thereby leading to an improved resource allocation and efficient overall wireless network.

[0043] At step 120, the method 100 includes training the DNN policy based on transitions of the base station and the updated local hyper parameter, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state, and the group reward. In an implementation, the transitions, such as the initial state and the resulting state of the base station are used to train the DNN policy. As a result, the training of the DNN policy based on the transitions of the base station and the updated local hyper parameter is used to enhance the decision-making capabilities of the base station in order to perform resource allocation in order to provide an efficient, reliable, and enhance overall network performance.

[0044] In accordance with an embodiment, the method 100 further comprises training the hyper DNN policy based on previous transitions. The training of the hyper DNN policy based on the previous transitions allows the optimization of the hyperparameters of the DNN policy dynamically in order to allow the base stations to adapt to the changing network conditions and improve the overall performance of the wireless network.

[0045] Advantageously, the method 100 for the base station associated with one or more terminal devices provides efficient and adaptive resource allocation through a deep neural network (DNN) policy. By determining an initial state at the first time and a resource allocation based on the DNN policy, the method ensures that resource allocation decisions are data-driven and optimized for current wireless network conditions. Moreover, causing the resource allocation to be executed and determining a resulting state at a second time allows for real-time feedback and adjustment of resource allocation strategies, enabling quick adaptation to change in the wireless network environment. Furthermore, by determining a local reward for the base station and receiving a neighbouring reward from a neighbouring base station, the method enhances the overall performance of the wireless network. The determination of the group reward based on local and neighbouring rewards, along with the exchange of hyper parameters and the agreement on a common reward statistic function with neighbouring base stations, helps mitigate network interference and improve overall wireless network efficiency by ensuring that base stations work together rather than in isolation. Additionally, updating the local hyper parameter based on the previous neighboring hyper parameter and the group reward using a hyper DNN policy allows continuous optimization of hyper parameters for optimal resource allocation policies. As a result, the method 100 provides a comprehensive and adaptive approach to resource allocation in wireless networks, leveraging advanced machine learning techniques to optimize performance, reduce interference, and enhance coordination among base stations, thereby providing a more efficient, reliable, and scalable wireless network.

[0046] The steps 102 to 120 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claim herein.

[0047] There is further provided a computer program product comprising program instructions for performing the method 100 when executed by one or more processors in the base station. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Elard Disk Drive (EIDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0048] FIG. 2 is a block diagram of a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a block diagram that includes a base station 202A, a communication network 210, a neighbouring base station 202B, and one or more terminal devices 212 associated with the base station 202A within a wireless network 200.

[0049] The base station 202A refers to an access point (AP) operating in the wireless network 200 and is responsible for allocating radio resources to the connected one or more terminal devices (e.g., user equipment, UEs) within the wireless network 200. Moreover, the neighboring base station 202B refers to another AP in the wireless network 200.

[0050] The controller (i.e., a first controller 204A) of the base station 202A is configured to determine an initial state (i.e., s‘i) at the first time (i.e., t) and further receive previous neighboring hyper parameter (i.e., Hj) from the neighboring base station 202B, such as through the controller (i.e., a second controller 204B) of the neighboring base station 202B. Examples of the controllers (i.e., the first controller 204A and the second controller 204B) of the base station 202A and the neighboring base station 202B may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.

[0051] A first memory 206A and a second memory 206B are used to store received neighboring rewards, group rewards, hyper parameters, and the like. Examples of implementation of the first memory 206A and the second memory 206B may include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Dynamic Random Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory. A first network interface 208A is used by the base station 202A and the second network interface 208B is used by the neighboring base station 202B to communicate with the first controller 204 A and the second controller 204B respectively. Examples of implementation of the first network interface 208A and the second network interface 208B may include but are not limited to a network interface, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.

[0052] The communication network 210 includes a medium (e.g., a communication channel) through which the base station 202A communicates with the neighboring base station 202B of the wireless network 200. Examples of the communication network 210 may include, but are not limited to, a cellular network (e.g., a 2G, a 3G, long-term evolution (LTE) 4G, a 5G, or 5G New Radio (NR) network, such as sub 6 GHz, cmWave, or mmWave communication network), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to-network (V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.

[0053] There is provided the base station 202 A associated with the one or more terminal devices 212. The base station 202 A is configured to provide a comprehensive and adaptive resource allocation in wireless networks by using the DNN policy. Moreover, by continuously optimizing the hyperparameters and coordinating with neighboring base stations, the base station 202A is configured to ensure an efficient, data-driven resource allocation with improved overall network performance, reduced interference, and a reliable wireless network (i.e., the wireless network 200).

[0054] In operation, the base station 202A is configured to determine an initial state at the first time. The determination of the initial state at the first time is to allow effective and efficient utilization of resources. Furthermore, the base station 202A is configured to determine a resource allocation (i.e., a)) for the one or more terminal devices 212 associated with the base station 202A based on the DNN policy. In an example, the base station 202A is configured to determine a resource allocation (i.e., a)) for a first terminal device 212A associated with the base station 202A based on the DNN policy. In another example, the base station 202A is configured to determine a resource allocation (i.e., a)) for a second terminal device 212B associated with the base station 202 A based on the DNN policy. In yet another example, the base station 202 A is configured to determine a resource allocation (i.e., a)) for nth terminal device 212N associated with the base station 202A based on the DNN policy. The DNN policy is used to process the initial state data and other relevant inputs to determine the optimal distribution of resources among the connected one or more terminal devices. As a result, the determination of the resource allocation for the one or more terminal devices associated with the base station based on the DNN policy is used to allow an accurate, reliable, and adaptive resource utilization. Furthermore, the base station 202A is configured to cause the resource allocation (i.e., a)) to be executed. In an implementation, by causing the execution of the resource allocation, the base station 202A is configured to dynamically adjust resource allocation based on current network conditions that enhance the overall efficiency, reduce device interference, and enhance the overall resource utilization within the wireless network (i.e., the wireless network 200). Furthermore, the base station 202A is configured to determine a resulting state (i.e., st+1i) at a second time (i.e., t+1) preceding the execution of the resource allocation. Moreover, the determination of the resulting state at the second time preceding the execution of the resource allocation is used to dynamically monitor the resources allocated to the one or more terminal devices 212 within the wireless network (or the wireless network 200) in order to identify the changing network conditions and take necessary required measures. Furthermore, the base station 202A is configured to determine a local reward (i.e., r‘i) for the base station 202A. Furthermore, the base station 202A is configured to receive a neighboring reward (i.e., r j) from the neighboring base station 202B. The transmission of neighboring reward from the neighboring base station is used to analyze the overall condition of the wireless network in order to reduce interference between base stations, leading to an optimized network environment with reduced interference, improved resource utilization, and an efficient wireless network. Furthermore, the base station 202A is configured to determine a group reward (i.e., R) based on the local reward (i.e., r‘i) and the received neighbouring reward (i.e., Tj). Furthermore, the base station 202A is configured to receive a previous neighbouring hyper parameter (i.e. ,Hj" ) from the neighbouring base station 202B. As a result, by sharing the previous neighbouring hyperparameters, the base stations are configured to coordinate the actions of each of the base stations that are included within the wireless network, which leads to an efficient and effective network-wide resource allocation. Furthermore, the base station 202A is configured to update a local hyper parameter (i.e., Hi) based on the previous neighboring hyper parameter (i.e., Hj") and the group reward (i.e., R). Moreover, updating the local hyper parameter (i.e., Hi) utilizes a hyper DNN policy. The updating of the previous neighboring hyper parameter and the group reward is to ensure that the resource allocation through the DNN policy is optimized locally as well as within the wireless network with reduced interference among the one or more terminal devices associated with the base station.

[0055] In accordance with an embodiment, the base station 202A is further configured to determine reward statistics (i.e., h(R)) for the group reward (i.e., R) and update the local hyper parameter (i.e., Hi) based on the group reward (i.e., R) includes updating the local hyper parameter (i.e., Hi) based on the reward statistics (i.e., h). The exchange of the indicator of the function for determining the reward statistics is used to allow the neighboring base stations to agree on a common function that evaluates the network condition and the performance of the neighboring base stations for coordinated and optimized resource allocation decisions across the wireless network, leading to reduced interference and optimized overall performance of the wireless network.

[0056] In accordance with an embodiment, the base station 202 A is further configured to update the local hyper parameter (i.e., Hi) based on the group reward (i.e., R) by updating the local hyper parameter (i.e., Hi) based on the k latest group reward (i.e., R). Moreover, k is a natural number. By updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the k latest group rewards are used to ensure that the base station and the associated one or more terminal devices are in compliance with the changing conditions of the network's environment thereby leading to an improved resource allocation and efficient overall wireless network.

[0057] In accordance with an embodiment, the base station 202A is further configured to transmit the local reward (i.e., r‘i) for the base station 202A to the neighbouring base station 202B, and transmit the previous local hyper parameter (i.e., H ) to the neighbouring base station 202B. In other words, the base station 202A is configured to transmit the previous local hyperparameter used in the DNN policy to allow the neighbouring base station 202B to incorporate the received data into their own decision-making processes, ensuring a synchronized and cooperative resource allocation with an improved wireless network (i.e., the wireless network 200).

[0058] Furthermore, the base station 202A is configured to train the DNN policy based on transitions (i.e., Ti) of the base station 202A and the updated local hyper parameter (i.e., Hi), where a transition (i.e., Ti) is a tuple comprising the initial state (i.e., s‘i), the resource allocation (i.e., a)), the resulting state (st+1i) and the group reward The training of the DNN policy based on the transitions of the base station and the updated local hyper parameter is used to enhance the decision-making capabilities of the base station in order to perform resource allocation in order to provide an efficient, reliable, and enhance overall network performance.

[0059] Advantageously, the base station 202A associated with the one or more terminal devices 212 is configured to provide an efficient and adaptive resource allocation through a deep neural network (DNN) policy. By determining an initial state at the first time and a resource allocation based on the DNN policy, the method ensures that resource allocation decisions are data- driven and optimized for current wireless network conditions. Moreover, causing the resource allocation to be executed and determining a resulting state at a second time allows for real-time feedback and adjustment of resource allocation strategies, enabling quick adaptation to change in the wireless network environment. Furthermore, by determining a local reward for the base station and receiving a neighbouring reward from a neighbouring base station, the method enhances the overall performance of the wireless network. The determination of the group reward based on local and neighbouring rewards, along with the exchange of hyper parameters and reward statistics with neighbouring base stations, helps mitigate network interference and improve overall wireless network efficiency by ensuring that base stations work together rather than in isolation. Additionally, updating the local hyper parameter based on the previous neighbouring hyper parameter and the group reward using a hyper DNN policy allows continuous optimization of hyper parameters for optimal resource allocation policies. As a result, the base station 202A is configured to provide a comprehensive and adaptive approach to resource allocation in wireless networks, leveraging advanced machine learning techniques to optimize performance, reduce interference, and enhance coordination among base stations, thereby providing a more efficient, reliable, and scalable wireless network.

[0060] FIG. 3 is a diagram that illustrates an architecture and a message exchange between a base station and a neighbouring base station of the same interference group, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIG. 2. With reference to FIG. 3, there is shown a diagram 300 of the architecture and the message exchange between the base station 202A and a second base station (or the neighbouring base station 202B) of the same interference group. The base station 202A is operated in a first local environment 302A, a first DNN hyper 304A, a first receiver or transmitter 306A, a first DNN policy 308A, and the first controller 204A. Similarly, the neighbouring base station 202B is operated in a second local environment 302B, a second DNN hyper 304B, a second receiver or transmitter 306B, a second DNN policy 308B, and the second controller 204B.

[0061] In an implementation, each base station, such as the base station 202A and the second base station (or the neighbouring base station 202B) is configured to send an initialization handshake to the one or more neighbouring base stations of the interference group (e.g., m, = {lD_hyper, ID_hi;H-1}). Moreover, such initialization handshake includes an ID of a hyper parameter, ID of a function, or an algorithm (i.e., h , which is selected from a predefined list and is used to compute the reward statistics, such as at operation 31O.Thereafter, the base station 202A and the second base station (or the neighbouring base station 202B) agree on a common hyper parameter ID and common function (i.e., h). Additionally, the initial handshake includes an initial hyperparameter (i.e., H f ). Moreover, each time when any of the base station joins the interference group, then, in that case, the initialization of handshake is performed. The DNN policies, such as the first DNN policy 308A and the second DNN policy 308B are continuously optimized. At each timeslot, the controller (i.e., the first controller 204A and the second controller 204B) are configured to make decisions based on the local state of the one or more associated terminal devices. The base station 202A is configured to determine the new local state and the local reward as a result of its action and exchanges the local reward with the second base station (or the neighboring base station 202B), such as at operation 312 before saving the transition information (i.e., local state, local action, group reward, and new local state). Moreover, the first controller 204A is configured to execute and compute the reward statistics (i.e., h(RK)), where RKis the group reward vector of K transitions at each timeslot (i.e., k timeslots). Moreover, the first controller 204A is configured to provide h(RK) and past neighbor hyper-parameters (i.e., HZ;) to the DNN-hyper (i.e., the first DNN hyper 304A) in order to compute the updated local hyper-parameter (i.e., H . Thereafter, the first controller 204A is configured to update the parameters 0, of the DNN-policy based on K transitions and H, and send the hyper parameters (i.e., H to the neighboring base station 204B and receive H_i;such as at operation 314. Finally, at each K x L timeslot, the first controller 204A is configured to update the parameters , of the DNN-hyper (i.e., the first DNN hyper 304A) based on K x L transitions. As a result, the architecture is configured to continuously optimize the overall network's performance through collaborative, adaptive decision-making based on local observations and shared information, leading to an improved network efficiency and responsiveness to changing conditions within the network 200.

[0062] FIG. 4 is a diagram that illustrates an exemplary scenario of a base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 4 is described in conjunction with elements from FIGs 1 to 3. With reference to FIG. 4, there is shown a diagram 400 illustrates the operations for resource allocation within the wireless network 200 including a first terminal device 402A, a second terminal device 402B, and a third terminal device 402C associated with the base station 202A. Similarly, a first terminal device 404A, a second terminal device 404B, and a third terminal device 404C is associated with the neighbouring base station 202B.

[0063] FIG. 5 is a diagram that illustrates another exemplary scenario of a base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 5 is described in conjunction with elements from FIGs 1 to 4. With reference to FIG. 5, there is shown a diagram 500 illustrating the operations for resource allocation within the wireless network 200 that including a third base station 504A, the base station 202A, and a second base station 504B. A first terminal device 502A is associated with a third base station 504A, a second terminal device 502B is associated with the base station 202A, and a third terminal device 502C is associated with the second base station 504B, such as at operation 506A, 506B, and 506C respectively.

[0064] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.

Claims

CLAIMS1. A method (100) for a base station (202A), the base station (202A) being associated with one or more terminal devices, the method (100) comprising: determining an initial state at a first time, determining a resource allocation for the one or more terminal devices associated with the base station (202A) based on a Deep Neural Network, DNN, policy, causing the resource allocation to be executed, determining a resulting state at a second time preceding the execution of the resource allocation, determining a local reward for the base station (202A), receiving a neighbouring reward from a neighbouring base station (202B), determining a group reward based on the local reward and the received neighbouring reward, receiving a previous neighbouring hyper parameter from the neighbouring base station (202B), updating a local hyper parameter based on the previous neighbouring hyper parameter and the group reward, wherein updating the local hyper parameter utilizes a hyper Deep Neural Network, DNN, policy and training the DNN policy based on transitions of the base station (202A) and the updated local hyper parameter, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state, and the group reward.

2. The method (100) according to claim 1, wherein the method (100) further comprises determining reward statistics for the group reward, and wherein updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the reward statistics.

3. The method (100) according to any preceding claim, wherein updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the k latest group rewards, wherein k is a natural number.

4. The method (100) according to any preceding claim, wherein the method (100) further comprises training the hyper DNN policy based on previous transitions.

5. The method (100) according to any preceding claim, wherein the method (100) further comprises exchanging an identifier of the hyper DNN policy with the neighboring base station (202B).

6. The method (100) according to claim 2 and 5, wherein the method (100) further comprises exchanging an indicator of a function for determining the reward statistics.

7. The method (100) according to claim 5 or 6, wherein the method (100) further comprises exchanging an initial hyper parameter with the neighbouring base station (202B).

8. The method (100) according to any preceding claim, wherein the method (100) further comprises: transmitting the local reward for the base station (202A) to the neighbouring base station (202B), and transmitting the previous local hyper parameter to the neighbouring base station (202B).

9. A computer program product comprising program instructions for performing the method (100) according to any preceding claim, when executed by one or more processors in a Base Station (202A).

10. A base station (202A) associated with one or more terminal devices, the base station (202A) configured to: determine an initial state at a first time, determine a resource allocation for the one or more terminal devices associated with the base station (202A) based on a Deep Neural Network, DNN, policy, cause the resource allocation to be executed, determine a resulting state at a second time preceding the execution of the resource allocation, determine a local reward for the base station (202A), receive a neighbouring reward from a neighbouring base station (202B), determine a group reward based on the local reward and the received neighbouring reward, receive a previous neighbouring hyper parameter from the neighbouring base station (202B), update a local hyper parameter based on the previous neighbouring hyper parameter and the group reward, wherein updating the local hyper parameter utilizes a hyper Deep Neural Network, DNN, policy, and train the DNN policy based on transitions of the base station (202A) and the updated local hyper parameter, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state, and the group reward.

11. The base station (202A) according to claim 10, wherein the base station (202A) is further configured to determine reward statistics for the group reward, and wherein updating the local hyper parameter based on the group reward includes updating the local hyper parameter based on the reward statistics.

12. The base station (202A) according to claim 10 or 11, wherein the base station (202A) is further configured to update the local hyper parameter based on the group reward by updating the local hyper parameter based on the k latest group rewards, wherein k is a natural number.

13. The base station (202A) according to claims 10, 11, or 12, wherein the base station (202A) is further configured to: transmit the local reward for the base station (202A) to the neighbouring base station (202B), and transmit the previous local hyper parameter to the neighbouring base station (202B).

Citation Information

Patent Citations

  • Joint distributed learning of signaling and policies for radio resource allocation

    WO2024110047A1