Method and base station for resource allocation in dynamic wireless environment with continual-learning
The method uses a DNN policy for base stations to adapt resource allocation to changing environments by exchanging information with neighbors, reducing interference, and improving network performance through collaborative learning.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-05
AI Technical Summary
Existing wireless networks face challenges in adapting decentralized base station radio resource allocation policies to dynamically changing environments due to high interference and communication overhead, leading to inaccurate and unreliable resource allocation.
A method for a base station using a Deep Neural Network (DNN) policy to determine resource allocation, exchange information with neighboring stations, and train the DNN policy based on local and neighboring rewards to adapt to changing conditions, reducing interference and improving network performance through collaborative learning.
The method enables efficient, adaptive resource allocation in dynamic wireless environments by optimizing DNN policies, mitigating device interference, and enhancing network performance through continuous learning without complete retraining.
Smart Images

Figure EP2024073814_05032026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND BASE STATION FOR RESOURCE ALLOCATION IN DYNAMIC WIRELESS ENVIRONMENT WITH CONTINUAI^LE ARNING
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to the field of wireless communication networks and more specifically, to a method for a base station associated with a one or more terminal devices and a base station associated with the one or more terminal devices, such as for resource allocation in a dynamic wireless environment with continual- learning.
[0004] BACKGROUND
[0005] In existing wireless networks, a base station is connected to a large number of devices to allocate radio resources, ensuring effective communication within the wireless network. Moreover, radio resource allocation policies are used to make these allocation decisions based on the available network-related information. However, due to the dynamically changing conditions of the wireless networks, these resource allocation policies may often face challenges due to which the resource allocation policies are updated by the mobile network operator manually for future allocations. Additionally, the adoption of decentralized base station policies, where each base station can take multiple actions based on local information received from the connected devices, results in high interference and degradation in overall network performance.
[0006] Conventionally, in wireless communication networks, resource allocation policies do not consider the changing conditions of the environment within the wireless network, resulting in inaccurate and unreliable resource allocation. Certain attempts have been taken to ensure the implementation of a reliable resource allocation policies, such as by using reinforcement learning, Meta-Gradient RL (Meta-GRL), and the like. However, such attempts often failed due to various reasons, including low adaptation rates, high device interference among base stations, and communication overhead. Thus, there exists a technical problem of how to train decentralized base station radio resource allocation policies with low communication overhead that can be adapted to changing environments.
[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional base stations and conventional methods for resource allocation in a dynamic wireless environment with continual- learning.
[0008] SUMMARY
[0009] The present disclosure provides a method for a base station and the base station associated with one or more terminal devices for resource allocation in a dynamic wireless environment with continual- learning. The present disclosure provides a solution to the existing problem of how to train decentralized base station radio resource allocation policies with low communication overhead that can be adapted to changing environments. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the base station and the method for base station associated with one or more terminal devices, such as for resource allocation in a dynamic wireless environment with continual- learning.
[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.
[0011] In one aspect, the present disclosure provides a method for a base station. The base station being associated with one or more terminal devices, the method comprising determining an initial state at a first time, determining a resource allocation for the one or more terminal devices associated with the base station based on a Deep Neural Network (DNN) policy, causing the resource allocation to be executed, determining a resulting state at a second time preceding the execution of the resource allocation, determining a local reward for the base station, determining local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward. Furthermore, the method includes receiving neighbouring importance values for current transition and historical transitions of a neighbouring base station, determining global importance values based on the local importance values and the neighbouring importance values, determining which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions, receiving neighbouring rewards for previous transitions of the neighbouring base station, and training the DNN policy based on previous transitions of the base station and the received neighbouring rewards.
[0012] Advantageously, the method for a base station associated with one or more terminal devices provides efficient and adaptive resource allocation in a dynamic wireless environment through optimized deep neural network (DNN) policies and continual learning. By determining the initial state at a first time and a resulting state and reward at a second time, the base station can effectively monitor and evaluate resource allocation, allowing real-time adaptation to changing network conditions. Exchanging information, such as rewards, with neighbouring base stations within the wireless network enables the base station to mitigate device interference and improve resource allocation across the network. Additionally, training the DNN policy based on previous transitions of the base station and the received neighbouring rewards ensures continuous learning, which adapts to new environments and changes in network conditions without requiring complete retraining from scratch. Thus, the method allows for dynamic, efficient resource allocation and improved network performance through collaborative learning, enhancing the overall performance of the wireless network.
[0013] In another aspect, the present disclosure provides a base station associated with one or more terminal devices, the base station being configured to determine an initial state at a first time determine a resource allocation for the one or more terminal devices associated with the base station based on a DNN policy, cause the resource allocation to be executed, determine a resulting state at a second time preceding the execution of the resource allocation, determine a local reward for the base station, determine local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward, receive neighbouring importance values for a current transition and historical transitions of a neighbouring base station, determine global importance values based on the local importance values and the neighbouring importance values determine which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions, receive neighbouring rewards for previous transitions of the neighbouring base station, and train the DNN policy based on previous transitions of the base station and the received neighbouring rewards.
[0014] The base station achieves all the advantages and technical effects of the method of the present disclosure.
[0015] It is to be appreciated that all the aforementioned implementation forms can be combined.
[0016] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.
[0017] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.
[0018] BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.
[0020] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:
[0021] FIG. 1 is a flowchart of a method for a base station associated with one or more terminal devices;
[0022] FIG. 2 is a block diagram of a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure;
[0023] FIG. 3 is a diagram that illustrates an architecture and a message exchange between a base station and a neighbouring base station of the same interference group, in accordance with an embodiment of the present disclosure; and
[0024] FIG. 4 is a diagram that illustrates an exemplary scenario of base stations operating in a wireless network, in accordance with an embodiment of the present disclosure.
[0025] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0026] DETAILED DESCRIPTION OF EMBODIMENTS
[0027] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0028] FIG. 1 is a flowchart of a method for a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure. With reference to FIG. 1, there is shown a flowchart of a method 100 that includes steps 102 to 122. The base station is configured to execute the method 100.
[0029] There is provided the method 100 for the base station associated with the one or more terminal devices. The method 100 is used to provide a comprehensive and adaptive resource allocation in a dynamic wireless network environment with continual- learning base station radio station policies with low communication overhead. Moreover, the method 100 is used to ensure an efficient, data-driven resource allocation with improved overall network performance, reduced device interference, and an improved wireless network.
[0030] At step 102, the method 100 for the base station associated with the one or more terminal devices includes determining an initial state at a first time. In an implementation, a controller of the base station is configured to collect relevant data, such as the number of connected devices, network conditions, traffic demands, and the like to understand the current network conditions of the base station. Moreover, by aggregating and analysing the collected data, the method 100 is used to determine the initial state at the first time. In addition, the initial state refers to a state of the base station that is used to determine the changing network conditions within the wireless network. As a result, the determination of the initial state at the first time is to allow effective and efficient utilization of resources.
[0031] At step 104, the method 100 includes determining a resource allocation for the one or more terminal devices associated with the base station based on a Deep Neural Network (DNN) policy. In an implementation, the DNN policy refers to the decisions that are derived from a deep neural network. The DNN policy processes data, such as network conditions, device requirements, channel states, and the like to determine optimal actions, for example, resource allocation (i.e., a‘i), scheduling, power control, and the like in order to ensure enhanced and improved network performance for optimized resource allocation. The DNN policy is used to process the initial state data and other relevant inputs to determine the optimal distribution of resources among the connected one or more terminal devices. As a result, the determination of the resource allocation for the one or more terminal devices associated with the base station based on the DNN policy is used to allow an accurate, reliable, and adaptive resource utilization.
[0032] At step 106, the method 100 includes causing the resource allocation to be executed. In an implementation, the base station is configured to cause the resource allocation to be executed to the one or more terminal devices based on the DNN policy that includes the determination of bandwidth, power, time slots, and the like. As a result, by causing the execution of the resource allocation, the base station is configured to dynamically adjust resource allocation based on current network conditions that enhance the overall efficiency, reduce device interference, and enhance the overall resource utilization within the wireless network.
[0033] At step 108, the method 100 includes determining a resulting state at a second time preceding the execution of the resource allocation. In an implementation, the method 100 is used to collect data, such as device throughput, latency, signal strength, network conditions, and the like in order to determine the resulting state (i.e., st+1i) at the second time preceding the execution of the resource allocation. Moreover, the resulting state refers to a state that reflects the changes and outcomes caused by the execution of the resource allocation. As a result, the determination of the resulting state at the second time preceding the execution of the resource allocation is used to dynamically monitor the resources allocated to the one or more terminal devices within the wireless network in order to identify the changing network conditions and take necessary required measures.
[0034] At step 110, the method 100 includes determining a local reward for the base station. Moreover, after the determination of the resulting state, the local reward (i.e., r‘i ) for the base station is determined. In an implementation, the local reward for the base station is determined based on the summation of the rewards (i.e., rf G R) of the one or more terminal devices associated with the base station.
[0035] At step 112, the method 100 includes determining local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward. In an implementation, the base station is configured to determine the current transition, which includes the initial state, the resource allocation action taken, the resulting state, and the local reward along with the historical transitions stored in a memory of the base station. Thereafter, the base station is configured to determine the local importance values for the current transition (i.e., T‘i) and historical transitions (i.e., TMi) , such as by applying local importance function. As a result, by determining the local importance value, the base station is further allowed to prioritize transitions that are more valuable for training the DNN policy model, thereby enhancing the overall performance of the wireless network by efficient and effective resource allocation. At step 114, the method 100 includes receiving neighbouring importance values for a current transition and historical transitions of the neighbouring base station. The neighbouring base station is configured to send the neighbouring importance values for the current transition and the historical transitions of the neighbouring base station to the base station in order to provide information about the changing network environment so that the base station can optimize the resource allocation accordingly.
[0036] In accordance with an embodiment, the method 100 further includes transmitting local rewards and the local importance values for the base station to the neighbouring base station. Moreover, the base station is configured to transmit the local rewards and the local importance values to allow the neighbouring base station to incorporate the received data into their own decisionmaking processes, ensuring a synchronized and cooperative resource allocation with an improved wireless network.
[0037] In accordance with an embodiment, the method 100 further includes transmitting the importance values for the current transition and the historical transitions to the neighbouring base station. The importance value function refers to a function that takes M + 1 transitions (i.e.,T as input and provides the importance values as a vector, for example, I; is composed of
[0038] M + 1 real numbers that represent the importance of each transition. Moreover, the importance value refers to a value that provides reward prioritization, such as by giving priority to transitions with high rewards or coverage prioritization by giving priority to transitions that have few neighbours. For example, two transitions are neighbouring if their distance (e.g., Euclidean) is lower than a fixed value. In addition, the importance value can be defined as a temporal difference (TD) prioritization that provides a priority to transitions that have poor prediction of their expected long-term rewards. As a result, by transmitting the importance values for the current transition and the historical transitions to the neighbouring base station, each of the base station can contribute to a collaborative, efficient, and adaptive decentralized resource allocation policy.
[0039] At step 116, the method 100 includes determining global importance values based on the local importance values and the neighbouring importance values. In other words, the method 100 is used to combine the local importance values and the neighbouring importance values received from neighbouring base stations. Moreover, such a combination is calculated through a weighted sum or another aggregation technique, without affecting the scope of the present disclosure. As a result, the determination of the global importance values is used to provide a balanced and network-wide optimization of resource allocation. In addition, the determination of the global importance values (e.g., F) is used to provide an enhanced coordination and reduction of the interference among the base stations are used to enhance the overall efficiency and performance of the wireless network.
[0040] In accordance with an embodiment, the method 100 further includes determining the global importance values as a summation of the local importance values and the neighbouring importance values. In other words, upon receiving the neighbouring importance values from the neighbouring base stations, the base station sums the local importance values with the received neighbouring importance values to determine the global importance value for each transition. By combining the local importance values with the neighbouring importance values, the base station is configured to prioritize the transitions that are deemed important not only locally but also by neighbouring base station. Moreover, such determination is used in managing interference, optimizing resource allocation, and improving overall network performance.
[0041] At step 118, the method 100 includes determining which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions. Firstly, the base station is configured to calculate the global importance values for the current transition and the historical transitions by summing up the local importance values and the neighbouring importance values. Moreover, once such global importance values are determined, then, the base station is configured to identify the transition with the lowest value. Such transition is then dropped from the memory and the remaining transitions are retained in order to ensure that only required transitions that can be further used for learning and decision-making is stored in order to provide an efficient, effective, and adaptive decentralized resource allocation policies. In accordance with an embodiment, the method 100 further includes determining if the first time exceeds the number of historical transitions M, t > M, and if so, determining the importance values and comparing importance values of transitions in order to determine which transition to drop, and if the first time does not exceed the number of historical transitions M, t < M, no transition is dropped. In an implementation, is the if t>M, then, in that case, the method 100 is used to calculate the importance values for all transitions stored in the memory and for the current transition, compares these values, and determines which transition to drop based on the lowest global importance value to ensure that only the most valuable transitions are retained. In another implementation, if it is not greater than M, then, in that case, the method 100 does not drop any transitions but stores the new transition in the memory. As a result, the method 100 is used to effectively manage the selective memory based on the number of historical transitions, ensuring efficient and adaptive operation in a decentralized resource allocation.
[0042] At step 120, the method 100 includes receiving neighbouring rewards for previous transitions of the neighbouring base station. The neighbouring rewards are received to enhance the accuracy and effectiveness of the learning process by incorporating broader and more comprehensive data in order to allow the base station to incorporate neighbouring rewards for previous transitions by contributing to an informed, accurate, and collaborative resource allocation in the wireless network.
[0043] In accordance with an embodiment, the method 100 further includes receiving neighbouring rewards for previous transitions of the neighbouring base station in a mini-batch of rewards. The method 100 further includes receiving neighbouring rewards for previous transitions of the neighbouring base station in a mini-batch of rewards, rather than receiving rewards individually to provide a collective set of reward values associated with multiple transitions, facilitating a more efficient data exchange process. Additionally, mini-batching of the rewards helps in synchronizing data collection across neighbouring base stations, leading to more consistent and accurate learning updates, thereby enhancing the efficiency, accuracy, and scalability of the resource allocation and learning processes within the network.
[0044] In accordance with an embodiment, the previous transitions include historical transitions and buffered transitions. Moreover, the buffered transitions are the B latest transitions, and the historical transitions are M transitions preceding the current transition. Furthermore, B and M are natural numbers, and determining which transition of the current transition and historical transitions that has the lowest global importance. In other words, the previous transitions include both historical transitions and buffered transitions. Moreover, the buffered transitions are the B most recent transitions, while historical transitions are the M transitions that occurred prior to the current transition.
[0045] At step 122, the method 100 includes training the DNN policy based on previous transitions of the base station and the received neighbouring rewards. In an implementation, the previous transitions, such as the buffered and the historical transitions of the base station are used to train the DNN policy. As a result, the training of the DNN policy based on the previous transitions of the base station and the received neighbouring rewards is used to enhance the decision-making capabilities of the base station in order to perform resource allocation in order to provide an efficient, reliable, and enhance overall network performance.
[0046] In accordance with an embodiment, the training of the DNN policy based on the previous transitions of the base station and the received neighbouring rewards comprises training based on transitions at time slots that are stored by both the base station and the neighbouring base station. In other words, the method 100 is used to identify the previous transitions that are stored by both the base station and the neighbouring base station that includes transitions from time slots that are present in the memory. Further, each of the base stations are trained using the transitions and corresponding rewards that are received from each other. As a result, the training of the DNN policy based on shared rewards and synched transitions helps in improving the overall learning accuracy and effectiveness of the resource allocation policy.
[0047] In accordance with an embodiment, the training of the DNN policy based on the previous transitions of the base station and the received neighbouring rewards includes determining a global reward and training based on the global reward. In an implementation, the method 100 is used to calculate the global reward for each transition by aggregating local rewards from the base station and rewards received from the neighbouring base station, such as by combining the rewards in order to obtain the global reward. Moreover, the training of the DNN policy based on global rewards provides a more comprehensive and accurate evaluation of each transition to enhance the learning process by incorporating a broader range of data, leading to more informed policy updates and better overall network performance.
[0048] Advantageously, the method 100 for the base station associated with the one or more terminal devices provides efficient and adaptive resource allocation in a dynamic wireless environment through optimized deep neural network (DNN) policies and continual learning. By determining the initial state at a first time and a resulting state and reward at a second time, the base station can effectively monitor and evaluate resource allocation, allowing real-time adaptation to changing network conditions. Exchanging information, such as rewards, with neighbouring base stations within the wireless network enables the base station to mitigate device interference and improve resource allocation across the network. Additionally, training the DNN policy based on previous transitions of the base station and the received neighbouring rewards ensures continuous learning, which adapts to new environments and changes in network conditions without requiring complete retraining from scratch. Thus, the method 100 allows for dynamic, efficient resource allocation and improved network performance through collaborative learning, enhancing the overall performance of the wireless network.
[0049] The steps 102 to 122 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claim herein.
[0050] There is further provided a computer program product comprising program instructions for performing the method 100 when executed by one or more processors in the base station. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Elard Disk Drive (EfDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.
[0051] FIG. 2 is a block diagram of a base station associated with one or more terminal devices, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a block diagram that includes a base station 202A, a communication network 210, a neighbouring base station 202B, and one or more terminal devices 212 associated with the base station 202A within a wireless network 200.
[0052] The base station 202A can be an access point (AP) operating in the wireless network 200 and is responsible for allocating radio resources to the connected one or more terminal devices (e.g., user equipment, UEs) within the wireless network 200. Moreover, the neighbouring base station 202B refers to another AP in the wireless network 200. In an implementation, the base station 202 A includes a first controller 204 A, a first network interface 208 A, a first memory 206 A, and a first FIFO buffer 214A. Similarly, the neighbouring base station 202B includes a second controller 204B, a second network interface 208B, a second memory 206B, and a second FIFO buffer 214B.
[0053] The controller (i.e., a first controller 204A) of the base station 202A is configured to determine an initial state (i.e., s)) at the first time (i.e., t) and further train the DNN policy based on the previous transitions of the base station 202A and the received neighbouring rewards. Examples of the controllers (i.e., the first controller 204A and the second controller 204B) of the base station 202A and the neighbouring base station 202B may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.
[0054] A first memory 206A and a second memory 206B are used to store historical transitions, and the like. Examples of implementation of the first memory 206A and the second memory 206B may include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Dynamic Random Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory. Furthermore, the first FIFO buffer 214A and the second buffer 214B are used to store the latest transitions. A first network interface 208A is used by the base station 202A and the second network interface 208B is used by the neighbouring base station 202B to communicate with the first controller 204A and the second controller 204B respectively. Examples of implementation of the first network interface 208 A and the second network interface 208B may include but are not limited to a network interface, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.
[0055] The communication network 210 includes a medium (e.g., a communication channel) through which the base station 202A communicates with the neighbouring base station 202B of the wireless network 200. Examples of the communication network 210 may include, but are not limited to, a cellular network (e.g., a 2G, a 3G, long-term evolution (LTE) 4G, a 5G, or 5G New Radio (NR) network, such as sub 6 GHz, cmWave, or mmWave communication network, or any communication networks in the future), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to-network (V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.
[0056] There is provided the base station 202 A associated with the one or more terminal devices 212. The base station 202 A is used to provide a comprehensive and adaptive resource allocation in a dynamic wireless network environment with continual- learning. Moreover, the base station 202A is used to ensure an efficient, data-driven resource allocation with improved overall network performance, reduced device interference, and an improved wireless network.
[0057] The base station 202A is configured to determine an initial state at a first time. In an implementation, a controller (i.e., the first controller 204A) of the base station 202A is configured to collect relevant data, such as the number of connected devices, network conditions, traffic demands, and the like to understand the current network conditions of the base station. Moreover, by aggregating and analysing the collected data, the base station 202A is used to determine the initial state at the first time. In addition, the initial state refers to a state of the base station 202A that is used to determine the changing network conditions within the wireless network 200. As a result, the determination of the initial state at the first time is to allow effective and efficient utilization of resources.
[0058] The base station 202 A is configured to determine a resource allocation for the one or more terminal devices 212 associated with the base station 202 A based on the DNN policy. The DNN policy is used to process the initial state data and other relevant inputs to determine the optimal distribution of resources among the connected one or more terminal devices 212. As a result, the determination of the resource allocation for the one or more terminal devices 212 associated with the base station 202 A based on the DNN policy is used to allow an accurate, reliable, and adaptive resource utilization.
[0059] The base station 202A is configured to cause the resource allocation to be executed. As a result, by causing the execution of the resource allocation, the base station 202A is configured to dynamically adjust resource allocation based on current network conditions that enhance the overall efficiency, reduce device interference, and enhance the overall resource utilization within the wireless network 200. The base station 202A is configured to determine a resulting state at a second time preceding the execution of the resource allocation. As a result, the determination of the resulting state at the second time preceding the execution of the resource allocation is used to dynamically monitor the resources allocated to the one or more terminal devices 212 within the wireless network 200 in order to identify the changing network conditions and take necessary required measures.
[0060] The base station 202A is configured to determine a local reward for the base station 202A. Moreover, after the determination of the resulting state, the local reward for the base station 202A is determined. In an implementation, the local reward for the base station 202A is determined based on the summation of the rewards (i.e., rf G R) of the one or more terminal devices associated with the base station 202A.
[0061] The base station 202A is configured to determine local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward. By determining the local importance value, the base station 202A is further allowed to prioritize transitions that are more valuable for training the DNN policy model, thereby enhancing the overall performance of the wireless network 200 by efficient and effective resource allocation.
[0062] The base station 202A is configured to receive neighbouring importance values for a current transition and historical transitions of a neighbouring base station 202B. The neighbouring base station 202B is configured to send the neighbouring importance values for the current transition and the historical transitions of the neighbouring base station 202B to the base station 202A in order to provide information about the changing network environment so that the base station 202A can optimize the resource allocation accordingly.
[0063] In accordance with an embodiment, the base station 202A is further configured to transmit local rewards and the local importance values for the base station to the neighbouring base station 202B. Moreover, the base station 202A is configured to transmit the local rewards that are used to train the DNN policy to allow the neighbouring base station 202B to incorporate the received data into their own decision-making processes, ensuring a synchronized and cooperative resource allocation with an improved wireless network 200.
[0064] In accordance with an embodiment, the base station 202A is further configured to transmit the importance values for the current transition and the historical transitions to the neighbouring base station 202B. The importance value function refers to a function that takes M + 1 transitions (i.e.,T as input and provides the importance values as a vector, for example, Ii;is composed of M + 1 real numbers that represent the importance of each transition. Moreover, the importance value refers to a value that provides reward prioritization, such as by giving priority to transitions with high rewards or coverage prioritization by giving priority to transitions that have few neighbours. For example, two transitions are neighbouring if their distance (e.g., Euclidean) is lower than a fixed value. In addition, the importance value can be defined as a temporal difference (TD) prioritization that provides a priority to transitions that have poor prediction of their expected long-term rewards. As a result, by transmitting the importance values for the current transition and the historical transitions to the neighbouring base station, each of the base station 202A and the neighbouring base station 202B can contribute to a collaborative, efficient, and adaptive decentralized resource allocation policy.
[0065] The base station 202A is configured to determine global importance values based on the local importance values and the neighbouring importance values. In other words, the base station 202A is used to combine the local importance values and the neighbouring importance values received from neighbouring base station 202B. Moreover, such a combination is calculated through a weighted sum or another aggregation technique, without affecting the scope of the present disclosure. As a result, the determination of the global importance values is used to provide a balanced and network- wide optimization of resource allocation. In addition, the determination of the global importance values is used to provide an enhanced coordination and reduction of the interference among the base stations are used to enhance the overall efficiency and performance of the wireless network 200.
[0066] In accordance with an embodiment, the base station 202A is further configured to determine the global importance values as a summation of the local importance values and the neighbouring importance values. By combining the local importance values with the neighbouring importance values, the base station 202A is configured to prioritize the transitions that are deemed important not only locally but also by neighbouring base station 202B. Moreover, such determination is used in managing interference, optimizing resource allocation, and improving overall network performance.
[0067] The base station 202A is configured to determine which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions. Such transition is then dropped from the memory (e . g . , the first memory 206 A) and the remaining transitions are retained in order to ensure that only required transitions that can be further used for learning and decision-making is stored in order to provide an efficient, effective, and adaptive decentralized resource allocation.
[0068] In accordance with an embodiment, the base station 202A is further configured to determine if the first time exceeds the number of historic transitions M, t > M, and if so, determine the importance values and comparing importance values of transitions in order to determine which transition to drop, and if the first time does not exceed the number of historical transitions M, t < M, no transition is dropped. As a result, the method 100 is used to effectively manage the selective memory based on the number of historical transitions, ensuring efficient and adaptive operation in a decentralized resource allocation.
[0069] The base station 202A is configured to receive neighbouring rewards for previous transitions of the neighbouring base station 202B. The neighbouring rewards are received to enhance the accuracy and effectiveness of the learning process by incorporating broader and more comprehensive data in order to allow the base station to incorporate neighbouring rewards for previous transitions by contributing to an informed, accurate, and collaborative resource allocation in the wireless network 200.
[0070] In accordance with an embodiment, the previous transitions includes historical transitions and buffered transitions. Moreover, the buffered transitions are the B latest transitions, and the historical transitions are M transitions preceding the current transition. Moreover, B and M are natural numbers, and determining which transition of the current transition and historical transitions that has the lowest global importance value comprises determining which transition of the current transition and historical transitions that has the lowest global importance value. The method 100 is used to determine which transition among the current transition, historical transitions, and buffered transitions has the lowest global importance value by evaluating the importance values of all these transitions and identifying the one with the lowest value. As a result, such transition can then be deprioritized or removed, allowing the base station 202A to prioritize transitions with high importance values, thereby improving the overall resource allocation.
[0071] The base station 202A is configured to train the DNN policy based on previous transitions of the base station 202A and the received neighbouring rewards. In an implementation, the previous transitions, such as the buffered and the historical transitions of the base station 202 A are used to train the DNN policy. As a result, the training of the DNN policy based on the previous transitions of the base station 202A and the received neighbouring rewards is used to enhance the decision-making capabilities of the base station 202A in order to perform resource allocation in order to provide an efficient, reliable, and enhance overall network performance.
[0072] In accordance with an embodiment, training the DNN policy based on the previous transitions of the base station 202A and the received neighbouring rewards comprises training based on transitions at time slots that are stored by both the base station 202A and the neighbouring base station 202B. As a result, the training of the DNN policy based on shared rewards and synched transitions helps in improving the overall learning accuracy and effectiveness of the resource allocation policy. Advantageously, the base station 202A associated with the one or more terminal devices 212 provides efficient and adaptive resource allocation in a dynamic wireless environment through optimized deep neural network (DNN) policies and continual learning. By determining the initial state at a first time and a resulting state and reward at a second time, the base station 202 A can effectively monitor and evaluate resource allocation, allowing real-time adaptation to changing network conditions. Exchanging information, such as rewards, with neighbouring base stations within the wireless network 200 enables the base station to mitigate device interference and improve resource allocation across the wireless network 200. Additionally, training the DNN policy based on previous transitions of the base stations and the received neighbouring rewards ensures continuous learning, which adapts to new environments and changes in network conditions without requiring complete retraining from scratch. Thus, the base station 202A is configured to allows for dynamic, efficient resource allocation and improved network performance through collaborative learning, enhancing the overall performance of the wireless network 200.
[0073] FIG. 3 is a diagram that illustrates an architecture and a message exchange between a base station and a neighbouring base station of the same interference group, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIG. 2. With reference to FIG. 3, there is shown a diagram 300 of the architecture and the message exchange between the base station 202A and the neighbouring base station 202B of the same interference group. The base station 202A is operated in a first local environment 302A and includes a first DNN policy 308A, a first receiver or transmitter 306A, a first FIFO buffer 304A, a first selective memory buffer 316A, and the first controller 204A. Similarly, the neighbouring base station 202B operates in a second local environment 302B and includes a second DNN policy 308B, a second receiver or transmitter 306B, a second FIFO buffer 304B, a second selective memory buffer 316B, and the second controller 204B.
[0074] In an implementation, at operation 310, each base station, such as the base station 202A and the neighbouring base station 202B is configured to send an initialization handshake message (i.e., m composed of the ID of its importance value function chosen from a predefined list, the sizes of its buffer, memory and mini-batch and an initial seed (e.g., m, = {ID3,memorysize,buffersize,minibatchsize,seed}). After that, the base station 202A is configured to determine importance value function (i. e. , J) and parameters that ensure synched transitions, such as size of selective memories, buffers and minibatches, mini-batch sampling method and the like. Moreover, the handshake initialization is performed every time any new base station joins or leaves the neighborhood. The controller (i.e., the first controller 204A), at a time-slot (i.e., t), is configured to take a decision (i.e., al) based on the local state (i.e., sf ) through the first DNN policy 308A and determine the new local state (i.e., sf+1) and the local reward (i.e., rf) as a result of the decision and saves the local transition (i.e., rf = (sf, af, rf, s j+1)) in the first FIFO buffer 304A. In an implementation, if t < M, then, in that case, the first controller 204A is configured to save the transition (i.e., xf) in the first selective memory buffer 316A. In another implementation, if t > M, then, in that case, the first selective memory buffer 316 A is full and the first controller 204 A is configured to decide whether to store the current transition (i.e., xf) or not. Thereafter, the first controller 204A is configured to compute (i.e., M + 1) local importance values (i.e., Fi) for the current local transition experience and for the M historical transitions stored in the first selective memory buffer 316 A. Therafter, the first controller 204 A of the base station 202 A is configured to send the importance value vector (i.e., If) to the neighboring base station 202B and receives the importance value vector for the neighbouring base station 202B, such as at operation 312. After receiving the importance value vectors, the first controller 204A is configured to compute the global importance value vector (i.e., F) and drops the transition that has the minimum global importance value (i.e., mini1). Moreover, if the dropped transition is in the first selective memory buffer 316A, then, in that case, the current transition takes its place. Furthermore, at each K timeslots, the base station 202A is configured to sample synchronized mini batches. For example, each base station (e.g., the base station 202A) selects, from a common pool of the FIFO buffer (i.e., the first FIFO buffer 304A) and the selective memory (i.e., the first selective memory buffer 316A), the same time-slot local transitions as the other base stations. Thereafter, the base station 202A is configured to exchange the respective rewards and after receiving all the rewards in mini-batches and for each sampled transition k, each base station (i.e., i) is configured to compute the associated global reward (i.e., Rk= jeJV- rk, where JVj is the neighboring BSs) , such as at operation 314. Finally, the base station 202 A is configured to use the sampled mini batch, after replacing local rewards by global rewards, to update the first DNN policy (0j) 308A. As a result, the architecture is configured to continuously optimize the overall network's performance through collaborative, adaptive decision-making based on local observations and shared information, leading to an improved network efficiency and responsiveness to changing conditions within the network 200 along with a stable DNN policy.
[0075] FIG. 4 is a diagram that illustrates an exemplary scenario of base stations operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 4 is described in conjunction with elements from FIGs. 1 to 3. With reference to FIG. 4, there is shown a diagram 400 illustrating the operations for resource allocation within the wireless network 200 including the base station 202A and the neighbouring base station 202B.
[0076] In an implementation, the base station (i.e., the base station 202A and the neighbouring base station 202B) defines a common seed as the summation of all seeds and a common size of buffers, memories, and minibatches as the minimum value (e.g., M = min{Mj}jeJV- )). Moreover, the base station (i.e., the base station 202A and the base station 202B) agrees on a common importance value function that can be defined as ?(TJ, = (r- , ... , r!I+1) = I;. Furthermore, the base station 202A and the base station 202B are configured to serve devices (e.g., K, devices and Kj devices) and each of the base station has access to the number of packets waiting to be downloaded in their respective channel gains (or the local state). At each time slot (t), the base station 202A, based on (i.e., sf ), selects a device that will occupy the channel and then monitors the local throughput (i.e., rf) and the updated state (i.e., sl+1) from the environment and stores the local transition (i.e., T? = (Sj,aj,rj,Sj+1)) in the FIFO buffer (e.g., the first FIFO buffer 304A). Similarly, the neighbouring base station 202B is configured to store the local transition (i.e., T?) in the second FIFO buffer 304B. Furthermore, at operation 402A, ift < M, then each base station (i.e., the first base station 202A and the neighbouring base station 202B) stores the local transition in the selective memory and if t > M, then the base station 202A is configured to compute the local importance value vector (i.e., I and the neighbouring base station 202B is configured to compute the local importance value vector (i.e., Ij). Both the base station 202A and the neighbouring base station 202B are configured to exchange the vectors, such as at operation 404A and operation 404B, and each one computes the global importance value vector (i.e., I = I , + Ij) and drops the transition with the lowest global value. Furthermore, at operation 402B, the base station 202A and the neighbouring base station 202B are configured to use the common seed and each one samples a mini-batch of transitions, such as at each K time-slots and further exchange the associated local rewards at operation 406A and 406B and update the DNN parameters based on global rewards,. As a result, the resource allocation in the wireless network 200 is optimized by enabling efficient coordination between base stations (i.e., the base station 202A and the neighbouring base station 202B) with minimized interference, and enhanced network performance.
[0077] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A method (100) for a base station, the base station (202A) being associated with one or more terminal devices (212), the method (100) comprising: determining an initial state at a first time, determining aresource allocation for the one or more terminal devices (212) associated with the base station based on a Deep Neural Network, DNN, policy, causing the resource allocation to be executed, determining a resulting state at a second time preceding the execution of the resource allocation, determining a local reward for the base station (202A), determining local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward, receiving neighbouring importance values for current transition and historical transitions of a neighbouring base station (202B), determining global importance values based on the local importance values and the neighbouring importance values determining which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions, receiving neighbouring rewards for previous transitions of the neighbouring base station (202B), and training the DNN policy based on previous transitions of the base station (202A) and the received neighbouring rewards.
2. The method (100) according to claim 1, wherein the method (100) further comprises transmitting local rewards and the local importance values for the base station (202A) to the neighbouring base station (202B).
3. The method (100) according to any preceding claim, wherein the method (100) further comprises determining the global importance values as a summation of the local importance values and the neighbouring importance values.
4. The method (100) according to any preceding claim, wherein the previous transitions comprises historical transitions and buffered transitions, wherein the buffered transitions are the B latest transitions, and the historical transitions are M transitions preceding the current transition, wherein B and M are natural numbers, and wherein determining which transition of the current transition and historical transitions that has the lowest global importance value comprises determining which transition of the current transition and historical transitions that has the lowest global importance value.
5. The method (100) according to any preceding claim, wherein the method (100) further comprises transmitting the importance value for the current transition and the historical transitions to the neighbouring base station (202B).
6. The method (100) according to any preceding claim, wherein training the DNN policy based on the previous transitions of the base station (202A) and the received neighbouring rewards comprises training based on transitions at time slots that are stored by both the base station (202A) and the neighbouring base station (202B).
7. The method (100) according to claim 5 or 6, wherein the method (100) further comprises receiving neighbouring rewards for previous transitions of the neighbouring base station (202B) in a mini-batch of rewards.
8. The method (100) according to any preceding claim, wherein training the DNN policy based on the previous transitions of the base station (202A) and the received neighbouring rewards includes determining a global reward and training based on the global reward.
9. The method (100) according to any preceding claim, wherein the method (100) further comprises determining if the first time exceeds the number of historic transitions M, t > M, and if so, determining the importance values and comparing importance values of transitions in order to determine which transition to drop, and if the first time does not exceed the number of historic transitions M, t < M, no transition is dropped.
10. A computer program product comprising program instructions for performing the method according to preceding claim, when executed by one or more processors in a base station (202A).
11. A base station (202A) associated with one or more terminal devices (212), the base station (202A) being configured to determine an initial state at a first time, determine a resource allocation for the one or more terminal devices (212) associated with the base station (202A) based on a Deep Neural Network, DNN, policy, cause the resource allocation to be executed, determine a resulting state at a second time preceding the execution of the resource allocation, determine a local reward for the base station (202A), determine local importance values for a current transition and historical transitions, where a transition is a tuple comprising the initial state, the resource allocation, the resulting state and the local reward, receive neighbouring importance values for a current transition and historical transitions of a neighbouring base station (202B), determine global importance values based on the local importance values and the neighbouring importance values determine which transition of the current transition and historical transitions that has the lowest global importance value and drop it and store the remaining transitions, receive neighbouring rewards for previous transitions of the neighbouring base station (202B), and train the DNN policy based on previous transitions of the base station (202A) and the received neighbouring rewards.
12. The base station (202A) according to claim 11, wherein the base station (202A) is further configured to transmit local rewards and the local importance values for the base station to the neighbouring base station (202B).
13. The base station (202A) according to claim 11 or 12, wherein the base station (202A) is further configured to determine the global importance values as a summation of the local importance values and the neighbouring importance values.
14. The base station (202A) according to claim 11, 12 or 13, wherein the previous transitions comprises historical transitions and buffered transitions, wherein the buffered transitions are the B latest transitions, and the historical transitions are M transitions preceding the current transition, wherein B and M are natural numbers, and wherein determining which transition of the current transition and historical transitions that has the lowest global importancevalue comprises determining which transition of the current transition and historical transitions that has the lowest global importance value.
15. The base station (202A) according to any of claims 11 to 14, wherein the base station (202A) is further configured to transmit the importance values for the current transition and the historical transitions to the neighbouring base station (202B).
16. The base station (202A) according to any of claims 11 to 15, wherein training of the DNN policy based on the previous transitions of the base station (202A) and the received neighbouring rewards comprises training based on transitions at time slots that are stored by both the base station (202A) and the neighbouring base station (202B).
17. The base station (202A) according to any of claims 11 to 16, wherein the base station (202A) is further configured to determine if the first time exceeds the number of historic transitions M, t > M, and if so determine the importance values and comparing importance values of transitions in order to determine which transition to drop, and if the first time does not exceed the number of historic transitions M, t < M, no transition is dropped.
Citation Information
Patent Citations
Joint distributed learning of signaling and policies for radio resource allocation
WO2024110047A1