First base station arranged in wireless network and method for controller of base station

The first base station uses RL with MAML to train decentralized AP resource allocation policies based on local and global environment identifiers, addressing interference and environmental variability in wireless networks, achieving efficient and adaptive resource allocation with reduced overhead.

WO2026012572A1PCT designated stage Publication Date: 2026-01-15HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/069306
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional radio resource allocation in wireless communication networks faces challenges in managing interference among multiple access points (APs) due to varying wireless environments, requiring manual tuning that is resource-intensive and time-consuming, and existing reinforcement learning methods like Model-Agnostic Meta-Learning (MAML) fail to provide effective policies for real-world networks with high communication overhead.

Method used

A first base station employs Reinforcement Learning (RL) with Model-Agnostic Meta-Learning (MAML) to train decentralized AP radio resource allocation policies by leveraging local and global environment identifiers, using signal measurements like SNR, RSSI, and RSRQ, and exchanging these identifiers between neighboring APs to adapt resource allocation models efficiently.

Benefits of technology

This approach reduces communication overhead, enhances interference management, and dynamically adapts to changing environments, ensuring efficient resource allocation with minimal manual intervention and maintaining high network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024069306_15012026_PF_FP_ABST
    Figure EP2024069306_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A first base station arranged to operate in a wireless network configured to initiate a connection with a second base station, determine a first local environment identifier for the base station, receive a second local environment identifier for the second base station, determine a global environment identifier based on the first local environment identifier for the base station and the second local environment identifier for the second base station, determine a model to be used based on the global environment identifier for the base station according to associations between environments and associated models, determine a local state from the first local environment, perform a local resource allocation action, determine a first local reward for the resource allocation action and a new local state for the base station, receive a second local reward for the second base station, Meta-train the model based on one or more transitions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] FIRST BASE STATION ARRANGED IN WIRELESS NETWORK AND METHOD FOR CONTROLLER OF

[0002] BASE STATION

[0003] TECHNICAL FIELD

[0004] The present disclosure relates generally to the field of wireless communication networks and more specifically, to a first base station arranged to operate in a wireless network and a method for a controller of the first base station, such as for resource allocation in a dynamic wireless environment with meta-leaming and identifier.

[0005] BACKGROUND

[0006] Typically, an Access Point (AP) is connected to multiple devices in a wireless communication networks to allocate radio resources to the devices and to facilitate communication within the network. Radio resource allocation strategies are essential for APs to effectively allocate and deallocate resources based on the available network-related information. Wireless communication networks often experience varying wireless environments, leading to challenges in maintaining efficient communication. Moreover, due to such variability, manual tuning of AP policies is required but is resource-intensive and timeconsuming.

[0007] Conventionally, in the wireless communication networks, radio resource allocation is often based on local information from connected devices, with each AP using state information from the associated devices. While such distributed policies are straightforward to implement and have significant limitations, particularly in managing interference among multiple APs. Moreover, high interference can degrade overall network performance, highlighting the need for algorithms that incorporate some level of communication between APs to manage interference effectively. Reinforcement Learning (RL), specifically Model-Agnostic Meta-Leaming (MAML), is used to address the above-mentioned limitations but often fails to provide policies suitable for real-world wireless communication networks. Thus, there exists a technical problem of how to train decentralized AP radio resource allocation policies with low communication overhead between neighboring APs that can be adapted to any wireless network.

[0008] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional controllers and conventional methods for resource allocation in a dynamic wireless environment with meta- leaming.

[0009] SUMMARY

[0010] The present disclosure provides a first base station arranged to operate in a wireless network and a method for a controller of the first base station of the wireless network. The present disclosure provides a solution to the existing problem of how to train decentralized AP radio resource allocation policies with low communication overhead between neighboring APs that can be adapted to any wireless network. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the base station and the method for the controller of the base station for resource allocation in a dynamic wireless environment with meta-leaming.

[0011] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.

[0012] In one aspect, the present disclosure provides a first base station arranged to operate in a wireless network. The first base station comprises a controller configured to initiate a connection with a second base station, determine a first local environment identifier for the base station, receive a second local environment identifier for the second base station, determine a global environment identifier based on the first local environment identifier for the first base station and the second local environment identifier for the second base station, determine a model to be used based on the global environment identifier for the base station according to associations between environments and associated models. Moreover, a model comprises an agnostic model or an adapted environment-specific model, determines a local state from the first local environment, performs a local resource allocation action based on the model to be used and the local state, determines a first local reward for the resource allocation action and a new local state for the base station, receive a second local reward for the second base station, Meta-train the model to be used based on one or more transitions, wherein a transition is a tuple comprising the local state, the local resource allocation action, the local and the received rewards and the new local state.

[0013] Advantageously, by leveraging local and global environment identifiers to guide resource allocation strategies. The first base station and the second base station determine local environment identifiers through signal measurements like signal-to-noise ratio (SNR), received signal strength indicator (RSSI), and reference signal received quality (RSRQ). Moreover, such local identifiers are exchanged between neighboring APs to compute a global environment identifier, which informs the selection of appropriate resource allocation models. The first base station employs Reinforcement Learning (RL), particularly Model- Agnostic Meta-Leaming (MAML), to train an agnostic model that can quickly adapt to new environments. By meta-training that model with transitions and rewards, the first base station is configured to enhance the resource allocation policies that reduce communication overhead, enhance interference management, and dynamically adapt to changing environments. The use of the learned agnostic model that can be quickly adapted to any environment by simple and fast fine-tuning allows the first base station to scale efficiently across various network conditions, ensuring efficient resource allocation with minimal manual intervention while maintaining high network performance.

[0014] In another aspect, the present disclosure provides a method for a controller in a first base station, the method comprising initiating a connection with a second base station, determining a first local environment identifier for the first base station, sending the first local environment identifier for the first base station to the second base station, receiving a second local environment identifier for the second base station, determining a global environment identifier based on the first local environment identifier for the first base station and the second local environment identifier for the second base station, maintaining an indexed memory. Moreover, the indexed memory stores associations between environments and associated models, wherein a model can be an agnostic model or an adapted environment-specific model, determining a model to be used based on the global environment identifier for the base station according to the indexed memory. Furthermore, the method includes determining a local state from the first local environment, determining a local resource allocation action based on the model to be used and the local state, executing the local resource allocation action, determining a first local reward and a new local state for the first base station, sending the first local reward for the first base station to the second base station, and receiving a second local reward for the second base station. Furthermore, the method includes Meta-training the model to be used based on one or more transitions. Moreover, a transition is a tuple comprising the local state the local resource allocation action, the local and the received rewards, and the new local state.

[0015] The method achieves all the advantages and technical effects of the first base station of the present disclosure.

[0016] It is to be appreciated that all the aforementioned implementation forms can be combined.

[0017] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.

[0018] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.

[0021] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:

[0022] FIG. 1 is a block diagram that illustrates a first base station arranged to operate in a wireless network, in accordance with an embodiment of the present disclosure;

[0023] FIG. 2 is a flowchart of a method for a controller in a first base station, in accordance with an embodiment of the present disclosure;

[0024] FIG. 3 is a diagram that illustrates signaling between a first base station and a second base station operating in a wireless network, in accordance with an embodiment of the present disclosure;

[0025] FIG. 4 is a diagram that illustrates communication between a first base station and a second base station operating in a wireless network, in accordance with an embodiment of the present disclosure;

[0026] FIG. 5 is a diagram that illustrates operations for resource allocation within a wireless network, in accordance with an embodiment of the present disclosure;

[0027] FIG. 6 is a diagram that illustrates an exemplary scenario of a first base station operating in a wireless network, in accordance with an embodiment of the present disclosure; and

[0028] FIG. 7 is a diagram that illustrates another exemplary scenario of a first base station operating in a wireless network, in accordance with an embodiment of the present disclosure.

[0029] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.

[0030] DETAILED DESCRIPTION OF EMBODIMENTS

[0031] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.

[0032] FIG. 1 is a block diagram that illustrates a first base station arranged to operate in a wireless network, in accordance with an embodiment of the present disclosure. With reference to FIG. 1 , there is shown a first base station 102A arranged to operate in the wireless network 100. The first base station 102A is further connected to a second base station 102B within the wireless network 100.

[0033] The first base station 102A refers to an access point (AP) operating in the wireless network 100 and is responsible for allocating radio resources to the connected devices (e.g., user equipment, UEs) within the wireless network 100. Moreover, the second base station 102B refers to another AP in the wireless network 100.

[0034] The controller of the first base station 102A (i.e., the first controller 104A) is configured to establish a connection with the second base station 102B, and the controller of the second base station 102B (i.e., the second controller 104B) is configured to establish the connection with the first base station 102A, such as through the first controller 104A. Examples of the controllers (i.e., the first controller 104A and the second controller 104B) of the first base station 102A and the second base station 102B may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.

[0035] The first memory 106A and the second memory 106B are used to store the associations between environments and associated models of the first base station 102A and the second base station 102B respectively. Examples of implementation of the first memory 106A and the second memory 106B may include, but are not limited to, Electrically Erasable Programmable Read- Only Memory (EEPROM), Dynamic Random Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory.

[0036] The first network interface 108A is used by the first base station 102A and the second network interface 108B is used by the second base station 102B to communicate with the first controller 104A and the second controller 104B respectively. Examples of implementation of the first network interface 108A and the second network interface 108B may include but are not limited to a network interface, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.

[0037] The communication network 110 includes a medium (e.g., a communication channel) through which the first base station 102A communicates with the second base station 102B of the wireless network 100. Examples of the communication network 110 may include, but are not limited to, a cellular network (e.g., a 2G, a 3G, long-term evolution (LTE) 4G, a 5G, or 5G New Radio (NR) network, such as sub 6 GHz, cmWave, or mmWave communication network), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to-network (V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.

[0038] There is provided the first base station 102A arranged to operate in the wireless network 100. Moreover, the first base station 102A is used for an efficient resource allocation in the wireless network 100, such as by maintaining an optimal network performance and low interference.

[0039] Moreover, the first base station 102A includes the controller (i.e., the first controller 104A) configured to initiate a connection with the second base station 102B. In an implementation, the connection is initiated by the controller in order to synchronize and coordinate resource allocation between the first base station 102A and the second base station 102B. As a result, the establishment of the direct communication between the first base station 102A and the second base station 102B allows an enhanced synchronization and coordination of resource allocation thereby reducing interference and enhancing the overall efficiency of the wireless network 100. In accordance with an embodiment, the controller is further configured to initiate the connection through a handshake protocol exchanging information on environment classes. In an implementation, the handshake protocol includes a protocol that includes an initial exchange of messages between the controllers of the first base station 102A and the second base station 102B both in order to understand network conditions and adjust the resource allocation policies accordingly. Therefore, the handshake protocol ensures a secure and efficient setup of the connection, facilitating seamless communication between the first base station 102A and the second base station 102B. As a result, the handshake protocol ensures that the controllers have up-to-date information about the environment classes, allowing for more precise and effective resource management with reduced communication overhead.

[0040] In accordance with an embodiment, initiating the connection includes initiating the agnostic model. In an implementation, the agnostic model utilizes machine learning algorithms to assess the current network conditions and optimize resource allocation strategies accordingly. By initiating the agnostic model during the initiation of the connection, the first base station 102A is configured to ensure that both the base stations (i.e., the first base station 102A and the second base station 102B) are prepared to handle dynamic network scenarios effectively.

[0041] Furthermore, the controller is configured to determine a first local environment identifier (ei) for the first base station 102A. The first local environment identifier refers to a unique identifier that depicts specific conditions of the first base station 102A operating in the wireless network 100. The controller is configured to determine the first environment identifier, such as by determining a signal strength, interference levels, and the number of connected devices. In an implementation, the first controller 104A is configured to determine the first local environment identifier for the first base station 102A. Moreover, the first local environment identifier is further utilized to provide an effective resource management in the wireless network 100. As a result, the first local environment identifier is used to provide precise and tailored policies for resource allocation with an improved network performance and reliability.

[0042] In accordance with an embodiment, the controller is further configured to determine the first local environment identifier by making measurements and determining the first local environment identifier based on the measurements. In other words, the controller (i.e., the first controller 104A of the first base station 102A) is configured to make measurements such as signal strength, interference levels, and the number of active devices connected with the first base station 102A. Thereafter, the controller is configured to determine the first local environment identifier based on the measurement. As a result, the determination of the first local environment leads to an improved network performance, reduced interference, and more efficient use of resources thereby enhancing the adaptability and resilience of the wireless network 100.

[0043] In accordance with an embodiment, the controller is further configured to send the first local environment identifier (ei) for the first base station 102A to the second base station 102B. The controller (i.e., the first controller 104A) is configured to collect and process the data to generate the first local environment identifier for the first base station 102A. Once the first local environment identifier is determined, the controller sends the determined first local environment identifier to the second base station 102B, such as by using a secure and efficient communication protocol. Moreover, such transmission enables the second base station 102B to receive and interpret the environmental conditions of the first base station 102A. As a result, the controller is allowed to understand the environmental context of the first base station 102A, leading to an informed and effective decisions that enhance the overall performance of the wireless network 100.

[0044] In accordance with an embodiment, an environment relates to a signal measurement and an environment identifier relates to a classification of the environment. In an implementation, the signal measurements are used to determine the first local environment identifier, such as by measuring the signal strength, interference levels, noise, and the like. Moreover, the environment identifier provides a clear and actionable classification of the environment that enables more effective resource allocation and interference mitigation that results in an adaptive, resilient, and high-performing wireless communication network 100.

[0045] In accordance with an embodiment, the signal measurement relates to signal-to-noise ratio (SNR), Reference Signal Received Quality, (RSRQ), Received Signal Strength Indicator, (RSSI), or Reference Signal Received Power, (RSRP). In an implementation, the signal measurement relates to the SNR. The SNR refers to a ratio of the received signal power to the noise power, indicating the quality of the received signal. In another implementation, the signal measurement relates to the RSRQ. Moreover, the RSRQ refers to the measurement of the signal quality, combining the received signal strength and the Signal-to- Interference-plus-Noise Ratio (SINR). In yet another implementation, the signal measurement relates to the RSSI. In addition, the RSSI refers to the total received power, including the desired signal, interference, and noise components. In another implementation, the signal measurement relates to the RSSP. Moreover, the RSSP refers to the average power of the received reference signals, which is a measure of the received signal strength. The signal measurements are used to determine the first local environment identifier by providing information about radio conditions experienced by the first base station 102A. The controller of the first base station 102A continuously monitors and measures the received signal quality and strength, such as by using SNR, RSRQ, RSSI, or RSRP. In an example, such signal measurements can be obtained from various reference signals or data transmissions received from the associated user equipment (UEs or devices) or neighboring base stations, such as the second base station 102B. As a result, the controller is configured to represent the local radio condition of the first base station 102A enabling collaborative adaptation of resource allocation strategies and improving the overall performance of the wireless network 100.

[0046] Furthermore, the controller is configured to receive a second local environment identifier (ej) for the second base station 102B. The controller of the first base station 102A establishes a communication link with the second base station 102B and through the established link, the second base station 102B transmits the second local environment identifier, which encapsulates the specific conditions and measurements of the operating environment of the second base station 102B. Therefore, the second local environment identifier is used by the controller of the first base station 102A to make an informed decision for optimizing the overall performance of the wireless network 100 and reducing the interference in the wireless network 100. As a result, the first base station 102A is configured to adjust the policies based on comprehensive environmental data that leads to an effective resource allocation, reduced interference, and enhanced overall network performance.

[0047] In accordance with an embodiment, the controller is further configured to receive second local environment identifiers from more than one second base station. In an example, the controller is configured to receive the second local environment identifier from the second base station 102B. Similarly, the controller is configured to receive the second local environment identifier from another second base station. The controller at the first base station 102A establishes communication links with one or more second base stations. Each of these second base stations transmits the second local environment identifiers respectively. The controller at the first base station 102A receives the second local environment identifiers and processes them to gain a comprehensive understanding of the environmental conditions across the wireless network 100. By receiving local environment identifiers from more than one second base station, the first base station 102A makes more informed decisions that account for a wider range of conditions with enhanced coordination and optimization of the resources and improved network performance with reduced network interference.

[0048] Furthermore, the controller is configured to determine a global environment identifier (e) based on the first local environment identifier (ei) for the first base station 102A and the second local environment identifier (ej) for the second base station 102B. The global identifier encapsulates the combined environmental conditions and characteristics of both the base stations (i.e., the first base station 102A and the second base station 102B), providing a comprehensive overview of the network environment. The determination of the global environment identifier allows the effective management of radio resources in dynamic network environments, leading to improved network performance, reliability, enhanced efficiency, reduced interference, and improved overall performance of the wireless network 100.

[0049] Furthermore, the controller is configured to determine a model to be used based on the global environment identifier for the first base station 102A and the second base station 102B according to associations between environments and associated models. Moreover, the model comprises an agnostic model or an adapted environment-specific model. In an implementation, different wireless environments may require different resource allocation strategies for optimal performance, such as different models based on the determined local environment identifiers. Moreover, an agnostic model is a general-purpose model that can work reasonably well in various environments and that can be quickly optimized for specific conditions. An adapted environment-specific model is a model that has been obtained by fine-tuning or adapting the agnostic model to perform optimally in a particular environment, represented by a specific global environment identifier. The controller of the first base station 102A receives the local environment identifier from the second base station 102B and further determines the global environment identifier, such as by using a predefined mapping function. Based on the determined global environment identifier, the controller fine-tunes the agnostic model to an adapted model. As a result, the ability to fine-tune the agnostic model based on the global environment identifier allows the controller to adapt its resource allocation strategy to various wireless conditions, optimizing the overall performance of the wireless network 100.

[0050] In accordance with an embodiment, the controller is further configured to receive a new environment identifier from the second base station, and in response thereto determine a new global environment identifier based on the new environment identifier and find a new resource allocation model based on the new global environment identifier. Firstly, the controller is configured to receive the new identifier from the second base station 102B. Thereafter, the controller is configured to determine the new global environment identifier based on the new environment identifier. After that, the controller is configured to adapt or finetune the agnostic resource allocation model based on the new global environment identifier. Moreover, the new global environment identifier allows an effective and coordinated resource allocation strategies that improve the overall performance of the wireless network 100 and reduce network interference. Therefore, finding a new resource allocation model, by fine- tuning the agnostic model, based on the new global environment ensures that the wireless network 100 adapts swiftly and efficiently to changes based on comprehensive signaling measurements resulting in a more resilient and high-performing wireless network 100, capable of maintaining optimal performance in a variety of scenarios.

[0051] In accordance with an embodiment, the controller is further configured to determine a change in a local environment, and in response thereto determine a new local environment identifier. Thereafter, the controller is configured to send the new environment identifier to the second base station 102B, determine a new global environment identifier based on the new environment identifier, and find a new resource allocation model based on the new global environment identifier. As the wireless network environments are dynamic, and changes in radio conditions, interference levels, or user mobility can occur over time, updating the local environment identifier and subsequently the global environment identifier is required to select the most appropriate resource allocation model for the corresponding changed conditions. The controller (e.g., the first controller 104A) continuously monitors the signal measurements (e.g., SNR, RSRQ, RSSI, RSRP) and neighboring base stations (e.g., the second base station 102B) If the signal measurements indicate a significant change in the local environment (e.g., crossing a predefined threshold), the controller determines a new local environment identifier based on the updated measurements. The controller sends the new local environment identifier to the second base station 102B and further determines a new global environment identifier based on the new environment identifier such as by using the mapping function. Therefore, the detection of changes in the local environment and updating the resource allocation model accordingly allow the controller to dynamically adapt to evolving wireless conditions, ensuring optimal performance. By promptly exchanging new environment identifiers and fast adapting the agnostic resource allocation model, the controller can quickly respond to changes in the wireless environment, minimizing performance degradation. Furthermore, the controller is configured to determine a local state from the first local environment. The local state represents the current conditions and status of the first base station 102A and associated devices that are essential for making informed resource allocation decisions. In an implementation, the controller analyses various parameters, such as signal measurements, user traffic patterns, and resource utilization, within the local environment to determine the local state that enables the controller to make resource allocation decisions tailored to the specific conditions of the first base station 102A with enhanced overall network performance.

[0052] Furthermore, the controller is configured to perform a local resource allocation action based on the model to be used and the local state. In an implementation, the controller is configured to perform a local resource allocation action for the first base station 102A based on the selected model (e.g., the agnostic model or the adapted model) and the determined local state. The resource allocation action aims to optimize the utilization of radio resources (e.g., scheduling, power control, modulation, and coding scheme selection) based on the current conditions represented by the model and local state. Moreover, the controller is configured to utilize the selected model, which has been trained or adapted for the specific environment, and applies it to the local state to determine the appropriate resource allocation action. As a result, by combining the selected model and local state allows for highly contextualized and optimized resource allocation decisions, improving overall network performance and user experience.

[0053] Furthermore, the controller is configured to determine a first local reward (n) for the resource allocation action and a new local state for the first base station 102A. After performing the resource allocation action, the controller determines the first local reward and a new local state for the first base station 102A. The local reward quantifies the performance or quality metric associated with the executed resource allocation action, which is required for training and adapting the agnostic model. The new local state represents the updated conditions after executing the action. The controller monitors the various conditions of the wireless network 100, such as throughput, latency, fairness, and the like. Thereafter, the controller is configured to update the local state based on the impact of the resource allocation action and enables the controller to learn and adapt the agnostic resource allocation model based on the outcomes of the executed actions, leading to continuous improvement and optimization.

[0054] In accordance with an embodiment, the controller is further configured to send the first local reward (n) for the first base station 102A to the second base station 102B. Moreover, the transmission of the first local reward for the first base station 102A to the second base station 102B allows collaborative learning and adaptation of the agnostic resource allocation models, considering the impact of interference and actions taken by neighboring base stations. The controller is configured to transmit the first local reward to the second base station 102B through a predefined signaling or communication channel to establish an effective and reliable communication between the first base station 102A and the second base station 102B thereby leading to an improved overall network performance.

[0055] Furthermore, the controller is configured to receive a second local reward (rj) for the second base station 102B. In an implementation, the first controller 104A of the first base station 102A receives the second local reward from the second base station 102B in order to provide collaborative learning and adaptation, such as through a predefined signaling or communication channel. As a result, the receiving of the second local reward for the second base station 102B enables the controller to optimize the resource allocation policies, such as by considering the effects of interference and actions taken by the second base station 102B

[0056] Furthermore, the controller is configured to Meta-train the model to be used based on one or more transitions. Moreover, a transition is a tuple comprising the local state, the local resource allocation action, the local and the received rewards (n,rj), and the new local state. The Meta-training of the agnostic model allows the controller to learn how to quickly adapt the agnostic resource allocation model based on the experiences and outcomes of executed actions, enabling continuous improvement and optimization. The controller is configured to collect and store transitions that can be further utilized to train the model to learn from real-world network conditions thereby allowing the controller to a robust and optimized resource allocation.

[0057] In accordance with an embodiment, Meta training includes adapting the agnostic model to an environment-specific model to be used based on the local environment identifiers (ei, ej) and the associated one or more transitions. The adaptation of the agnostic model to an environment-specific model allows the controller to fine-tune the resource allocation for the wireless network 100 with improved overall network performance. During meta-training, the controller uses the collected transitions and the corresponding local environment identifiers to adapt the agnostic model, creating an environment-specific model tailored for those conditions in order to achieve optimal resource allocation performance in various wireless environments, while maintaining a scalable and efficient approach.

[0058] In accordance with an embodiment, the meta training includes adapting the agnostic model to the environment-specific model to be used based on the global environment and one or more transitions. Moreover, the adaption of the agnostic model to an environment-specific model based on the global environment ensures that the resource allocation strategy is optimized for the overall network conditions, considering the impact of interference and actions taken by the second base station 102B. After that, during meta-training, the controller uses the collected transitions and the corresponding global environment identifier to adapt the agnostic model, creating an environment-specific model tailored for those global conditions. Therefore, by adapting the agnostic model to environment-specific models based on the global environment, the controller is configured to provide optimal resource allocation performance while considering the interference and actions of the second base station 102B, leading to improved overall network performance.

[0059] In accordance with an embodiment, the controller is further configured to update the agnostic model to be used with a metagradient step. The controller is configured to update the agnostic model itself with a meta-gradient step during meta-training to improve the general-purpose model's ability to adapt quickly to new environments during the fine-tuning phase. During meta- training, the controller is configured to compute meta-gradients based on the collected transitions and uses them to update the parameters of the agnostic model. As a result, by updating the agnostic model with meta-gradient steps, the controller is configured to reduce the fine-tuning efforts and improve the overall resource allocation performance across various conditions.

[0060] In accordance with an embodiment, the controller is further configured to maintain an indexed memory during the meta- training. Moreover, the indexed memory stores the associations between environments and associated or adapted models. The controller is configured to maintain an indexed memory that stores associations between environments (represented by environment identifiers) and the adapted resource allocation models. The indexed memory allows the controller to efficiently meta-train the agnostic model. In addition, the adapted models are also stored to collect trajectories that are further utilized to update the agnostic model and ensure a reliable meta-training of the agnostic model.

[0061] In accordance with an embodiment, during the meta-training the controller is further configured to store the meta-trained agnostic model in the indexed memory. The controller stores the meta-trained agnostic model in the indexed memory.

[0062] In accordance with an embodiment, during the meta-training the controller is further configured to maintain an indexed buffer storing associations between environments and associated one or more transitions. The controller is configured to maintain an indexed buffer that stores associations between environments (represented by environment identifiers) and the corresponding collected transitions (e.g., tuples of local state, resource allocation action, rewards, and new local state). The indexed buffer organizes the collected transition data in a structured manner, enabling efficient retrieval and utilization for meta-training and model adaptation. Moreover, the controller is configured to collect transitions during the operation and stores the transitions in the indexed buffer, associating each transition with the corresponding environment identifier, enabling effective meta-training and adaptation of resource allocation model. Advantageously, by leveraging local and global environment identifiers to guide resource allocation strategies. The first base station 102A and the second base station 102B determine local environment identifiers through signal measurements like signal-to-noise ratio (SNR), received signal strength indicator (RSSI), and reference signal received quality (RSRQ). Moreover, such local identifiers are exchanged between neighboring APs to compute a global environment identifier, which triggers the adaptation of the agnostic resource allocation model. The first base station 102A employs Reinforcement Learning (RL), particularly Model-Agnostic Meta-Leaming (MAML), to meta-train the agnostic model that can quickly adapt to new environments. By meta-training this model with transitions and rewards, the first base station 102A is configured to enhance the resource allocation policies that reduce communication overhead, enhance interference management, and dynamically adapt to changing environments.

[0063] FIG. 2 is a flowchart of a method for a controller in a first base station, in accordance with an embodiment of the present disclosure. FIG. 2 is described in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown a flowchart of method 200 that includes steps 202 to 224. The controller (i.e., the first controller 104A of the FIG. 1) is configured to execute the method 200.

[0064] There is provided the method 200 for the controller in the first base station 102A. Moreover, the first base station 102A is used for an efficient resource allocation in the wireless network 100, such as by maintaining an optimal network performance and low interference.

[0065] At step 202, the method 200 includes initiating a connection with the second base station 102B. At step 204, the method 200 includes determining the first local environment identifier for the first base station 102A, and at step 206, the method 200 includes sending the first local environment identifier for the first base station 102A to the second base station 102B. Furthermore, at step 208, the method 200 includes receiving a second local environment identifier for the second base station 102B and at step 210, the method 200 includes determining a global environment identifier based on the first local environment identifier for the first base station 102A and the second local environment identifier for the second base station 102B. At step 212, the method 200 includes maintaining an indexed memory. Moreover, the indexed memory stores associations between environments and associated models, and a model can be an agnostic model or an adapted environment-specific model. Thereafter, at step 214, the method 200 includes determining a model to be used based on the global environment identifier for the first base station 102A according to the indexed memory and further determining a local state from the first local environment, such as at step 216. Furthermore, at step 218, the method 200 includes determining a local resource allocation action based on the model to be used and the local state and at step 220, the method 200 includes executing the local resource allocation action. At step 222, the method 200 includes determining a first local reward and a new local state for the first base station 102A and at step 224, the method 200 includes sending the first local reward for the first base station 102A to the second base station 102B. At step 226, the method 200 includes receiving a second local reward for the second base station 102B. At step 228, the method 200 includes Meta-training the model to be used based on one or more transitions, wherein a transition is a tuple comprising the local state, the local resource allocation action, the local and the received rewards, and the new local state.

[0066] The steps 202 to 228 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.

[0067] There is further provided a computer program product comprising program instructions for performing the method 200 when executed by one or more processors in the base station. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Elard Disk Drive (EfDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0068] FIG. 3 is a diagram that illustrates signaling between a first base station and a second base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIG. 1. With reference to FIG. 3, there is shown a diagram 300 of the signaling between the first base station 102A and the second base station 102B operating in the wireless network 100. The first base station 102A includes a first local environment 302A, a first indexed buffer 304A, the first memory 106A, a first receiver or transmitter 306A, and the first controller 104A. Similarly, the second base station 102B includes a second local environment 302B, a second indexed buffer 304B, the second memory 106B, a second receiver or transmitter 306B, and the second controller 104B.

[0069] In an implementation, the first controller 104A is configured to identify the local environment (i.e., the first local environment 302, Ej) and send the corresponding first local Env ID (e to neighboring access points (Aps). Thereafter, the first controller 104A is configured to determine a global environment identifier (i.e., Env IDe) after receiving local Env IDs from other APs. Furthermore, the first controller 104A is configured to take actions (af) based on the policy provided by the first memory 106A (or a first indexed memory) and train the agnostic policy during the meta-training phase in order to adapt the agnostic policy to any environment in the fine-tuning phase. Moreover, the first memory 106A (or the indexed memory) is configured to keep track of multiple policies (with 'NN' parameters) required for meta-training and determine which policy will be used at each time slot. The first indexed buffer 304A is configured to store experience trajectories per global Env ID (e) in a structured manner. Similarly, the second controller 104B is configured to determine the second local environment identifier. The second controller 104B is configured to determine a second local reward for the resource allocation action and a new local state for the second base station 102B and Meta-train the model to be used based on one or more transitions that are stored in the second indexed buffer 304B. Moreover, the second memory 106B (or a second indexed memory) is configured to store associations between environments and the corresponding resource allocation models (i.e., agnostic or adapted). At operation 306, the first controller 104A is configured to initiate communication with the second base station 102B, such as by messaging for initialization. Thereafter, the second controller 104B of the second base station 102B is configured to send local environment identifiers, such as at operation 308. Furthermore, at operation 310, the first controller 104A is further configured to send the first local reward (i.e., n) for the first base station 102A to the second base station 102B. Finally, the resource allocation based on the determined local environment identifier is performed.

[0070] FIG. 4 is a diagram that illustrates communication between a first base station and a second base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 4 is described in conjunction with elements from FIG. 1, 2 and 3. With reference to FIG. 4, there is shown a diagram 400 of the communication between the first base station 102A and the second base station 102B operating in the wireless network 100.

[0071] In an implementation scenario, at operation 402, the first base station 102A is configured to determine the first local environment identifier and at operation 404, the first base station 102A is configured to perform an initiation of a handshake (mi) and exchange a handshake message (mi) in order to establish a connection between the first base station 102A and the second base station 102B, such as at operation 406. Furthermore, at operation 408, the second base station 102B is configured to receive the second local environment identifier from the second local environment 302B. Thereafter, at operation 410 and operation 412, the first base station 102A and the second base station 102B performed an exchange of local environment identifiers (ei, ej) to allow both the controllers to identify and exchange their unique environment identifiers (ei, ej) to establish their identities. Moreover, at operation 418 and at operation 420, the first base station 102A is configured to determine a first local reward for the resource allocation action. Similarly, at operation 414 and at operation 416, the second base station 102B is configured to determine a second local reward for the resource allocation action. At operation 422 and operation 424, the first base station 102A and the second base station 102B are configured to exchange rewards, such as through reward signals (n, rj). After that, the first base station 102A and the second base station 102B are configured to determine new local environment identifiers, such as at operation 426 and at operation 428. Furthermore, the first base station 102A and the second base station 102B are configured to notify new local environment identifiers (e;1) that are identified by the first base station 102A (e.g., at operation 430). Thereafter, at operation 432 and at operation 434, the first base station 102A and the second base station 102B are configured to exchange updated reward signals (ri, rj') based on the new environment identifier. Moreover, the APs may need to redefine or update their environment classes to accommodate the changes or new conditions in order to allow radio resource allocation based on new environment conditions to enhance the overall network performance. Therefore, at operation 436, the APs are configured to restart the meta-training process (mi1) in order to incorporate the new data and continue the collaborative learning.

[0072] FIG. 5 is a diagram that illustrates operations for resource allocation within a wireless network, in accordance with an embodiment of the present disclosure. FIG. 5 is described in conjunction with elements from FIGs 1 to 4. With reference to FIG. 5, there is shown a diagram 500 illustrating the operations for resource allocation within the wireless network 100.

[0073] In an exemplary scenario, at operation 502, the handshake signaling process is established in order to start or restart the meta- training of the agnostic models. Each of the base stations (i.e., the first base station 102A) or the access point (AP) defines the environment classes and the associated local environment identifier and initiates the associated local agnostic policy (i.e., 0°). In order to establish the start of meta-training, the first base station 102A is configured to send the initialization handshake message (i.e., m,) composed of the category type (e.g., SNR levels) to the second base station 102B along with the number of total classes (i.e., m, = {category _type, num_c lasses | C, | }). Moreover, if the first base station 102A decides to modify the environment classes then, in that case, the first base station 102A is configured to send the modified environment classes to the second base station 102B in order to restart the meta-training phase. At operation 504, the controller (i.e., the first controller 104A) of the first base station 102A is configured to receive the local environment identifier and exchange rewards in order to train or adapt the agnostic model, such as by meta-training the model to be used based on one or more transitions. The meta- training refers to a process of learning an optimal agnostic policy and monitoring when to restart meta-training. At each time interval, the controller of the second base station 102B is configured to identify the local environment identifier and send the local environment identifier to the neighbor base stations (i.e., the first base station 102A), if it is different from the previous one, and maps all the local IDs to a global environment identifier that is provided to the indexed buffer and memory. The indexed memory (i.e., the first memory 106A of the first base station 102A and the second memory 106B of the second base station 102B) is configured to send the parameters of the adapted policy if they are available, or otherwise the parameters of the agnostic model. Furthermore, the controller is configured to take a decision (i.e., action) based on the local state (i.e., the local monitored states) of associated devices. The controller is configured to determine a new local state and the corresponding local reward as a result of the action and exchanges the local reward with neighbors (i.e., the first base station 102A and the second base station 102B) before sending the transition information (i.e., local state, local action, global reward, and new local state) to the indexed buffer. At each time slot, the controller may execute any of the computation tasks, such as constructing the adapted policy (i.e., 0e) when there are enough trajectories {re} collected from the environment (i.e., e) using the agnostic parameters (i. e. , 0). Furthermore, the controller is configured to compute a meta-gradient (i. e. ,ge)~ when there are enough trajectories {ree} collected from the environment e when using (i.e., 0e) and update the agnostic policy when sufficient metagradients are calculated. Finally, the controller is configured to erase all the adapted parameters from the first memory 106A (i.e., the indexed memory) and all the trajectories from the indexed buffer. Moreover, the optimal agnostic parameters ( i. e., 0*) are obtained at the end of the meta-training phase. Moreover, this phase can be restarted (i.e., 0Z<- 0°) any time when the first base station 102A sends or receives a handshake message for initialization. At operation 512, the controller is configured to end the process of meta-training, and at operation 506, the controller is configured to perform fine-tuning of the agnostic model that provides fast adaptation to the environment (i.e., e) by fine-tuning (0* to 0*) and monitoring when to restart fine- tuning at operation 514 or restart meta-training at operation 510. The phase of fine-tuning starts after the end of the meta- training phase and restarts each time the global environment identifier changes. At the beginning of this phase, each base station (i.e., the first base station 102A and the second base station 102B) is configured to identify the global environment identifiers and exchange the rewards with the second base station 102B in order to adapt to the environment (i. e., e, i.e., fine-tune 0* to 0e*) using a few gradient steps. Furthermore, at operation 514, the operation 506 is performed again if the global environment identifier is changed. At operation 516, the controller is configured to end the process of fine-tuning and at operation 510, the controller is configured to message the first base station 102Ato restart the meta-training phase. At operation 508, the controller is configured to apply the adapted model (i.e., 0*), and monitoring when to fine-tune again at operation 518 or restart meta- training at operation 510. The controller is configured to apply the adapted parameters (i.e., 0*) on the environment ( i. e. , e ). As a result, the resource allocation with overall enhanced network performance can be performed by the first base station 102A operating in the wireless network 100.

[0074] FIG. 6 is a diagram that illustrates an exemplary scenario of a first base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 6 is described in conjunction with elements from FIGs. 1 to 5. With reference to FIG. 6, there is shown a diagram 600 illustrating the operations for resource allocation within the wireless network 100.

[0075] In an exemplary scenario, the local environment is defined by the local SNR level experienced by the first base station 102A (i.e., an AP (i) ) and is categorized into groups, such as very low, low, medium, high, and very high (i.e., EtE C,- = {very low, low, medium, high, very high}) with IDs, such as e;E {0, |C;|— 1}. Moreover, at operation 606, the first base station 102A is configured to initiate the handshake, such as by initiating a communication with the second base station 102B (i.e., mi={'SNR, 5}). At operation 608, the second base station 102B is configured to establish a communication link with the first base station 102A (i.e., m2={'SN , 4}). Thereafter, each of the base stations (i.e., the first base station 102A and the second base station 102B) or the AP is configured to identify the local environment from local signaling measurements: At operation 602, the first base station 102A is configured to identify its local environment from local signaling measurements (Vj), determining the local environment identifier (ei = I Div, ) = 3), which corresponds to "High SNR" from the set C, = {very low, low, medium, high, very high}, |Ct| = 5, and of type "SNR". Similarly, at operation 604, the second base station 102B is configured to identify the local environment (e2 = ID(u2) = 1), corresponding to "low" from the set C2= {very low, low, medium, high}, |C2| = 4, and of type "SNR". Thereafter, the first base station 102 A determine the global environment identifier after exchanging local environment identifiers (e.g., ei—3' and e2= ) with the second base station 102B, such as at operation 618 and at operation 620 respectively. At operations 614 and 616, the first base station 102A is configured to determine a first local reward for the resource allocation action and a new local state for the first base station 102A. Similarly, at operations 622 and 624, the second base station 102B is configured to determine a second local reward for the resource allocation action and a new local state for the second base station 102B. The global environment identifiers are derived from the combined local environment identifiers using functions fi and fi, such as at operation 610 and at operation 612. However, each base station (or AP) requires a local function to map the received and local environment identifiers to a global identifier. In an example the mapping function can be by the given below equation (1):

[0076] / (ei . ew) = Z =il e; 11 / =;+! |Cjl + eN(1)

[0077] In an example the first base station 102A has five possible categories, such as {very low, low, medium, high, very high (i.e., | | = 5), with type SNR and the second base station 102B has four categories, such as {very low, low, medium, high} (i.e., | C21 = 4). In such an example, the mapping function for each base station can be defined as the given below equation (2): fi(e1,e2') = f2(e1,e2') = 4et+ e2(2)

[0078] Where E1=”EIigh SNR”, ej= 3, e^fD^), e=f; (e1 ;e2), e2=ID(v2), e=f2(e1,e2), E2= "low”, e2=l. As a result, an efficient resource allocation within the wireless network 100 with enhanced communication and coordination between the first base station 102A and the second base station 102B through local and global environment identification and reward exchange is performed.

[0079] FIG. 7 is a diagram that illustrates another exemplary scenario of a first base station operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 7 is described in conjunction with elements from FIGs. 1 to 6. With reference to FIG. 7, there is shown a diagram 700 illustrating the operations for resource allocation within the wireless network 100.

[0080] In an exemplary scenario, the local environments are defined by the local SNR levels experienced by the base stations (i.e., APs) and are categorized into groups, such as {very low, low, medium, high, and very high} . At operation 712 and operation 714, the first base station 102A and the third base station 702 are configured to initiate handshake (i.e., m3={'SNR, 3} and mi={'SNR, 5}) in order to initiate the communication between both the base stations. Similarly, at operation 724 and 726 the second base station 102B and the third base station 702 are configured to initiate handshake (i.e., m3={'SNR, 3} and m2={'SNR, 4}) in order to initiate the communication between both the base stations. At operation 602, the first base station 102A is configured to identify its local environment from local signaling measurements (v; ), determining the local environment identifier (ei = ID(v1) = 3), which corresponds to "High " from the set Cj = {very low, low, medium, high, very high}, |Ct| = 5, and of type "SNR". Similarly, at operation 604, the second base station 102B is configured to identify the local environment (e2 = ID(V2) = 1), corresponding to "low" from the set C2= {very low, low, medium, high}, |C2| = 4, and of type "SNR". Additionally, the third local environment 704 with the third base station 702 identifies a local environment identifier from local signaling measurements (v3) at operation 706, determining the local environment identifier (e3 = H V3) = 0), which corresponds to "medium" from the set iC = {medium, high, very high}, |C3| = 3) and of type "SNR". Furthermore, at operation 716 and at operation 718, both the stations, such as the first base station 102A and the third base station 702 are configured to exchange the determined local environment identifiers (e.g., e3=0 and ei=3). Similarly, at operation 728 and at operation 730, both the stations, such as the second base station 102B and the third base station 702 are configured to exchange the determined local environment identifiers (e.g., e3=0 and e2=l). Each base station (or AP) requires a local mapping function to map the received and local environment identifiers. The mapping functions for each base station can be defined by the equation (3) given below: e= / (ei>e2>e3) = 12el+ 3e2+ e3(3)

[0081] . At operations 614 and 616, the first base station 102A is configured to determine a local reward for the resource allocation action and a new local state. Similarly, at operations 622 and 624, the second base station 102B determines a local reward and a new local state. At operations 626 and 628, the first base station 102A and the second base station 102B exchange local rewards between the base stations. Similarly, at operation 720 and 722, the first base station 102A and the third base station 702 are configured to exchange rewards for the resource allocation action (i.e., rs and n) and at operation 732 and 734, the second base station 102B and the third base station 702 are configured to exchange rewards for the resource allocation action (i.e., rs and n). As a result, an efficient resource allocation within the wireless network 100 with enhanced communication and coordination between the base stations is performed through local and global environment identification and reward exchange.

[0082] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.

Claims

CLAIMS1. A first base station (102A) arranged to operate in a wireless network (100), wherein the first base station (102A) comprises a controller configured to: initiate a connection with a second base station (102B), determine a first local environment identifier for the first base station (102A), receive a second local environment identifier for the second base station (102B), determine a global environment identifier based on the first local environment identifier for the first base station (102A) and the second local environment identifier for the second base station (102B), determine a model to be used based on the global environment identifier for the first base station (102A) according to associations between environments and associated models, wherein a model comprises an agnostic model or an adapted environment specific model, determine a local state from the first local environment, perform a local resource allocation action based on the model to be used and the local state, determine a first local reward for the resource allocation action and a new local state for the base station, receive a second local reward for the second base station,Meta-train the model to be used based on one or more transitions, wherein a transition is a tuple comprising the local state , the local resource allocation action, the local and the received rewards and the new local state.

2. The first base station (102A) according to claim 1, wherein the controller is further configured to send the first local environment identifier for the first base station (102A) to the second base station (102B).

3. The first base station (102A) according to claim 1 or 2, wherein the controller is further configured to send the first local reward for the first base station (102A) to the second base station (102B).

4. The first base station (102A) according to any one of claims 1 to 3, wherein the controller is further configured to maintain an indexed memory, wherein the indexed memory stores the associations between environments and associated models.

5. The first base station (102A) according to any one of claims 1 to 4, wherein the controller is further configured to store the Meta-trained agnostic model in the indexed memory.

6. The first base station (102A) according to any one of claims 1 to 5, wherein the controller is further configured to determine a change in a local environment, and in response thereto determine a new local environment identifier, send the new environment identifier to the second base station (102B), determine a new global environment identifier based on the new environment identifier, and select a new resource allocation model based on the new global environment identifier.

7. The first base station (102A) according to any preceding claim, wherein the controller is further configured to receive a new environment identifier from the second base station (102B), and in response thereto determine a new global environment identifier based on the new environment identifier, and select a new resource allocation model based on the new global environment identifier.

8. The first base station (102A) according to any preceding claim, wherein the controller is further configured to determine the first local environment identifier by making measurements and determining the first local environment identifier based on the measurements.

9. The first base station (102A) according to claim 8, wherein an environment relates to a signal measurement, and an environment identifier relates to a classification of the environment10. The first base station (102A) according to claim 9, wherein the signal measurement relates to signal-to-noise ratio, SNR, Reference Signal Received Quality, RSRQ, Received Signal Strength Indicator, RSSI, or Reference Signal Received Power, RSRP.

11. The first base station (102A) according to any preceding claim, wherein the controller is further configured to initiate the connection through a handshake protocol exchanging information on environment classes.

12. The first base station (102A) according to any preceding claim, wherein initiating the connection includes initiating the agnostic model.

13. The first base station (102A) according to any preceding claim, wherein the controller is further configured to maintain an indexed buffer storing associations between environments and associated one or more transitions.

14. The first base station (102A) according to claim 13, wherein Meta training includes adapting the agnostic model to environment-specific model to be used based on the local environment identifiers and the associated one or more transitions.

15. The first base station (102A) according to claim 14, wherein the controller is further configured to update the agnostic model to be used with a meta-gradient step.

16. The first base station (102A) according to any preceding claim, wherein Meta training includes adapting the agnostic model to environment-specific model to be used based on the global environment and one or more transitions.

17. The first base station (102A) according to any preceding claim, wherein the controller is further configured to receive second local environment identifiers from more than one second base station (102B).

18. A method (200) for a controller in a first base station (102A), the method (200) comprising: initiating a connection with a second base station (102B), determining a first local environment identifier for the first base station (102A), sending the first local environment identifier for the first base station (102A) to the second base station (102B), receiving a second local environment identifier for the second base station (102B), determining a global environment identifier based on the first local environment identifier for the first base station (102A) and the second local environment identifier for the second base station (102B), maintaining an indexed memory, wherein the indexed memory stores associations between environments and associated models, wherein a model can be an agnostic model or an adapted environment-specific model, determining a model to be used based on the global environment identifier for the base station according to the indexed memory, determining a local state from the first local environment, determining a local resource allocation action based on the model to be used and the local state,executing the local resource allocation action, determining a first local reward and a new local state for the first base station (102A), sending the first local reward for the first base station (102A) to the second base station, receiving a second local reward for the second base station (102B), Meta-training the model to be used based on one or more transitions, wherein a transition is a tuple comprising the local state, the local resource allocation action, the local and the received rewards and the new local state.

19. A computer program product comprising program instructions for performing the method (200) according to claim 18, when executed by one or more processors in a first base station (102A).