Network device and method to perform resource allocation tasks utilizing single multiple-parameter dnn
The network device employs a single multiple-parameter DNN to adapt dynamically to changing radio resources, minimizing interference and optimizing resource allocation in decentralized wireless networks, thereby improving network performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-18
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional wireless networks face challenges in dynamically adapting to changing radio resources due to high interference and excessive communication overhead, leading to inefficient resource allocation and reduced network performance.
A network device utilizing a single multiple-parameter Deep Neural Network (DNN) for decentralized resource allocation, which receives local and neighboring parameters, applies a mapping function to determine a common parameter, and performs resource allocation tasks, minimizing interference and communication overhead.
Enhances network performance by adapting to dynamic conditions, reducing interference, and optimizing resource allocation across decentralized wireless networks without the need for manual tuning or excessive communication.
Smart Images

Figure EP2024082724_21052026_PF_FP_ABST
Abstract
Description
[0001] NETWORK DEVICE AND METHOD TO PERFORM RESOURCE ALLOCATION TASKS UTILIZING SINGLE MULTIPLE-PARAMETER DNN
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to the field of wireless communication networks and more specifically, to a network device and a method for the network device configured to perform resource allocation tasks utilizing a single multipleparameter Deep Neural Network (DNN), such as by providing a multi-agent resource allocation in wireless networks with dynamic action spaces.
[0004] BACKGROUND
[0005] In existing wireless networks, network devices are responsible for allocating radio resources to a large number of connected devices, ensuring efficient communication within the network. Radio resource allocation strategies (or policies) are designed to allow network devices to make decisions based on available network-related information. While these policies may work well for a fixed or specific set of radio resources (e.g., defined resource blocks or power levels), wireless networks often experience dynamic variations in radio resources due to factors like changing channel conditions or frequency bands. Moreover, to handle such variations, mobile network operators often need to manually adjust the network device policies, which can be time-consuming and computationally expensive. Additionally, in decentralized systems, where each network device primarily relies on local information from its connected devices, there is a risk of high interference between network devices that leads to a degradation of overall network performance, as network devices may take actions that are locally optimal but harmful to the broader network.
[0006] Conventionally, in wireless communication networks, resource allocation policies do not effectively account for the dynamic nature of the environment, resulting in inaccurate and unreliable allocation of resources. While conventional decentralized network device policies rely on local information, they often fail to consider broader network- wide patterns that could improve resource allocation that leads to significant challenges, such as increased interference between network devices and reduced overall network performance. Certain attempts have been made to enhance the reliability of resource allocation policies, including the use of techniques, such as reinforcement learning (RL) but often face issues, such as slow adaptation to dynamic radio resources, high interference between neighbouring network devices, and excessive communication overhead. Thus, there exists a technical problem of how to train decentralized radio resource allocation policies with low communication overhead between neighboring network devices, which can be automatically and quickly adapted to the dynamic set of available resources.
[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional network devices and conventional methods for resource allocation utilizing Deep Neural Networks (DNN).
[0008] SUMMARY
[0009] The present disclosure provides a network device and a method for the network device configured to perform resource allocation tasks utilizing a single multiple-parameter Deep Neural Network (DNN), such as by providing a multi-agent resource allocation in wireless networks with dynamic action spaces. The present disclosure provides a solution to the existing problem of how to train decentralized radio resource allocation policies with low communication overhead between neighboring network devices, which can be automatically and quickly adapted to the dynamic set of available resources. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the network device and the method for network device associated with one or more terminal devices, such as for resource allocation in a dynamic wireless environment. One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.
[0010] In one aspect, the present disclosure provides a network device configured to perform resource allocation tasks utilizing a single multiple-parameter Deep Neural Network (DNN) and a resource allocation task is associated with an action space which is based on at least one parameter. Moreover, the network device is configured to receive an indication of a resource allocation task to be performed, determine a local parameter, receive neighbouring parameters from neighbouring network device(s), determine a common parameter by applying a mapping function to the determined local parameter and the received neighbouring parameters, and then perform the resource allocation task utilizing the single multiple-parameter DNN based on the common parameter. Moreover, the single multiple-parameter DNN has been trained based on the resource allocation task and all possible values for the common parameter.
[0011] Advantageously, the network device is configured to perform resource allocation through resource allocation policies in decentralized wireless networks with dynamic conditions. Unlike traditional systems that require manual tuning of policies or rely on a conventional reinforcement learning, which struggles to generalize across multiple action spaces. The network device is configured to handle multiple action spaces reducing the need for retraining and memory overhead. Additionally, the network device is configured to minimize the interference between the network devices by enabling the communication and cooperative policy learning with enhanced overall network performance without excessive communication overhead. As a result, the network device is configured to perform various wireless tasks, such as bandwidth or power allocation, and is adaptable to dynamic network environments, thereby enhancing the overall performance of the wireless network.
[0012] In another aspect, the present disclosure provides a method for a network device configured to perform resource allocation tasks utilizing a single multiple-parameter and a resource allocation task is associated with an action space which is based on at least one parameter. Moreover, the method comprises receiving an indication of a resource allocation task to be performed, determining a local parameter, receiving neighbouring parameters from neighbouring network device(s), determining a common parameter by applying a mapping function to the determined local parameter and the received neighbouring parameters), and then performing the resource allocation task utilizing the single multiple-parameter DNN, based on the common parameter, wherein the single multiple-parameter Deep Neural Network, DNN, has been trained based on the resource allocation task and all possible values for the common parameter.
[0013] The method achieves all the advantages and technical effects of the network device of the present disclosure.
[0014] It is to be appreciated that all the aforementioned implementation forms can be combined.
[0015] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.
[0016] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.
[0018] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:
[0019] FIG. 1 is a block diagram of a network device configured to perform resource allocation tasks utilizing a single multipleparameter Deep Neural network (DNN), in accordance with an embodiment of the present disclosure;
[0020] FIG. 2 is a flowchart of a method for performing resource allocation tasks utilizing a single multiple-parameter Deep Neural network (DNN), in accordance with an embodiment of the present disclosure;
[0021] FIG. 3 is a diagram that illustrates an architecture and communication of the network device with the neighbouring network devices, in accordance with an embodiment of the present disclosure;
[0022] FIG. 4A is a diagram that illustrates an exemplary scenario of the network device operating in a wireless network, in accordance with an embodiment of the present disclosure; and
[0023] FIG. 4B is a diagram that illustrates another exemplary scenario of the network device operating in a wireless network, in accordance with an embodiment of the present disclosure.
[0024] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0025] DETAILED DESCRIPTION OF EMBODIMENTS
[0026] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0027] FIG. 1 is a block diagram of a network device configured to perform resource allocation tasks utilizing a single multipleparameter Deep Neural network (DNN), in accordance with an embodiment of the present disclosure. With reference to FIG.
[0028] 1, there is shown a block diagram 100 that includes a network device 102, a communication network 110, and neighbouring network devices 112. Moreover, the network device 102 includes a controller 104, a memory 106, and a network interface 108.
[0029] The network device 102 and each of the neighbouring network devices 112, such as a first neighbouring network device 112A, a second neighbouring network device 112B, up to Nth neighbouring device 112N refers to N base stations operating in the wireless network and are responsible to provide multi-agent resource allocation in the wireless network with dynamic action spaces. In accordance with an embodiment, the network device 102 is a base station. The base station is responsible for allocating radio resources, such as bandwidth, power, and beamforming directions to ensure efficient communication for multiple users within its coverage area. As a result, the network device 102 is configured to enhance the ability of the wireless network to manage network traffic, provide high-quality service to users, and maintain efficient communication across the entire wireless infrastructure.
[0030] The controller 104 of the network device 102 is configured to perform resource allocation tasks utilizing the single multipleparameter DNN. Examples of the controller 104 may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.
[0031] The memory 106 is used to store data related to multi-agent resource allocation, and the like. Examples of implementation of the memory 106 may include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Dynamic Random Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory. The network interface 108 is used by the network device 102 to communicate with the controller 104 and examples of implementation of the network interface 108 may include but are not limited to a network interface, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.
[0032] The communication network 110 includes a medium (e.g., a communication channel) through which the network device 102 communicates with the neighbouring network devices 112 of the wireless network. Examples of the communication network 110 may include, but are not limited to, a cellular network (e.g., a 2G, a 3G, long-term evolution (LTE) 4G, a 5G, or 5G New Radio (NR) network, such as sub 6 GHz, cmWave, or mmWave communication network), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to-network ( V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.
[0033] There is provided the network device 102 that is configured to perform resource allocation tasks utilizing a single multipleparameter Deep Neural Network (DNN). Moreover, a resource allocation task is associated with an action space (A) which is based on at least one parameter (|i). The resource allocation task in the wireless network is associated with the action space (A) that represents the set of actions that the network device 102 can take when allocating resources. In an implementation, the at least one parameter (p) can be a bandwidth numerology index, power discretization level, resource requirement, and the like. Additionally, the network device 102 is configured to adapt to varying radio resources due to changes in channel conditions, frequency availability, and power requirements while minimizing communication overhead between the network devices. The network device 102 is configured to receive an indication of a resource allocation task to be performed. In an implementation, the network device 102 is configured to receive signals or instructions about specific resource allocation tasks that need to be addressed, such as distributing bandwidth among users or adjusting power levels based on network load. As a result, by receiving the indication of the resource allocation task to be performed, the network device 102 is configured to adapt to network change in order to optimize resource utilization and efficiently manage the task while maintaining scalability and network performance.
[0034] In accordance with an embodiment, the network device 102 is further configured to receive the indication of the resource allocation task to be performed through an initialization message comprising the indication of the resource allocation task to be performed. The indication of the resource allocation task allows the network device 102 to understand the specific resource allocation requirements, enabling the network device 102 to respond swiftly to the demands of the wireless network. As a result, by receiving the indication of the resource allocation task is used to initialize the resource allocation with an improved performance, especially in scenarios that require frequent reconfiguration of resources while reducing the overall communication overhead.
[0035] In accordance with an embodiment, the initialization message also comprises an indication of what parameters to be exchanged. In an implementation, the parameters to be exchanged are used to ensure smooth and efficient resource allocation in dynamic and decentralized wireless networks and the indication of what parameters to exchange include bandwidth settings, power levels, beam selection configurations, and the like.
[0036] In accordance with an embodiment, the initialization message also comprises an indication of the mapping function (f). The indication of the mapping function (f) is used to ensure proper coordination between the network devices during resource allocation. The mapping function (f) refers to a function that is used to harmonize local parameters across multiple network devices by mapping local radio resource parameters (e.g., bandwidth numerology indexes) into a common parameter. By including the mapping function in the initialization message, the network device 102 can optimize the resource allocation strategies, ensuring that each of the resources works cooperatively and avoids interference, particularly in decentralized environments where each network device manages its own local resources.
[0037] In accordance with an embodiment, the mapping function (f) is associated with the resource allocation task to be performed. For example, in the case of bandwidth allocation, the mapping function can be used to adjust bandwidth subdivisions of the network devices so that the network device 102 and the neighbouring network devices 112 utilize a common numerology for better coordination.
[0038] In accordance with an embodiment, all possible values for the common parameter (pc) are given by a predefined set of common parameter values ({pc}). In an implementation, all the possible values for the common parameter are used to provide a range of acceptable values that can be used by the network device 102 during resource allocation. For example, in a beamforming task, the common parameter could represent the selection of beamforming codebooks, and the predefined set may include all configurations that the network device 102 can further utilize. Therefore, by using all values for the common parameter that are given by a predefined set of common parameter values, the network device 102 is configured to ensure that all the network devices operate within a defined and compatible range thereby enhancing the coordination and further reducing the interference. In accordance with an embodiment, the initialization message also includes an indication of a predefined set of common parameter values ({pc}). The initialization message not only indicates the resource allocation task and mapping function but also includes the predefined set of common parameter values, which are used to identify the available range of common parameter values that can be used during the task. As a result, the initialization message that includes the indication of the predefined set of common parameter values is used to ensure smooth operation across the wireless network, as each network device (i.e., the network device 102) can adjust the resource allocation based on the same set of predefined parameters thereby ensuring an efficient communication and collaboration across the wireless environment.
[0039] Furthermore, the network device 102 is configured to determine a local parameter (pi). In an implementation, the local parameter refers to a parameter, such as bandwidth numerology index or power discretization level for the connected network device that serves as an initial input for performing a resource allocation task. Moreover, by determining the local parameter, the network device 102 is configured to ensure that the resource management is aligned with current network conditions, leading to efficient and optimized operations.
[0040]
[0041] bandwidth numerology indexes or the power discretization levels, the network device 102 is configured to make an informed decision to minimize interference and improve overall network performance. As a result, the overall performance of the wireless network is enhanced.
[0042]
[0043] In accordance with an embodiment, the network device 102 is further configured to determine a local reward (n), receive neighbouring rewards (r-i) from the neighbouring network devices 112, determine a group reward (R) based on the local reward (n) and the received neighbouring rewards (r-i), train one parameter-specific DNN based on the group reward (R) for each possible value of the common parameter (pc). Moreover, the specific-parameter DNN is specific to one value of the common parameter (pc) and train the single multiple-parameter DNN based on the specific-parameter DNNs, whereby the single multiple-parameter DNN will be trained for all values for the common parameter (pc) over time. In an implementation, the network device 102 is configured to determine the local reward based on the overall performance of its associated users and receives the neighbouring rewards from the neighbouring network devices 112. Moreover, such rewards are combined into a group reward (R), which reflects the overall network performance. Additionally, the network device 102 is configured to train a parameter-specific DNN for each value of the common parameter based on the group reward. As a result, the network device 102 is configured to ensure an efficient training and broad applicability of the multiple-parameter DNN across different resource allocation scenarios.
[0044]
[0045] In accordance with an embodiment, the network device 102 is further configured to determine the group reward (R) as the sum of the local reward (n) and the received neighbouring rewards (r-i). In an implementation, the neighbouring network devices 112, such as the first neighbouring network device 112A, the second neighbouring network device 112B, up to the nth neighbouring device 112N are configured to send the neighbouring rewards to the network device 102. Moreover, the determination of the group reward as the sum of the local reward and the received neighbouring reward is used to provide a comprehensive view of the overall performance of the wireless network thereby, encouraging cooperation between the network devices. As a result, by considering both local and neighbouring rewards for computing the group reward, the overall performance of the wireless network is identified in order to provide an enhanced balanced resource allocation with minimized interference by the neighbouring devices 112.
[0046] In accordance with an embodiment, the network device 102 is further configured to determine the local reward (n) as a summation of rewards of associated users to ensure that the resource allocation decisions reflect the overall satisfaction of all users, rather than focusing on individual metrics, leading to more equitable and efficient resource distribution across the network devices of the wireless network.
[0047] In accordance with an embodiment, the network device 102 is further configured to send the local parameter (pi) to neighbouring network devices. The transmission of the local parameter from the network device 102 to the neighbouring network devices 112 allows a coordination and optimization of the resource allocation across the wireless network thereby reducing the risk of interference and ensuring that the resources are distributed efficiently among all the network devices. In accordance with an embodiment, when the resource allocation task is to allocate bandwidth, the corresponding action space (A) is based on a number of associated users (A) and the parameter (p) and the parameter (p) is a numerology index. When the resource allocation task involves bandwidth allocation, the action space (A) is determined by both the number of connected devices (N) and the numerology index (p). The numerology index (p) is used to allocate the bandwidth as subdivided blocks across time and frequency domains. As a result, the network device 102 is configured to allow bandwidth resource allocation, such as based on user demand and network conditions.
[0048] In accordance with an embodiment, when the resource allocation task is to allocate power, the corresponding action space (A) is based on a number of associated users (A) and the parameter (p) and the parameter (p) is a discretization level. For power allocation tasks, the action space (A) is determined by the number of connected devices (N) and the discretization level (p), which defines the available power levels and further ensures that the power resources are distributed efficiently thereby balancing the overall requirement of each of the network devices of the wireless network while avoiding excess power consumption and minimizing interference.
[0049] Advantageously, the network device 102 is configured to perform resource allocation through resource allocation policies in decentralized wireless networks with dynamic conditions. Unlike traditional systems that require manual tuning of policies or rely on the conventional reinforcement learning, which struggles to generalize across multiple action spaces. The network device 102 is configured to handle multiple action spaces reducing the need for retraining and memory overhead. Additionally, the network device 102 is configured to minimize the interference between the network devices by enabling the communication and cooperative policy learning with enhanced overall network performance without excessive communication overhead. As a result, the network device 102 is configured to perform various wireless tasks, such as bandwidth or power allocation, and is adaptable to dynamic network environments, thereby enhancing the overall performance of the wireless network.
[0050] FIG. 2 is a flowchart of a method for performing resource allocation tasks utilizing a single multiple-parameter Deep Neural network (DNN), in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a flowchart of a method 200 that includes steps 202 to 210. The network device (i.e., the network device 102 of FIG. 1) is configured to execute the method 200.
[0051] There is provided the method 200 for the network device 102 configured to perform resource allocation tasks utilizing a single multiple-parameter DNN. Moreover, the resource allocation task is associated with an action space (A) which is based on at least one parameter. The resource allocation task in the wireless network is associated with the action space (A) that represents the set of actions that the network device 102 can take when allocating resources. In an implementation, the at least one parameter (p) can be a bandwidth numerology index, power discretization level, resource requirement, and the like. Additionally, the network device 102 is configured to adapt to varying radio resources due to changes in channel conditions, frequency availability, and power requirements while minimizing communication overhead between the network devices. At step 202, the method 200 includes receiving an indication of a resource allocation task to be performed. By receiving the indication of the resource allocation task to be performed, the network device 102 is configured to adapt to network change in order to optimize resource utilization and efficiently manage the task while maintaining scalability and network performance. At step 204, the method 200 includes determining a local parameter. By determining the local parameter, the network device 102 is configured to ensure that the resource management is aligned with current network conditions, leading to an efficient and optimized operations.
[0052] At step 206, the method 200 includes receiving neighbouring parameters from neighbouring network device(s). Moreover, the exchange of parameters allows for the cooperation between the network devices (e.g., the network device 102 and the neighbouring network devices 112). For example, when the neighbouring network devices 112 share the bandwidth numerology indexes or power discretization levels, the network device 102 is configured to make an informed decision to minimize interference and improve overall network performance. As a result, the overall performance of the wireless network is enhanced.
[0053] At step 208, the method 200 includes determining a common parameter by applying a mapping function (f) to the determined local parameter and the received neighbouring parameters and then performing the resource allocation task utilizing the single multiple-parameter DNN, based on the common parameter, at step 210. Moreover, the single multiple-parameter DNN, has been trained based on the resource allocation task and all possible values for the common parameter. In an implementation, the network device 102 is configured to dynamically adapt to changing network conditions and perform resource allocation more efficiently by utilizing a DNN that has been generalized across all values of the common parameters thereby eliminating the need for retraining.
[0054] In accordance with an embodiment, the method 200 further includes determining a local reward (n), receiving neighbouring rewards (r-i) from neighbouring network device(s), determining a group reward (R) based on the local reward (n) and the received neighbouring rewards (r-i), training one specific-parameter DNN based on the group reward (R) according to the common parameter during training. Moreover, the specific-parameter DNN is specific to one possible value of the common parameter and trains the single multiple-parameter DNN based on the specific-parameter DNNs, whereby the single multipleparameter DNN will be trained for all possible values for the common parameter over time. In an implementation, the network device 102 is configured to determine the local reward based on the overall performance of its associated users and receives the neighbouring rewards from the neighbouring network devices 112. Moreover, such rewards are combined into a group reward (R), which reflects the overall network performance. Additionally, the network device 102 is configured to train a specific-parameter DNN for each value of the common parameter based on the group reward. As a result, the network device 102 is configured to ensure an efficient training and broad applicability of the single multiple-parameter DNN across different resource allocation scenarios.
[0055] In accordance with an embodiment, the method 200 further includes repeating determining the local parameter, receiving neighbouring parameters from neighbouring network device(s), determining the common parameter by applying the mapping function (f), performing the resource allocation task utilizing the specific-parameter DNN based on the common parameter, receiving neighbouring rewards from neighbouring network device) s), sending the local reward to the neighbouring network device(s) and determining the group reward (R) based on the local reward and the received neighbouring rewards. At each k time steps train the specific-parameter DNNs. By repeating the process at defined intervals (k time steps), the network device 102 is configured to adapt to dynamic changes in the network environment, ensuring that resource allocation remains optimized in real time.
[0056] Advantageously, the method 200 is used to perform resource allocation through resource allocation policies in decentralized wireless networks with dynamic conditions. Unlike traditional systems that require manual tuning of policies or rely on standard reinforcement learning, which struggles to generalize across multiple action spaces. The method 200 is used to handle multiple action spaces that reduce the need for retraining and memory overhead. Additionally, the method 200 is used to minimize the interference between the network devices by enabling the communication and cooperative policy learning with enhanced overall network performance without excessive communication overhead. As a result, the method 200 is used to perform various wireless tasks, such as bandwidth or power allocation, and is adaptable to dynamic network environments, thereby enhancing the overall performance of the wireless network.
[0057] The steps 202 to 210 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claim herein.
[0058] There is further provided a computer program product comprising program instructions for performing the method 200 when executed by one or more processors in the network device 102. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Elard Disk Drive (EIDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.
[0059] FIG. 3 is a diagram that illustrates an architecture and communication of the network device with the neighbouring network devices, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIG. 1. With reference to FIG. 3, there is shown a diagram 300 of the architecture and communication of the network device 102 with the neighbouring network devices 112, such as the first neighbouring network device 112A. The network device 102 is operated in a first local environment 302A and includes a first controller 304A, a first experience buffer 308A, a first receiver or transmitter 306A, a first action space module 310A, and a first memory 312A. Similarly, the first neighbouring network device 112 A operates in a second local environment 302B and includes a second controller 304B, a second experience buffer 308B, a second receiver or transmitter 306B, a second action space module 310B, and a second memory 312B.
[0060] In an implementation, the first experience buffer 308A is configured to store the transitions generated by the specific-parameter DNN policies {0 / 1} and the second experience buffer 308B is configured to store the transitions generated by the specificparameter DNN policies {G1* } while the first memory 312A is configured to store the specific-parameter DNN policies {0 / 1} (i.e., during the training phase) and the single multiple-parameter DNN policy 0,, and the second memory 312B is configured to store the specific-parameter DNN policies
[0061]
[0062] (i e., during the training phase) and the single multiple-parameter DNN policy Gj. Moreover, the first action space module 310A is configured to map the local radio resource parameter / z;and the received neighboring ones
[0063]
[0064] to a common radio resource parameter / z and generate the action space c / l^ / z) based on [1, and the second action space module 310B is configured to map the local radio resource parameter n, and the received neighbouring ones fi_j to a common radio resource parameter [1 and generate the action space cZZ ( / z) based on / z. Additionally, the first controller 304A is configured to select local radio resource parameter / z;, send / z;to neighbouring BS and receive theirs (jU_;), take action af from c / Z,-(p). based on the local state s' , via the specific-parameter DNN policy 0^, exchange local rewards with neighboring BSs then compute the group reward, and train the specific-parameter DNN policies {0^} and the single multipleparameter DNN policy 0,. The second controller 304B is configured to select local radio resource parameter / Zy, send py to neighbouring BS and receive theirs (ji-j), take action a' from cZly(p), based on the local state s' , via the specific-parameter DNN policy Gj , exchange local rewards with neighboring BSs then compute the group reward, and train the specific-parameter DNN policies {9^} and the single multiple-parameter DNN policy Gj. In the initialization phase, a handshake signalling process establishes the start of learning, such as at operation 314. Furthermore, the network device 102 and the first neighbouring network device 112 A agree on a common task, such as bandwidth allocation for specific services, and its related common function that maps each combination of local radio resource parameters to a common one and exchange their sets of local radio resource parameters and define a common set as the intersection of these local sets. During the learning process, at each step (time slot), the first controller 304A is configured to select a local radio resource parameter and communicate any changes to the first neighbouring network device 112A, such as at operation 316. The action space module then maps all local radio resource parameters to a common one, which defines the action space. The first controller 304A is configured to make decisions based on the local state using its specific-parameter DNN policy to observe the new local state and reward, exchange rewards with the first neighbouring network device 112A such as at operation 318 and send transition information to the buffer. The learning phase for each network device 102 involves learning a single multiple-parameter DNN (multiple action space policy) that can perform as well as the specific-parameter DNNs (specific action space policies).
[0065] In the inference phase, the first controller 304A is configured to select a local radio resource parameter and communicate any changes to the first neighbouring network device 112A, such as at operation 316. The action space module (e.g., the first action space module 310A) is configured to map all local radio resource parameters to a common one, defining the action space and then making decisions based on the local state using its single multiple-parameter DNN policy. As a result, the network device 102 is configured to provide a single DNN for multiple dynamic action spaces, enabling efficient distributed learning and decision-making across the network devices of the wireless network with reduced resource utilization, interference, and bandwidth allocation.
[0066] FIG. 4A is a diagram that illustrates an exemplary scenario of the network device operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 4A is described in conjunction with elements from FIGs 1 to 3. With reference to FIG. 4A, there is shown a diagram 400A illustrating the operations for resource allocation within a wireless network 402 including the network device 102 and the first neighbouring network device 112A.
[0067] In an implementation scenario, the network device 102 is configured to allocate bandwidth to its associated users. Moreover, at the end of initialization, the network device 102 and the first neighbouring network device 112A agree on a common set of numerologies and a common mapping function. During the learning process, at each step (time slot), the controller (e.g., the first controller 304A of the network device 102) is configured to select a local numerology and send its index to neighbours if it is different from the previous one (e.g., at operation 404A). The same is done by the first neighboring device 112A (e.g., at operation 404B). Thereafter, the action space module maps all local numerologies to a common numerology that is provided to the generator in order to define the action space cZZ;(p) = {1,2, ... , N^i1, where / < is the number of blocks when operating with numerology p for the bandwidth and Ntis the number of associated users.. Therefore, each of the network devices is configured to take an action and observe the associated local reward and the new local state and exchange local rewards and each one computes the group reward for the transition. In another implementation scenario, the network device 102 is configured to select a transmission power level to each device of its associated users. Moreover, the power transmission is chosen from a set (i.e.,
[0068]
[0069] = {0, - pmax, - Pmax> ■■■ >Pmax})- where p,- > 2 is a discretization level and pmaxis the maximum power level. At the end of initialization, the network device 102 and the first neighboring device 112A agree on a "
[0070]
[0071] FIG. 4B is a diagram that illustrates another exemplary scenario of the network device operating in a wireless network, in accordance with an embodiment of the present disclosure. FIG. 4B is described in conjunction with elements from FIGs. 1 to 4A. With reference to FIG. 4B, there is shown a diagram 400B illustrating the operations for resource allocation within the wireless network 402 including the network device 102, the first neighbouring network device 112A, and the second neighbouring network device 112B.
[0072]
[0073] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A network device (102) configured to perform resource allocation tasks utilizing a single multiple-parameter Deep Neural Network (DNN), wherein a resource allocation task is associated with an action space (A) which is based on at least one parameter (p), wherein the network device (102) is configured to:receive an indication of a resource allocation task to be performed,determine a local parameter (pi),receive neighbouring parameters) from neighbouring network device(s) (112),determine a common parameter (pc) by applying a mapping function (f) to the determined local parameter (pi ) and the received neighbouring parameters), and thenperform the resource allocation task utilizing the single multiple-parameter DNN based on the common parameter (pc), wherein the single multiple-parameter DNN has been trained based on the resource allocation task and all possible values for the common parameter (pc).
2. The network device (102) according to claim 1, wherein the network device (102) is further configured to:determine a local reward (n),receive neighbouring rewards (r_;) from the neighbouring network devices,determine a group reward (R) based on the local reward (n) and the received neighbouring rewards (r_;)train one specific-parameter DNN based on the group reward (R) for each possible value of the common parameter (pc), wherein the specific-parameter DNN is specific to one possible value of the common parameter (pc), and train the single multiple-parameter DNN based on the specific-parameter DNNs, whereby the single multiple-parameter DNN will be trained for all possible values for the common parameter (pc) over time.
3. The network device (102) according to claim 2, wherein the Network device (102) is further configured to repeat: determining the local parameter (pi ),receiving neighbouring parameters) from neighbouring network devices (112),determining the common parameter (pc) by applying the mapping function (f),performing the resource allocation task utilizing the specific-parameter DNN, based on the common parameter (pc), receiving neighbouring rewards (r_;) from neighbouring network devices (112), anddetermining the group reward (R) based on the local reward (n ) and the received neighbouring rewards (r_;), train of the specific-parameter DNNs at each k time stepssend the local reward (n) to neighbouring network devices (112).
4. The Network device (102) according to claim 2 or 3, wherein the Network device (102) is further configured to: determine the group reward (R) as the sum of the local reward (n ) and the received neighbouring rewards (r_;).
5. The Network device (102) according to claim 2, 3 or 4, wherein the Network device (102) is further configured to determine the local reward (n) as a summation of rewards of associated users.
6. The Network device (102) according to any preceding claim, wherein the Network device (102) is further configured to receive the indication of the resource allocation task to be performed through an initialization message comprising the indication of the resource allocation task to be performed.
7. The Network device (102) according to claim 6, wherein the initialization message also comprises an indication of what parameters to be exchanged.
8. The Network device (102) according to claim 6 or 7, wherein the initialization message also comprises an indication of the mapping function (f).
9. The Network device (102) according to any preceding claim, wherein all possible values for the common parameter (pc) are given by a predefined set of common parameter values ({pc}).
10. The Network device (102) according to claim 9 and any of 6, 7 or 8, wherein the initialization message also comprises an indication of predefined set of common parameter values ({pc}).
11. The Network device (102) according to any preceding claim, wherein the Network device is further configured to send the local parameter (pi) to neighbouring network devices.
12. The Network device (102) according to any preceding claim, wherein the mapping function (f) is associated with the resource allocation task to be performed.
13. The Network device (102) according to any preceding claim, wherein when the resource allocation task is to allocate bandwidth, the corresponding action space (A) is based on a number of associated users (N) and the parameter (p), wherein the parameter (p) is a numerology index.
14. The Network device (102) according to any preceding claim, wherein when the resource allocation task is to allocate power, the corresponding action space (A) is based on a number of associated users (N) and the parameter (p), wherein the parameter (p) is a discretization level.
15. The network device (102) according to any preceding claim wherein the Network device (102) is a Base station.
16. A method (200) for a Network device (102) configured to perform resource allocation tasks utilizing a single multipleparameter DNN, wherein a resource allocation task is associated with an action space (A) which is based on at least one parameter (p), wherein the method (200) comprises:receiving an indication of a resource allocation task to be performed,determining a local parameter (pi),receiving neighbouring parameters) from neighbouring network device(s),determining a common parameter (pc) by applying a mapping function (f) to the determined local parameter (pi) and the received neighbouring parameters), and thenperforming the resource allocation task utilizing the single multiple-parameter DNN, DNN, based on the common parameter (pc), wherein the single multiple-parameter Deep Neural Network, DNN, has been trained based on the resource allocation task and all possible values for the common parameter (pc).
17. The method (200) according to claim 16, wherein the method (200) further comprises during training:determining a local reward (n),receiving neighbouring rewards (r_;) from neighbouring network device(s) (112),determining a group reward (R) based on the local reward (n) and the received neighbouring rewards (r_;) training one specific-parameter DNN based on the group reward (R) according to the common parameter (pc), wherein the specific-parameter DNN is specific to one possible value of the common parameter (pc), andtraining the single multiple-parameter DNN based on the specific-parameter DNNs, whereby the single multipleparameter DNN will be trained for all possible values for the common parameter (pc) over time.
18. The method (200) according to claim 17, wherein the method (200) further comprises repeating:determining the local parameter (pi ),receiving neighbouring parameters) from neighbouring network device(s),determining the common parameter (pc) by applying the mapping function (f),performing the resource allocation task utilizing the specific-parameter DNN, based on the common parameter (pc), receiving neighbouring rewards (r_;) from neighbouring network device(s) (112), anddetermining the group reward (R) based on the local reward (n ) and the received neighbouring rewards (r_;) at each k time steps train the specific-parameter DNNssend the local reward (n ) to neighbouring network device(s) (112).
19. A computer program product comprising program instructions for performing the method (200) according to claim 16, 17 or 18, when executed by one ormore processors in aNetwork device (102).15