Apparatus and method for multi-target radio resource allocation
By introducing preference modules and neural network models into wireless networks, and formulating global preference vectors and resource allocation strategies, the multi-objective optimization problem in existing technologies is solved, achieving efficient resource allocation under different service requirements and improving network performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2023-10-18
- Publication Date
- 2026-05-05
AI Technical Summary
Existing wireless network resource allocation strategies struggle to achieve a good balance among multiple performance metrics (such as rate, latency, reliability, energy efficiency, and network coverage), especially in scenarios that support different service requirements, such as enhanced mobile broadband and ultra-reliable low-latency communication. Existing methods often can only optimize a single objective, resulting in poor performance in multi-objective optimization.
By employing the preference module and policy module in network devices, and through global preference vectors and neural network models, resource allocation decisions are made, and parameters are updated in conjunction with the training module to achieve multi-objective optimization.
Under different business needs, it can provide optimal actions for multiple objectives, improve network performance, achieve flexibility and efficiency in resource allocation, and reduce the occurrence of suboptimal decisions.
Smart Images

Figure CN121986535A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to wireless communication networks, and more particularly to resource allocation within wireless communication networks. To allocate wireless resources to devices with different performance requirements, this disclosure proposes a network device and method for efficient multi-target wireless resource allocation for user devices (UEs) in a wireless network. Background Technology
[0002] In existing wireless networks, access points (APs) connect to a massive number of devices. One of the most important tasks of an AP is to allocate wireless resources to these connected devices so they can communicate with the outside world. To this end, APs rely on device state information (such as queue length, channel gain, etc.) to optimize their performance. This is typically achieved by optimizing one of the many important performance indicators (KPIs) of the wireless network, such as speed, latency, reliability, energy efficiency, and network coverage. Therefore, the operator's goal is to propose resource allocation strategies, which essentially map state information to a function of actions (resource allocation decisions) to ensure good performance on the target KPIs.
[0003] However, this approach is becoming increasingly outdated because, as already envisioned, next-generation wireless networks will be able to support a variety of services, such as enhanced mobile broadband (eMBB) applications and ultra-reliable and low-latency communications (URLLC). Importantly, these different services may require good performance across different KPIs. For example, eMBB services require high-throughput connections, while URLLC services require robust data exchange with stringent reliability and latency requirements. To support these services, it is clear that a key requirement for APs is the ability to perform well across more than one, and possibly even a combination of, KPIs / objectives. A convenient way to model priorities in the objective space is to introduce preferences (linear weights). These preferences describe the relative importance of different KPIs. Therefore, a new and challenging problem arises: how to design a resource allocation strategy based on the so-called multi-objective optimization paradigm by constructing strategies that perform well across a variety of different services (even combinations of them). Summary of the Invention
[0004] In view of the above challenges, this disclosure aims to provide a scheme for efficient multi-target wireless resource allocation in wireless networks. Specifically, the objective is to design a single resource allocation strategy that optimizes for arbitrary relative importance among multiple targets. Another objective is to provide a flexible scheme that can provide a single strategy across the entire preference space, thereby achieving optimal performance for a variety of different preference vectors.
[0005] These and other objectives are achieved by the solutions provided in this disclosure in the independent claims. Advantageous implementations are further specified in the dependent claims.
[0006] A first aspect of this disclosure provides a network device for allocating resources to one or more user equipments, the network device including a preference module and a policy module. The preference module is configured to: determine a global preference vector, wherein the global preference vector describes the overall weight of each network performance metric in a set of network performance metrics for one or more user equipments, and provide the global preference vector to the policy module. The policy module is configured to: obtain the global preference vector from the preference module and formulate resource allocation decisions based on the global preference vector.
[0007] This disclosure enables network devices to make resource allocation decisions based on a global preference vector designed to guide AP (i.e., network device) performance toward achieving good performance in desired KPIs.
[0008] In one implementation of the first aspect, the preference module is used to: obtain one or more preference vectors from one or more user equipments, wherein each preference vector describes the weight of each network performance indicator in the network performance indicator set of one of the one or more user equipments; generate a global preference vector based on the one or more preference vectors; or obtain a global preference vector from a network operator.
[0009] Optionally, individual device preferences can be received from each associated device. The network device then maps these device preferences to AP preferences (i.e., a global preference vector). In another implementation, the global preference vector can be obtained directly from the network operator, for example, by monitoring network conditions and daily / hourly traffic curves.
[0010] In one implementation of the first aspect, the strategy module is further configured to: obtain one or more local states of one or more user devices, and formulate resource allocation decisions based on the global preference vector and one or more local states of one or more user devices.
[0011] It is worth noting that at the beginning of each time slot, the policy module receives AP preferences from the preference module, and the policy module can also receive the AP status as input from its connected devices, such as channel gain and queue length. Based on this information, the policy module performs an action for that time slot, that is, makes a resource allocation decision.
[0012] In one implementation of the first aspect, the strategy module is further configured to: determine a first neural network model by inputting a global preference vector and one or more local states of one or more user devices into the first neural network model to obtain an output value set, and formulate resource allocation decisions based on the output value set.
[0013] The policy module can be parameterized as a neural network (e.g., a DNN) with one or more parameters, which can be the weights and biases of the neural network. Therefore, this disclosure proposes a trainable multi-objective policy function (e.g., a DNN) to provide resource allocation decisions.
[0014] In one implementation of the first aspect, the strategy module is further configured to: make resource allocation decisions by randomly selecting output values from the output value set according to the probability distribution of the output value set, or by selecting output values with preset values from the output value set.
[0015] Alternatively, the selection of resource allocation actions can be achieved, for example, randomly by sampling actions from a probability distribution; or deterministically by selecting the action with the maximum value.
[0016] In one implementation of the first aspect, the network device further includes a training module, which is used to: determine a second neural network model and determine whether to use the second neural network model to update the first neural network model and / or the second neural network model used by the strategy module.
[0017] This disclosure also proposes a training module for network devices to update the parameters of a policy module using the current policy (e.g., periodically) after a given number of time slots.
[0018] In one implementation of the first aspect, the preference module is further configured to: determine a global preference vector set, obtain a preference sample set by sampling the global preference vector set, and provide the preference sample set to the training module. The training module is further configured to: receive the preference sample set from the preference module.
[0019] During the training step, the preference module is responsible for providing the preference sample set to the training module, which is essential for improving the performance of the training algorithm.
[0020] In one implementation of the first aspect, the policy module is further configured to: provide one or more local states and one or more actions of one or more user devices to the training module, wherein each action is the result of a resource allocation decision, and the training module is further configured to: store one or more local states and at least one action in a buffer.
[0021] It is worth noting that at the end of each resource allocation step, the training module stores the local state and actions received from the policy module in the experience buffer.
[0022] In one implementation of the first aspect, the strategy module is further configured to: obtain one or more next local states of one or more user devices after making a resource allocation decision, and / or obtain at least one reward after making a resource allocation decision.
[0023] It is worth noting that the reward is a multi-objective reward accumulated by the wireless resource allocation strategy. This reward can be a vector, the size of which is equal to the size of the preference vector. For example, the size of the reward vector is the number of objectives.
[0024] In one implementation of the first aspect, the policy module is further configured to: provide one or more next local states and / or at least one reward of one or more user devices to the training module, and the training module is further configured to: store one or more next local states and / or at least one reward in a buffer.
[0025] It is worth noting that at the end of each resource allocation step, the training module stores the state transitions and actions received from the policy module, as well as the multi-objective rewards accumulated by the wireless resource allocation policy, in the experience buffer.
[0026] In one implementation of the first aspect, the training module is further configured to: generate a set of transformation tuples by sampling information stored in a buffer, train a first neural network model of the policy module using the set of transformation tuples and a set of preference samples, and / or train a second neural network model using the set of transformation tuples and the set of preference samples.
[0027] To train the multi-objective policy function, the training module samples a set of transformed tuples from the experience buffer and trains the multi-objective policy function based on the transformed tuple samples and preference vector samples from the preference module.
[0028] In one implementation of the first aspect, the training module is further configured to: update one or more parameters associated with the first neural network model and / or the second neural network using one or more loss functions.
[0029] It is possible that the network device could use actors and critics to train a loss function to update the parameters of the policy module, i.e., the weights.
[0030] In one implementation of the first aspect, the training module is further configured to: provide one or more parameters associated with the first neural network model to the policy module, and the policy module is further configured to: update the first neural network model based on one or more parameters.
[0031] The parameters of the policy module can be updated periodically with the help of the training module.
[0032] In one implementation of the first aspect, the first neural network model and the second neural network model form a reinforcement learning model.
[0033] In one implementation of the first aspect, the reinforcement learning model is based on the policy gradient method.
[0034] The network device runs a policy gradient method to update the policy model (DNN).
[0035] In one implementation of the first aspect, the preference module is used to: generate a global preference vector based on one or more preference vectors using at least one of the following solutions: negotiation solution, mean function, game theory solution, priority solution, or neural network.
[0036] Optionally, the preference module can multiply the preferences of each user device by a priority number (an integer greater than 0). Typically, the priority assigned to each user depends on the user's signed contracts. The global preference can then be the sum of the new preferences divided by the sum of the priorities. Alternatively, the preference module can use a function (e.g., a pre-trained neural network) to generate a global preference vector. For example, the preference module takes the user's preferences and their types as input (which can be generated from the user's contracts) and outputs a global preference vector.
[0037] In one implementation of the first aspect, the set of network performance metrics includes one or more of the following: rate, throughput, latency, reliability, energy efficiency, fairness, and network coverage.
[0038] The network performance metrics set may also include other network performance metrics.
[0039] A second aspect of this disclosure provides a user equipment for assisting in the allocation of resources to one or more user equipments, the user equipment being used to provide a preference vector associated with the user equipment to a network device, wherein the preference vector describes the weight of each network performance metric in a set of network performance metrics for one of the one or more user equipments.
[0040] This disclosure also proposes a user equipment for assisting resource allocation. Specifically, this user equipment is responsible for providing individual device preferences to network entities.
[0041] A third aspect of this disclosure provides a method for allocating resources to one or more user equipments, performed by a network device, wherein the method includes: determining a global preference vector, wherein the global preference vector describes the overall weight of each network performance metric in a set of network performance metrics for one or more user equipments, and making a resource allocation decision based on the global preference vector.
[0042] The method described in the third aspect can be implemented in a manner corresponding to the network device implementation described in the first aspect. The method and its implementation in the third aspect achieve the same advantages and effects as the network device and its implementation described in the first aspect.
[0043] A fourth aspect of this disclosure provides a method performed by a user equipment for assisting in the allocation of resources to one or more user equipments, wherein the method includes providing a preference vector associated with the user equipment to a network device, wherein the preference vector describes the weight of each network performance metric in a set of network performance metrics for one of the one or more user equipments.
[0044] The fifth aspect of this disclosure provides a computer program product including program code that, when implemented on a processor, performs the method described according to the third aspect and any implementation thereof, or the method described according to the fourth aspect and any implementation thereof.
[0045] It should be noted that all devices, elements, units, and apparatuses described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by the various entities described in this application, and the functions to be performed by the various entities described, are intended to indicate that each entity is suitable for or used to perform the corresponding steps and functions. Although the specific functions or steps performed by external entities are not reflected in the detailed description of the specific elements of the entities performing the specific steps or functions in the following description of specific embodiments, those skilled in the art will understand that these methods and functions can be implemented by corresponding software or hardware elements or any combination thereof. Attached Figure Description
[0046] The above aspects and implementations of this disclosure will be set forth in the following description of specific embodiments with reference to the accompanying drawings, wherein:
[0047] Figure 1 A network device according to an embodiment of this disclosure is shown;
[0048] Figure 2 A network device and its associated user equipment according to embodiments of this disclosure are shown;
[0049] Figure 3 A system block diagram is shown for multi-objective resource allocation during the training phase according to an embodiment of the present disclosure;
[0050] Figure 4 A system block diagram is shown for multi-objective resource allocation during the training phase according to an embodiment of the present disclosure;
[0051] Figure 5 A flowchart of the proposed system for multi-objective resource allocation is shown;
[0052] Figure 6 A user equipment according to an embodiment of this disclosure is shown;
[0053] Figure 7 Performance evaluations of different algorithms relative to two targets are shown according to embodiments of this disclosure;
[0054] Figure 8 A method according to an embodiment of this disclosure is shown;
[0055] Figure 9 A method according to an embodiment of this disclosure is shown. Detailed Implementation
[0056] Illustrative embodiments of network devices, user equipment, corresponding methods for resource allocation, and corresponding computer program products for resource allocation are described with reference to the accompanying drawings. While this description provides detailed examples of possible implementations, it should be noted that these details are exemplary only and in no way limit the scope of this application.
[0057] Furthermore, one embodiment / example may refer to multiple other embodiments / examples. For example, any descriptions (including but not limited to terms, elements, processes, explanations, and / or technical advantages) mentioned in one embodiment / example are applicable to multiple other embodiments / examples.
[0058] To address the issues discussed in the background section and inspired by the successful applications of Machine Learning (ML) in other fields, there has been a recent surge of interest in applying ML to wireless resource allocation. In the RL paradigm, an agent (AP) interacts with the environment (wireless network and its devices) and attempts to learn an optimal policy (resource allocation) through trial and error, aiming to optimize a single reward in the long run. A significant advantage of RL-based methods is that they do not require predefined models (learning models from data) and can adapt policies to dynamically changing environments. However, a major drawback of simple RL-based methods is their extremely slow convergence speed, which in practice means that the agent will make suboptimal decisions for a considerable period. The root cause of this problem is that the agent needs to learn a policy for every possible state, and in our setting, the number of possible states grows exponentially with the number of devices. This problem has been addressed by using Neural Networks (NNs), which can generalize in the state space and approximately learn the optimal policy; from a practical perspective, this means that neural networks can converge to a near-optimal policy relatively quickly. Nevertheless, handling multiple objectives in RL remains a challenging problem, even with neural networks, and most studies employ oversimplified approaches that result in suboptimal outcomes.
[0059] Resource allocation in wireless networks has been studied for a long time and remains a relevant problem due to the many unexplored aspects. For example, in specific scheduling problems, the two most well-known schedulers are the Proportional Fairness (PF) scheduler and the Max-Weight scheduler. PF is the standard scheduler in cellular systems, designed to fairly allocate service rates among devices and considered the gold standard for performance comparison. Max-Weight provides an efficient mechanism for achieving queue stability, where non-service traffic to connected devices waits in the queue for transmission. While these schemes are easy to implement, they optimize for only a single objective and perform poorly across the entire preference space.
[0060] In addition, there are also studies in the literature on the use of classic static offline optimization algorithms. The main idea is to build an offline model based on previously collected data and optimize resource allocation decisions based on this model. Some of these studies consider a single optimization objective, while others consider multiple objectives, assuming that the preference vector between the objectives is fixed and known throughout the optimization process.
[0061] To address the aforementioned challenges, this disclosure proposes a method that enables an AP to learn a single resource allocation strategy that optimizes for the relative importance of multiple objectives.
[0062] Figure 1 A network device 100 is shown for allocating resources to one or more user equipments 200. Specifically, the network device 100 may be located in or integrated into a network such as... Figure 2 In the AP or base station (BS) shown. It is worth noting that the network device 100 in this disclosure can also be referred to as an AP. Each user equipment 200 can be an associated or connected device of the AP, and user equipment can have different preferences, such as... Figure 2 As shown in the image.
[0063] Network device 100 may include processing circuitry (not shown) for performing, conducting, or initiating various operations of the network device 100 described herein. The processing circuitry may include hardware and software. Hardware may include analog circuitry, digital circuitry, or a combination of analog and digital circuitry. Digital circuitry may include components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or multi-purpose processors. Network device 100 may also include memory circuitry for storing one or more instructions that can be executed by a processor or processing circuitry (specifically, under software control). For example, the memory circuitry may include a non-transitory storage medium storing executable software code that, when executed by a processor or processing circuitry, causes the network device 100 to perform various operations. In one embodiment, the processing circuitry system includes one or more processors and non-transitory memory connected to the one or more processors. Non-transient memory can hold executable program code that, when executed by one or more processors, enables network device 100 to perform, conduct, or initiate the operations or methods described herein.
[0064] Network device 100 includes a preference module 110 and a policy module 120. The preference module 110 determines a global preference vector 111 and provides it to the policy module 120. Notably, the global preference vector 111 describes the overall weight of each network performance metric in the network performance metric set of one or more user equipments 200. The policy module 120 obtains the global preference vector 111 from the preference module 110 and formulates resource allocation decisions 121 based on the global preference vector 111.
[0065] This disclosure presents a network device 100 for efficient multi-target radio resource allocation in a wireless network. The network device 100 includes a preference module 110 and a policy module 120, and is located at an access point (AP) or a base station (BS). This network device enables the AP (or BS) to allocate radio resources using only a single policy, thereby providing optimal action for any relative importance among multiple targets.
[0066] At the beginning of each resource allocation step, the preference module 110 sets the global preference vector 111. Provided to strategy module 120, where It is AP to the target The preference, and satisfy .
[0067] It is worth noting that the global preference vector 111 can be determined internally by taking into account the preferences of its associated devices or preferences from the network operator.
[0068] According to one embodiment of this disclosure, the preference module 110 is further configured to obtain one or more preference vectors 201 from one or more user equipments 200. Specifically, each preference vector 201 describes the weight of each network performance metric in the network performance metric set of one of the one or more user equipments 200. The preference module 110 is further configured to generate a global preference vector 111 based on the one or more preference vectors 201.
[0069] This means that the preference module 110 receives from each associated device (That is, one or more user equipment 200) receive the AP's preference vector ,in It is equipment For the target The preference, and satisfy Then, the preference module maps these individual device preferences to global AP preferences. And send it to the policy module 120.
[0070] It's important to note that the global preference vector 111 cannot represent the individual preference vectors of each device, but it can provide an indication of the overall situation of individual preference vectors. For a simple example, given a pair of preference vectors... and , and ,or and The mean function can be used to obtain the global vector. .
[0071] Device preferences can be determined based on several factors, such as the applications currently in use, user profiles and priorities, and the device's capabilities. Executing a mapping function from device preferences to AP preferences should guide AP performance in a direction that yields good results within the desired KPIs.
[0072] According to one embodiment of this disclosure, the preference module 110 can be used to obtain a global preference vector 111 from a network operator, for example, by monitoring network conditions and daily / hourly traffic curves. The network operator can directly provide the global preference vector to the network device 100. In one possible scenario, the network operator can provide all preference vectors of the device to the network device 100.
[0073] According to one embodiment of this disclosure, the policy module 120 is further configured to acquire one or more local states of one or more user equipments 200; and
[0074] Resource allocation decisions are made based on a global preference vector 111 and one or more local states of one or more user devices 200 121.
[0075] In each time slot At the beginning, the policy module 120 receives AP preferences from the preference module 110. The policy module 120 can also receive the status of the AP from its connected devices. As inputs, for example, channel gain and queue length. Based on this information, policy module 120 performs actions for that time slot. This action, namely, the resource allocation decision 121, can be the allocation of any wireless resource, such as transmission power, device scheduling, precoding or beamforming when using multiple antennas, allocating resource blocks to devices, etc.
[0076] Strategy module 120 can be parameterized to have parameters Neural networks (e.g., DNNs), parameters It can be the weights and biases of a neural network.
[0077] According to one embodiment of this disclosure, the strategy module 120 is further configured to determine a first neural network model by inputting a global preference vector 111 and one or more local states of one or more user devices 200 into the first neural network model to obtain an output value set, and to formulate a resource allocation decision 121 based on the output value set.
[0078] In general, the DNN (i.e., the first neural network model) of policy module 120 can be represented as a function of probability distributions in the action space. This is represented as follows. Then, based on this DNN, the selection of resource allocation actions can be achieved, for example, randomly by sampling actions from a probability distribution; or deterministically by selecting the action with the maximum value.
[0079] According to one embodiment of this disclosure, the strategy module 120 can also be used to make a resource allocation decision 121 by randomly selecting an output value from the output value set according to the probability distribution of the output value set. Alternatively, the strategy module 120 can also be used to make a resource allocation decision 121 by selecting an output value with a preset value from the output value set.
[0080] In practice, this wireless resource allocation can be implemented by the corresponding physical layer and MAC layer modules of the AP.
[0081] It is worth noting the parameters of strategy module 120. It can be updated periodically with the help of training module 130. In this case, training module 130 also includes parameters. The second DNN is used to approximate other important RL functions.
[0082] According to one embodiment of this disclosure, the network device 100 further includes, for example, Figure 3 The training module 130 shown is shown.
[0083] Specifically, the training module 130 is used to determine the second neural network model and whether to use the second neural network model to update the first neural network model and / or the second neural network model used by the strategy module 120.
[0084] According to one embodiment of this disclosure, the preference module 110 is further configured to determine a global preference vector set 111, obtain a preference sample set by sampling the global preference vector set 111, and provide the preference sample set to the training module 130. Accordingly, the training module 130 is further configured to receive the preference sample set from the preference module 110.
[0085] During the training step, the preference module 110 is responsible for providing the preference sample set to the training module, which is necessary to improve the performance of the training algorithm.
[0086] According to one embodiment of this disclosure, the policy module 120 is further configured to provide one or more local states and one or more actions of one or more user devices 200 to the training module 130, wherein each action is the result of resource allocation decision 121. The training module 130 is further configured to:
[0087] Store one or more local states and at least one action in a buffer.
[0088] According to one embodiment of this disclosure, the policy module 120 is further configured to obtain one or more next local states of one or more user devices 200 after making a resource allocation decision 121, and / or obtain at least one reward after making a resource allocation decision 121.
[0089] It's important to note that the reward is a vector, and its size is equal to the size of the preference vector. For example, the size of the reward vector is the number of objectives.
[0090] According to one embodiment of this disclosure, the policy module 120 is further configured to provide one or more next local states and / or at least one reward of one or more user devices 200 to the training module 130. The training module 130 is further configured to store one or more next local states and / or at least one reward in a buffer.
[0091] The training module 130 has two main functions. First, at the end of each resource allocation step, the training module 130 stores the state transitions and actions received from the policy module 120, as well as the accumulated multi-objective reward of the wireless resource allocation policy. This information is crucial for training the algorithm. Specifically, the training module stores tuples... Stored in the experience buffer, the tuple includes the state at the start of the time slot. Resource allocation decisions adopted (i.e., action), multi-objective reward and resource allocation decisions The next state generated .
[0092] Secondly, after a given number of time slots, the training module 130 periodically updates the parameters of the policy module 120 using the current policy: for this purpose, the training module samples from the experience buffer. Transformed tuples and receive from preference module 110 A sample of preference vectors. Based on this input, the training module runs the policy gradient method to update the policy model (DNN).
[0093] According to one embodiment of this disclosure, the training module 130 is further configured to: generate a set of transformation tuples using information stored in the sampling buffer. Furthermore, the training module 130 can also be configured to train a first neural network model of the policy module 120 using the set of transformation tuples and the preference sample set. Optionally, the training module 130 can also be configured to train a second neural network model using the set of transformation tuples and the preference sample set.
[0094] More precisely, the actor-critic method is used to estimate the gradient in the direction of updating the policy. The actor is the policy DNN that must perform the optimal action given the state and preference vectors. In other words, for a given AP preference... What needs to be optimized is a single strategy. The critic is another neural network with parameters... Used to evaluate actors' behavior. Critics take action, state, and preference vectors as input and output Q-scores. The Q value is in the state. and preferences Next action Then, the sum of expected discounts on future rewards, which has the same dimension as the rewards. The discount factor is defined as follows: .
[0095] According to one embodiment of this disclosure, the training module 130 is further configured to use one or more loss functions to update one or more parameters associated with the first neural network model and / or the second neural network.
[0096] In this context, AP uses the actor and critic training loss functions defined below.
[0097] The loss of actor DNN:
[0098] .
[0099] The loss of critic DNN:
[0100] .
[0101] Finally, based on these losses, training module 130 updates the weights of policy module 120. And update the weights of training module 130 using gradient steps. .
[0102] According to one embodiment of this disclosure, the training module 130 is further configured to provide one or more parameters associated with the first neural network model to the policy module 120. The policy module 120 is further configured to update the first neural network model based on one or more parameters.
[0103] According to one embodiment of this disclosure, a first neural network model and a second neural network model form a reinforcement learning model.
[0104] According to one embodiment of this disclosure, the reinforcement learning model is based on the policy gradient method.
[0105] According to one embodiment of this disclosure, the preference module 110 is used to generate a global preference vector 111 based on one or more preference vectors 201 using at least one of the following solutions: negotiation solution, mean function, game theory solution, priority solution, or neural network.
[0106] Possibly, the preference module could be used to multiply the preferences of each user device by a priority number (an integer greater than 0). Typically, the priority assigned to each user depends on the contracts the user has entered into. The global preference (i.e., the global preference vector 111) would then be the sum of the new preferences divided by the sum of the priorities. Alternatively, the preference module could be a pre-trained function or neural network that takes user preferences and their types (generated from the user's contracts) as input. As output, the preference module generates the global preference.
[0107] According to one embodiment of this disclosure, the network performance metric set includes one or more of the following: rate, throughput, latency, reliability, energy efficiency, fairness, and network coverage. The network performance metric set may also include other network performance metrics.
[0108] Figure 3 and Figure 4 The proposed system for multi-objective resource allocation during the training phase is illustrated. Specifically, Figure 3 This illustrates one stage of the training phase as the experience buffer is filled. Figure 4 This illustrates one phase of the training phase when running the RL algorithm to update the policy weights.
[0109] In each resource allocation step (slot), the preference module 110 provides an AP preference vector to the policy module 120. The policy module 120 then makes a radio resource allocation decision based on this preference vector and the monitored state of its associated devices. The policy module also observes the new state as a result of its actions and sends transition information (state, action, and new state) to the training module. Upon receiving the information, the training module stores the transition along with the associated multi-objective reward (which is revealed after the action is performed) in its buffer.
[0110] After a given number of steps and having collected sufficient transformed data, the training module 130 runs the policy gradient RL algorithm by sampling the transformed data from its buffer and the preference vector from the preference module 110. At the end of this step, the parameters of the policy module 120 are updated.
[0111] Figure 5 A flowchart of a proposed scheme for multi-objective resource allocation according to an embodiment of the present disclosure is shown.
[0112] Figure 6A user equipment 200 is shown for assisting in the allocation of resources to one or more user equipments 200. Specifically, the user equipment 200 is an associated or connected device of network entity 100. It is worth noting that network device 100 can be, for example, Figure 1 or Figure 2 Network 100 is shown in the figure.
[0113] User equipment 200 may include processing circuitry (not shown) for performing, conducting, or initiating various operations of user equipment 200 as described herein. The processing circuitry may include hardware and software. Hardware may include analog circuitry, digital circuitry, or a combination of analog and digital circuitry. Digital circuitry may include components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or multi-purpose processors. User equipment 200 may also include memory circuitry for storing one or more instructions that can be executed by a processor or processing circuitry (specifically, under the control of software). For example, the memory circuitry may include a non-transitory storage medium storing executable software code that, when executed by a processor or processing circuitry, causes user equipment 200 to perform various operations. In one embodiment, the processing circuitry system includes one or more processors and non-transitory memory connected to one or more processors. The non-transitory memory may carry executable program code that, when executed by one or more processors, causes user equipment 200 to perform, conduct, or initiate the operations or methods described herein.
[0114] Specifically, user equipment 200 provides a preference vector 201 associated with user equipment 200 to network device 100. The preference vector 201 describes the weight of each network performance metric in the network performance metric set of one or more user equipments 200.
[0115] In a specific embodiment, the implementation of the proposed system in a cellular network environment is discussed, wherein the AP (i.e., such as Figure 1 The network device 100 shown must connect to its connected devices (i.e., such as...) Figure 6 One or more user equipment (200) shown in the diagram perform uplink scheduling. Therefore, the AP must decide which of its associated devices to schedule in each time slot. For the service model, it is assumed that in each time slot... There is traffic (by (Data packets in bits arrive at the device according to a random process) .
[0116] In detail, for the AP and its K connected or associated devices, the state ,in It includes the following three components:
[0117] The estimated rate;
[0118] The average received throughput over a period of time;
[0119] The head-of-line delay is estimated based on the previous transmission delay.
[0120] Further assume that the expected KPIs (two targets) include throughput and latency.
[0121] For policy decision-making, the channel condition for the AP during the time slot duration T is as follows: The time step t schedules device k. The number of bits transmitted is... Where G is the antenna gain, W is the transmission bandwidth, P is the transmission power, and n is the noise power. (for ).
[0122] The queue evolution at device k is given by the following equation: ,in It is the queue size at position t. It is the size of the data packet in bits. It represents the number of data packets arriving at position t in queue k.
[0123] Mapping preferences from user equipment to access points (APs) is a challenging task because objectives may conflict and user equipment may have different preferences. To address this issue, this embodiment employs a Nash bargaining solution to obtain global preferences.
[0124] Each goal (That is, KPI) modeled as having values There are a total of L participants (in this setting). Competitive sharing of resources. Ownership belongs to the participants. The amount of resources will be related to the target Associated weights.
[0125] Each with negotiating ability Participants (target) You will receive a share ,in These participants should agree on some "reasonable" axioms, such as the symmetry axiom: if two participants have the same value, then they should receive the same amount of resources.
[0126] Then, the following optimization problem can be obtained:
[0127] ,
[0128] The constraints are ,and .
[0129] The Nash negotiation solution is sufficient for this situation because it represents a unique solution satisfying four axioms: symmetry, independence of irrelevant alternatives, Pareto optimality (no one can gain a better payoff without harming others), and invariance to affine transformations. The Nash negotiation solution represents the amount of resources given to each objective, i.e. .
[0130] To conduct a performance evaluation, this paper discusses The system simulation is performed on one device, where the queue size for each device is 500 data packets. (bits). Transmission duration (TTI) is 1 millisecond, and scheduling duration is 5000 TTI.
[0131] To evaluate the performance of the proposed scheme, it will be compared with standard scheduling algorithms (proportional fairness, maximum weight, exponential proportional fairness delay, exponential proportional fairness buffer, and random algorithm).
[0132] Figure 7 The performance of each algorithm relative to two objectives is shown. It's important to note that latency is intentionally described with negative values, so higher values for each axis indicate a better solution. This figure illustrates how a single neural network, trained only once across the entire preference space, can capture the frontier of different behaviors based on input preferences. Each point from the system represents the proposed solution for a given preference vector. The performance, of which It is a weight associated with throughput. These are weights associated with latency. From left to right, From 0 to 1, From 1 to 0. Compared to a baseline designed to optimize a specific KPI, the proposed scheme achieves almost the same performance for that KPI, but significantly improves performance for other KPIs. For example, it can be seen that by selecting This scheme achieves nearly the same throughput as the maximum weighted scheme, but with only one-third of the latency. The proposed scheme can achieve near-optimal performance for each objective, but can also provide a wide range of optimal trade-offs for any possible preference.
[0133] Figure 8 A method 800 for allocating resources to one or more user equipments 200 according to an embodiment of this disclosure is illustrated. Method 800 is performed by a network device, particularly by... Figure 1 , Figure 2 or Figure 6 The network device 100 shown performs this. Method 800 includes a step 801 of determining a global preference vector, wherein the global preference vector describes the overall weight of each network performance metric in a set of network performance metrics for one or more user devices. Method 800 also includes a step 802 of making resource allocation decisions based on the global preference vector. Possibly, each user device 200 is... Figure 2 or Figure 6 The user equipment shown.
[0134] Figure 9 A method 900 for assisting in the allocation of resources to one or more user equipments 200, according to an embodiment of the present disclosure, is illustrated. The method 900 is performed by one of the one or more user equipments, in particular... Figure 2 or Figure 6 The network device 200 is shown. Method 900 includes step 901, providing a preference vector 201 associated with user equipment 200 to network device 100, wherein the preference vector 201 describes the weight of each network performance metric in a set of network performance metrics for one of the user equipments 200. Possibly, network device 100 is... Figure 1 , Figure 2 or Figure 6 The network device 100 shown is shown.
[0135] In summary, this disclosure enables an AP (or BS) to allocate radio resources using only a single policy, thereby providing optimal actions for arbitrary relative importance among multiple objectives. Specifically, this disclosure provides a trainable multi-objective policy function (e.g., a DNN) to provide resource allocation decisions. This disclosure proposes dedicated message interaction between the AP and devices to understand current device preferences. Furthermore, this disclosure enables mapping device preferences to a collective preference vector of the AP via a novel preference module. This is a flexible scheme capable of providing a single policy across the entire preference space, thus achieving optimal performance for a variety of different preference vectors. This scheme can be easily transferred to related tasks with different preferences. This disclosure also proposes using a policy gradient algorithm to update the policy model using samples from an experience buffer and a preference module. Therefore, this scheme alleviates the dependence on scalar reward design. Unlike recent multi-objective RL-based methods, this is a scalable approach that uses a single DNN (instead of a different DNN for each preference vector), thereby reducing the hardware requirements of the AP.
[0136] This disclosure has been described in conjunction with various embodiments as examples and implementations. However, based on a study of the drawings, this disclosure, and the appended claims, those skilled in the art will be able to understand and implement other variations in practicing the claimed embodiments of this disclosure. In the claims and the description, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" does not exclude a plurality. A single element or other unit may fulfill the function of several entities or items described in the claims. The enumeration of certain measures in dissimilar dependent claims does not imply that combinations of these measures cannot be used in advantageous implementations.
[0137] Furthermore, any method according to embodiments of this disclosure can be implemented in a computer program having code components, which, when run by a processing component, causes the processing component to perform method steps. The computer program is included in a computer-readable medium of the computer program product. The computer-readable medium can substantially include any memory, such as read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), flash memory, electrically erasable PROM (EEPROM), or a hard disk drive.
[0138] Furthermore, those skilled in the art will recognize that the proposed network device 100 or user equipment 200 and corresponding computer program products include the necessary communication capabilities, in the form of, for example, functions, devices, units, elements, etc., for executing the scheme. Other examples of such devices, units, elements, and functions include: processors, memories, buffers, control logic, encoders, decoders, rate matchers, de-rate matchers, mapping units, multipliers, decision units, selection units, switches, interleavers, deinterleavers, modulators, demodulators, inputs, outputs, antennas, amplifiers, receiving units, transmitting units, DSPs, trellis-coded modulation (TCM) encoders, TCM decoders, power supply units, power feeders, communication interfaces, communication protocols, etc., which are suitably arranged together to execute the scheme.
[0139] Specifically, one or more processors in network device 100 or user equipment 200 may include, for example, a Central Processing Unit (CPU), a processing unit, processing circuitry, a processor, an Application Specific Integrated Circuit (ASIC), a microprocessor, or one or more instances of other processing logic capable of interpreting and executing instructions. Therefore, the expression "processor" can refer to a processing circuitry that includes multiple processing circuits, such as any, some, or all of the processing circuits described above. The processing circuitry can also perform data processing functions for inputting, outputting, and processing data, including data buffering and device control functions such as call processing control, user interface control, etc.
Claims
1. A network device (100) for allocating resources to one or more user equipments (200), characterized in that, The network device (100) includes a preference module (110) and a policy module (120). The preference module (110) is used for: A global preference vector (111) is determined, wherein the global preference vector (111) describes the overall weight of each network performance metric in the network performance metric set of the one or more user equipments (200). The global preference vector (111) is provided to the policy module (120). The strategy module (120) is used for: The global preference vector (111) is obtained from the preference module (110). Resource allocation decisions (121) are made based on the global preference vector (111).
2. The network device (100) according to claim 1, characterized in that, The preference module (110) is used for: One or more preference vectors (201) are obtained from the one or more user equipments (200), wherein each preference vector (201) describes the weight of each network performance metric in the network performance metric set of one of the one or more user equipments (200), and the global preference vector (111) is generated based on the one or more preference vectors (201), or Obtain the global preference vector (111) from the network operator.
3. The network device (100) according to claim 1 or 2, characterized in that, The strategy module (120) is also used for: Obtain one or more local states of the one or more user equipments (200). The resource allocation decision (121) is made based on the global preference vector (111) and the one or more local states of the one or more user devices (200).
4. The network device (100) according to claim 3, characterized in that, The strategy module (120) is also used for: Determine the first neural network model. By inputting the global preference vector (111) and the one or more local states of the one or more user devices (200) into the first neural network model, an output value set is obtained. The resource allocation decision is made based on the output value set (121).
5. The network device (100) according to claim 4, characterized in that, The strategy module (120) is also used for: The resource allocation decision is made by randomly selecting output values from the output value set according to the probability distribution of the output value set (121), or The resource allocation decision is made by selecting an output value with a preset value from the set of output values (121).
6. The network device (100) according to any one of claims 1 to 5, characterized in that, The network device (100) further includes a training module (130), which is used for: Determine the second neural network model. Determine whether to use the second neural network model to update the first neural network model and / or the second neural network model used by the policy module (120).
7. The network device (100) according to claim 6, characterized in that, The preference module (110) is also used for: Determine the set of global preference vectors (111). The preference sample set (112) is obtained by sampling the global preference vector (111) set. The preference sample set (112) is provided to the training module (130). The training module (130) is also used for: The preference sample set (112) is received from the preference module (110).
8. The network device (100) according to claim 6 or 7 and claim 4 or 5, characterized in that, The strategy module (120) is also used for: The training module (130) provides one or more local states and one or more actions of the one or more user devices (200), wherein each action is the result of the resource allocation decision (121). The training module (130) is also used for: Store the one or more local states and at least one action in a buffer.
9. The network device (100) according to any one of claims 2 to 8, characterized in that, The strategy module (120) is also used for: After making the resource allocation decision (121), obtain one or more next local states of the one or more user equipments (200), and / or After making the resource allocation decision (121), at least one reward is obtained.
10. The network device (100) according to any one of claims 9 and 6 to 8, characterized in that, The strategy module (120) is also used for: The training module (130) provides the training module (130) with the one or more next local states of the one or more user devices (200) and / or the at least one reward. The training module (130) is also used for: The one or more next local states and / or at least one reward are stored in the buffer.
11. The network device (100) according to claim 10, characterized in that, The training module (130) is also used for: A set of transformed tuples is generated by sampling the information stored in the buffer. The first neural network model of the policy module (120) is trained using the set of transformation tuples and the set of preference samples, and / or The second neural network model is trained using the set of transformed tuples and the set of preference samples.
12. The network device (100) according to claim 11, characterized in that, The training module (130) is also used for: Use one or more loss functions to update one or more parameters associated with the first neural network model and / or the second neural network.
13. The network device (100) according to claim 12, characterized in that, The training module (130) is also used for: The one or more parameters associated with the first neural network model are provided to the policy module (120). The strategy module (120) is also used for: The first neural network model is updated based on one or more of the parameters.
14. The network device (100) according to any one of claims 4 to 13, characterized in that, The first neural network model and the second neural network model form a reinforcement learning model.
15. The network device (100) according to any one of claims 7 to 14, characterized in that, The reinforcement learning model is based on the policy gradient method.
16. The network device (100) according to any one of claims 2 to 15, characterized in that, The preference module (110) is used for: Based on the one or more preference vectors (201), the global preference vector (111) is generated using at least one of the following solutions: negotiation solution, mean function, game theory solution, priority solution, or neural network.
17. The network device (100) according to any one of claims 1 to 16, characterized in that, The network performance metrics set includes one or more of the following: rate, throughput, latency, reliability, energy efficiency, fairness, and network coverage.
18. A user equipment (200) for assisting in allocating resources to one or more user equipments (200), characterized in that, The user equipment is used for: A preference vector (201) associated with the user equipment (200) is provided to the network device (100), wherein the preference vector (201) describes the weight of each network performance metric in the network performance metric set of one of the user equipments (200).
19. A method for allocating resources to one or more user equipments (200), characterized in that, The method includes: A global preference vector (111) is determined, wherein the global preference vector (111) describes the overall weight of each network performance metric in the network performance metric set of the one or more user equipments (200). Resource allocation decisions (121) are made based on the global preference vector (111).
20. A method for assisting in allocating resources to one or more user equipments (200), characterized in that, The method includes: A preference vector (201) associated with a user equipment (200) is provided to a network device (100), wherein the preference vector (201) describes the weight of each network performance metric in the network performance metric set of one of the user equipments (200).
21. A computer program product, characterized in that, Includes program code that, when implemented on a processor, executes the method according to claim 19 or 20.