Power allocation method for communication system based on multi-agent cooperation
Through the deep reinforcement learning method of multi-agent cooperation, the power allocation in high- and low-frequency collaborative networking is optimized, which solves the problems of interference suppression and energy efficiency in ultra-densely deployed 6G networks and achieves a balance between system capacity and energy efficiency.
Patent Information
- Application Number
- CN202410738605.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-06-07
AI Technical Summary
Existing technologies have difficulty effectively suppressing interference and achieving joint optimization of system capacity and energy efficiency in ultra-densely deployed 6G networks. This is especially true in the power allocation process of high-frequency base stations. Existing algorithms require global information and have high computational complexity, making them unable to adapt to complex environments.
A deep reinforcement learning method based on multi-agent cooperation is adopted. By constructing a high- and low-frequency collaborative networking scenario, each high-frequency base station acts as an agent, and a two-layer deep reinforcement learning algorithm is used for sub-band selection and power allocation. Combined with a multi-head attention mechanism, the system capacity and energy efficiency are optimized, taking into account the interference and energy consumption between agents.
It achieves efficient power allocation in complex environments, optimizes system capacity and energy efficiency, reduces interference, improves spectrum utilization efficiency, and achieves a balance in energy consumption control.
Smart Images

Figure CN118678452B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a communication system power allocation method based on multi-agent cooperation. Background Art
[0002] Targeting commercial use in 2030, 6G's application scenarios will expand from the three typical 5G scenarios to eight categories, encompassing ten potential key technology areas. These will encompass richer and more intelligent business needs, such as immersive reality, sensory connectivity, and digital twins. 4G networks, previously based on a single sub-6GHz frequency band, no longer meet communication needs. Therefore, 5G communication standards are evolving toward millimeter wave (mmWave) frequency bands, employing a collaborative networking approach for high and low frequency band deployment. This model will be more widely and deeply applied and developed as 6G achieves full coverage in all scenarios: retaining low-frequency bands with high reliability and better transmission characteristics to ensure stable wide-area network coverage, while high-frequency bands will be used to enhance system capacity in hotspots. The sub-6GHz low-frequency band (450-6000MHz) has wide coverage but narrow bandwidth, while the high-frequency mmWave band (24.25-52.6GHz) has large bandwidth but limited coverage. Therefore, the high data traffic demands in hotspots require ultra-dense deployment of high-frequency base stations. However, ultra-dense deployment of high-frequency base stations exacerbates interference and increases system energy consumption, resulting in reduced system capacity and energy efficiency. Therefore, for ultra-dense deployment of high-frequency base stations, effectively controlling downlink transmit power is crucial for increasing system capacity, significantly improving spectrum utilization efficiency. At the same time, attention should also be paid to the dramatic increase in energy consumption caused by dense base station deployment. Energy consumption control factors must be fully considered during system planning, seeking the optimal balance between base station density, transmit power configuration, system capacity, and energy efficiency.
[0003] Therefore, a reasonable power allocation strategy is crucial for 6G networks. Providing stable connections through low-frequency base stations and precisely controlling their transmit power can effectively mitigate interference caused by ultra-dense deployments and improve system capacity. Furthermore, reasonable power allocation can control base station energy consumption and achieve efficient energy utilization.
[0004] Among traditional power allocation algorithms, existing techniques have proposed improving system capacity through power allocation. These methods use an ideal fractional programming (Ideal FP) approach to transform a non-convex problem into a convex one, simplifying the power allocation decision-making process. However, this Ideal FP approach assumes instantaneous access to global channel state information (CSI), which is challenging and inadequate for the complex environments of densely deployed base stations. Others have proposed optimizing power allocation using a particle swarm optimization algorithm, establishing a relationship between interference and power allocation to improve system capacity while also considering energy consumption. However, this approach suffers from high computational complexity and is difficult to apply to dense network environments. Another approach involves jointly optimizing channel and power allocation, using a hypergraph model to accurately characterize the complex cumulative interference in the network's power allocation process. These traditional algorithms rely on global information for power allocation, making them inefficient in dense networks and difficult to constrain in terms of capacity and energy consumption. In contrast, the rapid development of deep reinforcement learning (DRL) technology in recent years has enabled it to explore optimal decisions based on partial information rather than global information and combined with empirical evidence, leading to its widespread application in densely deployed communication networks. In this context, each base station is not only the executor of power allocation decisions but can also be considered an instance of an intelligent agent, which must independently respond based on local information. A base station (agent) must be able to adapt in real time to the changing needs of the specific user group it serves, while also taking into account the activities of other surrounding base stations (multi-agents) to optimize its power output. Currently, there are methods that use the relationship between system capacity and signal-to-interference-and-noise ratio (SINR) to identify the impact of power allocation on system capacity. This approach accurately achieves capacity improvement requirements and is applicable to various complex interference scenarios. It is currently widely used to address capacity improvement in ultra-dense networks through DRL. Some existing technologies fail to incorporate constraints on base station downlink transmit power during the power allocation process, rendering power allocation meaningless. Other existing technologies use the emerging Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach to address power allocation. These multiple agents learn decisions in a distributed manner from past experience, but do not consider interference mitigation between agents. Some existing technologies consider the balance between capacity and energy consumption during power allocation but fail to account for the impact of interference externalities between base stations. Some studies have also improved system capacity through power allocation, achieving a trade-off between capacity improvement and energy efficiency, but have not considered decision-making cooperation among agents. Some studies on power allocation have considered multi-objective optimization of system capacity and energy efficiency in ultra-dense networks, but the downlink power allocation process has not considered the interference connections between multiple agents.Another approach uses a deep Q-network (DQN) to jointly handle subband selection and transmit power control to improve system capacity, taking into account the impact of interference. However, its single deep reinforcement learning space is the Cartesian product of available subbands and transmit power. The number of state-action pairs visited for convergence during training does not scale well with the number of subbands, resulting in inefficiency in power allocation strategies to improve system capacity in dense deployment scenarios. Another approach controls the transmit power of intelligent agents to maximize downlink transmission rates, but the allocation process fails to consider energy consumption. Another approach uses multi-agent power allocation, but the allocation process fails to account for the significant energy consumption in densely deployed communication systems. Some research uses a two-layer deep reinforcement learning algorithm, DQN and Deep Deterministic Policy Gradient (DDPG), to jointly optimize subband selection and power allocation to improve system capacity. This approach effectively mitigates interference in dense deployment scenarios, but the multi-agent power allocation process is complex and lacks a balance between capacity and energy consumption.
[0005] In summary, in the power allocation process of ultra-densely deployed network scenarios, existing technologies rarely consider achieving joint optimization of system capacity and energy efficiency while effectively suppressing interference. Summary of the Invention
[0006] The present invention provides a communication system power allocation method based on multi-agent cooperation, which is mainly applied to downlink power control between high-frequency base stations (HF-BS). Each HF-BS optimizes power allocation through multi-agent cooperative game to maximize system capacity and energy efficiency, while controlling inter-cell interference.
[0007] The technical solution adopted by the present invention to solve its technical problems is:
[0008] A communication system power allocation method based on multi-agent cooperation, comprising the following steps:
[0009] Build a high- and low-frequency collaborative networking scenario that integrates low-frequency base stations (LF-BS) below 6 GHz and millimeter-wave high-frequency base stations (HF-BS), with each HF-BS acting as an intelligent agent instance.
[0010] Construct the state space, action space, and reward function of the intelligent agent. Integrate the concept of cooperative game into the construction of the reward function. Define the reward function as a constraint relationship model between system capacity, energy efficiency, and interference penalty. The reward function sets the optimization goal.
[0011] After setting the optimization target for the reward function, a two-layer deep reinforcement learning algorithm architecture is used for subband selection and power allocation. The upper layer uses the deep Q network (DQN) to select subbands, and the lower layer uses the deep deterministic policy gradient (DDPG) for power allocation. A multi-head attention mechanism is added between the first and second fully connected layers of the DDPG critic network in the lower layer to selectively capture the state and action information of other intelligent agents and weightedly fuse this information with its own features.
[0012] Furthermore, the state space is composed of the number of currently connected users Channel state h t , Subband selection for user connection and the current transmit power At time t, the state space S is expressed as:
[0013]
[0014] Furthermore, the action space A is expressed as:
[0015]
[0016] where k n is the subband to be power controlled by the nth HF-BS at the current time t, Represents the power adjustment amount of the nth HF-BS at the current time t, and its value range is
[0017] Furthermore, the reward function consists of system capacity, energy efficiency, and interference penalty. The optimization goal is to maximize system capacity and energy efficiency. A multi-objective optimization method is used to integrate multiple objectives into a comprehensive goal for solution. The reward function is defined as:
[0018]
[0019] in, represents the reward function, α represents the balance parameter between energy efficiency and system capacity, and its value range is 0≤α≤1; EE n represents the energy efficiency of the nth HF-BS; B tot is the total bandwidth; E tot is the total power consumption of the HF-BS when transmitting at full power; is the interference penalty of the nth HF-BS; is the capacity on subband k; B / K represents K carrier subbands with a total bandwidth of B.
[0020] Furthermore, the optimization objective is expressed as:
[0021]
[0022] Constraints:
[0023] (1)
[0024] (2)
[0025] Among them, p max With p min They represent the maximum transmission power and minimum transmission power of HF-BS, E total is the total system power consumption, E max Is the maximum power consumption limit of the system; is the transmission power of subband k; n∈N={1,2,…N}, N is the number of HF-BSs within a certain range; is the total system capacity; EE total Represents the total energy efficiency of the system; E n is the power consumption of the nth HF-BS; C t Represents the total system capacity at time t.
[0026] Furthermore, the total system capacity for:
[0027]
[0028] Where: K is the total number of carrier subbands, M is the total number of users, is the capacity of the HF-BS and user m on subband k, expressed as:
[0029]
[0030] is the signal-to-noise ratio;
[0031] The power consumption of HF-BS includes zero-load normally open power p A and transmission power And the power p in low power sleep mode S :
[0032]
[0033] Q n is the dynamic switching factor of HF-BS;
[0034] Energy Efficiency (EE) n Defined as:
[0035]
[0036] The beneficial effects of the present invention include:
[0037] The present invention proposes an adaptive power allocation method based on collaborative cooperation among multiple agents, which can optimize power allocation in dense networks. First, the concept of cooperative game is incorporated into the design of the reward function for agent learning, and a constraint relationship model between system capacity, energy efficiency and interference penalty is constructed. The externality impact of the agent is taken into account when allocating power to the agent, guiding the agent to reduce interference with other agents in the decision-making process, and cooperate to improve system capacity and energy efficiency. Secondly, after setting the optimization target for the reward function, the DQN and DDPG two-layer architecture is used for sub-band selection and power allocation, and a multi-head attention mechanism is introduced in the Critic network to selectively capture the status and action information of other agents, and weightedly fuse this external information with its own features. This enables the Critic network to more accurately evaluate the potential benefits of the action, providing the Actor network with more globally effective strategy support, thereby not only optimizing individual decisions, but also fully considering the improvement of overall network performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a structural diagram of a high- and low-frequency collaborative networking scenario;
[0039] Figure 2 It is a two-layer deep reinforcement learning framework with an attention mechanism.
[0040] Figure 3 This is the working diagram of the attention mechanism between the fully connected layers of the Critic network;
[0041] Figure 4 It is the flow chart of sub-band selection and power allocation. DETAILED DESCRIPTION
[0042] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0043] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] Example 1
[0045] 1 System Model
[0046] The network environment deployment solution of the present invention integrates low-frequency base stations (LF-BS) below 6GHz and millimeter-wave high-frequency base stations (HF-BS) to achieve high- and low-frequency collaborative networking. In this scenario, the network consists of a macro cell and multiple micro cells. The low-frequency base station is responsible for providing wide-area coverage of the macro cell control plane to ensure stable network connection and management. At the same time, the high-frequency base station focuses on the micro cell, providing transmission of high-speed data services. Since the LF-BS is responsible for wide-area coverage and ensures reliable transmission of the system, it needs to operate continuously. The present invention mainly considers the dynamic adaptive adjustment of the downlink transmission power of the HF-BS in the micro cell, so that limited power resources can be fully utilized while expanding the system capacity and reducing interference.
[0047] The network scenario is covered by n HF-BSs, n∈N={1,2,…N}, where N is the number of HF-BSs within a certain range. Each HF-BS applies to K carrier subbands with a total bandwidth of B. The transmission link between the user and the base station uses orthogonal channels. Therefore, during the downlink transmission process, there is no inter-cell downlink interference within the coverage area of each HF-BS microcell. In the network scenario, it is assumed that the system operation is completely synchronous, and time is divided into time slots of fixed length T. Given the relative scarcity of available spectrum resources, N is usually much larger than the number of subbands. Each link selects a subband at the beginning of each time slot. The LF-BS collects network information within the macrocell within the coverage area, such as signal-to-interference-and-noise ratio, transmission power, etc.
[0048] At time t, the user can access the LF-BS in the area and establish a dual connection with a HF-BS in the coverage area. At this time, the HF-BS is in the active state and its transmission power is The number of users connected to the nth HF-BS is represented by Indicates that M is the total number of users. The total number of user connections is as follows:
[0049]
[0050] definition As the downlink channel gain from the nth HF-BS to the mth user on the kth sub-band in time slot t, this gain combines the effects of large-scale and small-scale fading:
[0051]
[0052] in is the large-scale fading including path loss and shadow fading, It is a small-scale fading. Considering the slow fading characteristics of large-scale fading and the small path loss in a small range of HF-BS, it is assumed that all sub-bands have the same large-scale fading. The Jakes fading model is used to describe The small-scale fading of each channel follows a first-order complex Gauss-Markov process, and the small-scale fading of the sub-band within the frequency band is:
[0053]
[0054] Where ρ represents the correlation between two fading blocks, ρ = J0(2πf d T), where J0(.) depends on the maximum Doppler frequency f d Bessel functions of the first kind; and the fading variation of the current time slot are independent and identically distributed complex Gaussian random variables with zero mean and unit variance.
[0055] The downlink signal to interference and noise ratio (SINR) of the nth HF-BS in time slot t can be calculated using To express:
[0056]
[0057] in and They represent the channel gains of the nth and n'th HF-BS and the mth user on the kth subband respectively. and denote the total transmission power of the nth and n'th HF-BS on the kth sub-band respectively. is the dynamic switching factor of HF-BS, σ 2 Represents noise. Using binary variables To indicate whether the HF-BS allocates the kth subband to user m in time slot t, if subband k is selected, then and Indicates whether user m is connected to the nth HF-BS, also represented by 0 and 1.
[0058] The performance of the system is analyzed by the HF-BS capacity, which is the capacity of the nth HF-BS and user m on subband k. It can be expressed by the following formula:
[0059]
[0060] Therefore, the capacity of the nth HF-BS can be defined as:
[0061]
[0062] The fundamental purpose of this solution is to increase the overall system capacity by controlling the downlink power of the base station and reducing the downlink inter-cell interference, while also considering the energy consumption in this scenario to achieve a compromise between energy efficiency and system capacity. The transmit power of the nth HF-BS on all subbands at time t can be defined as:
[0063]
[0064] If there is no user request or few user requests within the HF-BS coverage area, the system will be set to sleep. is the dynamic switching factor of the nth HF-BS. When the HF-BS enters the low-power sleep mode, some functions are turned off to reduce energy consumption.
[0065] Therefore, the power consumption of each HF-BS can be expressed as two categories: the first category is the zero-load always-on power p when the base station is in the active state. A and the transmission power when there is a transmission task The second category is the sleep power p of the base station in low-power sleep mode. S .
[0066] In order to simplify the analysis process of the problem, the switching power consumption consumed by the base station switching state is not considered here, and p A With p s They can all be considered as constants, and their values do not affect the trend of system performance changes, but only the specific value of energy consumption. In this case, the power consumption model of the nth HF-BS can be expressed as follows:
[0067]
[0068] Energy efficiency is expressed as:
[0069]
[0070] Therefore, in the high- and low-frequency collaborative networking scenario, downlink power control is performed on densely deployed high-frequency micro base stations to rationally allocate power resources in different subbands, maximizing system capacity while ensuring power consumption is within a tolerable range. The optimization goal of this invention is to increase system capacity by controlling the downlink power of the base station, reducing interference, and taking energy consumption into consideration to achieve a balance between energy efficiency and system capacity. The optimization problem can be expressed as:
[0071]
[0072] Among them, p max With p min They represent the maximum transmission power and minimum transmission power of HF-BS, E total is the total system power consumption, E max It is the maximum power consumption limit of the system.
[0073] 2 Deep Reinforcement Learning for Multi-Agent Cooperative Games
[0074] A deep reinforcement learning algorithm is used to solve the power allocation problem in a multi-agent cooperative game. Each HF-BS agent perceives the environment through a deep learning network and makes decisions using reinforcement learning strategies, enabling the entire system to achieve adaptive power control and resource allocation in complex environments.
[0075] 2.1 Multi-agent cooperative game learning scheme
[0076] Deep reinforcement learning combines the decision-making capabilities of reinforcement learning with the perception capabilities of deep learning. It consists of five basic elements: agent, environment, state, action, and reward. To allow distributed execution, each HF-BS transmitter operates as an independent part by treating other agents as part of the local environment.
[0077] Construct the state space, action space, and reward function.
[0078] (1) State space S: Each HF-BS needs to know the number of users it currently serves and the channel state observed at this time t. It also needs to know its own transmit power at this time. So the state space S is the number of users connected at this time. Channel state h t , subband selection connected to the user And the transmission power at this time The state S at time t can be expressed as:
[0079] (2) Action space A: At time t, each HF-BS agent needs to know the transmit power that should be controlled at this time. For base station n, it is necessary to select the frequency band to be power controlled at this time and the size of the power change on it. therefore The sign represents the increase or decrease of power. And is determined by the total capacity of the system. If the reward of agent n Increase, the action selection makes A positive value increases the transmission power, and vice versa.
[0080] (3) Reward function R: In the power allocation process, the optimization strategy of a single agent is insufficient to cope with the complexity of the entire network. Cooperative game theory emphasizes optimizing the entire network through mutual cooperation and strategy coordination among agents in a multi-agent system. The decision of each agent not only affects itself but also the performance of other agents in the network. Therefore, when designing the reward function, it is necessary not only to consider the performance optimization of a single agent but also to focus on cooperation and coordination among agents to maximize the overall efficiency of the network.
[0081] HF-BS downlink power allocation, interference and energy efficiency control are defined as the following cooperative game:
[0082] F=(S,A,R) (11)
[0083] It is the set of utility functions associated with the agent and its strategy, i.e., the reward function, which consists of system capacity, energy efficiency, and interference penalty. This study aims to maximize both system capacity and energy efficiency simultaneously. However, these two goals are often contradictory. We can use a multi-objective optimization method to integrate multiple goals into a comprehensive goal for solution. Since the unit of energy efficiency is bits / J and the unit of system capacity is bps, in order to keep the units consistent in the calculation process, the system capacity function is multiplied by (B tot is the total bandwidth, E tot is the total power consumption of the agent when transmitting at full power. In addition, a parameter α (0≤α≤1) is added to balance energy efficiency and system capacity. During peak demand periods, priority is given to system capacity to support more services; during low-peak periods, the focus is on improving energy efficiency to reduce energy consumption. Increasing α tends to increase system capacity. When α is 1, the goal is to maximize system capacity; when α is 0, the goal shifts to maximizing energy efficiency. At the same time, the total bandwidth B tot The increase of will increase the weight of system capacity in the objective function. Conversely, when power is sufficient, the objective function will pay more attention to energy efficiency.
[0084] Therefore, the reward function is defined as:
[0085]
[0086] The penalty term for an agent depends on the externality impact on its interfering neighbors. For the penalty, we first calculate the agent n’s interference in time t. externalities, is the set of receivers affected by interference.
[0087]
[0088] in There is no agent n pairs of subbands during time slot t The system capacity of j when the interference is .
[0089] The goal of deep reinforcement learning is to learn a policy that maximizes the cumulative discounted reward at time t. Therefore, the reward of agent n is defined as:
[0090]
[0091] 2.2 DQN and DDPG Joint Subband Selection and Power Allocation Architecture
[0092] We previously optimized the reward function design based on cooperative game theory and defined the optimization goal. This section will address this goal through deep reinforcement learning techniques, using intelligent technology to achieve efficient dynamic adjustment and optimization in multi-agent systems.
[0093] Specifically, this solution uses a two-layer deep reinforcement learning algorithm to optimize power allocation. The upper layer uses a deep Q network to select subbands, and the lower layer DDPG performs power allocation. An attention mechanism is added between the first and second fully connected layers of the lower layer Critic network to improve the assessment of environmental status and actions.
[0094] The learning action value function of DQN is Q(s,a). Let π(a|s) be the probability of taking action a in the current state s. The expected cumulative reward of the Q function when taking action a in state s is:
[0095]
[0096] Assume that for a given state s, The most advantageous action a * , the optimal strategy π * If (a|s) is equal to 1, then the optimal Q function satisfies the Bellman equation:
[0097]
[0098] in is the expected reward for taking action a in state s, and is the transition probability from state s to the next state s' through action a, Pr represents the transition probability. γ∈(0,1) is the discount factor for the relative importance of future rewards to the current policy.
[0099] DQN uses a deep neural network to approximate the Q function q(s,a;ψ), where ψ represents the network parameters. It stores past experiences in an experience replay pool, represented as D in the form of e=(s,a',s'). Small values for the maximum memory size |D| can lead to overfitting, while large values can slow down learning. Create a model with parameter ψ targetThe target network predicts the target value in the following mean squared Bellman error:
[0100]
[0101] Where the target y(r',s')=r'+γmax a' q(s',a';ψ target ). Represents the state transition (s,a,r',s') sampled from the experience replay pool D; (y(r',s')-q(s,a;ψ)) 2 : represents the mean square error between the Q value q(s,a;ψ) and the target Q value y(r',s'); q(s,a;ψ): represents the Q value predicted by DQN based on the current parameter ψ; r' represents the immediate reward obtained after performing action a in state s.
[0102] Each iteration samples a batch from the experience replay, calculates the mean squared gradient of the Bellman error, and updates the neural network parameters. To minimize the Bellman error (17), ψ is updated by sampling a random mini-batch B from D and running gradient descent:
[0103]
[0104] After each iteration, ψ is updated with train During training, an ε-greedy strategy is used for exploration, randomly selecting actions with a certain probability ε.
[0105] To overcome the limitations of deep Q-learning in handling continuous actions, the underlying continuous power control uses a deep deterministic policy gradient algorithm. An actor network and a critic network are trained simultaneously. The critic network, defined by φ, is iteratively trained to represent the action-value function. The critic network is then used to train the actor network, defined by θ, which parameterizes the deterministic policy. The deterministic policy is defined as μ:S→A. For a given state, the action is determined by a = μ(s;θ).
[0106] Therefore, the target strategy μ * Satisfies Bellman properties:
[0107]
[0108] Similar to the deep Q network, the critic network is trained by minimizing the mean square Bellmann error. Compared with the deep Q network, the single output of the critic network can give an estimate of the Q function value for a given state and action output. The goal in DDPG is y critic (r',s')=r'+γq(s',μ(s';θ);φ target ).
[0109] The power control action is continuous, and q(s,a;φ) is differentiable with respect to the action, so the policy parameters only need the following gradients to be updated:
[0110]
[0111]
[0112]
[0113] 2.3 Critic Network with Multi-Head Attention
[0114] Leveraging the aforementioned reward function strategy, the deep reinforcement learning framework optimizes power allocation decisions through interactions between agents. This solution introduces an attention mechanism within the critic network, enabling each agent to selectively focus on specific agent information as needed, rather than simply averaging the influence of all agents. This allows for dynamic weighting of local information. This allows agents to consider both their individual optimal strategies and the overall performance optimization of the network.
[0115] To calculate the state-value function of the nth agent, the Critic network receives the states s∈S and actions a∈A of all agents in this time slot, where q(s, a; φ) is a function that calculates the action-state value of agent n and the contribution from other agents:
[0116] q(s,a;φ)=f n (l n (s n ,a n ),x n ) (twenty one)
[0117] l n is the first layer of the fully connected network of Critic, f n It is the last two layers, and the input is the current state s of the agent n and action a n , x n Represents contributions from other agents, in the form of:
[0118]
[0119] V is a shared matrix within the attention, which linearly transforms the contribution of other agents through V, and is activated by the RELU nonlinearity represented by h, and finally multiplied by the attention weight α n' ,α n' The differences between agents are analyzed using an attention query key system and a softmax nonlinear activation in the form of:
[0120]
[0121] Among them, ξ n' =l n' (s n' ,a n' ). W q Will n Converted into "query", W k Will n' Converted into a "key" and then scaled to match the dimensions of the two matrices to prevent gradients from vanishing.
[0122]
[0123] By rationally designing the state space, action space, and reward function, and combining it with deep reinforcement learning methods, the present invention can effectively optimize the power allocation of the communication system, improve system capacity and energy efficiency, and achieve efficient and intelligent network resource management.
[0124] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A communication system power allocation method based on multi-agent cooperation, its characteristic steps include: Build a high- and low-frequency collaborative networking scenario that integrates low-frequency base stations (LF-BS) below 6 GHz and millimeter-wave high-frequency base stations (HF-BS), with each HF-BS acting as an intelligent agent instance. Construct the state space, action space, and reward function of the intelligent agent. Integrate the concept of cooperative game into the construction of the reward function. Define the reward function as a constraint relationship model between system capacity, energy efficiency, and interference penalty. The reward function sets the optimization goal. After setting the optimization target for the reward function, a two-layer deep reinforcement learning algorithm architecture is used for subband selection and power allocation. The upper layer uses the deep Q network (DQN) to select subbands, and the lower layer uses the deep deterministic policy gradient (DDPG) for power allocation. A multi-head attention mechanism is added between the first and second fully connected layers of the DDPG critic network in the lower layer to selectively capture the state and action information of other intelligent agents and weightedly fuse this information with its own features.
2. The communication system power allocation method based on multi-agent cooperation according to claim 1, characterized in that: The state space is determined by the number of currently connected users. Channel state h t , Subband selection for user connection and the current transmit power At time t, the state space is expressed as:
3. The communication system power allocation method based on multi-agent cooperation according to claim 2, characterized in that: The action space is expressed as: where k n is the subband to be power controlled by the nth HF-BS at the current time t, Represents the power adjustment amount of the nth HF-BS at the current time t, and its value range is 4. The communication system power allocation method based on multi-agent cooperation according to claim 3, characterized in that: The reward function consists of system capacity, energy efficiency, and interference penalty. The optimization goal is to maximize system capacity and energy efficiency. A multi-objective optimization method is used to integrate multiple objectives into a comprehensive goal for solution. The reward function is defined as: in, represents the reward function, α represents the balance parameter between energy efficiency and system capacity, and its value range is 0≤α≤1; EE n represents the energy efficiency of the nth HF-BS; B tot is the total bandwidth; E tot is the total power consumption of the HF-BS when transmitting at full power; is the interference penalty of the nth HF-BS; is the capacity on subband k; B / K represents K carrier subbands with a total bandwidth of B.
5. The communication system power allocation method based on multi-agent cooperation according to claim 4, characterized in that: The optimization objective is expressed as: Constraints: (1) (2) Among them, p max With p min They represent the maximum transmission power and minimum transmission power of HF-BS, E total is the total system power consumption, E max Is the maximum power consumption limit of the system; is the transmission power of subband k; n∈N={1,2,…N}, N is the number of HF-BSs within a certain range; is the total system capacity; EE total Represents the total energy efficiency of the system; E n is the power consumption of the nth HF-BS; C t Represents the total system capacity at time t.
6. The communication system power allocation method based on multi-agent cooperation according to claim 5, characterized in that: Total system capacity for: Where: K is the total number of carrier subbands, M is the total number of users, is the capacity of the HF-BS and user m on subband k, expressed as: is the signal-to-noise ratio; The power consumption of HF-BS includes zero-load normally open power p A and transmission power And the power p in low power sleep mode S : Q n is the dynamic switching factor of HF-BS; Energy Efficiency (EE) n Defined as:
Citation Information
Patent Citations
Multi-agent deep reinforcement learning strategy optimization method based on attention mechanism
CN113392935A
TACS network resource allocation method based on deep reinforcement learning
CN116405904A