5G / 5G-A network dynamic power distribution method based on continuous near-end strategy optimization
By optimizing agents based on PPO deep reinforcement learning and attention mechanisms, the accuracy and energy consumption issues of dynamic power allocation in 5G/5G-A networks are solved, realizing a high-precision, low-energy dynamic power allocation strategy suitable for multi-node heterogeneous reliability scenarios.
Patent Information
- Application Number
- CN202511456049.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-12-12
AI Technical Summary
Existing dynamic power allocation methods for 5G/5G-A networks suffer from insufficient allocation accuracy, excessive learning burden, and high environmental uncertainty when facing the requirements of multi-node heterogeneous reliability and low energy consumption. It is difficult to achieve high-precision dynamic power allocation in complex channel environments.
We employ a deep reinforcement learning approach based on continuous proximal policy optimization (PPO) combined with an attention mechanism to construct an agent that optimizes power resource allocation. By solving high-dimensional feature problems through a continuous action space, we enhance the feature processing capability of multiple nodes, improve allocation accuracy, and reduce energy consumption.
It achieves high-precision dynamic power allocation in complex channel environments, meets the reliability requirements of multi-node heterogeneity, reduces system energy consumption, and improves the learning speed and stability of allocation strategies.
Smart Images

Figure CN121126533A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of mobile communication and relates to a 5G / 5G-A network dynamic power distribution method based on continuous proximal policy optimization (PPO). BACKGROUND
[0002] With the deep integration of communication technology and industry application, 5G technology provides communication foundation for industrial control and vehicle networking fields. 5G-A, as an evolution and enhancement form of 5G, further improves the performance upper limit of the communication system. 5G / 5G-A has the characteristics of ultra-high reliability and low latency, which provides the possibility to meet the real-time and low-energy consumption requirements of communication. The transmission capability of the 5G / 5G-A network node depends on the transmission power. In order to ensure that the node transmits the data in the buffer in time within the specified latency deadline using the minimum power, the base station needs to obtain and analyze complex channel state information and make power distribution decisions, thereby reducing energy consumption and improving reliability.
[0003] Power distribution is a key technology in 5G / 5G-A. In order to optimize the energy consumption and reliability of the communication system, how to realize dynamic power distribution based on a deep reinforcement learning framework has become a core problem to be solved. The power distribution methods include D3QN, DQN, etc. In these methods, the agent is usually deployed at the base station side, learns the resource allocation strategy autonomously by perceiving the environment state, and realizes the joint optimization of energy consumption and reliability of the communication system. However, most of the above methods make distribution decisions based on discrete action space, which leads to quantization error of continuous resources such as power distribution, limiting the distribution accuracy. At the same time, in order to improve the distribution accuracy, the high dimension of the action space also increases the learning burden of the agent, affecting the convergence speed. In addition, different nodes have differentiated reliability requirements, and the power distribution strategy is difficult to balance global and individual constraints. Finally, the mobility of nodes in the industrial environment will cause the channel gain to change rapidly, further increasing the uncertainty of the environment, and posing a severe challenge to the deep reinforcement learning algorithm. Therefore, how to invent a dynamic power distribution method suitable for 5G / 5G-A network, which still has high-precision power distribution performance in the real-time change of the environment, meets the requirements of multi-node heterogeneous reliability and low energy consumption, has become an important challenge. SUMMARY
[0004] Therefore, the present application aims to provide a 5G / 5G-A network dynamic power allocation method based on continuous proximal policy optimization, which is suitable for 5G / 5G-A network scenarios with single base station and multiple nodes, and can allocate power resources under the constraint of reliability to ensure the real-time performance and low energy consumption of the communication system. The present application constructs a deep reinforcement learning agent based on PPO, and formulates the problem of minimizing system energy consumption under the constraint of reliability as a Markov decision process. In order to solve this optimization problem, the present application introduces an attention mechanism to enhance the processing capacity of multiple node features, meets the heterogeneous reliability requirements of multiple nodes, effectively overcomes the high-dimensional feature problem through continuous action space, speeds up the learning speed of the allocation strategy, improves the power allocation accuracy, and achieves the goal of low energy consumption.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions: A 5G / 5G-A network dynamic power allocation method based on continuous proximal policy optimization, the method comprising: S1, obtaining the relevant parameter information of the 5G / 5G-A network, and establishing an attention mechanism improved continuous proximal policy optimization (PPO) network; S2, establishing state space, action space and reward function according to the optimization target of reliability constraint and minimum system energy consumption; S3, updating the packet loss amount, backlog amount of data buffer queue and channel gain of each node, and the attention mechanism improved continuous proximal policy optimization algorithm explores the continuous action space according to the state and outputs the optimal action directly representing the node power value; the system feeds back the action value and timely reward value, and saves the state, reward value, action value, logarithmic probability and state value to the replay area; S4, training the attention mechanism improved continuous proximal policy optimization (PPO) network, calculating the loss function through the logarithmic probability, state value and entropy, and then updating the parameters by using the back propagation and gradient descent mechanism until convergence to obtain the trained network; and the trained network is used for real-time network dynamic power allocation.
[0006] Further, in step S1, it is assumed that in a single base station uplink scenario, it includes network slices and mobile nodes and a base station providing power allocation service for the nodes; The base station collects the channel state information between the nodes in different slices and itself for power allocation decision, and sends power adjustment commands to the nodes, and divides the subframe into multiple time slots by configuring the subcarrier spacing; The reliability requirements of the nodes under different slices are different, and the nodes occupy physical resource blocks, and the bandwidth of each physical resource block is The nodes transmit data to the base station using the allocated spectrum resources; Each node maintains a queue in which data to be sent to the base station is buffered; the indices for slices, mobile nodes, time slots, and queues are defined as follows: , , and ; Introducing user mobility and considering real-time changing channel gain, the rate at which a user transmits data in the uplink can be expressed as:
[0007] in, For transmission rate, This refers to the bandwidth used by the node. For transmission power, For channel gain, This is channel noise.
[0008] Furthermore, the attention mechanism-improved continuous proximal policy optimization (PPO) network established in step S1 is as follows: It includes an Actor network and a Critic network. The Actor network consists of a Linear input layer, a fully connected layer, and a Linear output layer, with Tanh activation functions used between layers, and channel information as multi-dimensional feature input. The Critic network consists of three Linear layers with an output dimension of 1. The number of nodes is considered in relation to the network's structure. The output dimension of the Actor network is determined as follows: At the same time, the input feature dimension is determined to be 8. The Actor network maintains a power allocation strategy; the Critic network evaluates the state, obtains the state value, and uses it in the pruning objective function; the allocation of node transmit power is completed at the beginning of each time slot, and the power allocation of each node is independent of each other; the attention mechanism allows the agent to pay attention to multiple feature dimensions of channel information at the same time, the node features are mapped to a high-dimensional feature space through a fully connected layer, and the acquired new features are used as input to the Actor network.
[0009] Furthermore, in step S2, the process of establishing the optimization objective and its constraints based on minimizing system energy consumption is as follows: First, we establish an optimization objective based on minimizing system energy consumption, expressed as:
[0010] in, Indicates the total number of time slots. Represents a node The transmission power, N represents the number of nodes, the total energy consumption of the system throughout the time slot is embodied as the sum of the transmission power value of each time slot; Then, the reliability constraint condition is established:
[0011] Wherein, the data in the buffer queue exceeding the delay deadline is determined as a packet loss, is the statistical packet loss amount, is the reliability requirement, is the data buffered into the queue; The node exchange position specification constraint is established:
[0012] The power constraint is established:
[0013] The delay constraint is established:
[0014] In the case of limited bandwidth, the total delay is represented as , is the queue delay, is the delay deadline of different nodes.
[0015] Further, step S2 also includes the initialization establishment process of the state space, action space and reward function of the system: The state space of the system is established is:
[0016] Wherein, the network information of each user is a sub-state is represented as:
[0017] Wherein, each sub-state includes traffic value , channel gain , buffer queue data backlog , buffer queue length , transmission rate , power , power surplus and packet loss amount ; For the action space, the continuous action space of the node transmission power is all possible allocation methods under the upper and lower power constraints, and for multiple nodes, multiple continuous power values are selected as actions, then the overall action space is represented as:
[0018] where each node's continuous action takes value ; The established reward function is expressed as:
[0019] where, is the difference between the allocated power and the set maximum power upper limit, representing the power surplus amount; is the weight of the node for the power surplus; is the penalty of the node's packet loss amount, is the weight of the packet loss item; is the penalty of the node's insufficient transmission capacity, is the corresponding weight.
[0020] Further, in step S3, for the backlog of the data buffer queue, each time slot starts, each user maintains the data buffer queue stores the upcoming data packet , and controls the transmission of the data packet according to the transmission power control parameter provided by the base station, and the backlog is expressed as:
[0021] where, represents the amount of data transmitted; For the packet loss amount, the data packet in the data buffer queue that exceeds the maximum deadline is discarded within each time slot, which is expressed as:
[0022] where, is the packet loss amount, is the first data packet of the data buffer queue; is an indicator function, if the queue length of the current time slot is greater than the maximum deadline, then , otherwise ; For the channel gain, its real-time change is jointly determined by the large-scale fading and the small-scale fading, and can be expressed as:
[0023] where the large-scale fading is expressed as:
[0024] where the free space path loss is a fixed value, the distance between the base station and the node changes between time slots, is a Gaussian random variable; the modulus of the small-scale fading It follows a Rayleigh distribution.
[0025] Furthermore, step S3 also includes an action selection process. In the action selection stage, the Actor network observes the channel state information, generates a multivariate normal distribution based on the mean of the output actions, samples this multivariate normal distribution to obtain the action value representing the power value, and its selection strategy is expressed as follows:
[0026] in, As a strategy, The mean of the action. Standard deviation; Based on the action value representing the power value, the system provides timely rewards and saves the status, timely reward value, action value, log probability, and status value to the replay area.
[0027] Furthermore, in step S4, the training process for the attention mechanism-improved continuous proximal policy optimization PPO network is as follows: The objective function for clipping proximal constraints is determined as follows:
[0028] in, It is the ratio of the new strategy to the old strategy. It is a hyperparameter that controls the degree of policy updates. It is an advantage function that represents how much better or worse the current action is compared to the average action of the strategy. Determine the loss function, select an objective function that incorporates pruning, and use the mean squared error of the cumulative reward and state value, along with changes in entropy, to guide the algorithm's convergence. The loss function... The calculation formula is:
[0029] in, It is a cumulative reward; It is the mean squared error loss of state value and reward; It is the entropy of the network with respect to the normal distribution of actions; , These represent the weighting coefficients for mean squared error loss and entropy, respectively. Every few steps, the gradient is calculated through backpropagation, and the parameters are updated using the optimizer until the network converges.
[0030] The beneficial effects of this invention are as follows: 1) For the uplink of a single-base station multi-node network, dynamic power allocation is used to ensure node reliability while minimizing power consumption. An improved continuous PPO method based on an attention mechanism enhances the ability to capture multi-node channel characteristics. The continuous action space overcomes the problem of high-dimensional action space, thus accelerating training speed and achieving a dynamic power resource allocation method. This invention uses an improved continuous PPO method based on an attention mechanism to minimize power allocation under reliability constraints, thereby reducing total energy consumption.
[0031] 2) Based on the deep reinforcement learning framework, the continuous PPO with improved attention mechanism directly corresponds the action value selected by the policy to the accurate transmission power value required by the node through the continuous action space, which reduces the dimensionality of the action space and improves the power allocation accuracy. This power allocation method is suitable for scenarios with multi-node heterogeneous reliability requirements.
[0032] 3) This invention overcomes the convergence difficulty caused by the high-dimensional action space in power resource allocation, improves the convergence speed and stability of power allocation under complex channel conditions, and reduces system energy consumption while meeting the reliability requirements of multi-node heterogeneity.
[0033] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Fig. 1 This is a schematic diagram of the structure of the Proximal Continuous Policy Optimization (PPO) network based on the improved attention mechanism of a deep learning framework, according to an embodiment of the present invention. Fig. 2 This is a schematic diagram illustrating the execution flow of the dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to an embodiment of the present invention. Detailed Implementation
[0035] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0036] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0037] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0038] Please see Figs. 1-2 This is a dynamic power allocation method for 5G / 5G-A networks based on continuous near-end strategy optimization.
[0039] The method specifically includes the following steps: S1: Obtain relevant parameter information of 5G / 5G-A network, construct intelligent agent, and build a continuous PPO algorithm with improved attention mechanism; S2: Establish the state space, action space, and reward function based on reliability constraints and the optimization objective of minimizing system energy consumption; S3: Update the packet loss amount, data buffer queue backlog, and channel gain of each node. The continuous near-end policy optimization algorithm with improved attention mechanism explores the continuous action space and outputs the optimal action that directly represents the node power value. The system provides timely reward value based on the action value and saves the state, reward value, action value, log probability, and state value to the replay area. S4: Train the agent and calculate the loss function, then use backpropagation and gradient descent mechanisms to update the parameters until convergence, thus obtaining the trained agent.
[0040] In step S1 of this embodiment, constructing a continuous PPO with an improved attention mechanism specifically includes the following steps: S101: In a single-base station cell uplink scenario, the system consists of a base station providing power allocation services to nodes. A network slice and The system consists of several mobile nodes. The base station collects channel state information between itself and nodes in different slices to make power allocation decisions and sends power adjustment commands to the nodes to ensure the uplink service quality and reliability. Furthermore, the reliability requirements of nodes differ across slices. This embodiment sets the system to be time-slot based, dividing subframes into multiple time slots by configuring subcarrier spacing. Additionally, nodes occupy physical resource blocks, each resource block having… The bandwidth is Hz, and data is transmitted to the base station using the allocated spectrum resources; each node maintains a queue in which data to be sent to the base station is buffered; the indices of slices, mobile nodes, time slots, and queues are defined as follows: , , and ; S102: Introducing user mobility and considering real-time changing channel gain facilitates base station scheduling decisions. In the uplink of this system, using Shannon's formula, the rate at which users transmit data can be expressed as:
[0041] in, For transmission rate, This refers to the bandwidth used by the node. For transmission power, For channel gain, Channel noise; S103: The distance between the base station and the node varies in each time slot, causing changes in path loss. The channel gain can be expressed as:
[0042]
[0043] in, For free space path loss, For base stations and nodes The distance between them It is a Gaussian random variable; S104: As Fig. 1As shown, the attention mechanism-improved continuous proximal policy optimization includes an Actor network and a Critic network. The Actor network consists of a Linear input layer, a fully connected layer, and a Linear output layer, with Tanh activation functions used between layers, taking channel information as multi-dimensional feature input. The Critic network consists of three Linear layers with an output dimension of 1. The number of nodes... The output dimension of the Actor network is determined as follows: At the same time, the input feature dimension is determined to be 8. The Actor network maintains a power allocation strategy; the Critic network evaluates the state, obtains the state value, and uses it in the pruning objective function; the allocation of node transmit power is completed at the beginning of each time slot, and the power allocation of each node is independent of each other; the attention mechanism allows the agent to pay attention to multiple feature dimensions of channel information at the same time, and the node features are mapped to a high-dimensional feature space through a fully connected layer, and the acquired new features are used as input to the Actor network. Step S2 in this embodiment specifically includes the following steps: S201: Construct an optimization objective based on minimizing system energy consumption, expressed as:
[0044] in, Indicates the total number of time slots. Represents a node The transmission power, This represents the number of nodes, and the total energy consumption over the entire time slot of the system is represented by the sum of the transmit power values of each time slot. The constraints are as follows: To meet reliability constraints, ensure that data buffer backlog does not become too large, and reduce packet loss:
[0045] In this process, data in the buffer queue that has exceeded the delay deadline is considered lost. For the statistical data on packet loss, For reliability requirements, To buffer the data being queued; Ensure that the node changes position within the specified range:
[0046] The range of constrained power selection is between the minimum power and the maximum power:
[0047] S202: Under limited bandwidth conditions, the total delay is expressed as... :
[0048] in, Indicates queue delay. For different nodes, the delay deadline is set. S203: Establish the system's state space for:
[0049] Among them, each user's network information is a sub-state. It can be represented as:
[0050] This includes flow rate values. Channel gain Data backlog in buffer queues Buffer queue length Transmission rate ,power Power surplus and packet loss ; S204: The continuous action space for node transmit power represents all possible allocation methods under power upper and lower limit constraints. For multiple nodes, the agent will select multiple consecutive power values. The overall action space can be represented as:
[0051] The continuous action value of each node is... ; S205: The reward function ensures that the optimization process gradually converges to the desired solution, improving the efficiency and accuracy of problem solving. The reward function is constructed as follows:
[0052] in, The difference between the allocated power and the set maximum power limit represents the power margin. The weight of a node for power surplus; Penalty for the amount of packet loss at a node. The weight of the packet loss item; As a penalty for insufficient node transmission capacity, For the corresponding weights; Step S3 in this embodiment specifically includes the following steps: S301: At the start of each time slot, each user maintains a data buffer queue. All will store upcoming data packets Meanwhile, the transmission of data packets is controlled according to the transmit power control parameters provided by the BS. The backlog in the data buffer queue can be expressed as:
[0053] in, Indicates the amount of data sent.
[0054] S302: This system discards data packets that have exceeded the maximum deadline in the data buffer queue within each time slot, which can be represented as:
[0055] in, For packet loss, This is the first data packet in the data buffer queue; It is an indicator function that checks if the queue length of the current time slot is greater than the maximum deadline. ,otherwise ; S303: During the action selection phase, the Actor network observes channel state information, generates a multivariate normal distribution based on the mean of the output actions, and samples this multivariate normal distribution to obtain the action value representing the power value. The Actor network maintains a strategy, which can be expressed as:
[0056] in, As a strategy, The mean of the action. The standard deviation is denoted as .
[0057] Step S4 in this embodiment specifically includes the following steps: S401: Proximal constraints are used to stabilize the training process in continuous PPO with improved attention mechanisms; the objective function for pruning proximal constraints is:
[0058] in, It is the ratio of the new strategy to the old strategy. It is a hyperparameter that controls the degree of policy updates. It is an advantage function that represents how much better or worse the current action is compared to the average action of the strategy. S402: Improved attention mechanisms in continuous PPO The function selection, combined with the objective function of pruning, and the changes in the mean square error and entropy of the cumulative reward and state value guide the convergence of the algorithm. The calculation formula is:
[0059] in, It's a cumulative reward. It is the mean squared error loss of state value and reward, which helps to improve the accuracy of value assessment. It is the entropy of the agent's actions according to the normal distribution. This term encourages the policy to explore randomly and avoids the policy from getting trapped in local optima. S403: Every few steps, the gradient will be calculated through backpropagation, and the parameters will be updated using the optimizer until the network converges.
[0060] Example 2 This embodiment Fig. 2 The flowchart of the dynamic power allocation method for 5G / 5G-A networks based on continuous power point allocation (PPO) of the present invention is shown. Specifically, it includes: V1: Dynamic power allocation for 5G / 5G-A networks begins; V2: Obtain channel status and traffic data in 5G / 5G-A networks; V3: Establish an improved PPO with an attention mechanism and initialize network parameters; V4: Initializes the network's state space, action space, and reward function; V5: Updates packet loss, queue backlog, and channel gain for each node; V6: Based on the current state, calculate the mean and Gaussian distribution of the actions, and sample the optimal action; V7: Rewards based on action feedback; V8: Saves state, reward value, action value, log probability, and state value to the replay area; V9: Calculates the loss function using log probability, state value, and entropy; V10: Calculate gradients and update network weights; V11: Updates the ActorCritic network parameters every few steps; V12: Determine if convergence has occurred. If not, return to V5; otherwise, proceed to V13. V13: Obtain the power allocation result; V14: End.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization, characterized in that: The method includes: S1. Obtain relevant parameter information of 5G / 5G-A network and establish a continuous near-end strategy to optimize PPO network with attention mechanism improvement; S2. Establish the state space, action space, and reward function based on reliability constraints and the optimization objective of minimizing system energy consumption; S3. Update the packet loss amount, data buffer queue backlog, and channel gain of each node. The attention mechanism-improved continuous near-end policy optimization algorithm explores the continuous action space based on the state and outputs the optimal action that directly represents the node power value. The system provides timely reward values based on the action values and saves the state, reward value, action value, log probability, and state value to the replay area. S4. Train the PPO network with improved attention mechanism and continuous proximal policy. Calculate the loss function using log probability, state value, and entropy. Then update the parameters using backpropagation and gradient descent until convergence to obtain the trained network. Use this network for real-time network dynamic power allocation.
2. The dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 1, characterized in that: In step S1, a single-base station cell uplink scenario is defined, including... A network slice and One mobile node and one base station that provides power distribution services to the nodes; The base station collects channel state information between itself and nodes in different slices to make power allocation decisions, and sends power adjustment commands to the nodes. It divides the subframe into multiple time slots by configuring the subcarrier spacing. The reliability requirements of nodes differ across different slices. Nodes occupy physical resource blocks, and the bandwidth of each physical resource block is... The nodes use the allocated spectrum resources to transmit data to the base station; Each node maintains a queue in which data to be sent to the base station is buffered; The indices for slices, movable nodes, time slots, and queues are defined as follows: , , and ; Introducing user mobility and considering real-time changing channel gain, the rate at which a user transmits data in the uplink can be expressed as: in, For transmission rate, This refers to the bandwidth used by the node. For transmission power, For channel gain, This is channel noise.
3. The dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 2, characterized in that: The attention mechanism-improved continuous proximal policy optimization PPO network established in step S1 is as follows: It includes an Actor network and a Critic network. The Actor network consists of a Linear input layer, a fully connected layer, and a Linear output layer, with Tanh activation functions used between layers, and channel information as multi-dimensional feature input. The Critic network consists of three Linear layers with an output dimension of 1. The number of nodes is considered in relation to the network's structure. The output dimension of the Actor network is determined as follows: At the same time, the input feature dimension is determined to be 8. The Actor network maintains a power allocation strategy; the Critic network evaluates the state, obtains the state value, and uses it in the pruning objective function; the allocation of node transmit power is completed at the beginning of each time slot, and the power allocation of each node is independent of each other; the attention mechanism allows the agent to pay attention to multiple feature dimensions of channel information at the same time, the node features are mapped to a high-dimensional feature space through a fully connected layer, and the acquired new features are used as input to the Actor network.
4. The dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 3, characterized in that: In step S2, the process of establishing the optimization objective and its constraints based on minimizing system energy consumption is as follows: First, we establish an optimization objective based on minimizing system energy consumption, expressed as: in, Indicates the total number of time slots. Represents a node The transmission power, This represents the number of nodes, and the total energy consumption over the entire time slot of the system is represented by the sum of the transmit power values of each time slot. Then, establish reliability constraints: Data in the buffer queue that has exceeded the delay deadline is considered lost. For the statistical data on packet loss, For reliability requirements, To buffer the data being queued; Establish node exchange location specification constraints: Establish power constraints: Establish delay constraints: With limited bandwidth, the total latency is expressed as , Indicates queue delay. These are the delay deadlines for different nodes.
5. A dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 4, characterized in that: Step S2 also includes the initialization process for the system's state space, action space, and reward function: Establish the system's state space for: Among them, each user's network information is a sub-state. Represented as: Each sub-state includes a flow value. Channel gain Data backlog in buffer queues Buffer queue length Transmission rate ,power Power surplus and packet loss ; For the action space, the continuous action space of node transmit power is all possible allocation methods under the upper and lower power limits. For multiple nodes, multiple consecutive power values are selected as actions, and the overall action space is represented as follows: The continuous action value of each node is... ; The established reward function is expressed as follows: in, The difference between the allocated power and the set maximum power limit represents the power margin. The weight of a node for power surplus; Penalty for the amount of packet loss at a node. The weight of the packet loss item; As a penalty for insufficient node transmission capacity, For the corresponding weights.
6. A dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 5, characterized in that: In step S3, regarding the backlog of the data buffer queue, at the beginning of each time slot, each user maintains a data buffer queue. All for the upcoming data packets If data packets are stored and transmitted according to the transmit power control parameters provided by the base station, the backlog is expressed as: in, Indicates the amount of data sent; Regarding packet loss, packets that have exceeded the maximum deadline and are buffered in the data buffer queue within each time slot are discarded, as shown below: in, For packet loss, This is the first data packet in the data buffer queue; It is an indicator function that checks if the queue length of the current time slot is greater than the maximum deadline. ,otherwise ; The real-time variation of channel gain is determined by both large-scale fading and small-scale fading, as expressed as: Large-scale fading is represented as: Among them, free space path loss For fixed values, base stations and nodes Distance between It varies between time slots. For Gaussian random variables; the mode of small-scale fading It follows a Rayleigh distribution.
7. A dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 6, characterized in that: Step S3 also includes an action selection process. In the action selection phase, the Actor network observes the channel state information, generates a multivariate normal distribution based on the mean of the output actions, samples this multivariate normal distribution to obtain the action value representing the power value, and its selection strategy is expressed as follows: in, As a strategy, The mean of the action. Standard deviation; Based on the action value representing the power value, the system provides timely rewards and saves the status, timely reward value, action value, log probability, and status value to the replay area.
8. A dynamic power allocation method for 5G / 5G-A networks based on continuous near-end policy optimization according to claim 7, characterized in that: In step S4, the training process for the attention mechanism-improved continuous proximal policy optimization PPO network is as follows: The objective function for clipping proximal constraints is determined as follows: in, It is the ratio of the new strategy to the old strategy. It is a hyperparameter that controls the degree of policy updates. It is an advantage function that represents how much better or worse the current action is compared to the average action of the strategy. Determine the loss function, select an objective function that incorporates pruning, and use the mean squared error of the cumulative reward and state value, along with changes in entropy, to guide the algorithm's convergence. The loss function... The calculation formula is: in, It is a cumulative reward; It is the mean squared error loss of state value and reward; It is the entropy of the network with respect to the normal distribution of actions; , These represent the weighting coefficients for the mean squared error loss and the entropy, respectively. Every few steps, the gradient is calculated through backpropagation, and the parameters are updated using the optimizer until the network converges.