A high-energy-efficiency power allocation method under dynamic traffic based on deep reinforcement learning
Patent Information
- Application Number
- CN202610947706.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
现有技术仅根据可达数据率公式计算吞吐量,无法应用于动态流量场景
[0115](1)本发明引入深度强化学习并设置了内外循环的训练方式,考虑了时隙间队列状态的变化对决策的影响,使actor网络能够根据用户信道状态及流量队列积压信息进行决策,提升了系统的长期能效性能。
Smart Images

Figure CN122846360A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication, specifically relating to a high-efficiency power allocation method for dynamic traffic based on deep reinforcement learning. Background Technology
[0002] With the widespread application of massive MIMO and multi-base station cooperation technologies, the transmission rate of wireless communication systems has significantly improved, while the overall energy consumption of wireless networks has also continued to increase. Improving system energy efficiency (EE), which is the number of bits that can be transmitted per joule of energy, can effectively reduce power consumption while ensuring service quality. Therefore, energy efficiency research has become one of the key directions in the field of communications.
[0003] On the base station side, optimizing the allocation strategy of radio resources to improve system energy efficiency has become an important technical means. For example, optimizing key aspects such as precoding, user access, and transmit power allocation can reduce the power consumption of the base station, thereby improving system energy efficiency. In the process of energy efficiency optimization, it is usually necessary to comprehensively consider the trade-offs between various performance indicators, such as achieving a balance between system throughput, energy consumption, and quality of service constraints.
[0004] In existing technologies, energy efficiency optimization problems are typically modeled as fractional programming problems, and optimization results are obtained by iteratively updating relevant parameters. However, this type of method usually involves multiple iterations during the solution process, which is computationally complex and time-consuming, thus limiting its application in practical communication systems.
[0005] In order to reduce the computation time of complex algorithms, deep learning methods have been increasingly used in recent years to optimize wireless resource allocation in order to improve the instantaneous or long-term energy efficiency of the system. For example, in terms of improving instantaneous energy efficiency, existing technologies have adopted multiple deep neural networks (DNNs) trained by unsupervised learning to execute power allocation, user access and base station sleep strategies respectively [1]; or, a graph neural network (GNN) based on channel pseudo-inverse Taylor expansion is constructed [2] and the GNN is trained by unsupervised learning to execute precoding strategies; furthermore, DNNs are trained by deep reinforcement learning to execute antenna selection and precoding strategies [3], or DNNs are trained by deep reinforcement learning to execute beamforming weight strategies [4]. In terms of improving long-term energy efficiency, DNNs trained by deep reinforcement learning are used to execute user access strategies [5].
[0006] Although the above methods reduce computational complexity and improve system energy efficiency to some extent, they are only applicable to scenarios where the base station is operating at full load, i.e., the user data backlog is infinite and the throughput can be directly calculated using the reachable data rate formula, ignoring the dynamic change of load traffic that is common in communication systems.
[0007] In scenarios with low traffic load, base stations can serve users based on the data queues requested by users, thereby saving power and improving energy efficiency when the backlog of traffic data is small. In this scenario, the throughput of the base station is determined by the user's data backlog, power allocation decisions, and wireless channel conditions. Existing technologies calculate throughput solely based on the achievable data rate formula, which cannot be applied to dynamic traffic scenarios. Furthermore, the user's data backlog and channel state in one time slot will affect the decision-making in the next time slot, creating a coupling between the state and decision-making between time slots. Maximizing the energy efficiency of only one time slot cannot achieve optimal long-term energy efficiency.
[0008] The references are as follows:
[0009] [1]Z. Ma, N. Zhang, Y. Gao, G. Han, J. Pu and Z. Hao, "Deep-Learning-Based System-Level Energy Efficiency Maximization in Ultra-Dense Micro CellNetworks," IEEE Wireless Commun. Lett., vol. 14, no. 2, pp. 435-439, Feb.2025.
[0010] [2] J. Guo and C. Yang, "A Model-Based GNN for Learning Precoding," IEEE Trans. Wireless Commun., vol. 23, no. 7, pp. 6983-6999, July 2024
[0011] [3] A. Farhadi, R. Hatami, M. Robat Mili, C. Masouros and M. Bennis, "A Meta-Learning Approach for Energy-Efficient Resource Allocation and Antenna Selection in STAR-BD-RIS Aided Wireless Networks," IEEE WirelessCommun. Lett., vol. 14, no. 5, pp. 1421-1425, May 2025
[0012] [4] W. Li, W. Ni, H. Tian and M. Hua, "Deep Reinforcement Learning for Energy-Efficient Beamforming Design in Cell-Free Networks," IEEE WCNCW, 2021.
[0013] [5] J. Moon, S. Kim, H. Ju and B. Shim, "Energy-Efficient UserAssociation in mmWave / THz Ultra-Dense Network via Multi-Agent DeepReinforcement Learning," IEEE Trans. Green Commun. Networking, vol. 7, no.2, pp. 692-706, June 2023. Summary of the Invention
[0014] Existing technologies, when optimizing energy efficiency, do not fully consider the dynamic changes in traffic, making it difficult to achieve efficient wireless resource allocation. Furthermore, existing technologies decouple long-term energy efficiency optimization into optimizing the energy efficiency of each time slot, which does not achieve optimal long-term energy efficiency. To address these issues, this invention provides a high-efficiency power allocation method based on deep reinforcement learning under dynamic traffic conditions.
[0015] Includes the following steps:
[0016] Step 1: In the downlink transmission scenario of a multi-antenna multi-user system, enable the actor network to interact with the transmission scenario. Each time slot is used as one inner loop. Each interactive time slot generates an empirical quadruple and saves all quadruples as training samples.
[0017] The transmission scenario includes base stations, equipped with One antenna serves the community. One user;
[0018] Step 101, in the time slot The base station collects data on channel status, traffic queues, and new requests from all users within the cell. The values are 1, 2, ... ;
[0019] Channel state matrix Traffic queue ; Newly requested data ;
[0020] Step 102: The base station precodes the transmitted data according to the channel states of all users to obtain the equivalent channel between users. And calculate the queue backlog based on the user traffic queue and the data of new requests. ;
[0021] Equivalent Channel The element for , ; For users Normalized precoding vector; Among them, users queue backlog ;
[0022] Step 103: Utilize the equivalent channel With queue backlog Model the graph and resolve the queue backlog. Zhang Cheng's diagonal matrix and equivalent channel The splicing yields a multidimensional matrix. ;
[0023] Base station sends to The data streams of each user are pre-encoded, therefore there is There are 10 data vertices; simultaneously, each user receives data, therefore, there exists One user vertex; forming two types of vertices in the graph;
[0024] Multidimensional matrix As an edge feature of the graph, its first two dimensions of index The resulting slice is the data vertex. With user apex Edge features between them ;
[0025] Step 104: Convert the multidimensional matrix As a state, input the actor network to obtain actions. , i.e., slot Power allocation vector for all users;
[0026] The actor network adopts a dual-network structure, corresponding to a main network and a target network. The initial values of the target network parameters are directly copied from the main network parameters.
[0027] The specific steps are as follows:
[0028] Step 4.1, for the multidimensional matrix The queue backlog and equivalent channel in the input are scaled and then fed into the main actor network.
[0029] Step 4.2: Using the scaled queue backlog and equivalent channel, construct the initial hidden representation on the edges of the main actor network;
[0030] First, regarding the connection data vertices With user apex The edges are initially hidden as vectors. ;
[0031] Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog is used as a vector The third dimension value, if and only if When the queue is full, the value of the third dimension is queue backlog; otherwise, the value of the third dimension is 0.
[0032] Step 4.3: Update the hidden representations on all edges of the main actor network using the initial hidden representations;
[0033] The updated formula is as follows:
[0034]
[0035]
[0036]
[0037]
[0038] in, The number of layers in the GNN used by the actor network. For the first After the layer update, the first Hidden representation on each edge; , , , and For different training weights, It is an activation function. Apply LayerNorm to the input;
[0039] Step 4.4: Merge the hidden representations on all edges updated in the last layer to obtain... Multidimensional matrix ;
[0040] To hide the dimensions of the representation;
[0041] Step 4.5, for the multidimensional matrix Pooling is performed on the second dimension, and through... The readout layer of the dimension is obtained using Softmax. Normalizing the dimensional vector yields ;
[0042] Step 4.6, convert the multidimensional matrix Pooling is performed on the first two dimensions, and then through... The readout layer is dimensional, and the range of values is obtained through a scaling function. scaling factor between ;
[0043] Step 4.7, based on the base station's maximum transmit power Using normalized vectors and scaling factor Calculate the power allocation vector output by the main actor network:
[0044]
[0045] Step 105: Based on the current action Calculate time slots Reward value ;
[0046] Specifically, it includes:
[0047] Step 5.1, according to the action Equivalent channel Calculate users Transmittable data volume :
[0048] ,
[0049] in For bandwidth, For the duration of a single time slot, and Actions Chinese users Power of other users power, The equivalent channel representing the useful signal. Indicates user For users The equivalent interference channel; Noise power;
[0050] Step 5.2, according to the user queue backlog With the amount of data that can be transmitted Calculate the actual amount of data served. :
[0051]
[0052] Step 5.3: Calculate the current timeslot using the actual data volume of all users' services. System throughput And further combined with energy consumption Calculate time slots Energy efficiency bonus ;
[0053] ;
[0054] ;
[0055] Among them, energy consumption The value is the product of the total power consumption of the base station and the duration of a single timeslot; For Dinkelbach parameters;
[0056] Step 5.4, utilizing time slots Energy efficiency bonus Calculate the queue penalty term as follows and overall reward value ;
[0057] ;
[0058] ;
[0059] in These are the weighting coefficients;
[0060] Step 106, Proceed to the next time slot +1, the base station collects the channel status of all users. And calculate the equivalent channel. and queue backlog The next state is obtained by splicing. ;
[0061] According to time slot The amount of data requested by the user Calculate the updated queue backlog ;
[0062] Among them, users queue information The calculation method is as follows:
[0063]
[0064] Step 107: Divide the time slot The generated quadruple It is stored as a training sample in the Replay buffer array;
[0065] Step 108: Convert the multidimensional matrix As a state, input the main actor network to obtain actions. And further calculate the reward value. Update the next time slot The state of +2 yields a quadruple. The training sample is then stored in the replay buffer, and the loop is repeated until the time slot is reached. This terminates the current inner loop.
[0066] Step 2: In the current inner loop, using the training samples, train the actor network and the critic network based on the deep deterministic policy gradient reinforcement learning algorithm.
[0067] The critic network employs a dual-network structure, consisting of a main network and a target network. The initial values of the target network parameters are directly copied from the main network parameters.
[0068] Specifically as follows:
[0069] Step 201, starting from the current inner loop In the next time slot, the interval In each time slot, several training samples are randomly sampled to update the parameters of the main critic network and the target critic network;
[0070] Step I, training samples middle As a state, the input to the main critic network outputs the Q value. .
[0071] Specifically:
[0072] Step 2.1, for each sample middle The queue backlog and equivalent channel in the input are scaled and then fed into the main critic network.
[0073] Step 2.2, using Scaled queue backlog and equivalent channel and each sample Actions in Construct the initial hidden representations on the edges of the main critic network;
[0074] The critic network uses the same GNN as the actor network, and the initial hidden representation is:
[0075] First, let's define the connection data vertices. With user apex The initial hidden edge is represented as a vector. ;
[0076] Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog as The third dimension value, if and only if When the queue is full, the third dimension value is queue backlog; otherwise, it is 0.
[0077] Next, it will be assigned to the user. power As a vector The fourth dimension value, if and only if At that time, the value of the fourth dimension is Otherwise, it is 0;
[0078] Step 2.3: Using the constructed initial hidden representation, the main critic network updates the hidden representation on all edges using the same update formula as the main actor network.
[0079] Step 2.4: Merge the hidden representations on all edges updated in the last layer to obtain... Multidimensional matrix , For the hidden layer dimension;
[0080] Step 2.5, Pooling is performed on the first two dimensions, and through The readout layer of the dimension obtains the Q-value of the main critic network output. .
[0081] Step II: Calculate the target value using the target critic network and the target actor network. ;
[0082] Specifically:
[0083] First, the sample In Input the target actor network to obtain the estimated action ;
[0084] Then, take each sample The estimated action Input the target critic network to obtain the estimated Q value. ;
[0085] Next, the target value is calculated according to the Bellman equation:
[0086] ;
[0087] This indicates that the average is calculated over all sampled training samples; Discount factor
[0088] Step III, Calculate the target value With the output of the main critic network The mean squared error between the two values is used as the loss function for training the main critic network, thereby updating the parameters of the main critic network. ;
[0089] The main critic network is trained to minimize the mean squared error between its output and the target value, and its parameters are updated accordingly. The mean squared error is the loss function, calculated as follows:
[0090]
[0091] Step IV: After updating the main critic network parameters, update the target critic network parameters using a soft update method. ;
[0092]
[0093] in , where is the soft update coefficient;
[0094] Step 202: Starting from the current inner loop... In the next time slot, the interval In each time slot, several training samples are randomly sampled again to update the parameters of the main actor network and the main critic network, and then the corresponding target networks are updated.
[0095] Specifically:
[0096] First, the main critic network is updated again using the new training samples. and the parameters of the target critic network ;
[0097] Then, the training samples In The power allocation vector is obtained by inputting the main actor network. ;
[0098] Next, and Input the main critic network, train the actor network to maximize the corresponding Q-value. And update its parameters, with the negative value of Q being the loss function, calculated as follows;
[0099]
[0100] Finally, each time the main actor network parameters are updated... Then, update the target actor network parameters using a soft update method. :
[0101]
[0102] Step 3: Return to Step 1 to perform the next inner loop, until the set number of iterations is reached. End the current phase of the loop.
[0103] Step 4: Utilize the current stage The average energy efficiency is calculated in the second internal cycle, and the parameters are updated using a moving average. ;
[0104] The average energy efficiency of each internal cycle is calculated as follows:
[0105]
[0106] in and The first The inner loop number Throughput and energy consumption per time slot;
[0107] The Dinkelbach parameters are updated and calculated as follows:
[0108]
[0109] in It is a smoothing factor;
[0110] Step 5: Utilize the updated parameters Proceed to the next stage of the cycle, that is, repeat. The cycle continues until the average energy efficiency converges.
[0111] Step 6: Using the trained master actor network, perform power allocation based on the equivalent channel and user queue backlog for each time slot, and calculate the total throughput across multiple time slots. Energy consumption Then, energy efficiency can be calculated.
[0112] The calculation method is as follows:
[0113]
[0114] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0115] (1) This invention introduces deep reinforcement learning and sets up an inner and outer loop training method, taking into account the impact of changes in queue state between time slots on decision-making, enabling the actor network to make decisions based on user channel state and traffic queue backlog information, thereby improving the long-term energy efficiency performance of the system.
[0116] (2) In the application stage, the actor network trained by this invention can directly output power allocation decisions based on the current channel and queue states, without performing multiple iterative calculations in the traditional fractional programming method, which significantly reduces the computational complexity and processing latency of online decision-making. Since the influence of time slot states on decision-making is considered, the resulting long-term energy efficiency is also higher than that of existing methods. Attached Figure Description
[0117] Figure 1 This is a flowchart of a high-efficiency power allocation method for dynamic traffic based on deep reinforcement learning, according to the present invention.
[0118] Figure 2 This is a comparison chart of the long-term energy efficiency and queue stability satisfaction rate of the present invention and existing methods in the test sample under the condition of changing user data packet request rate (reflecting traffic load).
[0119] Figure 3 This is a comparison chart of the long-term energy efficiency and queue stability satisfaction rate of the present invention and existing methods in the test sample under the condition of varying number of users served by the base station. Detailed Implementation
[0120] The embodiments of the present invention will now be described in complete and detailed manner with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0121] like Figure 1 As shown, the specific steps are as follows:
[0122] Step 1: In the downlink transmission scenario of a multi-antenna multi-user system, enable the actor network to interact with the transmission scenario. Each time slot is used as one inner loop. Each interactive time slot generates an empirical quadruple and saves all quadruples as training samples.
[0123] In the downlink transmission scenario described above, a base station is located within a cell and is equipped with... A uniform linear antenna array serves the cell. One user; each user has a single antenna;
[0124] Step 101, in the time slot The base station collects data on channel status, traffic queues, and new requests from all users within the cell. The values are 1, 2, ... ;
[0125] In the time slot , The channel state matrix from each user to the base station antenna is as follows: ;
[0126] for A dimensional vector representing a user to base station The channel vector of the root antenna.
[0127] The current traffic queue for each user is ;
[0128] No. Queue information for each user Indicates the cutoff time slot ,user Data that has been requested but not yet served;
[0129] The data of the current new requests from each user is ;
[0130] , For time slots user Number of requested data packets, The number of bits contained in the data packet;
[0131] Step 102: The base station, according to the time slot The channel states of all users are used to pre-encode the transmitted data to obtain the equivalent channel between users. And calculate the queue backlog based on the user traffic queue and the data of new requests. ;
[0132] Equivalent Channel The element for , ; For users The normalized precoding vector, ; Among them, users queue backlog ;
[0133] Step 103: Utilize the equivalent channel With queue backlog Model the graph and resolve the queue backlog. Zhang Cheng's diagonal matrix and equivalent channel The splicing yields a multidimensional matrix. ;
[0134] Base station sends to The data streams of each user are pre-encoded, therefore there is There are 10 data vertices; simultaneously, each user receives data, therefore, there exists One user vertex; forming two types of vertices in the graph;
[0135] Queue backlog Convert to a diagonal matrix, where the diagonal elements represent the queue backlog of each user, and the off-diagonal elements are zero. Then, combine this diagonal matrix with the equivalent channel... The splicing yields a multidimensional matrix. This matrix is used as the edge feature of the graph, with its first two dimensions being indices. The resulting slice is the data vertex. With user apex Edge features between them ;
[0136] Step 104: Convert the multidimensional matrix As a state, input the actor network to obtain actions. ;
[0137] In other words, a time slot Power allocation vector for all users The actor network is a GNN that incorporates an attention mechanism, learning a power allocation vector based on the input state. ;
[0138]
[0139] in, For users The power allocation vector;
[0140] The actor network adopts a dual-network structure, corresponding to a main network and a target network. The initial values of the target network parameters are obtained by directly copying the parameters of the main network.
[0141] The specific steps are as follows:
[0142] Step 4.1, for the multidimensional matrix The queue backlog and equivalent channel in the input are scaled and then fed into the main actor network.
[0143] First, calculate the average amount of data M to be transmitted per time slot based on the user's average packet request rate (packets per second);
[0144] Then, the multidimensional matrix The queue backlog is divided by the data size M to obtain the scaled queue backlog.
[0145] Next, the average large-scale fading value is calculated based on the channel state collected by the base station. The reciprocal of the large-scale fading is multiplied by the equivalent channel, and the result is used as the scaled equivalent channel.
[0146] Step 4.2: Using the scaled queue backlog and equivalent channel, construct the initial hidden representation on the edges of the main actor network;
[0147] First, regarding the connection data vertices With user apex The edges are initially hidden as vectors. ;
[0148] Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog is used as a vector The third dimension value, if and only if When the queue is full, the value of the third dimension is queue backlog; otherwise, the value of the third dimension is 0.
[0149] Step 4.3: Update the hidden representations on all edges of the main actor network using the initial hidden representations;
[0150] The updated formula is as follows:
[0151]
[0152]
[0153]
[0154]
[0155] in, The number of layers in the GNN used by the actor network. For the first After the layer update, the first Hidden representation on each edge; , , , and For different training weights, It is an activation function. Apply LayerNorm to the input;
[0156] Step 4.4: Merge the hidden representations on all edges updated in the last layer of the main actor network to obtain a multidimensional matrix. ;
[0157] Multidimensional matrix for dimension, To hide the dimensions of the representation;
[0158] Step 4.5, for the multidimensional matrix Pooling is performed on the second dimension, and through... The readout layer of the dimension is obtained using Softmax. Normalizing the dimensional vector yields ;
[0159] Step 4.6, convert the multidimensional matrix Pooling is performed on the first two dimensions, and then through... The readout layer is dimensional, and the range of values is obtained through a scaling function. scaling factor between ;
[0160] Step 4.7, based on the base station's maximum transmit power Using normalized vectors and scaling factor Calculate the power allocation vector output by the main actor network:
[0161]
[0162] Step 105: Based on the current action Calculate time slots Reward value ;
[0163] Specifically, it includes:
[0164] Step 5.1, according to the action Equivalent channel Calculate users Transmittable data volume :
[0165] ,
[0166] in For bandwidth, For the duration of a single time slot, and Actions Chinese users Power of other users power, The equivalent channel representing the useful signal. Indicates user For users The equivalent interference channel; Noise power;
[0167] Step 5.2, according to the user queue backlog With the amount of data that can be transmitted Calculate the actual amount of data served. :
[0168]
[0169] Step 5.3: Calculate the current timeslot using the actual data volume of all users' services. System throughput And further combined with energy consumption Calculate time slots Energy efficiency bonus ;
[0170] ;
[0171] ;
[0172] Among them, energy consumption The value is the product of the total power consumption of the base station and the duration of a single timeslot; For Dinkelbach parameters;
[0173] Step 5.4, utilizing time slots Energy efficiency bonus Calculate the queue penalty term as follows and overall reward value ;
[0174] ;
[0175] ;
[0176] in As weighting coefficients, energy efficiency rewards and queue penalties need to be normalized separately;
[0177] Step 106, Proceed to the next time slot +1, the base station collects the channel status of all users. And calculate the equivalent channel. and queue backlog The next state is obtained by splicing. ;
[0178] According to time slot The amount of data requested by the user Calculate the updated queue backlog ;
[0179] Among them, users queue information The calculation method is as follows:
[0180]
[0181] Step 107: Divide the time slot The generated quadruple As a training sample, it is stored in the Replay buffer array;
[0182] State After inputting the actor network, the output is the action. After interacting with the wireless environment, it enters the next state. , including the next time slot Equivalent channel and updated queue; actor network output action The current time slot also needs to be calculated. Reward value Then the tuple The tuple is stored as a training sample in the Replay Buffer. One lesson learned from the interaction between the actor network and the environment;
[0183] Step 108: Convert the multidimensional matrix As a state, input the main actor network to obtain actions. And further calculate the reward value. Update the next time slot The state of +2 yields a quadruple. The training sample is then stored in the replay buffer, and the loop is repeated until the time slot is reached. This terminates the current inner loop.
[0184] Step 2: In the current inner loop, using the training samples, train the actor network and the critic network based on the deep deterministic policy gradient reinforcement learning algorithm.
[0185] The actor network and critic network are trained using inner loops, with each inner loop's actor network interacting with the wireless environment. In each time slot, while the actor network interacts with the environment, a certain number of training samples are continuously sampled from the Replay Buffer. The actor network and the critic network are trained based on a deep deterministic policy gradient reinforcement learning algorithm. The actor network then adjusts its training based on the state of the samples. Learning power allocation and the status With action Input the critic network; the critic network adjusts its state accordingly. Corresponding actions Learn the Q-value, which is used to evaluate the state. Next action The quality;
[0186] Similar to the actor network, the critic network is also a GNN that incorporates an attention mechanism. The critic network adopts a dual-network structure, consisting of a main network and a target network. The initial values of the target network parameters are obtained by directly copying the parameters of the main network.
[0187] Specifically as follows:
[0188] Step 201, starting from the current inner loop In the next time slot, the interval In each time slot, several training samples are randomly sampled to update the parameters of the main critic network and the target critic network;
[0189] Step I, training samples middle As a state, the input to the main critic network outputs the Q value. .
[0190] Step 2.1, for each sample middle The queue backlog and equivalent channel in the input are scaled and then fed into the main critic network.
[0191] Step 2.2, using Scaled queue backlog and equivalent channel and each sample Actions in Construct the initial hidden representations on the edges of the main critic network;
[0192] The critic network uses the same GNN as the actor network, and the initial hidden representation is:
[0193] First, let's define the connection data vertices. With user apex The initial hidden edge is represented as a vector. ;
[0194] Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog as The third dimension value, if and only if When the queue is full, the third dimension value is queue backlog; otherwise, it is 0.
[0195] Next, it will be assigned to the user. power As a vector The fourth dimension value, if and only if At that time, the value of the fourth dimension is Otherwise, it is 0;
[0196] Step 2.3: Using the constructed initial hidden representation, the main critic network updates the hidden representations on all edges;
[0197] The update formula for each layer of the critic network is the same as the update formula for each layer of the main actor network.
[0198] Step 2.4: Merge the hidden representations on all edges updated in the last layer to obtain... Multidimensional matrix , For the hidden layer dimension;
[0199] Step 2.5, Pooling is performed on the first two dimensions, and through The readout layer of the dimension obtains the Q-value of the main critic network output. .
[0200] Step II: Calculate the target value using the target critic network and the target actor network. ;
[0201] First, each sample In Follow steps 4.1-4.7 to input the target actor network to obtain the estimated action. ;
[0202] Then, take each sample The estimated action Follow steps 8.1-8.5 to input the target critic network and obtain the estimated Q value. ;
[0203] Next, the target value is calculated according to the Bellman equation:
[0204] ;
[0205] in The target actor network input is the state. The action output at that time Input to the target critic network is The Q value output at that time. This indicates that the average is calculated over all sampled training samples; Discount factor;
[0206] Step III, Calculate the target value The mean squared error between the main critic network's output Q-value and the output Q-value is used as the loss function to train the main critic network, thereby updating the main critic network's parameters. ;
[0207] The main critic network is trained to minimize the mean squared error between its output and the target value, and its parameters are updated accordingly. The mean squared error is the loss function, calculated as follows:
[0208]
[0209] Step IV: After updating the main critic network parameters, update the target critic network using a soft update method. ;
[0210]
[0211] in , where is the soft update coefficient.
[0212] Step 202: Starting from the current inner loop... In the next time slot, the interval In each time slot, several training samples are randomly sampled again. The parameters of the main actor network and the main critic network are updated using the sampled training samples. After updating the main actor network and the main critic network, their corresponding target networks are updated.
[0213] Specifically:
[0214] First, using the method described in step 201, update the main critic network again with the new training samples. and the parameters of the target critic network ;
[0215] Then, each sampled training sample In Follow steps 4.1-4.7 to input the power allocation vector into the main actor network.
[0216] Next, each training sample In and Input the main critic network to obtain the corresponding Q value. Then, the main actor network is trained to maximize the estimate of the main critic network. The loss function used to update the main actor network is as follows:
[0217]
[0218] Finally, each time the main actor network is updated... Then, update the target actor network parameters using a soft update method. ;
[0219]
[0220] Step 3: After the current inner loop finishes, return to Step 1 to start the next inner loop, until the set time is reached. Next, the current phase of the loop ends.
[0221] Step 4: Utilize the current stage The average energy efficiency is calculated in the second internal cycle, and the parameters are updated using a moving average. This is used to calculate the reward value for the next stage of the cycle;
[0222] An outer loop is used to continuously update the Dinkelbach parameters based on Dinkelbach fractional programming. ,Every After each inner loop completes, the statistics are... The average throughput and energy consumption of the actor network interacting with the environment within each inner loop are calculated, energy efficiency is determined, and updates are performed. Used to calculate reward values when the actor network interacts with the environment. ; The average energy efficiency of each internal cycle is calculated as follows:
[0223]
[0224] in and The first The inner loop number Throughput and energy consumption per time slot;
[0225] The Dinkelbach parameters are updated and calculated as follows:
[0226]
[0227] in It is a smoothing factor;
[0228] Step 5: Utilize the updated parameters Proceed to the next stage of the cycle, that is, repeat. The cycle continues until the average energy efficiency converges.
[0229] Step Six: During the testing phase, the trained master actor network performs power allocation inference based on the equivalent channel and user queue backlog for each time slot, and calculates the total throughput across multiple time slots. Energy consumption Then, energy efficiency can be calculated.
[0230] The calculation method is as follows:
[0231]
[0232] Example:
[0233] The embodiments of the present invention include the following eight steps:
[0234] Step 1: Collect data on channels, traffic queues, and new requests for each user within the cell at the base station.
[0235] The examples are generated according to a first-order autoregressive model. Small-scale time-varying channel for individual users , No. The first time slot The small-scale channel for each user is:
[0236]
[0237] in For random noise, the mean is 0 and the variance is... Generated by a complex Gaussian distribution, The channel correlation coefficient. For users Doppler shift, It is a zeroth-order Bessel function of the first kind.
[0238] Generated according to the 3GPP NLOS urban macro model Large-scale fading factor of individual user channels The corresponding path loss, and the path loss model is as follows:
[0239]
[0240] in The distance between the user and the base station, carrier frequency GHz.
[0241] user The channel between the base station and the station is:
[0242]
[0243] The base station and user heights are 25 m and 1.5 m respectively, the cell diameter is 500 m, and the cell edge signal-to-noise ratio is 10 dB.
[0244] In this embodiment of the invention, the system simulation conditions are as follows: , Duration of a single time slot The time is 1ms. The user data packet request rate is... Packets per second. The bandwidth normalized size of the data packets is 0.025 bits / Hz, bandwidth. kHz, per data packet The size is 375 bits. Each time slot Number of requested packets Following a Poisson distribution, the requested intervals should follow an exponential distribution. (User) The data requested is The queue is .
[0245] Step two: The base station performs precoding based on the channel state information to obtain the equivalent channel between users. And calculate the queue backlog based on the user traffic queue and the data of new requests. ;
[0246] In this embodiment, the normalized precoding vector is calculated using regularized zero-forcing precoding.
[0247] Step 3, based on the precoded equivalent channel With queue backlog Modeling a graph with users as vertices; managing queue backlogs. Convert to a diagonal matrix, where the diagonal elements represent the queue backlog of each user, and the off-diagonal elements are zero. Then, combine this diagonal matrix with the equivalent channel... The splicing yields a multidimensional matrix. This matrix is used as the edge feature of the graph, with its first two dimensions being indices. The resulting slice is the vertex. With vertex Edge features between them ;
[0248] Step four, put The state is input into the actor network to obtain the action, i.e., the power allocation vector. ;
[0249] In this embodiment, the actor network is updated three times, meaning it has three layers. The dimensions of the hidden representation change from the initial 3 to 32, 32, and 32 respectively. Layer number l=0 represents the initial hidden representation of the input.
[0250] Combining the hidden representations of all edges updated in the last layer of the actor network yields a... Multidimensional matrix ,right Pooling is performed on the second dimension, and then through a... The readout layer of the dimension is obtained using Softmax. Normalizing the dimensional vector yields ;Will Pooling is performed on the first two dimensions, and then through another... The readout layer is dimensional, and the range of values is obtained through a scaling function. scaling factor between ;set up Given the base station's maximum transmit power, the power allocation vector output by the actor network is: In this embodiment, the maximum transmit power of the base station is set to 10 W.
[0251] Step 5, convert the state space After inputting an actor network, the actor network outputs actions. After interacting with the wireless environment, it enters the next state. , including the next time slot Equivalent channel and updated queue; actor network output action After interacting with the environment, the current time slot also needs to be calculated. Reward value Then the tuple Storing tuples as training samples in the Replay Buffer allows them to be processed. This can be considered as an experience gained from the interaction between the actor network and the environment.
[0252] The Replay Buffer is a buffer array used to store experiences generated by the interaction between the actor network and the environment. When the number of experience groups exceeds the capacity of the Replay Buffer, the oldest experience is deleted and the newly generated experience is added. In this embodiment, the capacity of the Replay Buffer is set to 200,000.
[0253] in This is the Dinkelbach parameter, initially set to 200.
[0254] In this embodiment, the weighting coefficient Energy efficiency rewards and queue penalties were normalized using the Welford algorithm for online normalization.
[0255] The base station power consumption model used in this embodiment is:
[0256]
[0257] in, This represents the total power consumption of the base station. For transmission power, The circuit power consumption of each antenna, For signal processing power consumption, For the fixed power consumption required by the base station control signaling and processor, For local oscillator power consumption, To reflect the power amplifier efficiency, power supply, and cooling coefficients, the power consumption of each antenna circuit, signal processing power consumption, fixed power consumption, and local oscillator power consumption are 1.0 W, 7.54 W, 10 W, and 0.2 W, respectively. Take 4.7. Energy consumption. The calculation method is as follows:
[0258]
[0259] Step 6: Train the actor network and critic network in the inner loop;
[0260] In each inner loop, the actor network interacts with the wireless environment. In each time slot, while the actor network interacts with the environment, a certain number of training samples are continuously sampled from the Replay Buffer to train the actor network and the critic network based on a deep deterministic policy gradient reinforcement learning algorithm. Each time, 128 samples are sampled.
[0261] The critic network is also a GNN that incorporates an attention mechanism. Both the actor network and the critic network correspond to a main network and a target network, respectively.
[0262] In this embodiment of the invention, the critic network is updated three times, i.e., it has three layers. The dimensions of the hidden representation change from the initial 4 to 32, 32, and 32 respectively. Layer number l=0 represents the initial hidden representation of the input. The hidden representations of all edges updated in the last layer are combined to obtain a single hidden representation. Multidimensional matrix ,Will The first two dimensions are pooled, and then through a... The readout layer of the dimension obtains the output Q value.
[0263] This embodiment uses a dual-delay deep deterministic policy gradient algorithm to train the network, employing two critic networks to learn the Q-value and one actor network to learn the power allocation decision.
[0264] set up ( ) represent the Q-values output by the main critic network and the target critic network, respectively. Each... Each time slot samples 128 samples from the Replay Buffer. For each sampled sample, the target value is calculated using the smaller Q-value from the outputs of the two target critic networks, as follows:
[0265]
[0266] in As a discount factor, For the target actor network, the input state is Output power.
[0267] Train the main critic network to minimize the mean squared error between its output and the target value:
[0268]
[0269] in For sampling, Indicates the number of samples. It is a set of experiences in the sampled data.
[0270] Every The actor network is trained in each time slot to maximize the Q-value estimated by one of the main critic networks, using the following loss function:
[0271]
[0272] After each update of the main actor network and the main critic network, all target network parameters are updated using a soft update method, as follows:
[0273]
[0274] in , The parameters representing the target network, The parameters of the main network.
[0275] In this embodiment, the actor network in each inner loop interacts with the wireless environment. Each time slot. To calculate the energy consumption and throughput of each time slot, first estimate the number of time slots required for the queue to reach stability. The estimation formula is as follows:
[0276]
[0277] in (Packets per second) represents the average request rate of user data packets. Take 10, Take 1.2. When the number of time slots is greater than After reaching 500, the energy consumption and throughput of each time slot are calculated.
[0278] Step seven, repeat. After the inner loop of this phase ends, the next inner loop begins, and the outer loop updates the Dinkelbach parameters. ;
[0279] In this embodiment, each After the inner loop finishes, the next inner loop begins.
[0280] Step 8: After the average energy efficiency converges, stop the cycle and enter the testing phase.
[0281] During the testing phase, power allocation decisions were inferred solely from the state of each time slot using the trained actor network. Queue stability was evaluated as the ratio of the system's average throughput per second to the average data request rate per user.
[0282] In this embodiment of the invention, the Adam algorithm is used to optimize the neural network weights during the training phase, with both the actor and critic networks having a learning rate of 0.0005. Each inner loop contains channels with 1000 time slots. At the beginning of each inner loop, the queue is initialized, assuming an average user service rate... Then calculate the flow intensity. The initial queue backlog is .
[0283] In this embodiment of the invention, 100 sets of samples are used to test the actor network during the testing phase. Each set of samples includes channels with 1000 time slots. The queue initialization method for each set of samples is the same as described above. When the number of time slots is greater than... After reaching 500, the system begins to calculate the energy consumption and throughput of each time slot. After testing each set of samples, the total throughput and total energy consumption are calculated, and then the system energy efficiency is calculated.
[0284] In this embodiment of the invention, the above-mentioned GNN and training and testing framework are built using Python 3.7.10 and PyTorch 1.9.0.
[0285] The neural network trained according to the embodiment is compared with the performance of other commonly used methods on the test set. The performance includes long-term energy efficiency and queue stability satisfaction rate. The queue stability satisfaction rate is the ratio between the average data rate of the system serving users and the data rate of user requests. When the queue is stable, the satisfaction rate should be about 100%.
[0286] The comparison results are as follows Figure 2 and 3 As shown, the comparison schemes include:
[0287] Option 1: Use a graph neural network without an attention mechanism as both the actor and critic networks. Its structure is implemented by parameter sharing (reference [6]).
[0288] Option 2: A numerical optimization algorithm using Lyapunov drift theory (reference [7]). Its power allocation decision is obtained by solving a convex optimization problem.
[0289] Option 3: Use equal power allocation decision for users. When the queue backlog is 0, set the user power to zero.
[0290] Option 4: Use user-equal power allocation decision-making to transmit at full power.
[0291] Reference [6]: B. Zhao, J. Guo, and C. Yang, “Understanding thePerformance of Learning Precoding Policies With Graph and ConvolutionalNeural Networks,” IEEE Trans. Commun., vol. 72, no. 9, pp. 5657–5673, Sept.2024.
[0292] Reference [7]: D. Bethanabhotla, G. Caire, and MJ Neely, “Wiflix: Adaptive video streaming in massive MUMIMO wireless networks,” IEEE Trans.Wireless Commun., vol. 15, no. 6, pp. 4088–4103, June 2016.
[0293] like Figure 2 As shown, when user data packet request rates vary, the long-term energy efficiency achieved by the power allocation decision determined according to this invention exceeds that of schemes one to four. Furthermore, the power allocation decision determined by this invention enables queue stability.
[0294] like Figure 3 As shown, when the number of users served by the base station varies, the long-term energy efficiency achieved by the power allocation decision determined according to this invention exceeds that of schemes one to four. Furthermore, the power allocation decision determined by this invention enables queue stability.
[0295] Because the trainable weights of the GNN with the introduced attention mechanism are shared across all edges, the trained neural network can be directly applied to systems with different numbers of users without retraining. Furthermore, during testing, the power allocation decision determined by this invention is directly obtained by the actor network based on the channel and queue backlog status, requiring significantly less computation time than Scheme Two.
[0296] The above provides a detailed description of the high-efficiency power allocation method for dynamic traffic based on deep reinforcement learning proposed in this invention, and elucidates the principles and implementation methods of this invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A high-efficiency power allocation method for dynamic traffic based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: In the downlink transmission scenario of a multi-antenna multi-user system, enable the actor network to interact with the transmission scenario. Each time slot is used as one inner loop. Each interactive time slot generates an empirical quadruple and saves all quadruples as training samples. Step 2: In the current inner loop, using the training samples, train the actor network and the critic network based on the deep deterministic policy gradient reinforcement learning algorithm. Both the actor network and the critic network adopt a dual-network structure, each including a main network and a target network. The initial values of the target network parameters are directly copied from the main network parameters. Specifically, this includes: from the current inner loop. In the next time slot, the interval In each time slot, several training samples are randomly sampled to update the parameters of the main critic network and the target critic network; at intervals... In each time slot, several training samples are randomly sampled again to update the parameters of the main actor network and the main critic network, and then the corresponding target networks are updated. Step 3: Return to Step 1 to perform the next inner loop, until the set number of iterations is reached. End the current phase of the loop; Step 4: Utilize the current stage The average energy efficiency is calculated in the second internal cycle, and the parameters are updated using a moving average. ; The average energy efficiency of each internal cycle is calculated as follows: in and The first The inner loop number Throughput and energy consumption per time slot; The Dinkelbach parameters are updated and calculated as follows: in It is a smoothing factor; Step 5: Utilize the updated parameters Proceed to the next stage of the cycle, that is, repeat. The cycle continues until the average energy efficiency converges. Step 6: Using the trained master actor network, perform power allocation based on the equivalent channel and user queue backlog for each time slot, and calculate the total throughput across multiple time slots. Energy consumption Then, energy efficiency is calculated; The calculation method is as follows: .
2. The method as described in claim 1, characterized in that, Step one specifically involves: Step 101, in the time slot The base station collects channel status, traffic queues, and new request data from all users within the cell, and precodes this data to obtain the equivalent channel between users. Simultaneously calculate queue backlog ; Step 102: Utilize the equivalent channel With queue backlog Model the graph and resolve the queue backlog. Zhang Cheng's diagonal matrix and equivalent channel The splicing yields a multidimensional matrix. ; The vertices in the modeling graph include Data vertices and Individual user vertices; multidimensional matrix As an edge feature of the graph, its first two dimensions of index The resulting slice is the data vertex. With user apex Edge features between them ; Step 103: Convert the multidimensional matrix As a state, input the actor network to obtain actions. And further calculate the time slots Reward value ; Step 104, Proceed to the next time slot +1, the base station collects the channel status of all users. And calculate the equivalent channel. and queue backlog The next state is obtained by splicing. ; Step 105: Divide the time slot The generated quadruple It is stored as a training sample in the Replay buffer array; Step 106: Convert the multidimensional matrix As a state, input the main actor network to obtain actions. And further calculate the reward value. Update the next time slot The state of +2 yields a quadruple. The training sample is then stored in the replay buffer, and the loop is repeated until the time slot is reached. This terminates the current inner loop.
3. The method as described in claim 2, characterized in that, The channel state matrix Traffic queue ; Newly requested data ; Equivalent Channel The element for , ; For users Normalized precoding vector; Among them, users queue backlog .
4. The method as described in claim 3, characterized in that, The action is obtained in step 103. The details are as follows: Step 4.1, for the multidimensional matrix The queue backlog and equivalent channel in the input are scaled and then fed into the main actor network. Step 4.2: Using the scaled queue backlog and equivalent channel, construct the initial hidden representation on the edges of the main actor network; First, regarding the connection data vertices With user apex The edges are initially hidden as vectors. ; Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog is used as a vector The third dimension value, if and only if When the queue is full, the value of the third dimension is queue backlog; otherwise, the value of the third dimension is 0. Step 4.3: Update the hidden representations on all edges of the main actor network using the initial hidden representations; The updated formula is as follows: in, The number of layers in the GNN used by the actor network. For the first After the layer update, the first Hidden representation on each edge; , , , and For different training weights, It is an activation function. Apply LayerNorm to the input; Step 4.4: Merge the hidden representations on all edges updated in the last layer to obtain... Multidimensional matrix ; To hide the dimensions of the representation; Step 4.5, for the multidimensional matrix Pooling is performed on the second dimension, and through... The readout layer of the dimension is obtained using Softmax. Normalizing the dimensional vector yields ; Step 4.6, convert the multidimensional matrix Pooling is performed on the first two dimensions, and then through... The readout layer is dimensional, and the range of values is obtained through a scaling function. scaling factor between ; Step 4.7, based on the base station's maximum transmit power Using normalized vectors and scaling factor Calculate the power allocation vector output by the main actor network: .
5. The method as described in claim 2, characterized in that, Calculate the time slot in step 103 Reward value Specifically as follows: Step 5.1, according to the action Equivalent channel Calculate users Transmittable data volume : , in For bandwidth, The duration of a single time slot, and Actions Chinese users Power of other users power, The equivalent channel representing the useful signal. Indicates user For users The equivalent interference channel; Noise power; Step 5.2, according to the user queue backlog With the amount of data that can be transmitted Calculate the actual amount of data served. : Step 5.3: Calculate the current timeslot using the actual data volume of all users' services. System throughput And further combined with energy consumption Calculate time slots Energy efficiency bonus ; ; ; Among them, energy consumption The value is the product of the total power consumption of the base station and the duration of a single timeslot; For Dinkelbach parameters; Step 5.4, utilizing time slots Energy efficiency bonus Calculate the queue penalty term as follows and overall reward value ; ; ; in These are the weighting coefficients.
6. The method as described in claim 2, characterized in that, In step 104, according to the time slot The amount of data requested by the user Calculate the updated queue backlog ; Among them, users queue information The calculation method is as follows: 。 7. The method as described in claim 2, characterized in that, Step two is as follows: Step I, training samples middle As a state, the input to the main critic network outputs the Q value. ; Step II: Calculate the target value using the target critic network and the target actor network. ; Specifically: First, the sample In Input the target actor network to obtain the estimated action ; Then, take each sample The estimated action Input the target critic network to obtain the estimated Q value. ; Next, the target value is calculated according to the Bellman equation: ; This indicates that the average is calculated over all sampled training samples; Discount factor; Step III, Calculate the target value With the output of the main critic network The mean squared error between the two values is used as the loss function for training the main critic network, thereby updating the parameters of the main critic network. ; The main critic network is trained to minimize the mean squared error between its output and the target value, and its parameters are updated accordingly. The mean squared error is the loss function, calculated as follows: Step IV: After updating the main critic network parameters, update the target critic network parameters using a soft update method. ; in , where is the soft update coefficient.
8. The method as described in claim 7, characterized in that, Step I specifically involves: Step 2.1, for each sample middle The queue backlog and equivalent channel in the input are scaled and then fed into the main critic network. Step 2.2, using Scaled queue backlog and equivalent channel and each sample Actions in Construct the initial hidden representations on the edges of the main critic network; The critic network uses the same GNN as the actor network, and the initial hidden representation is: First, let's define the connection data vertices. With user apex The initial hidden edge is represented as a vector. ; Then, the real and imaginary parts of the scaled equivalent channel corresponding to that edge are respectively used as vectors. The first two dimension values will determine the user The scaled queue backlog as The third dimension value, if and only if When the queue is full, the third dimension value is queue backlog; otherwise, it is 0. Next, it will be assigned to the user. power As a vector The fourth dimension value, if and only if At that time, the value of the fourth dimension is Otherwise, it is 0; Step 2.3: Using the constructed initial hidden representation, the main critic network updates the hidden representation on all edges using the same update formula as the main actor network. Step 2.4: Merge the hidden representations on all edges updated in the last layer to obtain... Multidimensional matrix , For the hidden layer dimension; Step 2.5, Pooling is performed on the first two dimensions, and through The readout layer of the dimension obtains the Q-value of the main critic network output. .
9. The method as described in claim 7, characterized in that, The interval In each time slot, the parameters of the main actor network and the main critic network are updated, thereby updating their respective target networks; specifically: First, the main critic network is updated again using the new training samples. and the parameters of the target critic network ; Then, the training samples In The power allocation vector is obtained by inputting the main actor network. ; Next, and Input the main critic network, train the actor network to maximize the corresponding Q-value. And update its parameters, with the negative value of Q being the loss function, calculated as follows; Finally, each time the main actor network parameters are updated... Then, update the target actor network parameters using a soft update method. : 。