Energy harvesting D2D communication system long-term energy efficiency optimization method and system based on reinforcement learning and Lyapunov optimization
By combining reinforcement learning with Lyapunov optimization, the non-stationarity of energy and task arrival in D2D energy harvesting communication systems is addressed, achieving long-term optimal energy efficiency and system stability in non-stationary environments. An adaptive dual-deep Q-network structure with Dropout mechanism is used to control the stability of the training network, improving the robustness and adaptability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEAT UNIV OF SCI & TECH
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-05
AI Technical Summary
Existing power control methods fail to effectively address the non-stationarity of energy and task arrival in energy harvesting D2D communication systems, leading to decreased energy utilization and system instability. Traditional reinforcement learning algorithms are prone to getting stuck in local optima or slow convergence when training in non-stationary environments, making it difficult to balance maximizing energy efficiency with queue stability.
A wireless communication system model with multiple device nodes is constructed using a reinforcement learning and Lyapunov optimization approach. Power control is performed through the training network, and the objective is optimized by combining the Lyapunov drift term to achieve optimal long-term energy efficiency and system stability of the energy harvesting D2D communication system. An adaptive dual-deep Q-network structure with Dropout mechanism is used to control the stability of the training network.
The system achieves optimal long-term energy efficiency and system stability in a non-stationary environment, exhibiting good robustness and adaptability. It can maintain high energy efficiency and system stability even under conditions where both energy and mission arrival are non-stationary.
Smart Images

Figure CN121985377A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power control and energy efficiency optimization technology, and more specifically to a long-term energy efficiency optimization method and system for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization. Background Technology
[0002] With the rapid development of the Internet of Things (IoT) and 5G / 6G communication technologies, device-to-device (D2D) communication has become an important technical means to improve spectrum utilization and reduce communication latency. Meanwhile, the introduction of energy harvesting (EH) technology enables terminal devices to obtain energy from environmental energy sources such as solar, wind, and radio frequency energy, providing a new direction for building green and low-carbon communication systems. However, EH-D2D systems often exhibit significant non-stationarity in actual deployments. Energy sources such as solar and wind power are affected by factors such as weather, obstruction, and time of day, while user service loads also change constantly with application scenarios. This dual non-stationarity causes traditional power control methods based on steady-state assumptions to experience significant performance degradation when the environment changes, easily leading to energy depletion, service backlog, and a lack of robustness required for engineering operation.
[0003] Existing power control methods are typically based on static energy models or short-term optimal strategies, which fail to effectively address the non-stationary characteristics of energy and task arrival, easily leading to decreased energy utilization or system instability. Meanwhile, traditional reinforcement learning algorithms are prone to getting stuck in local optima or experiencing slow convergence when training in non-stationary environments, making it difficult to balance maximizing energy efficiency with queue stability.
[0004] Therefore, how to propose a long-term energy efficiency optimization method and system for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization, and achieve optimal long-term energy efficiency and system stability in non-stationary environments, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a long-term energy efficiency optimization method and system for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization, so as to achieve optimal long-term energy efficiency and system stability of energy harvesting D2D communication systems under non-stationary environments. To achieve the above objectives, the present invention adopts the following technical solution: A long-term energy efficiency optimization method for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization includes: Construct a wireless communication system model containing multiple device nodes and define the system state; Input the current system state into the training network to output power control actions; After executing a power control action, the system receives environmental feedback as a reward and then transitions to the next state; Based on the current system state, power control actions, environmental feedback rewards, and the next state, construct the current interaction sample and store it in the experience memory bank. Randomly sample sample data from the experience memory bank to update the parameters of the training network. Calculate the Lyapunov drift term in the queue, introduce the drift term into the optimization objective, and update the optimized value after introduction to the environmental feedback reward of the training network; The iterative training process continues until the network converges, obtaining the optimal power control strategy, and performing joint optimization of long-term energy efficiency and system stability constraints.
[0006] Optionally, each of the device nodes can harvest energy from three energy sources: solar, wind, and radio frequency energy.
[0007] Optionally, the system status includes node energy status, data queue status, and channel condition information.
[0008] Optionally, the training network is an n-layer neural network with Dropout units set between the hidden layers. It takes the system state as input, outputs the action value of each action in the action set, and selects the action with the highest action value as the power control action.
[0009] Optionally, the action value is used to characterize the long-term cumulative return obtained by selecting a specific power action in a given state, and is defined as: ; in, System status, For the value of the action, , are the parameters of the neural network. For discount rate, Yes Approximate to .
[0010] Optionally, the Dropout unit adaptively perturbs the Q-value distribution by randomly masking neurons. The calculation is as follows: ; in, In order to perform the action, For random sampling value.
[0011] Optionally, the step of randomly sampling sample data from the experience memory bank to update the parameters of the training network includes: ; in, For predicted values, To define rewards based on actual observations for .
[0012] Optionally, it also includes: optimizing using a loss function that minimizes the difference between the predicted Q-value and the target Q-value, and synchronizing parameters through the target network. ; Parameter update: ; in, It's the learning rate. .
[0013] Optionally, the calculation of the Lyapunov drift term in the queue, introducing the drift term into the optimization objective, and updating the optimized value after introduction to the environmental feedback reward of the training network includes:
[0014] in, For a set of queues, For Lyapunov functions, For Lyapunov drift function, For the expectation, For virtual queues, For transmission rate, Energy consumption; Incorporating the drift term into the optimization objective In the middle, As a reward for training the network, long-term energy efficiency and queue stability are jointly controlled.
[0015] Optionally, a long-term energy efficiency optimization system for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization includes: System model building module: used to build a wireless communication system model containing multiple device nodes and define the system state; Power control action module: used to input the current system state into the training network to output power control actions; Environmental feedback reward module: After executing a power control action, the system obtains environmental feedback rewards and moves to the next state; The parameter update module is used to construct the current interaction sample based on the current system state, power control action, environmental feedback reward, and next state, and store it in the experience memory bank. It also randomly samples sample data from the experience memory bank to update the parameters of the training network. Lyapunov optimization module: used to calculate queue Lyapunov drift terms, introduce the drift terms into the optimization objective, and update the optimized value after introduction to the environment feedback reward of the training network; Joint optimization module: Used for iterative training until the network converges, to obtain the optimal power control strategy, and to perform joint optimization of long-term energy efficiency and system stability constraints.
[0016] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a long-term energy efficiency optimization method and system for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization, which has the following beneficial effects: This invention discloses a power control and long-term energy efficiency optimization method for energy harvesting-based device-to-device (EH-D2D) systems applicable to non-stationary energy environments. Firstly, it utilizes the Poisson process to determine the energy arrival time... and the mission arrives Modeling and establishing energy queues and data queue The dynamic evolution model; in terms of the channel, the time-varying channel gain is calculated based on path loss and Rayleigh fading characteristics. Subsequently, through Lyapunov optimization, a joint energy queue was established. Data queue and virtual queues Constructing Lyapunov functions Deriving the Lyapunov drift term using Lyapunov functions. The system's energy efficiency is optimized while maintaining system stability by embedding the reward function into the reinforcement learning mechanism. The reinforcement learning part employs an adaptive dual-deep Q-network structure based on the Dropout mechanism, using the state vector as input and achieving an adaptive balance between exploration and exploitation through randomization of the Q-value distribution. The agent dynamically adjusts the transmission power of the EH-D2D system based on the optimal action output by the training network, and achieves parameter update stability through the target network. Through multiple rounds of training iterations, the agent gradually obtains the optimal power control strategy, achieving long-term optimal energy efficiency and queue stability in an energy harvesting environment. This invention maintains high energy efficiency and system stability even under non-stationary conditions of energy and task arrival, demonstrating good robustness, adaptability, and engineering application value. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This invention provides a schematic flowchart of a long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization.
[0019] Figure 2 A schematic diagram comparing the cumulative distribution functions of different algorithms provided in this invention.
[0020] Figure 3 This is a schematic diagram comparing queue compression using different algorithms provided by the present invention.
[0021] Figure 4 The diagram illustrates a long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization, provided by this invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] This invention discloses a long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization, comprising: Construct a wireless communication system model containing multiple device nodes and define the system state; Input the current system state into the training network to output power control actions; After executing a power control action, the system receives environmental feedback as a reward and then transitions to the next state; Based on the current system state, power control actions, environmental feedback rewards, and the next state, construct the current interaction sample and store it in the experience memory bank. Randomly sample sample data from the experience memory bank to update the parameters of the training network. Calculate the Lyapunov drift term in the queue, introduce the drift term into the optimization objective, and update the optimized value after introduction to the environmental feedback reward of the training network; The iterative training process continues until the network converges, obtaining the optimal power control strategy, and performing joint optimization of long-term energy efficiency and system stability constraints.
[0024] In a specific implementation, a long-term energy efficiency optimization method for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization is provided for energy harvesting device-to-device communication environments where both energy arrival and service arrival have non-stationary characteristics. Figure 4 As shown, it specifically includes: Step 1: Construct a wireless communication system model containing multiple device nodes; each device node can harvest energy from three energy sources: solar, wind, and radio frequency energy; define the system state. This includes node energy status, data queue status, and channel condition information; Step 2: Regarding the current state By training the network to control output power The target network is used as an auxiliary evaluation network to stabilize the training process.
[0025] The training network is a 4-layer neural network, with Dropout units placed between the hidden layers to... As input, output the action value of each action in the action set. , and select The biggest action as ,in, These are the parameters of the neural network.
[0026] The action value is used to characterize the long-term cumulative return that can be obtained by choosing a specific power action in a given state, and is specifically defined as: ; in, For discount rate, Yes The target network structure is identical to the training network, and its parameters are approximations. Through periodic replication The target Q-value is obtained and used to calculate the target Q-value to stabilize the training process, as follows: The goal of training is to minimize the loss function. No target network, target value Depends on the optimization This leads to instability in the optimization process. A structure with consistent parameters but independent parameters is adopted. Calculate the target value and order Only periodically (e.g., every C steps) from Synchronization ensures that during gradient descent within step C... Maintaining a constant gradient allows the optimization algorithm to converge along a stable gradient direction, avoiding oscillations.
[0027] Step 3: Execute the power control action. Afterwards, the system receives environmental feedback rewards. and transition to the next state. ; Step 4: Transfer the current interaction sample ( , , , Stored in the experience memory bank for subsequent training; Step 5: Randomly sample data from the experience memory bank ( , , , The parameters of the trained network are updated using the following formula: ; Where, on the left side of the formula For pure prediction, the right side of the formula represents the reward, which is partially based on actual observations. Therefore, the right side is more credible, and the right side is defined as... The loss function that minimizes the difference between the predicted Q-value and the target Q-value is used for optimization, and parameter synchronization is performed through the target network to achieve training stability. The specific formula is as follows: ; Parameter update is ,in, It is the learning rate; .
[0028] Step 6: Calculate the queue Lyapunov drift. To characterize changes in system stability, it is specifically expressed as: ; ; ; Incorporating the drift term into the optimization objective In the middle, As a reward for training the network This enables joint control of long-term energy efficiency and queue stability.
[0029] Step 7: Repeat steps 2 to 6 until the network converges, obtain the optimal power control strategy, and achieve joint optimization of long-term energy efficiency and system stability constraints.
[0030] Furthermore, in step 1, the energy queue of the sending node is defined. It evolves dynamically according to the following relationship: ; in, This is the upper limit of the energy queue. For energy consumption, The energy collected.
[0031] Data queue of sending node Update according to the following relationships: ; in, For transmission rate, This is newly arrived data.
[0032] The channel gain is calculated using the path loss and Rayleigh fading model, and the specific formula is as follows: ; Based on a combination of energy, data, and channel information, the node state is defined as follows: .
[0033] Furthermore, in step 2, a Dropout mechanism is introduced into the training network to achieve adaptive perturbation of the Q-value distribution by randomly masking neurons, so as to balance exploration and exploitation. The calculation is as follows: ; in, In order to perform the action, For random sampling value.
[0034] Furthermore, by constructing a Lyapunov function comprising an energy queue, a data queue, and a virtual queue, and calculating a drift term, the reward of the reinforcement learning is adjusted in conjunction with the drift term, thereby achieving adaptive robust control over non-stationary environments. The specific process is as follows: First, construct the Lyapunov function. Deriving the Lyapunov drift term using Lyapunov functions This ensures system stability, and the specific calculations are as follows: ; ; ; ; ; in, For a set of queues, For Lyapunov functions, For Lyapunov drift function, For the expectation, Let B be the average power consumption, and let B be a finite constant, which satisfies: ; The above equation shows that the Lyapunov drift expression for all queues has an upper bound, so the stability of the system can be ensured by minimizing the drift.
[0035] Then, the Lyapunov drift term is embedded into the reinforcement learning reward function to achieve the goal of maintaining system stability while optimizing energy efficiency. The specific expression is as follows: ; ; Where T is the total length of the time slot. As a reward.
[0036] Finally, after embedding the Lyapunov drift term into the reward function of reinforcement learning, the agent continuously updates the Q-network parameters based on the reward signal and stores the experience samples in the experience memory bank for subsequent training.
[0037] Furthermore, in step 7, the agent iteratively updates the policy parameters during continuous state observation, action selection, environmental interaction, and network training, achieving convergence by minimizing the error between the predicted Q-value and the target Q-value. When the training process reaches stability, the optimal power control strategy is obtained, achieving long-term optimal energy efficiency and system stability guarantee for the energy harvesting device-to-device communication system in non-stationary environments.
[0038] In a specific implementation, a long-term energy efficiency optimization method for energy harvesting D2D communication systems based on reinforcement learning and Lyapunov optimization is provided, which aims to improve the robustness of EH-D2D systems under non-stationary energy environments. Figure 1 As shown, the specific steps include: Step 11: Consider a wireless communication system model containing multiple device nodes; define the system state. This includes node energy status, data queue status, and channel condition information; .
[0039] Step 22: Establish an energy arrival and task arrival model based on the Poisson distribution to describe the randomness of external energy input and data generation processes, and establish the dynamic evolution relationship between the energy queue and the data queue accordingly, as follows: ; ; ; in, and The Poisson parameters that govern energy arrival and mission arrival, respectively. This is the maximum capacity of the energy queue. For the energy consumed, This refers to the transmission rate.
[0040] Step 33: Calculate the time-varying channel gain between devices based on path loss and Rayleigh fading characteristics to provide dynamic channel information for subsequent power control and learning processes; ; in, This indicates the distance between the two devices in the EH-D2D pair. That is the path loss coefficient. This represents Rayleigh fading modeled as a zero-mean Gaussian random distribution.
[0041] Step 4: Combine the energy state, data queue state, and channel state to form the system state vector. As input for reinforcement learning, it is used to describe the real-time operating state of the system.
[0042] Step 5: Integrated Energy Queue Data queue and virtual queues A Lyapunov function is established to provide theoretical constraints for system stability, as shown below: The specific update is as follows ,in, Average power consumption The energy consumed.
[0043] .
[0044] Step 6: Derive the drift term using the Lyapunov function. This is then embedded into the reward function of reinforcement learning to achieve a joint balance between energy efficiency optimization and queue stability. The reward function is defined as: ; in, Indicates energy efficiency. This is the balance coefficient.
[0045] Step 7: The agent employs an adaptive dual-deep Q-network (LAD3) structure based on the Dropout mechanism, achieving stable convergence through alternating updates between the training and target networks. The agent selects a power action based on the current state. The network parameters are continuously optimized based on system feedback, as detailed below: ; in, This represents the Q value of a random sample.
[0046] Step 8: After multiple rounds of iterative training, the agent obtains the optimal power control strategy, achieving long-term optimal energy efficiency and stable queue control of the energy harvesting system.
[0047] In a specific embodiment, an EH-D2D communication scenario is used. The scenario includes one energy source and 10 EH-D2D pairs. Each EH-D2D pair has two D2D user equipments, one transmitting and one receiving, and all equipment has energy harvesting capabilities. The distance between each pair of equipment is 20 to 50 meters. The channel is characterized by both path fading and Rayleigh fading. This example uses a combination of Double Deep Q-Network (DDQN) and Dropout, setting the dropout rate to 0.5 and the noise power to... This includes the following steps: Step 21: Construct an EH-D2D system model under a non-stationary environment. Because energy and the mission arrival process vary over time and are uncertain, the system as a whole exhibits non-stationary characteristics. Step 22: Establish a dynamic model of energy arrival and task generation based on the Poisson distribution. The energy arrival process describes the random energy input in the external environment, while the task arrival process reflects the change in communication requirements over time. Step 23: Calculate the time-varying channel gain between each device based on the path loss model and Rayleigh fading characteristics. This model can describe the real-time changes in the wireless channel; Step 24: Combine the energy queue state, data queue state, and channel state to form the input state vector for reinforcement learning. This state vector reflects the complete state of the system in each time slot and is used to guide the agent's action selection; Step 25: Integrated Energy Queue Data queue and virtual queues The dynamic changes of these three factors are used to construct a Lyapunov function to quantify system stability. This function plays a role in balancing energy efficiency and queue constraints during the optimization process. Step 26: According to the Lyapunov function Derive the Lyapunov drift term of the system And embed the drift amount into the reward function of reinforcement learning. In this way, the intelligent agent can simultaneously focus on improving energy efficiency and maintaining system stability during the learning process; Step 27: Embed the Lyapunov drift term into the reinforcement learning reward function so that the reward value reflects both energy efficiency and system stability; Step 28: The agent continuously performs state observation, action selection, environmental interaction and network update in consecutive time slots. After multiple rounds of training iterations, it achieves convergence and obtains the optimal power control strategy, realizing the long-term optimal energy efficiency and system stability guarantee of the energy harvesting D2D communication system in non-stationary energy environments.
[0048] The results are as follows Figure 2 The figure shows the cumulative distribution function (CDF) performance of the four algorithms. As can be seen from the figure, the LAD3 algorithm has a significant advantage in maintaining energy efficiency, with a 95% probability of keeping the EE above 8000 bit / J, while the corresponding probabilities for DDQN, DDQN+Dropout, and ECLB algorithms are only 60%, 84%, and 75%, respectively. This indicates that LAD3 can achieve high-efficiency transmission more stably in the face of energy fluctuations and uneven load environments. The results show that LAD3 not only performs outstandingly in the high-end region of energy efficiency distribution but also exhibits highly consistent performance in multiple independent experiments, demonstrating its excellent robustness and stability in non-stationary energy environments.
[0049] Figure 3 This paper showcases the queue backlog performance of four algorithms under extreme operating conditions. These extreme conditions primarily involve scenarios with severe energy shortages and high transmission loads, aiming to verify the robustness and stability of each algorithm under resource-constrained conditions. The results show that the proposed LAD3 algorithm exhibits superior stability and energy efficiency distribution across different operating samples. Its average queue backlog is reduced by approximately 26.4%, 12.2%, and 24.6% compared to the DDQN, DDQN+Dropout, and ECLB algorithms, respectively, indicating that LAD3 can more effectively maintain queue balance and energy utilization efficiency when dealing with high loads and energy fluctuations. Further independent experiments verify the reliability and consistency of this performance improvement, fully demonstrating that the LAD3 algorithm has stronger adaptability and robustness in non-stationary energy environments.
[0050] In a specific implementation, a long-term energy efficiency optimization system for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization includes: System model building module: used to build a wireless communication system model containing multiple device nodes and define the system state; Power control action module: used to input the current system state into the training network to output power control actions; Environmental feedback reward module: After executing a power control action, the system obtains environmental feedback rewards and moves to the next state; The parameter update module is used to construct the current interaction sample based on the current system state, power control action, environmental feedback reward, and next state, and store it in the experience memory bank. It also randomly samples sample data from the experience memory bank to update the parameters of the training network. Lyapunov optimization module: used to calculate queue Lyapunov drift terms, introduce the drift terms into the optimization objective, and update the optimized value after introduction to the environment feedback reward of the training network; Joint optimization module: Used for iterative training until the network converges, to obtain the optimal power control strategy, and to perform joint optimization of long-term energy efficiency and system stability constraints.
[0051] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0052] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization, characterized in that, include: Construct a wireless communication system model containing multiple device nodes and define the system state; Input the current system state into the training network to output power control actions; After executing a power control action, the system receives environmental feedback as a reward and then transitions to the next state; Based on the current system state, power control actions, environmental feedback rewards, and the next state, construct the current interaction sample and store it in the experience memory bank. Randomly sample sample data from the experience memory bank to update the parameters of the training network. Calculate the Lyapunov drift term in the queue, introduce the drift term into the optimization objective, and update the optimized value after introduction to the environmental feedback reward of the training network; The iterative training process continues until the network converges, obtaining the optimal power control strategy, and performing joint optimization of long-term energy efficiency and system stability constraints.
2. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 1, characterized in that, Each of the device nodes harvests energy from three energy sources: solar, wind, and radio frequency energy.
3. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 1, characterized in that, The system status includes node energy status, data queue status, and channel condition information.
4. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 1, characterized in that, The training network is an n-layer neural network with Dropout units set between the hidden layers. It takes the system state as input, outputs the action value of each action in the action set, and selects the action with the highest action value as the power control action.
5. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 4, characterized in that, The action value is used to characterize the long-term cumulative return obtained by choosing a specific power action in a given state, and is defined as follows: ; in, System status, For the value of the action, , are the parameters of the neural network. For discount rate, Yes Approximate to .
6. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 4, characterized in that, The Dropout unit adaptively perturbs the Q-value distribution by randomly masking neurons. The calculation is as follows: ; in, In order to perform the action, For random sampling value.
7. The long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization according to claim 1, characterized in that, The step of randomly sampling sample data from the experience memory bank to update the parameters of the training network includes: ; in, For predicted values, To define rewards based on actual observations for .
8. A long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 7, characterized in that, It also includes: optimizing using a loss function that minimizes the difference between the predicted Q-value and the target Q-value, and synchronizing parameters through the target network. ; Parameter update: ; in, It's the learning rate. .
9. A long-term energy efficiency optimization method for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization as described in claim 1, characterized in that, The computation queue Lyapunov drift term is incorporated into the optimization objective, and the optimized value after incorporation is updated to the environmental feedback reward of the training network, including: in, For a set of queues, For Lyapunov functions, For Lyapunov drift function, For the expectation, For virtual queues, For transmission rate, Energy consumption; Incorporating the drift term into the optimization objective In the middle, As a reward for training the network, long-term energy efficiency and queue stability are jointly controlled.
10. A long-term energy efficiency optimization system for an energy harvesting D2D communication system based on reinforcement learning and Lyapunov optimization, characterized in that, include: System model building module: used to build a wireless communication system model containing multiple device nodes and define the system state; Power control action module: used to input the current system state into the training network to output power control actions; Environmental feedback reward module: After executing a power control action, the system obtains environmental feedback rewards and moves to the next state; The parameter update module is used to construct the current interaction sample based on the current system state, power control action, environmental feedback reward, and next state, and store it in the experience memory bank. It also randomly samples sample data from the experience memory bank to update the parameters of the training network. Lyapunov optimization module: used to calculate queue Lyapunov drift terms, introduce the drift terms into the optimization objective, and update the optimized value after introduction to the environment feedback reward of the training network; Joint optimization module: Used for iterative training until the network converges, to obtain the optimal power control strategy, and to perform joint optimization of long-term energy efficiency and system stability constraints.