A power control method for energy harvesting wireless communication system

By combining the Bellman equation and reinforcement learning framework, a power control method with low computational complexity is designed, which solves the problem of high power control complexity in energy-collecting wireless communication systems, and achieves theoretical optimal approximation of throughput performance, which is suitable for IoT scenarios with limited computing resources.

CN119854923BActive Publication Date: 2025-06-06ZHEJIANG GONGSHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510318380.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-06
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In energy-harvesting wireless communication systems, how to design an online power control algorithm with low computing complexity can achieve the optimal performance of throughput performance indicator approximation theory under multiple energy arrival distributions, especially in the Internet of Things scenario where computing resources are limited.

Method used

By combining the power allocation formula based on Bellman equation inference with the reinforcement learning framework, a transmitter power control method with extremely low computational complexity is designed. The method includes a preparation phase and a real-time control phase, utilizing state value neural network and dynamic equal-volume ratio, calculate the power consumption in real time and update the parameters to achieve an approximate optimal power allocation.

Benefits of technology

It realizes the control of transmitter power at extremely low computing complexity, and can achieve theoretical optimal performance of throughput performance index approximation of wireless communication systems under various energy arrival distributions, reducing computing overhead and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854923B_ABST
    Figure CN119854923B_ABST
Patent Text Reader

Abstract

The present invention discloses a power control method for an energy collection wireless communication system, which belongs to the field of communication technology, and includes: reasoning a power allocation formula based on the Bellman equation and an empirical state value function; defining a state value neural network parameter in a reinforcement learning framework as a calculation parameter in the empirical state value function; calculating the consumed power based on the battery power, channel coefficient, dynamic average capacity ratio, state value neural network parameter and power allocation formula of the current time slot; sending and observing the data volume according to the consumed power of the current time slot, recording the battery power of the next time slot, generating an empirical data item including the battery power and data volume of the current time slot and the battery power of the next time slot; updating the differential reward and updating the state value neural network parameter according to the empirical data item; calculating the ratio of the charged battery energy in the current time slot to the maximum chargeable battery energy and updating the dynamic average capacity ratio. In this way, the calculation complexity can be reduced while optimizing the communication performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technology, and in particular relates to a power control method used in an energy harvesting wireless communication system. Background Art

[0002] With the exponential expansion of the scale of the Internet of Things, a large number of low-power wireless devices are widely used in multiple scenarios such as remote node access and environmental monitoring (sensor data collection such as temperature, humidity, and light). In order to meet the mobility requirements of devices, most terminals rely on batteries with limited capacity, which leads to bottlenecks in their operating life. Especially when the devices are deployed in areas where battery replacement is difficult (such as inside industrial equipment and remote areas) or when ultra-large-scale IoT nodes need to be maintained, frequent battery replacement will lead to high operation and maintenance costs. In this context, energy harvesting technology is regarded as a key method to break through energy constraints due to its sustainability and environmentally friendly characteristics. By integrating environmental energy harvesting modules and energy storage devices, IoT devices can autonomously collect energy from the surrounding environment (solar energy, electromagnetic radiation, mechanical vibration, etc.), thereby significantly extending their life cycle. However, the randomness and instability of environmental energy make the dynamic optimization of energy collection-storage-distribution a core challenge in the design of wireless communication systems. In addition, the fading characteristics of wireless channels further exacerbate the complexity of the problem. The unpredictability of channel states requires that energy allocation strategies must have real-time adaptive capabilities to cope with dynamically changing channel conditions and energy supply.

[0003] System throughput is a key indicator for measuring the performance of wireless communication systems. Its goal is to maximize the average amount of data transmitted per unit time by wireless devices, thereby improving communication efficiency and overall system performance. In the case of limited, intermittent and fluctuating energy supply, how to maximize the throughput of wireless devices by optimizing energy use is a key issue that needs to be addressed in actual systems. Under the offline power control model, this key issue can be solved by a directional water injection algorithm to obtain the optimal offline strategy, which can theoretically achieve the best performance. However, there is a large gap between offline control and actual application scenarios, and it is difficult to directly apply it to actual systems. Under the online power control model, this key issue is usually modeled as a Markov decision process (MDP), and the solution to its optimal power control strategy can be attributed to solving the Bellman equation corresponding to the model. The closed-form solution of the Bellman equation is usually difficult to obtain. Existing studies often use dynamic programming algorithms based on the Bellman equation (such as value iteration and policy iteration) to obtain approximate numerical solutions. However, such methods have two limitations: (1) Both closed-form solutions and dynamic programming algorithms require precise model knowledge; (2) Even if dynamic programming algorithms are used, their time complexity still grows exponentially with the growth of the state and action space dimensions, and the computing power requirements are extremely high.

[0004] Online power control based on reinforcement learning algorithms can avoid the reliance on precise model knowledge to a certain extent and reduce the algorithm complexity through the adaptive potential of interactive learning. However, the power control strategy designed based on the general reinforcement learning algorithm still needs to estimate a large number of parameters to ensure its performance, resulting in its high complexity. In addition, the existing power control algorithms based on reinforcement learning generally lack performance comparison with the theoretical optimal strategy (i.e., the MDP model optimal strategy). Therefore, in the IoT scenario with limited computing resources, how to design a lightweight online power control algorithm with both low computing overhead and high throughput guarantee is a technical problem that needs to be solved in the field of energy harvesting wireless communications. Summary of the invention

[0005] In view of the above, the object of the present invention is to provide a power control method for an energy harvesting wireless communication system, aiming to design a transmitter power control method based on a customized reinforcement learning framework with extremely low complexity based on domain knowledge, so as to achieve the throughput performance index of the energy harvesting wireless communication system approaching the theoretical optimal performance.

[0006] To achieve the above-mentioned object of the invention, an embodiment provides a power control method for an energy harvesting wireless communication system, comprising the following steps:

[0007] Preparation stage: Based on the empirical state value function and Bellman equation, the power allocation formula of the energy harvesting wireless communication system is obtained; based on the empirical state value function, the state value neural network in the reinforcement learning framework is defined and the training parameters are initialized, where the state value neural network parameters are defined as the calculation parameters in the empirical state value function;

[0008] Real-time control stage: Calculate the power consumption of the current time slot based on the battery power, channel coefficient, dynamic average capacity ratio, state value neural network parameters and power allocation formula of the current time slot; The transmitter sends data on the fading channel according to the power consumption of the current time slot, observes the amount of data sent in the current time slot, and observes the battery power of the next time slot before the end of the current time slot, and generates an experience data item including the battery power of the current time slot, the amount of data in the current time slot, and the battery power of the next time slot; Calculate and update the differential reward based on the experience data item, and update the state value neural network parameters based on the differential reward;

[0009] The ratio of the charged battery energy in the current time slot to the maximum chargeable battery energy is calculated, and the dynamic average capacity ratio is updated based on the ratio.

[0010] Preferably, the power allocation formula of the energy harvesting wireless communication system is obtained based on the empirical state value function and the Bellman equation, including:

[0011] Combined with the empirical state value function, the Bellman equation in any situation is approximated as the Bellman equation under the worst-case Bernoulli distribution. The optimization problem in the Bellman equation under the Bernoulli distribution is constructed, and the optimal solution obtained by solving it is the power allocation formula for the energy harvesting wireless communication system.

[0012] Preferably, a state value neural network in a reinforcement learning framework is defined based on the empirical state value function and training parameters are initialized, wherein the state value neural network parameters are defined as calculation parameters in the empirical state value function, including:

[0013] Initializing the reinforcement learning time slot , average reward , dynamic average capacity ratio , the learning rate parameter , the learning rate parameter , where the symbol Indicates assignment, , , ,as well as These are all initial state values ​​and will be optimized during subsequent training;

[0014] Initialize an empty first-in-first-out queue for storing experience data, called the experience pool queue;

[0015] Among them, the experience state value function is expressed as:

[0016] ;

[0017] in, Indicates the battery level. represents the equivalent linear power control coefficient, represents the equivalent channel coefficient, Indicates the state value, symbol Meaning is defined as;

[0018] When defining the state-value neural network in the reinforcement learning framework based on the empirical state-value function, the battery charge As the input of the state-value neural network, and As the parameters of the state-value neural network, they are initialized as and , state value As the output of the state-value neural network.

[0019] Preferably, the power consumption of the current time slot is calculated based on the battery power, channel coefficient, dynamic average capacity ratio, calculation parameters and power allocation formula of the current time slot, including:

[0020] Set the current time slot Power consumption Assigned to ,in, Indicates the current time slot The battery charge random variable at time , Indicates the current time slot The channel coefficient random variable at time , The calculation formula for power consumption, i.e., power control strategy, is defined as:

[0021] ;

[0022] express The battery charge value at the time of power control decision, express The channel coefficient value at the power control decision time is assumed to be ,but .

[0023] Preferably, the current time slot is observed The amount of data sent Assigned to , and the empirical data item is ,in, The battery power of the next time slot is pushed into the experience pool queue. If the experience pool queue contains experience data items exceeding the preset maximum length, , then the earliest experience data in the queue is popped out.

[0024] Preferably, the differential reward is calculated and updated based on the experience data item, including:

[0025] Calculate the current time slot The difference reward , calculated as:

[0026] ;

[0027] in, Value taken from , Value taken from , Value taken from , Represents single empirical data The corresponding differential reward is Representing state value neural network based on input The state value of Representing state value neural network based on input The state value of

[0028] Based on differential rewards The average reward Updated to ,in, represents the updated average reward.

[0029] Preferably, updating the state value neural network parameters according to the differential reward includes:

[0030] Randomly draw from the experience pool Item of empirical data, assuming The empirical data of the item is , , based on the differential reward for each experience data , update the state value neural network parameters:

[0031] ;

[0032] ;

[0033] in, and Represents the updated parameters, It means assigning the value on the right to the left and using it as the condition for the previous update calculation.

[0034] Preferably, the method further comprises: clipping the updated parameter value according to the boundary conditions of the parameter value to ensure that the parameter value is within the boundary.

[0035] Preferably, calculating the ratio of the charged battery energy in the current time slot to the maximum chargeable battery energy, and updating the dynamic average capacity ratio based on the ratio, includes:

[0036] ;

[0037] ;

[0038] in, Represents the ratio of the current time slot, Indicates the maximum chargeable battery energy. Indicates the updated dynamic average capacity ratio.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The power control method of the present invention can control the transmitter power with extremely low computational complexity under fading channel conditions, and achieve throughput performance indicators approaching the theoretical optimal performance under various energy arrival distributions, which is specifically derived from the following technical features:

[0041] Simplified value estimation: The state value neural network defined by the power allocation formula in the present invention only includes the state value neural network parameters and the average reward parameters when performing state estimation, a total of 3 parameters, which greatly simplifies the value estimation process and effectively solves the problem of high computational complexity of state value estimation in the prior art;

[0042] Simplify the decision-making process: The power allocation formula based on dynamic average capacity ratio and state value neural network parameters adopted in the present invention greatly simplifies the calculation process of searching for the optimal action or approximating the optimal strategy in reinforcement learning decision-making, thereby greatly reducing the calculation overhead at the optimal decision-making level. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 The invention is a flowchart of a power control method in an energy harvesting wireless communication system provided by an embodiment. DETAILED DESCRIPTION

[0045] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0046] The technical concept of the present invention is: the dynamic programming algorithm based on the Markov decision process can solve the optimal strategy, but the complexity is extremely high and requires accurate model knowledge; the method based on general reinforcement learning or deep reinforcement learning can adapt to the actual scene through value estimation, but because its performance depends on the number of parameters, it is difficult to achieve a reasonable compromise between performance and computing power on a device with low power consumption requirements, and the existing algorithm has no comprehensive or partial verification of the optimal performance even without computing power constraints. To this end, the embodiment of the present invention provides a power control method for an energy harvesting wireless communication system, which cleverly combines the power allocation formula based on the Bellman equation reasoning with reinforcement learning to achieve power control in the energy harvesting wireless communication system with extremely low computing overhead, and achieve the throughput performance index of the energy harvesting wireless communication system close to the theoretical optimal performance. The whole method includes a preparation stage and a real-time control stage, wherein the preparation stage includes the initialization of the previous value function, the power allocation formula and the reinforcement learning framework, and the real-time control stage includes real-time output of power consumption based on reinforcement learning, and control of the transmitter to transmit data based on power consumption.

[0047] The embodiment of the present invention considers a wireless transmitter equipped with an energy harvester and a rechargeable battery. The energy harvester collects ambient energy and stores it in the rechargeable battery to power the transmitter. The transmitter completes data transmission on a point-to-point fading channel by consuming battery energy. The entire transmitter system is a discrete time system. The capacity of the rechargeable battery is , current time slot The observable state of the system includes the battery level and channel coefficients , current time slot When the system control quantity is time slot Energy consumed by the internal transmitter , which means power consumption. and power consumption Under the condition of The amount of data transmitted within (i.e., the system reward value) is .

[0048] Based on this, Figure 1 As shown, an embodiment provides a power control method for an energy harvesting wireless communication system, comprising the following steps:

[0049] S1, the power allocation formula of the energy harvesting wireless communication system is obtained based on the empirical state value function and Bellman equation.

[0050] In energy harvesting wireless communication systems, according to existing literature, if the energy arrival random sequence is independent and identically distributed, then the Bernoulli distribution is the worst case among the energy arrival distributions with the same effective mean, that is, the system throughput reaches the lowest. Moreover, both theoretical and experimental results show that the optimal or near-optimal strategy under the Bernoulli distribution exhibits good performance under all energy distributions with the same effective mean. Therefore, if the state value estimation based on the state value neural network is accurate, and the effective mean of the energy charged into the battery can be approximated as the dynamic mean capacity ratio and the maximum chargeable battery energy The product of , then the Bellman equation in any case can be approximated as the Bellman equation under the worst-case Bernoulli distribution. Combined with the empirical state-value function, that is, let the theoretical value of the Bellman equation be equal to the empirical value calculated based on the state-value function, then:

[0051] ;

[0052] in, is the optimal reward value, which is the optimal throughput indicator in the present invention, is the optimal state value function, is the battery status, i.e. the battery charge, is the channel coefficient, It uses energy. is the reward value function, Indicates the battery status is The state value of Indicates the battery status is The status value.

[0053] Proposition: The optimization problem in the Bellman equation under Bernoulli distribution:

[0054] ;

[0055] Solving the optimization problem, the optimal solution is:

[0056] ;

[0057] in, , , .

[0058] Proof: The objective function of the Bellman equation To perform the derivation:

[0059] ;

[0060] Its first and second derivatives are:

[0061] ;

[0062] ;

[0063] Obviously, for all , .therefore, In the interval is strictly monotonically decreasing. , the unique zero point of its unbounded constraint can be solved as:

[0064] ;

[0065] Combination After classifying and discussing the monotonicity of the objective function, the optimal power allocation power is obtained as follows:

[0066] .

[0067] S2, based on the empirical state value function, defines the state value neural network in the reinforcement learning framework and initializes the training parameters, where the state value neural network parameters are defined as the calculation parameters in the empirical state value function.

[0068] Initializing the reinforcement learning time slot , average reward , dynamic average capacity ratio , the learning rate parameter , the learning rate parameter , where the symbol Indicates assignment, , , ,as well as are all initial state values, which will be optimized in the subsequent training process. Specifically, the average reward is Initialized to 0, dynamic average capacity ratio Initialized to 0, the learning rate parameter and Initialized to ; At the same time, initialize an empty first-in-first-out queue to store experience data, called the experience pool queue. The maximum length of this queue is .

[0069] Among them, the experience state value function is expressed as:

[0070] ;

[0071] in, Indicates the battery level. represents the equivalent linear power control coefficient, represents the equivalent channel coefficient, Indicates the state value, symbol Meaning is defined as;

[0072] When defining the state-value neural network in the reinforcement learning framework based on the empirical state-value function, the battery charge As the input of the state-value neural network, and As the parameters of the state-value neural network, they are initialized as and Specifically, the initial value of the parameter is and , state value As the output of the state-value neural network.

[0073] S3, calculates the power consumption of the current time slot based on the battery power, channel coefficient, dynamic average capacity ratio, state value neural network parameters and power allocation formula of the current time slot.

[0074] In the embodiment, the current time slot Power consumption Assigned to ,in, Indicates the current time slot The battery charge random variable at time , Indicates the current time slot The channel coefficient random variable at time , The calculation formula for power consumption, i.e., power control strategy, is defined as:

[0075] ;

[0076] express The battery charge value at the time of power control decision, express The channel coefficient value at the power control decision time is assumed to be ,but .

[0077] S4, the transmitter sends data on the fading channel according to the power consumption of the current time slot, observes the amount of data sent in the current time slot, and observes the battery power of the next time slot before the end of the current time slot, and generates an empirical data item including the battery power of the current time slot, the amount of data in the current time slot, and the battery power of the next time slot.

[0078] In the embodiment, the transmitter is based on the current time slot Power consumption Send data on a fading channel and observe the current time slot The amount of data sent Assigned to , and in the time slot The battery level of the next time slot is observed before the end , then the empirical data item is , the experience data item is pushed into the experience pool queue. If the experience data item contained in the experience pool queue exceeds the preset maximum length , then the earliest experience data in the queue is popped out.

[0079] S5, calculates and updates the differential reward based on the empirical data item, and updates the state value neural network parameters based on the differential reward.

[0080] In the embodiment, the current time slot is calculated The difference reward , calculated as:

[0081] ;

[0082] in, Value taken from , Value taken from , Value taken from , Represents single empirical data The corresponding differential reward is Representing state value neural network based on input The state value of Representing state value neural network based on input The state value of

[0083] Differential Rewards The average reward Updated to ,in, represents the updated average reward.

[0084] In the embodiment, randomly selected from the experience pool (For example, N can be 64) empirical data, assuming The empirical data of the item is , , based on the differential reward for each experience data , update the state value neural network parameters:

[0085] ;

[0086] ;

[0087] in, and Represents the updated parameters, It means assigning the value on the right to the left and using it as the condition for the previous update calculation.

[0088] At the same time, the updated parameter value is clipped according to the boundary conditions of the parameter value to ensure that the parameter value is within the boundary.

[0089] S6, calculating the ratio of the battery energy charged in the current time slot to the maximum battery energy that can be charged, and updating the dynamic average capacity ratio based on the ratio.

[0090] In the embodiment, the ratio of the battery energy charged in the current time slot to the maximum battery energy that can be charged is calculated. ,Right now:

[0091] ;

[0092] Update the dynamic average capacity ratio :

[0093] ;

[0094] in, Represents the ratio of the current time slot, Indicates the maximum chargeable battery energy. Indicates the updated dynamic average capacity ratio.

[0095] S7, increment time slot variable , and jump to step S3.

[0096] In order to verify the effect of the power control method of the present invention, simulation experiments were also carried out under the conditions of Rayleigh fading channel and three energy arrival distributions (Bernoulli distribution, uniform distribution and exponential distribution), and compared with the MDP optimal strategy (i.e. the theoretical optimal strategy). Table 1-3 shows the simulation performance of the algorithm of the present invention and the MDP optimal strategy under various parameter combinations, that is, the calculation system throughput, also known as the reward value, is calculated as follows: , For time slot The amount of data sent, the unit of the reward value is nats / s / Hz, which means the amount of data sent within a unit bandwidth (1Hz) within a unit time (1s). Because the reward value function is calculated by the ln function, the unit of the amount of data sent is nats (similar to bit). Among them, the nominal mean-to-capacity ratio (NMCR) parameter is the ratio of the mean of the energy arrival distribution to the battery capacity, and is specifically defined as:

[0097] ;

[0098] in, represents the energy arrival distribution, represents the arrival energy, represents the mean of the arrival energies subject to Q;

[0099] The nominal signal-to-noise ratio (NSNR) parameter is the energy arrival distribution over the battery capacity. The equivalent decibel value of the effective mean under the limit is specifically defined as:

[0100] ;

[0101] in, .

[0102] Table 1: Simulation performance under Bernoulli distribution

[0103]

[0104] Table 2: Simulation performance under uniform distribution

[0105]

[0106] Table 3: Simulation performance under exponential distribution

[0107]

[0108] Table 4 summarizes the results of Tables 1-3, and gives the maximum additive gap and minimum multiplication factor of the method of the present invention and the MDP optimal strategy in the nominal signal-to-noise ratio range of 0-30dB under different energy arrival distributions and different NMCR conditions. and multiplication factors They are defined as:

[0109] ;

[0110] ;

[0111] in, is the average reward value of the present invention as the average throughput, is the average throughput of the MDP optimal strategy. The results show that under various energy arrival distributions and various parameter conditions, the present invention exhibits performance close to the MDP optimal strategy, with the maximum absolute performance gap being approximately 0.003-0.038 (nats / s / Hz) and the maximum relative performance gap being approximately 0.2%-1.1%.

[0112] Table 4: Performance comparison between the proposed algorithm and the MDP optimal strategy

[0113]

[0114] Table 5 gives a comparison of the technical features of the present invention and the reinforcement learning algorithms in the existing literature. As shown in the table, the method SARSA (from Reinforcement learning exploration algorithms for energy harvesting communications systems) uses the traditional table method to design the reinforcement learning algorithm, which is only applicable to models with a small number of states, specifically 12 states in this paper. In order to achieve better results in value estimation and decision-making, DQN (from Action-bounding for reinforcement learning in energy harvesting communication systems, action boundary setting method for reinforcement learning in energy harvesting communication systems), Actor-Critic (An actor-critic reinforcement learning approach for energy harvesting communications systems, an actor-critic reinforcement learning method for energy harvesting communication systems), PPO (Autonomous management of energy-harvesting iot nodes using deep reinforcement learning, an autonomous management method of energy-harvesting IoT nodes based on deep reinforcement learning), and DDPG (Deep deterministic policy gradient (DDPG)-based energy harvesting wireless communications) in the table are processed using multi-layer neural networks, in which the number of parameters of the neural network is up to about 17,000 at most, and the minimum is 105. In contrast, the strategy proposed in the present invention only requires 4 parameters, and the algorithm complexity difference is significant.

[0115] Table 5: Comparison and summary of technical features of the present invention and related literature

[0116]

[0117] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A power control method for an energy harvesting wireless communication system, characterized in that: The following steps are involved: Preparation stage: Based on the empirical state value function and Bellman equation, the power allocation formula of the energy harvesting wireless communication system is obtained; Based on the experience state value function, the state value neural network in the reinforcement learning framework is defined and the training parameters are initialized, where the state value neural network parameters are defined as the calculation parameters in the experience state value function, including: Initializing the reinforcement learning time slot , average reward , dynamic average capacity ratio , the learning rate parameter , the learning rate parameter , where the symbol Indicates assignment, , , ,as well as These are all initial state values ​​and will be optimized during subsequent training; Initialize an empty first-in-first-out queue for storing experience data, called the experience pool queue; Among them, the experience state value function is expressed as: ; in, Indicates the battery level. represents the equivalent linear power control coefficient, represents the equivalent channel coefficient, Indicates the state value, symbol Meaning is defined as; When defining the state-value neural network in the reinforcement learning framework based on the empirical state-value function, the battery charge As the input of the state-value neural network, and As the parameters of the state-value neural network, they are initialized as and , state value As the output of the state-value neural network; Real-time control stage: Calculate the power consumption of the current time slot based on the battery power, channel coefficient, dynamic average capacity ratio, state value neural network parameters and power allocation formula of the current time slot, including: Set the current time slot Power consumption Assigned to ,in, Indicates the current time slot The battery charge random variable at time , Indicates the current time slot The channel coefficient random variable at time , The calculation formula for power consumption, i.e., power control strategy, is defined as: ; express The battery charge value at the time of power control decision, express The channel coefficient value at the power control decision time is assumed to be ,but ; The transmitter sends data on the fading channel according to the power consumption of the current time slot, observes the amount of data sent in the current time slot, and observes the battery power of the next time slot before the end of the current time slot, and generates an experience data item including the battery power of the current time slot, the amount of data in the current time slot, and the battery power of the next time slot; calculates and updates the differential reward based on the experience data item, and updates the state value neural network parameters based on the differential reward; The ratio of the charged battery energy in the current time slot to the maximum chargeable battery energy is calculated, and the dynamic average capacity ratio is updated based on the ratio.

2. The power control method for an energy harvesting wireless communication system according to claim 1, characterized in that: Based on the empirical state value function and Bellman equation, the power allocation formula of the energy harvesting wireless communication system is obtained, including: Combined with the empirical state value function, the Bellman equation in any situation is approximated as the Bellman equation under the worst-case Bernoulli distribution. The optimization problem in the Bellman equation under the Bernoulli distribution is constructed, and the optimal solution obtained by solving it is the power allocation formula for the energy harvesting wireless communication system.

3. The power control method for an energy harvesting wireless communication system according to claim 1, characterized in that: Observe the current time slot The amount of data sent Assigned to , and the empirical data item is ,in, The battery power of the next time slot is pushed into the experience pool queue. If the experience pool queue contains experience data items exceeding the preset maximum length, , then the earliest experience data in the queue is popped out.

4. The power control method for an energy harvesting wireless communication system according to claim 3, characterized in that: Calculate and update the differential reward based on the experience data items, including: Calculate the current time slot The difference reward , calculated as: ; in, Value taken from , Value taken from , Value taken from , Represents single empirical data The corresponding differential reward is Representing state value neural network based on input The state value of Representing state value neural network based on input The state value of Based on differential rewards The average reward Updated to ,in, represents the updated average reward.

5. The power control method for an energy harvesting wireless communication system according to claim 4, characterized in that: Update the state value neural network parameters according to the differential reward, including: Randomly draw from the experience pool Item of empirical data, assuming The empirical data of the item is , , based on the differential reward for each experience data , update the state value neural network parameters: ; ; in, and represents the updated parameters, It means assigning the value on the right to the left and using it as the condition for the previous update calculation.

6. The power control method for an energy harvesting wireless communication system according to claim 5, characterized in that: Also includes: The updated parameter value is clipped according to the boundary conditions of the parameter value to ensure that the parameter value is within the boundary.

7. The power control method for an energy harvesting wireless communication system according to claim 5, characterized in that: Calculate the ratio of the charged battery energy in the current time slot to the maximum chargeable battery energy, and update the dynamic average capacity ratio based on the ratio, including: ; ; in, Represents the ratio of the current time slot, Indicates the maximum chargeable battery energy. Indicates the updated dynamic average capacity ratio.

Citation Information

Patent Citations

  • Adaptive modulation and power control system based on energy collection and optimization method thereof

    CN111491358A

  • Power control method under actual charging and discharging characteristics of battery in energy harvesting wireless communication system

    CN113141646A