An adaptive power control method for energy harvesting communication systems

By adaptively adjusting the transmit power using a predictive model, the problem of high-complexity power control in energy harvesting communication systems is solved, achieving near-optimal power control with low complexity, performance loss of less than 2%, and significant reduction in computation time.

CN122093911BActive Publication Date: 2026-07-24ZHEJIANG GONGSHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG GONGSHANG UNIVERSITY
Filing Date
2026-04-22
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing power control methods for energy harvesting communication systems are computationally complex, making it difficult to adapt to complex time-varying energy harvesting and channel conditions, and lacking effective low-complexity solutions under non-independent and identically distributed models.

Method used

An adaptive power control method based on a prediction model is adopted. By predicting the equivalent linear strategy slope and equivalent channel quality metric of the next time slot, the transmit power is adaptively adjusted using a four-dimensional parameter vector, including the dynamic average capacity ratio of energy arrival in the current time slot, the minimum-average charging energy ratio, the equivalent linear strategy slope of the power control strategy in the next time slot, and the equivalent channel quality metric of the next time slot. This reduces computational complexity and achieves near-optimal control.

Benefits of technology

It significantly reduces computational complexity, achieves near-optimal power control under different energy arrival and channel conditions, with performance loss of less than 2% and computation time far lower than existing reinforcement learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093911B_ABST
    Figure CN122093911B_ABST
Patent Text Reader

Abstract

The application discloses a kind of adaptive power control methods for energy collection communication system, belong to communication technical field, comprising: in each time slot, four-dimensional parameter vector required for power control is predicted using updated prediction model, including the dynamic average of current time slot energy arrival ratio, current time slot minimum-average charging energy ratio, the equivalent linear strategy slope of next time slot power control strategy, and next time slot equivalent channel quality metric;Based on the system state of current time slot starting time and the power control strategy described by the four-dimensional parameter vector prediction value, the energy consumption value of current time slot is calculated;Instruction communication system completes corresponding communication task in current time slot according to the energy consumption value, and observes the utility value of completing communication task.The method can adaptively adjust the transmitting power according to the system state, to realize approximately optimal power control under specified performance indicators, and has lower calculation time overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, and specifically relates to an adaptive power control method for energy harvesting communication systems. Background Technology

[0002] Driven by the rapid development of IoT technology and the increasingly urgent need for green communication technologies in the context of global warming, the design of energy harvesting communication systems that utilize environmental energy has become one of the core research and development directions in the field of communications in recent years. Given the random intermittent nature of environmental energy sources, the capacity constraints of energy storage devices in communication systems, and the time-varying characteristics of wireless communication channels, designing power control algorithms that combine low complexity and high performance is one of the core challenges in the field of energy harvesting communications. Online power control in point-to-point energy harvesting communication under wireless fading channels is a fundamental problem in this challenge. To date, the following types of solutions have been developed:

[0003] Reference 1 (Amirnavaei F, Dong M. Online power control optimization for wireless transmission with energy harvesting and storage[J]. IEEE Transactions on Wireless Communications, 2016, 15(7): 4888-4901) proposes a power control method based on Lyapunov optimization, the core of which is a closed-loop strategy. Although this method has low computational complexity, it optimizes a relaxed approximation problem rather than the original problem itself. Therefore, the obtained strategy differs significantly in structure from the known theoretically optimal strategy in some cases, making it difficult to effectively approximate the optimal performance in all scenarios.

[0004] Reference 2 (Ku ML and Lin T J. Neural-network-based power control prediction for solar-powered energy harvesting communications[J]. IEEE Internet of Things Journal, 2021, 8(16): 12983-12998) proposes a power control method based on an optimal offline power control strategy and neural network power control prediction, which can essentially be regarded as a predictor trained based on the optimal offline strategy. In the solar energy scenario, the neural network constructed by this method has more than 30,000 parameters, resulting in high computational complexity. Simulation results show that the performance gap between it and the theoretical optimal performance is no more than 4%. However, since the objectives of the optimal offline strategy and the optimal online control are fundamentally different, the ability of a neural network trained entirely on the former to approximate the theoretical optimal performance in actual online scenarios is highly uncertain.

[0005] The mainstream approach to solving online power control in energy harvesting communications is the design of optimal online power control strategies based on Markov decision processes (MDPs). The core method involves solving the Bellman equation associated with a given MDP model. Under typical energy harvesting and channel models, closed-form solutions to this equation are often difficult to obtain, or even impossible to find. Classical numerical methods for solving the Bellman equation include value iteration and policy iteration, but these algorithms, when applied to power control problems, not only have extremely high computational complexity but also require precise model knowledge, making them difficult to adapt to adaptive power control under complex time-varying energy harvesting and channel conditions. Reinforcement learning-based online power control can approximate the Bellman equation through interactive learning without model knowledge, and to some extent reduces computational complexity. However, power control methods based on general reinforcement learning algorithms still have high computational complexity, and existing schemes lack rigorous performance verification. Patent document 3 (CN202510318380.1) proposes a customized reinforcement learning power control algorithm that integrates domain knowledge. Under the condition that energy arrival and channel fading are both independent and identically distributed, the algorithm achieves near-theoretical optimal performance within a nominal signal-to-noise ratio range of 0-30 dB (with a performance loss of approximately 1% compared to the optimal strategy). Comparative analysis in Patent Document 3 shows that the computational complexity of this algorithm is significantly lower than that of power control methods based on general reinforcement learning. This advantage stems from its closed-loop strategy and the special state value network containing only two parameters.

[0006] However, the method proposed in Patent Document 3 still has three shortcomings:

[0007] (1) Due to the use of a reinforcement learning framework, the algorithm complexity is still relatively high;

[0008] (2) The strategy function is designed based on the scenario where energy reaches its maximum fluctuation. There is still room for performance improvement in scenarios where energy reaches its minimum fluctuation.

[0009] (3) No effective low-complexity solution is given for general models that are not independent and identically distributed.

[0010] Therefore, how to overcome the limitations of reinforcement learning frameworks and design a low-complexity, high-performance general adaptive power control method is a key research problem in power control design for energy harvesting communication. Summary of the Invention

[0011] In view of the above, the purpose of this invention is to provide an adaptive power control method for an energy harvesting communication system, which can adaptively adjust the transmit power according to the system state under different energy arrival characteristics (i.e., the output energy of the energy harvester) and channel conditions (including but not limited to independent identically distributed processes and Markov processes) to achieve near-optimal power control under specified performance indicators, and has low computational time overhead.

[0012] To achieve the above-mentioned objectives, an embodiment provides an adaptive power control method for an energy harvesting communication system, comprising the following steps:

[0013] In each time slot, the updated prediction model is used to predict the four-dimensional parameter vector required for power control based on the historical system state up to the start of the current time slot. This includes the dynamic average capacity ratio of energy arrival in the current time slot, the minimum-average charging energy ratio of the current time slot, the equivalent linear strategy slope of the power control strategy in the next time slot, and the equivalent channel quality metric for the next time slot.

[0014] Based on the system state at the start of the current time slot and the power control strategy characterized by the predicted values ​​of the four-dimensional parameter vector, calculate the energy consumption value of the current time slot;

[0015] The command communication system completes the corresponding communication task in the current time slot according to the energy consumption value, and observes the utility value of completing the communication task.

[0016] Preferably, the updated prediction model predicts the four-dimensional parameter vector required for power control based on the historical system state up to the start of the current time slot, including:

[0017] The historical system states up to the start of the current time slot are constructed as the context features of the current time slot, and the updated prediction model is then used based on these context features. The four-dimensional parameter vector required for predictive power control; where, This indicates that the input is a contextual feature. And the parameters are Predictive models;

[0018] The context features For a fixed length of A sequence, where each element of the sequence is in a given set. The value is taken from the top; where, If it is a non-negative integer, when When the sequence is an empty sequence This indicates that historical system status information is not used;

[0019] Specifically, the historical system states up to the start of the current time slot are constructed as the context features of the current time slot, including:

[0020] The current time slot context features are calculated based on the system state at the start of the current time slot and the context features of the previous time slot, including:

[0021] in, For the current time slot context features, the function Generate functions for contextual features. For the context features of the previous time slot, This refers to the system state at the start of the current time slot, including the initial stored energy of the current time slot. Current time slot channel quality metric Other available state information related to power control decisions ,symbol Indicates the definition of an operation. The context feature sequence of the previous time slot The Middle One element, For a set of values ​​from the system state to a set The mapping.

[0022] Preferably, in each time slot, based on samples Update prediction model parameters ,in, For the context features of the previous time slot, This is the ratio of the energy charged in the previous time slot to the available charging energy capacity. The slope of the equivalent linear strategy for the current time slot power control strategy. As a measure of the current time slot channel quality, This represents the energy consumption value for the current time slot. This is the utility value of the current time slot.

[0023] Preferably, the prediction model employs a two-level prediction architecture to predict the four-dimensional parameter vector required for power control, wherein the two-level prediction architecture is defined as follows:

[0024]

[0025] in, This indicates that the input is the current time slot context feature. And the parameters are The prediction model, core prediction function The input is And the parameters are The output is a six-dimensional vector. , This is the predicted value of the current slot dynamic average capacity ratio. This is the predicted value of the ratio of the standard deviation of the current time slot charging energy to the available charging energy capacity. This is the predicted slope of the equivalent linear policy for the next time slot policy. This is a predicted value for the expected channel quality metric in the next time slot. This is the predicted energy consumption value for the next time slot. This is a predicted value for the expected utility of the next time slot;

[0026] Prediction post-processing function The input is Output six-dimensional vector The parameter is the weighting coefficient. The output is That is, the entire prediction model The output, where, and Calculate using the following formula:

[0027]

[0028]

[0029] in, The minimum-to-average charging energy ratio for the current time slot. As the equivalent channel quality metric for the next time slot, To find the minimum value of the function, the function Represents utility function Consuming energy and Under the condition of current time slot channel quality metric The inverse function value.

[0030] Preferably, when context features The set of values ​​that can be taken is a finite set. Then the core prediction function Defined as:

[0031]

[0032] Among them, symbols This indicates the definition of an operation, parameters. , , , , ,as well as Each of the six sub-parameters;

[0033] Based on samples Update the core prediction function parameters Specifically:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040] Among them, symbols This indicates an assignment operation. The learning rate hyperparameter is preset. , , , , ,as well as These are respectively represented as features based on the context of the previous time slot. Implemented adaptive power control data The six parameters have been updated.

[0041] Preferably, the prediction model Neural networks, support vector machines, linear regression, random forests, gradient boosting trees, or Bayesian methods, all based on samples, can be used. Update prediction model parameters ;

[0042] When using a neural network, update the parameters. Loss function used for:

[0043]

[0044] in, , , , , The prediction model is based on the context features of the previous time slot. The predicted dynamic average capacity ratio of the previous time slot, the minimum-mean charging energy ratio of the previous time slot, the equivalent linear strategy slope of the current time slot, and the equivalent channel quality metric of the current time slot; weight parameters. Weight parameters loss function , , , The loss function is selected from a variety of options, including the mean squared error loss function, the mean absolute error loss function, or the Huber loss function which includes adjustable hyperparameters; loss function For including quantile hyperparameters The quantile loss function.

[0045] Preferably, the ratio of charging energy in the previous time slot to available charging energy capacity. Calculated in the following way:

[0046]

[0047] in, To begin storing energy for the current time slot, The remaining available energy from the previous time slot, This refers to the energy capacity of the energy storage element.

[0048] Preferably, the slope of the equivalent linear strategy of the current time slot power control strategy Calculated in the following way:

[0049]

[0050] in, For the current time slot equivalent channel quality metric, This is the current time slot power control strategy.

[0051] Preferably, based on the system state at the start time of the current time slot. and the power control strategy characterized by the predicted values ​​of the four-dimensional parameter vector. Calculate the energy consumption value of the current time slot. :

[0052]

[0053] in, Let be the power decision function, and its parameters are: The input is and The output is a suggested energy consumption value. ; For the power correction function, it will Mapped to system state From the determined set of feasible energy consumption, the final executable energy consumption value is obtained. .

[0054] Preferably, the power decision function Defined as:

[0055]

[0056] in, For the equation In the closed interval The above is about energy consumption The only root, and They are respectively and The corresponding function value at that time.

[0057]

[0058] in, Representation function about The partial derivative, when hour, ;

[0059] in, and To find the minimum and maximum values ​​of the function, This refers to the energy capacity of the energy storage element.

[0060] Preferably, the method further includes: initializing the parameter vector of the prediction model. Previous time slot context features Remaining available energy in the previous time slot and the current time slot equivalent channel quality metric ;

[0061] Updating the parameter vector Then, the context features of the previous time slot will also be included. Update to current slot context features Use the remaining available energy from the previous time slot Updated to the remaining available energy in the current time slot. The current time slot equivalent channel quality metric Update to the next time slot equivalent channel quality metric .

[0062] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0063] 1. This invention completely breaks away from the classical reinforcement learning paradigm that relies on value function estimation by predicting the equivalent linear policy slope and the equivalent channel quality metric for the next time slot policy. This invention fundamentally avoids computationally intensive operations such as value function fitting, significantly reducing computational complexity. Experimental results show that the method of this invention not only has long-term average utility comparable to the current best reinforcement learning power control methods, but also has a much lower time cost than reinforcement learning algorithms.

[0064] A performance comparison of the method of this invention and the method of patent CN202510318380.1 when both energy arrival and channel fading are independently and identically distributed shows that:

[0065] (1) Compared with the performance of the optimal strategy of MDP, the maximum and average relative utility loss of the method of this invention are 2.37% and 0.37%, respectively, while the maximum and average relative utility loss of the method of patent CN202510318380.1 are 3.03% and 0.48%, respectively;

[0066] (2) Under the same test environment, the average computation time of the method of the present invention is 2.376 seconds per 100,000 decisions, while the average computation time of the method of patent CN202510318380.1 is 55.606 seconds per 100,000 decisions.

[0067] As described in the background section, the method in patent CN202510318380.1 is a customized reinforcement learning power control algorithm with performance close to the theoretical optimum and computational complexity significantly lower than existing power control methods based on general reinforcement learning.

[0068] 2. This invention distinguishes energy harvesting scenarios with different levels of fluctuation by introducing the minimum-to-average charging energy ratio of energy arrival. Based on this minimum-to-average charging energy ratio, as well as the dynamic average capacity ratio, the slope of the equivalent linear strategy in the next time slot, and the equivalent channel quality metric in the next time slot, a worst-case optimal power allocation method is designed, which effectively improves the long-term average utility of the system under different levels of energy arrival fluctuation.

[0069] Experimental results show that when the energy reaches a relatively low-fluctuation uniform and exponential distribution, the relative utility loss of the method of this invention compared to the optimal strategy of MDP is significantly lower than that of the method in patent CN202510318380.1. Under a uniform distribution, the maximum and average relative utility losses of the method of this invention are 2.37% and 0.45%, respectively, while the maximum and average relative utility losses of the method in patent CN202510318380.1 are 3.03% and 0.64%, respectively. Under an exponential distribution, the maximum and average relative utility losses of the method of this invention are 1.97% and 0.46%, respectively, while the maximum and average relative utility losses of the method in patent CN202510318380.1 are 2.35% and 0.69%, respectively.

[0070] 3. This invention uses a context-based online learning prediction model to estimate the four-dimensional parameter vector (dynamic average capacity ratio, minimum-average charging energy ratio, slope of the equivalent linear strategy in the next time slot, and equivalent channel quality metric in the next time slot) required for the power control strategy in real time, thereby effectively addressing various general energy arrival and channel fading scenarios in an adaptive manner.

[0071] Experimental results show that the method of the present invention still has near-optimal adaptive control performance under the general energy arrival and channel fading model, and its maximum relative utility loss compared with the optimal MDP strategy is less than 2%. Attached Figure Description

[0072] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0073] Figure 1 This is a flowchart of an adaptive power control method for an energy harvesting communication system provided in the embodiments. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.

[0075] In this invention, the operation of the energy harvesting communication system is modeled as a process based on discrete, equal-length periodic time slots. The performance-energy consumption relationship of the system within a single time slot is defined by a utility function of the following form:

[0076]

[0077] in, and Let represent the energy consumption and channel quality metric within the time slot, respectively. This utility function satisfies the following property: when When fixed, Regarding variables It is increasing, strictly concave, and continuously differentiable; when When fixed, Regarding variables It shows a strict increasing trend.

[0078] The energy harvesting communication system is equipped with an energy harvester and an energy capacity of [missing information]. The energy storage element follows a collection-storage-use working mode. Specifically, in each time slot, the ambient energy obtained by the energy collector must first be stored in the energy storage element; the system is powered by the energy storage element, and the energy newly stored in the energy storage element can only be used by the system from the next time slot.

[0079] The embodiment uses an energy harvesting wireless communication transmitter as an example to provide a detailed description of the adaptive power control method of the aforementioned energy harvesting communication system.

[0080] The energy harvesting wireless communication transmitter (hereinafter referred to as the transmitter) includes an energy harvester, an energy storage element (such as a rechargeable battery or supercapacitor), and a processing unit, wherein the processing unit is configured to execute the adaptive power control method of this embodiment. The transmitter operates based on a discrete and equal-length periodic time slot structure. The energy capacity of the energy storage element is... The transmitter achieves self-powering by harvesting energy from the environment and follows the harvest-store-use operating mode described above. Its energy is mainly consumed in the wireless signal transmission process of the radio frequency front end.

[0081] The transmitter is equipped with a single antenna and communicates with a receiver through a flat-fading Gaussian channel of a given bandwidth. The channel gain of this channel remains constant within a single time slot, but varies according to a certain statistical law between time slots. This embodiment aims to maximize the long-term average spectral efficiency of the transmitter. Accordingly, according to the Gaussian channel capacity formula, the efficiency of the channel is defined as follows: when the energy consumed is... The channel quality metric (i.e., the channel signal-to-noise ratio) is... At that time, the utility that the transmitter can obtain for:

[0082]

[0083] Among them, symbols Indicates the definition of an operation. This represents the unit of spectral efficiency.

[0084] At the beginning of each time slot, the processing unit can obtain the real-time system state information at the start of the current time slot, denoted as... ,in, To begin storing energy for the current time slot, As a measure of the current time slot channel quality, This includes other available state information related to power control decisions. In this embodiment, Specifically, it refers to the cumulative output energy of the energy harvester in the previous time slot, i.e., the "energy arrival" in this time slot.

[0085] like Figure 1 As shown in the embodiment, an adaptive power control method for an energy harvesting communication system includes the following steps:

[0086] S1, in each time slot, the updated prediction model is used to predict the four-dimensional parameter vector required for power control based on the historical system state up to the start of the current time slot. This includes the dynamic average capacity ratio of energy arrival in the current time slot, the minimum-average charging energy ratio of the current time slot, the equivalent linear policy slope of the power control strategy for the next time slot, and the equivalent channel quality metric for the next time slot.

[0087] In the embodiment, the parameter vector of the prediction model, the context features of the previous time slot, the remaining available energy of the previous time slot, and the equivalent channel quality metric of the current time slot are first initialized.

[0088] Among them, each time slot context feature Defined as a fixed-length sequence, denoted as , among which, the k element In a given set Take the value from, that is Length of context feature sequence It is a fixed non-negative integer, which is a preset hyperparameter. When the sequence is an empty sequence This indicates that historical system status information is not used. , It is a positive integer.

[0089] Predictive Model Based on the current time slot context features As input, this feature contains historical and current information relevant to the prediction task. When the system is in a contextualized state... When the state is as described in the prediction model For use based on The four-dimensional parameter vector required to predict the current time slot power control strategy The four components of this vector are the current time slot (defined by context features) The following parameters are used for the time slot starting from the observation time and its next time slot:

[0090] (1) Current time slot dynamic average capacity ratio (abbreviated as dynamic average capacity ratio) The dynamic average capacity ratio is defined as the ratio of the expected value of the charging energy actually charged and stored in the energy storage element to the available charging energy capacity, wherein the available charging energy capacity is defined as the difference between the energy capacity of the energy storage element and the energy currently stored in the energy storage element.

[0091] (2) Minimum-average charging energy ratio in the current time slot The minimum-average charge energy ratio is defined as the ratio of the minimum possible value to the expected value of the charge energy actually charged and stored in the energy storage element.

[0092] (3) Slope of the equivalent linear policy of the next time slot policy ;

[0093] (4) Equivalent channel quality metric for the next time slot .

[0094] It should be noted that the charging energy mentioned in this invention refers to the energy that is ultimately stored in the energy storage element and can be used for subsequent discharge.

[0095] In the embodiment, the prediction model A two-level prediction architecture is adopted, specifically defined as follows:

[0096]

[0097] in, The core prediction function has the following parameters. For prediction models Parameters to be optimized; For the prediction post-processing function, its parameters These are the preset weighting coefficient hyperparameters.

[0098] Core prediction function The input is context features The parameters are The output is a six-dimensional vector. ,in, This is the predicted value of the ratio of the standard deviation of the current time slot charging energy to the available charging energy capacity. This is a predicted value for the expected channel quality metric in the next time slot. This is the predicted energy consumption value for the next time slot. This is the predicted value of the expected utility for the next time slot.

[0099] Prediction post-processing function The input is Output six-dimensional vector The parameter is the weighting coefficient. The output is That is, the entire prediction model The output of, where, and Calculate using the following formula:

[0100] (2)

[0101] (3)

[0102] in, The minimum-to-average charging energy ratio for the current time slot. As the equivalent channel quality metric for the next time slot, To find the minimum value of the function, the function Representing utility function exist and Under what conditions regarding The inverse function value.

[0103] In this embodiment, the context features of the previous time slot are also initialized. Remaining available energy in the previous time slot and the current time slot equivalent channel quality metric These variables are also updated during the current time slot calculation process for use in the calculation of subsequent time slots.

[0104] In the embodiments, the prediction model defined above is based on the input current time slot context features. The four-dimensional parameter vector required for predictive power control is derived from the current time slot context features, which are constructed based on the historical system states up to the start of the current time slot, as shown in the following formula:

[0105] (4)

[0106] Among them, the function Generate functions for contextual features. The context feature sequence of the previous time slot The Middle One element, For a set of values ​​from the system state to a set The mapping is specifically defined as:

[0107] (5)

[0108] Component mapping For a sequence of incrementing positive real number partition points Multi-level mappings are specifically defined as:

[0109] (6)

[0110] Among them, intermediate variables The value is 1 and it is an intermediate variable. Values ,Right now intermediate variables The value is 2 and it is an intermediate variable. Values ,Right now , j Indicates the split point index. express The j One dividing point, This represents the total number of dividing points. Represents the context features of the previous time slot After element With elements Cascade.

[0111] S2, based on the system state at the start of the current time slot and the power control strategy characterized by the predicted values ​​of the four-dimensional parameter vector, calculate the energy consumption value of the current time slot.

[0112] In the embodiment, based on system state and predict four-dimensional parameter vector Through power control strategy Calculate the energy consumption value of the current time slot. :

[0113] (7)

[0114] in, Let be the power decision function, and its parameters are: The input is and The output is a suggested energy consumption value. ; For the power correction function, it will Mapped to system state From the determined set of feasible energy consumption, the optimal energy consumption value for final execution is obtained. Among them, the power decision function Based on the dynamic equal capacity ratio The minimum-to-average charging energy ratio is The approximate solution to the following optimization problem under the given conditions yields:

[0115] (8)

[0116] in, Represents the mathematical expectation. This represents the cumulative output energy of the energy harvester within the current time slot.

[0117] In one implementation, the power correction function can be defined as follows: ,in, This represents the maximum discharge energy of the energy storage element within a single time slot.

[0118] Under normal circumstances: based on the dynamic equal capacity ratio... The minimum-to-average charging energy ratio is Under the given conditions, the worst-case scenario for the cumulative output energy of the energy harvester within the current time slot is approximately a two-point distribution, which is proportional to... Values With probability Values ,in, Therefore, the optimization problem defined by formula (8) can be approximately transformed into the following optimization problem:

[0119] (9)

[0120] The optimization problem shown in formula (9) is a convex optimization problem. By solving its KKT conditions, the optimal power decision value can be derived, i.e., formula (10):

[0121] (10)

[0122] in, For the equation In the closed interval The above is about energy consumption The only root, and They are respectively and The corresponding function value at that time.

[0123] (11)

[0124] in, Representation function about The partial derivative of .

[0125] The power decision function is calculated directly based on formula (10). The value of needs to be found by solving the equation. In the closed interval The above is about energy consumption The only root This can be achieved through numerical iteration methods. When the function When a specific mathematical form is used, a pre-derived analytical quadratic formula can also be used for calculation.

[0126] In the embodiment, due to The equation Regarding variables It has a closed-form solution. The power decision function can be obtained by deriving this closed-form solution. The explicit expression is:

[0127] (12)

[0128] in, and To find the minimum and maximum values ​​of the function, Let be the energy capacity of the energy storage element. Formula (12) allows the optimal power decision to be calculated directly without any iteration, thereby significantly improving the real-time running efficiency of the algorithm.

[0129] S3, the command communication system completes the corresponding communication task in the current time slot according to the energy consumption value, and observes the utility value of completing the communication task.

[0130] In the embodiments, in each time slot, data is also based on samples. Update prediction model parameters ,in, This is the ratio of the energy charged in the previous time slot to the available charging energy capacity. The slope of the equivalent linear strategy for the current time slot power control strategy. The energy consumption of the current time slot is calculated using a power control strategy. According to the energy consumption of the current time slot The utility value of completing the corresponding communication task is observed when the corresponding communication task is completed within the current time slot.

[0131] Among them, energy storage based on time slots And the remaining available energy in the previous time slot Calculate the ratio of the charging energy in the previous time slot to the available charging energy capacity. :

[0132] (13)

[0133] Based on system state Predicted four-dimensional parameter vector and the current time slot equivalent channel quality metric Calculate the current time slot power control strategy Equivalent linear strategy slope :

[0134] (14)

[0135] In the embodiment, when context features The set of values ​​that can be taken is a finite set. Core prediction function Specifically defined as:

[0136] (1)

[0137] Among them, parameters , , , , ,as well as They are respectively The six sub-parameters.

[0138] At this point, based on the above samples Update the core prediction function parameters Specifically:

[0139] (15)

[0140] (16)

[0141] (17)

[0142] (18)

[0143] (19)

[0144] (20)

[0145] Among them, symbols This indicates an assignment operation. The learning rate hyperparameter is preset. , , , , ,as well as These are respectively represented as features based on the context of the previous time slot. Implemented adaptive power control data The six parameters have been updated.

[0146] In the embodiment, when updating the parameter vector Then, the context features of the previous time slot will also be included. Update to current slot context features Use the remaining available energy from the previous time slot Updated to the remaining available energy in the current time slot. The current time slot equivalent channel quality metric Update to the next time slot equivalent channel quality metric .

[0147] After the update, proceed to step S1 to calculate the next time slot.

[0148] Example 1

[0149] To verify the effectiveness of the method described above, simulation experiments were conducted in the following three scenarios:

[0150] Scenario 1: Energy arrival and channel quality metrics are both independent and identically distributed.

[0151] Scenario 2: Energy arrival is a Markov process, and the channel quality metric is independent and identically distributed.

[0152] Scenario 3: Both energy arrival and channel quality metrics are Markov processes.

[0153] The experiment uses the optimal strategy based on MDP theory as the performance benchmark and compares its performance with the method proposed in patent CN202510318380.1 in the background technology under scenario 1. The specific models and parameters involved in the simulation experiment are as follows:

[0154] (1) Energy arrival model.

[0155] Scenario 1: Independent and identically distributed (Bernoulli distribution, exponential distribution, uniform distribution).

[0156] First, define three basic parameters: Average Capacity Ratio (MCR): the available charging energy capacity is... Dynamic average capacity ratio under certain conditions; Nominal average capacity ratio (NMCR): the ratio of expected energy arrival to the energy capacity of the energy storage element. The ratio; nominal signal-to-noise ratio (NSNR): The unit is decibel (dB).

[0157] Based on the definitions of MCR and NSNR, it can be deduced that:

[0158] (twenty one)

[0159] Based on the above parameters and formulas, the three energy arrival distributions are specifically expressed as follows:

[0160] Bernoulli distribution: it is based on probability Values With probability Values , and its .

[0161] Exponential distribution: its probability density function is , , and its ,in, .

[0162] Uniform distribution: its probability density function is , , and its ,in, .

[0163] Scenario 2 and Scenario 3: A value taken from The first-order time-sigma Markov process has the following transition probability matrix: ,in, .

[0164] (2) Channel quality measurement model.

[0165] Scenario 1 and Scenario 2: Exponential distribution model (corresponding to the memoryless Rayleigh fading channel model), its probability density function is... , .

[0166] Scenario 3: A first-order time-straight Markov process model, whose value set is as follows: The transition probability matrix is .

[0167] (3) Communication system parameters:

[0168] Scenario 1: Energy capacity of energy storage devices The maximum discharge energy is given by formula (21). .

[0169] Scenario 2 and Scenario 3: Energy capacity of energy storage components Maximum discharge energy .

[0170] (4) Algorithm hyperparameters:

[0171] Scenario 1: Context Feature Length Learning rate Weighting coefficient .

[0172] Scenario 2 and Scenario 3: Context Feature Length Learning rate Weighting coefficient .

[0173] As shown in Table 1, the general conditions optimized by the method in Scenario 1 and Patent CN202510318380.1 are as follows ( Under nominal signal-to-noise ratio (SNR), both methods achieve near-optimal performance (with a maximum utility loss of approximately 1% relative to the optimal MDP strategy). However, under more comprehensive conditions including negative SNR (…),… Under the nominal signal-to-noise ratio (SNR), the overall performance of the method of this invention is significantly better than that of the method in patent CN202510318380.1. Specifically, the maximum utility loss of the method of this invention relative to the optimal MDP strategy is less than 2.4%, while the maximum utility loss of the method in patent CN202510318380.1 exceeds 3%. In terms of energy arrival distribution, the method of this invention outperforms the method under exponential and uniform distributions, but is slightly inferior to the method in patent CN202510318380.1 under the Bernoulli distribution. Furthermore, as shown in Table 2, the measured runtime results under the same test environment (CPU: Intel i9-12900H; Operating system: Ubuntu 20.04; Programming language: Python 3.10) show that the runtime of the method of this invention is significantly shorter than that of the method in patent CN202510318380.1, thus its actual computational overhead is lower.

[0174] Table 1. Comparison of utility loss (percentage) of the method of this invention and the method of patent CN202510318380.1 relative to the optimal strategy of MDP in scenario 1.

[0175]

[0176] Table 2 Comparison of computation time between the method of this invention and the method of patent CN202510318380.1 in scenario 1

[0177]

[0178] As shown in Table 3, in more complex scenarios (Scenario 2 and Scenario 3) where energy arrival or channel fading is a Markov process, the method of the present invention can still maintain excellent performance. Its maximum utility loss relative to the optimal MDP strategy is less than 2%, which fully verifies the universal near-optimal performance of the method of the present invention.

[0179] Table 3. Utility loss (percentage) of the method of the present invention relative to the optimal MDP strategy in scenarios 2 and 3.

[0180]

[0181] Example 2

[0182] Example 2 also provides another implementation of the method of the present invention, the core difference of which lies in the prediction model. The implementation method is defined more broadly. Furthermore, in this embodiment, the context feature generation function `context` is defined using... The function is defined as an identity function, meaning it does not perform any processing on the system state and directly uses the state before (and including) the start time of the current time slot. The historical system state serves as the context feature of the current time slot.

[0183] Specifically, prediction models It is defined as a machine learning model that can learn from data and perform specific mappings. Its essential function is to learn from the contextual features of the input. Output a four-dimensional parameter vector Predictive model parameters It can be obtained through training and can be updated online using real-time data during system operation, thereby achieving adaptive optimization.

[0184] To illustrate more clearly, the prediction model can be used to implement the above. The algorithms include, but are not limited to: neural networks, support vector machines, linear regression, random forests, gradient boosting trees, and Bayesian methods. The training process is the parameter update process described in Example 1, i.e., based on samples... Update prediction model parameters This can be an online update of a single sample in each time slot, or a periodic batch update based on historical sample data. The loss function used when employing a neural network approach can be illustrated as follows:

[0185] (twenty two)

[0186] in, , , , , The prediction model is based on the context features of the previous time slot. The predicted dynamic average capacity ratio of the previous time slot, the minimum-mean charging energy ratio of the previous time slot, the equivalent linear strategy slope of the current time slot, and the equivalent channel quality metric of the current time slot; weight parameters. Weight parameters loss function , , , A variety of loss functions can be selected, including but not limited to the mean squared error loss function, the mean absolute error loss function, or the Huber loss function which includes adjustable hyperparameters; loss function For including quantile hyperparameters Quantile loss function:

[0187] (twenty three)

[0188] The remaining steps in this embodiment are the same as those in the above method, and all follow the same procedure. Figure 1 The overall process is shown below. By adopting a broader model definition, this embodiment provides greater flexibility for practical system implementation, allowing for trade-offs between algorithm complexity, prediction accuracy, and computational resources based on specific scenarios.

[0189] Example 3

[0190] This embodiment 3, based on embodiment 1, further considers the energy loss that occurs in the energy storage element during actual discharge. This loss can typically be represented by a discharge efficiency coefficient. This is used to characterize the ratio of the actual usable energy at the load end to the energy drawn from the energy storage element.

[0191] Specifically, the utility function described in Example 1 only needs to be modified as follows:

[0192] (twenty four)

[0193] After the above modifications, the objective function and constraints of the power control optimization problem remain essentially unchanged in form; therefore, the method described in Example 1 is still fully applicable. The only difference lies in the power decision function. The explicit expression differs. Specifically, in this embodiment, the explicit expression for the optimal power decision value is:

[0194] (25).

[0195] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive power control method for an energy harvesting communication system, characterized in that, Includes the following steps: In each time slot, the updated prediction model is used to predict the four-dimensional parameter vector required for power control based on the historical system state up to the start of the current time slot. This includes the dynamic average capacity ratio of energy arrival in the current time slot, the minimum-average charging energy ratio of the current time slot, the equivalent linear strategy slope of the power control strategy in the next time slot, and the equivalent channel quality metric for the next time slot. Based on the system state at the start of the current time slot and the power control strategy characterized by the predicted values ​​of the four-dimensional parameter vector, calculate the energy consumption value of the current time slot; The command communication system completes the corresponding communication task in the current time slot according to the energy consumption value, and observes the utility value of completing the communication task; Specifically, the updated prediction model uses historical system states up to the start of the current time slot to predict the four-dimensional parameter vector required for power control, including: The historical system states up to the start of the current time slot are constructed as the context features of the current time slot, and an updated prediction model is used based on these context features. The four-dimensional parameter vector required for predictive power control; where, This indicates that the input is a contextual feature. And the parameters are Predictive models; The context features For a fixed length of A sequence, where each element of the sequence is in a given set. The value is taken from the top; where, If it is a non-negative integer, when When the sequence is an empty sequence This indicates that historical system status information is not used; Specifically, the historical system states up to the start of the current time slot are constructed as the context features of the current time slot, including: The current time slot context features are calculated based on the system state at the start of the current time slot and the context features of the previous time slot, including: in, For the current time slot context features, the function Generate functions for contextual features. For the context features of the previous time slot, This refers to the system state at the start of the current time slot, including the initial stored energy of the current time slot. Current time slot channel quality metric Other available state information related to power control decisions ,symbol Indicates the definition of an operation. The context feature sequence of the previous time slot The Middle One element, For a set of values ​​from the system state to a set The mapping.

2. The adaptive power control method for an energy harvesting communication system according to claim 1, characterized in that, In each time slot, based on samples Update prediction model parameters ,in, For the context features of the previous time slot, This is the ratio of the energy charged in the previous time slot to the available charging energy capacity. The slope of the equivalent linear strategy for the current time slot power control strategy. As a measure of the current time slot channel quality, This represents the energy consumption value for the current time slot. This is the utility value for the current time slot.

3. The adaptive power control method for an energy harvesting communication system according to claim 2, characterized in that, The prediction model employs a two-stage prediction architecture to predict the four-dimensional parameter vector required for power control, wherein the two-stage prediction architecture is defined as follows: in, This indicates that the input is the current time slot context feature. And the parameters are The prediction model, core prediction function The input is And the parameters are The output is a six-dimensional vector. , This is the predicted value of the current slot dynamic average capacity ratio. This is the predicted value of the ratio of the standard deviation of the current time slot charging energy to the available charging energy capacity. This is the predicted slope of the equivalent linear policy for the next time slot policy. This is a predicted value for the expected channel quality metric in the next time slot. This is the predicted energy consumption for the next time slot. This is a predicted value for the expected utility of the next time slot; Prediction post-processing function The input is Output six-dimensional vector The parameter is the weighting coefficient. The output is That is, the entire prediction model The output of, where, and Calculate using the following formula: in, The minimum-to-average charging energy ratio for the current time slot. As the equivalent channel quality metric for the next time slot, To find the minimum value of the function, the function Represents utility function Consuming energy and Under the condition of current time slot channel quality metric The inverse function value.

4. The adaptive power control method for an energy harvesting communication system according to claim 3, characterized in that, When context features The set of values ​​that can be taken is a finite set. Then the core prediction function Defined as: Among them, symbols This indicates the definition of an operation, parameters. , , , , ,as well as Each of the six sub-parameters; Based on samples Update the core prediction function parameters Specifically: Among them, symbols This indicates an assignment operation. The learning rate hyperparameter is preset. , , , , ,as well as These are respectively represented as features based on the context of the previous time slot. Implemented adaptive power control data The six parameters have been updated.

5. The adaptive power control method for an energy harvesting communication system according to claim 2, characterized in that, The prediction model Neural networks, support vector machines, linear regression, random forests, gradient boosting trees, or Bayesian methods, all based on samples, can be used. Update prediction model parameters ; When using a neural network, update the parameters. Loss function used for: in, , , , , The prediction model is based on the context features of the previous time slot. The predicted dynamic average capacity ratio of the previous time slot, the minimum-mean charging energy ratio of the previous time slot, the equivalent linear strategy slope of the current time slot, and the equivalent channel quality metric of the current time slot; weight parameters. Weight parameters loss function , , , The loss function is selected from a variety of options, including the mean squared error loss function, the mean absolute error loss function, or the Huber loss function which includes adjustable hyperparameters; loss function For including quantile hyperparameters The quantile loss function.

6. The adaptive power control method for an energy harvesting communication system according to claim 2, characterized in that, The ratio of charging energy in the previous time slot to available charging energy capacity Calculated in the following way: in, To begin storing energy for the current time slot, The remaining available energy from the previous time slot, This refers to the energy capacity of the energy storage element.

7. The adaptive power control method for an energy harvesting communication system according to claim 2, characterized in that, The slope of the equivalent linear strategy of the current time-slot power control strategy Calculated in the following way: in, For the current time slot equivalent channel quality metric, This is the current time slot power control strategy.

8. The adaptive power control method for an energy harvesting communication system according to claim 1, characterized in that, Based on the system state at the start of the current time slot and the power control strategy characterized by the predicted values ​​of the four-dimensional parameter vector, calculate the energy consumption value of the current time slot. : in, The system state includes the initial stored energy of the current time slot. Current time slot channel quality metric Other available state information related to power control decisions , For power control strategy, Let be the power decision function, and its parameters are: The input is and The output is a suggested energy consumption value. ; For the power correction function, it will Mapped to system state From the determined set of feasible energy consumption, the final executable energy consumption value is obtained. .

9. The adaptive power control method for an energy harvesting communication system according to claim 8, characterized in that, Power decision function Defined as: in, For the equation In the closed interval The above is about energy consumption The only root, and They are respectively and The corresponding function value at that time. in, Representation function about The partial derivative, when hour, ; in, and To find the minimum and maximum values ​​of the function, This refers to the energy capacity of the energy storage element.

Citation Information

Patent Citations

  • Adaptive modulation and power control system based on energy collection and optimization method thereof

    CN111491358A

  • Power control method for energy harvesting wireless communication system

    CN119854923A