Wireless sensor data transmission power control method and system

The energy and data transmission strategies of wireless sensors are optimized by a flexible actor-critic algorithm, which solves the energy and data transmission balance problem in wireless sensor networks, improves transmission efficiency and reduces costs.

CN118695219BActive Publication Date: 2025-09-26ZHEJIANG NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410742529.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-09-26
Estimated Expiration
2044-06-11

AI Technical Summary

Technical Problem

Due to the dynamic nature of the environment and the randomness of energy harvesting, wireless sensor networks find it difficult to effectively balance data transmission and energy consumption, resulting in increased network lifespan and operating costs.

Method used

The flexible actor-critic (SAC) algorithm is combined with the Markov decision process to optimize the energy and data transmission strategies of wireless sensors. An energy collection, data collection and transmission model is established, and a reward function is designed to control the transmission power.

Benefits of technology

In the case of unknown future information, the energy usage strategy of wireless sensors is optimized, data transmission efficiency is improved, network life is extended and operation costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118695219B_ABST
    Figure CN118695219B_ABST
Patent Text Reader

Abstract

The present invention aims to solve the problem that existing wireless sensors have poor strategic allocation between changing environmental energy and transmission data, making it difficult to maximize the long-term average data collection throughput. The patent of the present invention discloses a wireless sensor data transmission power control method and system, which expresses the data transmission problem of wireless sensors as a Markov decision process and then solves it using the flexible actor-critic algorithm (SAC) method. When the future information of energy, data, and channel status is unknown, the algorithm can learn and optimize the energy usage strategy by tracking the specific energy, data size, and channel status information received in each time slot. The action sequence generated by this strategy can maximize the long-term data throughput of the sensor node, and the optimal value can be obtained by limiting the conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data transmission control, and in particular to a method and system for controlling data transmission power of a wireless sensor. Background Art

[0002] In recent years, advancements in low-power and miniaturized electronics have led to a dramatic increase in the use of the Internet of Things (IoT). These widely deployed wireless devices are typically powered by batteries, which often have limited capacity, creating a significant performance bottleneck for implementing large-scale wireless communication networks. In particular, in certain environments, regular manual battery replacement and wired charging are costly, cumbersome, dangerous, and even impossible in some cases, such as implanted sensors. This significantly limits the lifespan of wireless networks and increases network operating costs.

[0003] Energy harvesting (EH) technology enables wireless sensor networks (WSNs) to be self-sustainable in order to maintain their long-term key performance indicators, such as data throughput, as shown in the patent application publication number CN117858188A. Due to the high dynamism and complexity of the environment, a stable energy supply cannot be provided, and an efficient learning algorithm is required to enable the system to adapt to the energy supply of the environment. Wireless communication nodes with EH capabilities have the potential to achieve self-sustainability and permanent operation. Communication nodes can collect piezoelectric, radio frequency (RF) energy, or use natural energy such as solar energy, thermal energy, wind energy, etc. to charge the battery in an environmentally friendly way, and then use the collected energy to transmit data, such as Figure 1 However, the arrival of natural energy is highly random, and the environment in which EH nodes operate is dynamically changing. Consequently, the energy harvesting process exhibits both randomness and dynamic characteristics. These uncertainties can affect the critical performance of sensor networks. Therefore, a control method and system are needed that can balance wireless sensor data transmission and energy consumption. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies of the prior art and to provide a method and system for controlling data transmission power of a wireless sensor.

[0005] In order to solve the above problems, the present invention adopts the following technical solution: a wireless sensor data transmission power control method, characterized by comprising the following steps:

[0006] Step 1: Setting up a communication connection between a wireless sensor and a data receiving node, wherein the wireless sensor includes a storage module for data storage and a battery for energy storage;

[0007] Step 2: The wireless sensor collects data from the outside world and stores it in its own memory. At the same time, the wireless sensor also obtains energy from the outside world to charge the battery.

[0008] Step 3: Calculate the amount of data and energy stored in the wireless sensor at set intervals;

[0009] Step 4: Establish an optimization problem of data transmission and energy consumption. The goal is to maximize the data transmission volume of wireless sensors within a set time period under the premise of energy sustainability.

[0010] Step 5: Formulate the optimization problem in step 4 as a Markov decision process, and then use the flexible actor-critic algorithm to solve the process to obtain the data transmission strategy of the wireless sensor at the current moment;

[0011] Step 6: The wireless sensor executes the data transmission strategy and transmits data to the data receiving node.

[0012] Furthermore, the process of expressing the optimization problem as a Markov decision in step 5 includes the following steps:

[0013] Step 51: Based on the Markov decision process, a model is built for the energy collection and data transmission process of the wireless sensor. The model is expressed as MDP M = {S, A, P, R, γ}, where S represents all possible states s in the state space; A represents all possible actions a in the action space; P represents the state transition probability, which is the state s of the wireless sensor at time t. t Next, perform action a t After that, the state transition occurs and transfers to state s t+1 The probability of; R represents the set reward value; the parameter γ represents the discount factor, which is the importance of future rewards;

[0014] Step 52: Building the Energy Harvesting Model Data Collection Model and data transmission model R i , and based on the energy collection model, data collection model and data transmission model, define the problem statement of the wireless sensor data transmission strategy; where the energy collection model represents the amount of electricity stored in the battery at the beginning of the current time slot i, the data collection model represents the amount of data stored by the wireless sensor at the beginning of the current time slot i, and the data transmission model represents the data transmission rate in time slot i;

[0015] Step 53: Set the reward function of the reward value R;

[0016] Step 54: Setting up a model algorithm based on the flexible actor-critic algorithm;

[0017] Step 55: Building and training a model based on the model algorithm;

[0018] Step 56: Run the model to obtain the data transmission strategy of the wireless sensor at the current moment, and control the output power of the wireless sensor's transmission data.

[0019] The energy harvesting model in step 52 Using the harvest-store-use (HSU) architecture, it is expressed as formula (1):

[0020]

[0021] in is the energy harvesting model of the wireless sensor, which represents the amount of electricity at the beginning of the current i-th time slot; τp i-1 is the energy consumed in the previous time slot i-1; p i-1 is the wireless sensor transmission power of the previous time slot i-1; E i-1 is the energy collected in the battery in the previous time slot i-1; E max The upper limit of the battery capacity of the wireless sensor.

[0022] The data collection model in step 52 The harvest-store-use (HSU) architecture is adopted, which is expressed as formula (2):

[0023]

[0024] in represents the amount of data stored in the wireless sensor storage module at the beginning of the current i-th time slot; τR i-1 The amount of data sent in the previous time slot i-1; R i is the transmission rate of the i-th time slot; D i-1 is the amount of data obtained in the i-1th time slot; D max The upper limit of the data cache capacity of the storage module.

[0025] The data transmission model R in step 52 i , expressed as formula (3):

[0026]

[0027] Among them, R i represents the data transmission rate in time slot i; h i represents the channel gain; σ 2 represents the variance of the noise.

[0028] The problem of wireless sensor data transmission strategy in step 52 can be expressed as formula (4):

[0029]

[0030] Among them, R i represents the data transmission rate in time slot i; h i represents the channel gain; σ 2 represents the variance of the noise; τ represents a time interval; N represents a natural number.

[0031] The reward function of the reward value R in step 53 includes the single-step data transmission amount r1, the action energy excess penalty r2, the action data excess penalty r3 and the energy overflow penalty r4; wherein the single-step data transmission amount r1 represents the amount of data transmitted by the wireless sensor in a single-step action; the action energy excess penalty r2 is related to the distance between the action energy and the boundary value of the set energy range, and the action energy represents the energy consumed by the wireless sensor to transmit a set amount of data during the action; the action data excess penalty r3 is related to the amount of data transmitted by the wireless sensor during the action and the amount of data remaining in the cache; the energy overflow penalty r4 is related to the energy overflowed by the wireless sensor and the amount of untransmitted data stored by the wireless sensor when the energy overflows.

[0032] In step 54, the model algorithm is set based on the flexible actor-critic algorithm, as shown in formula (9):

[0033]

[0034] Among them, the content expressed by formula (9) is that the mean value of the constraint entropy is greater than Under the premise of maximizing expected return; r(s t ,a t ) means in state s t Next take action a t Instant rewards; represents the mean of the cumulative returns obtained under strategy π; Medium (s t ,a t )~ρ π Represents a state-action pair (s t , a t ) is taken from the state-action distribution ρ π , which means that we need to calculate the state-action distribution ρ caused by the policy π π The mean of the negative log-probability of the action, where represents the expected value; π t (a t |s t ) indicates that in the strategy π t Next, in state s t When selecting action a t probability; is the set target value.

[0035] Formula (9) is simplified by the Lagrange multiplier method, and the loss function of the coefficient α is obtained as follows:

[0036]

[0037] Among them, s t ~R represents state s t is sampled from the experience replay pool R; a t ~π(s t ) indicates action a t is from policy π, in state s t Downsampled.

[0038] A wireless sensor data transmission power control system, comprising:

[0039] A data receiving node and a wireless sensor, wherein the wireless sensor includes a memory for storing data and a battery for storing electricity;

[0040] A collection module, used to collect the remaining energy and data volume in the wireless sensor;

[0041] A training module is used to obtain wireless sensor data and train models. The training module is used to establish an optimization problem. The goal of the optimization problem is to maximize the data transmission volume of wireless sensors within a set time while maintaining energy sustainability. A Markov decision process is established for the optimization problem, and then the flexible actor-critic algorithm is used to solve the process.

[0042] The beneficial effects of the present invention are:

[0043] By formulating the problem as a Markov decision process and then solving it using the flexible actor-critic algorithm (SAC) method, and through some reinforcement learning techniques and careful design of the reward function, a policy model is obtained. The final policy can learn and optimize the energy usage strategy by tracking the specific energy, data size, and channel state information received in each time slot when the future information of energy, data, and channel state is unknown, so as to achieve better results. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Models for energy and data harvesting and data transmission for wireless sensors;

[0045] Figure 2 This is a diagram of the SAC algorithm architecture of Example 1;

[0046] Figure 3 The long-term average throughput varies with the number of iterations in the case of unlimited data in Example 1;

[0047] Figure 4The long-term average throughput varies with the number of iterations in the case of limited data in Example 1. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0049] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. Therefore, the figures only show components related to the present invention and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0050] Example 1:

[0051] like Figure 2 As shown, a wireless sensor data transmission power control method includes the following steps:

[0052] Step 1: Setting up a communication connection between a wireless sensor and a data receiving node, wherein the wireless sensor includes a storage module for data storage and a battery for energy storage;

[0053] Step 2: The wireless sensor collects data from the outside world and stores it in its own memory. At the same time, the wireless sensor also obtains energy from the outside world to charge the battery.

[0054] Step 3: Calculate the amount of data and energy stored in the wireless sensor at set intervals;

[0055] Step 4: Establish an optimization problem of data transmission and energy consumption. The goal is to maximize the data transmission volume of wireless sensors within a set time period under the premise of energy sustainability.

[0056] Step 5: Formulate the optimization problem in step 4 as a Markov decision process, and then use the flexible actor-critic algorithm to solve the process to obtain the data transmission strategy of the wireless sensor at the current moment;

[0057] Step 6: The wireless sensor executes the data transmission strategy and transmits data to the data receiving node.

[0058] The process of expressing the optimization problem as a Markov decision in step 5 includes the following steps:

[0059] Step 51: Based on the Markov decision process, a model is built for the energy collection and data transmission process of the wireless sensor. The model is expressed as MDP M = {S, A, P, R, γ}, where S represents all possible states s in the state space; A represents all possible actions a in the action space; P represents the state transition probability, which is the state s of the wireless sensor at time t. t Next, perform action a t After that, the state transition occurs and transfers to state s t+1 The probability of R being the set reward value; the parameter γ is the discount factor, which is the importance attached to future rewards;

[0060] Step 52: Building the Energy Harvesting Model Data Collection Model and data transmission model R i , and based on the energy collection model, data collection model and data transmission model, define the problem statement of the wireless sensor data transmission strategy; where the energy collection model represents the amount of electricity stored in the battery at the beginning of the current time slot i, the data collection model represents the amount of data stored by the wireless sensor at the beginning of the current time slot i, and the data transmission model represents the data transmission rate in time slot i;

[0061] Step 53: Setting a reward function for the reward value R; the reward function includes a single-step data transmission amount r1, an action energy excess penalty r2, an action data excess penalty r3, and an energy overflow penalty r4; wherein the single-step data transmission amount r1 represents the amount of data transmitted by the wireless sensor in a single-step action; the action energy excess penalty r2 is related to the distance between the action energy and the boundary value of the set energy range; the action energy represents the energy consumed by the wireless sensor to transmit the set amount of data in the action; the action data excess penalty r3 is related to the amount of data transmitted by the wireless sensor in the action and the amount of data remaining in the cache; the energy overflow penalty r4 is related to the energy overflowed by the wireless sensor and the amount of untransmitted data stored by the wireless sensor when the energy overflowed;

[0062] Step 54: Setting up a model algorithm based on the flexible actor-critic algorithm;

[0063] Step 55: Building and training a model based on the model algorithm;

[0064] Step 56: Run the model to obtain the data transmission strategy of the wireless sensor at the current moment, and control the output power of the wireless sensor's transmission data.

[0065] In step 51, a state s in the state space S is expressed as in Indicates the remaining power of the wireless sensor; h iIndicates the channel gain state of the wireless sensor; E i R represents the random arrival energy when wireless sensors transmit data, that is, the energy consumed in transmitting data; Indicates the amount of remaining data in the wireless sensor; D i Represents the random arrival data volume of wireless sensor transmission data. According to the Markov characteristic, the state s t Inheriting all the historical state information before time t, the wireless sensor can make the next power allocation decision based on this historical state information to form a new state s t+1 .

[0066] The action a contained in the action space A in step 51 is the transmission power p of the wireless sensor at the corresponding time sequence. i , represents the power of data transmission; it should be noted that at a certain moment, the transmission power p of the wireless sensor i is a constant, because there is a nonlinear relationship between the transmission power and the remaining data amount of the wireless sensor. Therefore, when the transmission power is constant, the wireless sensor has the highest energy utilization efficiency.

[0067] The state transition probability P in step 51 characterizes the system dynamics over the time slot. Since both the state space S and the action space A are continuous and infinite, the space of the state transition probability P is also continuous and infinite. Therefore, the optimal power allocation strategy for wireless sensors cannot be derived using traditional offline optimization techniques.

[0068] Since the wireless sensor continuously collects renewable energy from the environment when it is working and stores it in its battery for subsequent use, the energy collection model in step 52 The harvest-store-use (HSU) architecture is adopted, which is expressed as formula (1):

[0069]

[0070] in is the energy harvesting model of the wireless sensor, which represents the amount of electricity at the beginning of the current i-th time slot; τp i-1 is the energy consumed in the previous time slot i-1; p i-1 is the wireless sensor transmission power of the previous time slot i-1; E i-1 is the energy collected in the battery in the previous time slot i-1; E max The upper limit of the battery capacity of the wireless sensor.

[0071] Similarly, the data collection model The harvest-store-use (HSU) architecture is also adopted, which is expressed as formula (2):

[0072]

[0073] in represents the amount of data stored in the wireless sensor storage module at the beginning of the current i-th time slot; τR i-1 The amount of data sent in the previous time slot i-1; R i is the transmission rate of the i-th time slot; D i-1 is the amount of data obtained in the i-1th time slot; D max The upper limit of the data cache capacity of the storage module.

[0074] Assume that the noise is independent and identically distributed zero-mean additive white Gaussian noise (AWGN) with variance σ 2 ; In addition, the transmission power p i In each time interval τ, the energy collected by the wireless sensor is only used to transmit data to the data receiving node. The amount of data transmitted in each time slot of the wireless sensor is the transmission power p. i and channel gain h i Jointly determined, so the data transmission model R i , expressed as formula (3):

[0075]

[0076] Among them, R i represents the data transmission rate in time slot i; h i represents the channel gain; σ 2 represents the variance of the noise.

[0077] The purpose of this model is to maximize the long-term average data throughput from the sensor to the data receiver, so the problem of wireless sensor data transmission strategy can be expressed as formula (4):

[0078]

[0079] Where τ represents a time interval and N represents a natural number. The transmission power p of the wireless sensor i It is also affected by the amount of power and data, expressed as

[0080]

[0081] The reward R in step 53 primarily represents the amount of data r1 transmitted by the wireless sensor in each time slot, as the strategy aims to maximize the amount of data transmitted by the wireless sensor within a time period. Furthermore, an action violation penalty and an energy overflow penalty r4 are set as negative rewards. The action violation penalty consists of an action energy limit penalty r2 and an action data limit penalty r3. The negative reward is set to achieve algorithm convergence and maximize the amount of data transmitted by the wireless sensor r1. The action energy limit penalty r2 and action data limit penalty r3 are set because the actions of the wireless sensor are limited by both battery power and data volume. Therefore, it is very easy for the wireless sensor to exceed the action energy constraint during the action. Furthermore, to ensure smooth model training, the action energy limit penalty r2 is set. The energy overflow penalty r4 is also multiplied by a set value representing the current amount of remaining data. The greater the amount of remaining data, the more accelerated the data transmission needs to be, and the greater the penalty for energy overflow.

[0082] The flexible actor-critic algorithm (SAC) in step 54 is a reinforcement learning algorithm that works well in the field of continuous control. The SAC algorithm combines the advantages of policy gradient and value function, and balances the problem of exploration and utilization by maximizing the information entropy of the environment. The network architecture of the SAC algorithm is as follows: Figure 2 As shown in Figure 1, it consists of one actor network, two critic networks, and two critic target networks. The actor network receives the input state and outputs the action and the corresponding entropy. The critic network evaluates the Q value of the current state-action pair. The critic target network provides a more stable Q value target, and its parameters are gradually updated from the main Q network through a delayed update method.

[0083] In addition, the idea of ​​maximum entropy reinforcement learning (RL) is to maximize the cumulative reward while making the strategy more random. An entropy regularization term is added to the reinforcement learning goal to control the randomness of the strategy. In summary, the optimal strategy π* is expressed as formula (5):

[0084]

[0085] Among them, π* represents the optimal strategy; r(s t ,a t ) means in state s t Next take action a t The immediate reward of Where X is a random variable, and its probability density function is p(x); α is a regularization coefficient used to control the importance of entropy. The larger α is, the stronger the exploration ability of the wireless sensor is, which helps to find a better strategy and reduce the possibility of the strategy falling into a poor local optimum; π(·|s t ) means in state s t Next, the probability distribution of actions selected according to the strategy π; represents the mathematical expectation, represents the expected value according to the policy π.

[0086] The policy network is considered as an "actor" that outputs actions in the continuous action space, while the Q network is considered as a "critic" that evaluates the value of the actions output by the policy network. In the initial stage, the Q network does not know whether the actions output by the policy network are good or bad, and needs to use the temporal difference (TD) method to learn the value of these actions. The function Q(s) of the state action value is t ,a t ) indicates that from the current state s t Start executing action a t The expected reward that can be obtained at the end of the round is used to measure the quality of performing an action in a given state. State value function V(s t ) is defined as t Under this condition, if we continue to execute according to the strategy π, we can obtain the expected cumulative reward value; the state value function V(s t ) represents the degree of execution of the strategy in a given state and is used to evaluate the value of the state s. The state value function V(s t ) is expressed as formula (6):

[0087]

[0088] Where V(s t ) represents state s t The state value of π(a t |s t ) means that under the strategy π, in state s t When selecting action a t probability; Indicates about action a t Calculate the expected value according to the probability distribution of strategy π; Q(s t ,a t ) is the state-action-value function.

[0089] In the calculation of the Critic, which is the Q-value network, the SAC algorithm uses a target network and experience replay mechanism, similar to DQN. Based on the idea of ​​Double DQN, the SAC algorithm uses two Q networks representing the action value function. Each time a Q network is used, a network with a smaller Q value is selected to alleviate the problem of overestimation of the Q value. In order to make the training parameter update process more stable, SAC also uses the target network mechanism, that is, to generate two target Q networks. The two Q networks are updated using soft updates. One-to-one correspondence. When calculating the loss function of the Q network, use Calculate the Q target estimate part of the latter term in the loss function. The loss function L of any Q network Q (ω) is formula (7):

[0090]

[0091] Among them, (s t ,a t ,r t ,s t+1 )~R represents the historical data taken from the experience replay pool R; a t+1 ~π θ (s t+1 ) indicates action a t+1 is from the strategy π θ In state s t+1 Downsampled; Q ω represents the estimated value of any Q network under parameter ω, Q ω (s t ,a t ) means in state s t Execute a t Action This state action pair (s t ,a t ) is the estimated value of Q under the parameter ω. Represents the estimated value of the target Q network. It should be noted that in this example, there are two target networks corresponding to the Q network of the action value function, which are represented as and Similarly, in the process of updating the Q network using the estimated value of the target network,

[0092] Select The smaller value of the two reduces the estimation bias; γ represents the discount factor; Indicates state s t +1 state value; π(a t+1 |s t+1 ) means that under the strategy π, in state s t+1 When selecting action a t +1 probability; r t Indicates immediate reward.

[0093] For a continuous action space environment, the SAC algorithm's strategy outputs the mean and standard deviation of a Gaussian distribution, but the process of sampling actions based on a Gaussian distribution is not differentiable. Here, we use the reparameterization technique, first sampling from a standard normal distribution, then multiplying the sampled value by the standard deviation and adding the mean, so that it can be considered as sampling from the strategy Gaussian distribution, and this is differentiable for the strategy function. Specifically, according to the current state s t , the Actor network obtains the Gaussian distribution parameter μ(s t ,θ) and σ(s t ,θ). Output action a t =μ(s t ,θ)+σ(s t ,θ)·∈ t , where ∈ t ~N(0,1)) is the noise sampled from the standard normal distribution, which is expressed as a t =f θ (∈ t ,s t ). This can be considered as sampling from the policy Gaussian distribution, and it is differentiable with respect to the policy function. In this embodiment, Ornstein-Uhlenbeck noise suitable for continuous action control is also used to increase the exploration ability of the action. The loss function of the policy is shown in (8):

[0094]

[0095] Among them, s t ~R represents state s t is sampled from the experience replay pool R; ∈ t ~N represents noise∈ t is sampled from a standard normal distribution N; the strategy π θ The θ in θ represents the parameters of the policy network; π θ (f θ (∈ t ,s t )|s t ) represents the policy network π in s t Output f in state θ The probability of the function, f θ Defined as f θ (∈ t ,s t )=μ(s t ,θ)+σ(s t ,θ)·∈ t , where μ and σ are the strategy π θ In state s t The mean and variance of the Gaussian distribution of the output; Q ωj (st ,f θ (∈ t ,s t )) represents the two action value functions Q network in state s t Output f θ The expected value reward that the function can obtain.

[0096] In the flexible actor-critic algorithm, the choice of the entropy coefficient α is very important. In this example, in a state s where the optimal action a of the wireless sensor is uncertain, the wireless sensor's exploration of the environment is insufficient. At this time, the entropy value should be larger to increase the wireless sensor's exploration of the environment. After the environment has been fully explored, a smaller entropy value is needed to seek the optimal action. The exploration of the environment can be understood as the rate at which the wireless sensor collects data and obtains power. In order to automatically adjust the entropy regularization term, SAC rewrites the reinforcement learning objective into a constrained optimization problem as shown in Equation (9):

[0097]

[0098] Among them, the content expressed by formula (9) is that the mean value of the constraint entropy is greater than Under the premise of maximizing expected return; r(s t ,a t ) means in state s t Next take action a t Instant rewards; represents the mean of the cumulative returns obtained under strategy π; Medium (s t ,a t )~ρ π Represents a state-action pair (s t , a t ) is taken from the state-action distribution ρ π , which means that we need to calculate the state-action distribution ρ caused by the policy π π The mean of the negative log-probability of the action, where represents the expected value; π t (a t |s t ) indicates that in the strategy π t Next, in state s t When selecting action a t probability; is the set target value.

[0099] Next, we can obtain the loss function of α by simplification. The simplification process is shown in Equations (10) to (12):

[0100] By using the Lagrange multiplier method, the constrained optimization problem is converted into an unconstrained optimization problem. Introduce the Lagrange multiplier, in this case the coefficient α is used as the multiplier; rewrite its objective function (9) into the new Lagrange objective function

[0101]

[0102] By expanding the Lagrangian objective function above You can get:

[0103]

[0104] In order to make the entropy of the policy as close as possible to the target entropy Minimize the loss function with respect to the coefficient α of the constrained entropy:

[0105]

[0106] Among them, s t ~R represents state s t is sampled from the experience replay pool R; a t ~π(s t ) indicates action a t is from policy π, in state s t Downsampled.

[0107] The process of training the model in step 55 first requires initializing the critic network with random network parameters ω1, ω2 and θ respectively. and Actor Network π θ , copy the same parameters Initialize the target network separately and Then, the experience replay pool size, OU noise parameter, sampling batch size N, and discount factor γ are initialized respectively; finally, the battery power of the wireless sensor represented by the sending node S is initialized respectively. Amount of data in cache Channel gain h0, collected energy E0, collected data D0.

[0108] The procedural steps for training the model in step 55 are as follows:

[0109] 1) for sequence i=1→Ido

[0110] 2) Initial state

[0111] 3) for time step t = 1 → T do

[0112] 4) t Input to the Actor strategy network and get action at (Data transmission power p t )

[0113] 5) Sending node S performs action a t , get reward r t (r t The environment state becomes s, which is the sum of data transmission amount r1, action energy over-limit penalty r2, action data over-limit penalty r3, and energy overflow penalty r4. t+1

[0114] 6) The obtained (s t ,a t ,r t ,s t+1 ) Deposit into experience replay pool R

[0115] 7) If the amount of data in the current experience replay pool is greater than N

[0116] 8) for the number of training rounds k = 1 → K do

[0117] 9) Sample N tuples {(s i ,a i ,r i ,s i+1 )} i=1,...,N

[0118] 10) For each tuple, use the target critic network calculate

[0119]

[0120] 11) Perform the following gradient updates on the parameters of both Critic networks: for j=1,2 minimum

[0121] Optimized loss function

[0122]

[0123] 12) Reparameterization, sampling action Then, the following loss function is used to update the current Actor network parameters to improve the wireless sensor strategy so that it can select a better transmission power:

[0124]

[0125] 13) Update the entropy coefficient α according to the following loss function to balance the agent's exploration and reward

[0126] Encourage improvement

[0127]

[0128] 14) Soft update the target critic network parameters so that the target network parameters slowly catch up with the critic network to ensure training stability:

[0129]

[0130] 15)end for

[0131] 16)end if

[0132] 17)end for

[0133] 18)end for

[0134] like Figure 3 、 4 As shown in Figure 2, in order to verify the effectiveness of the above model algorithm, a simulated scenario is established for simulation training. The first consideration in establishing the scenario is whether the data collected by the wireless sensor is limited or unlimited.

[0135] In the first scenario, the data is assumed to be infinite, that is, only the random arrival of energy is considered. This makes sense in scenarios where there is always data to be sent. At the same time, the offline problem in this scenario becomes a convex problem. Here, offline means that the arrival of energy and channel state information in the entire sequence are completely known at the beginning. The SCS solver is used to find the numerical optimal solution of the offline problem and compare it with the algorithm execution result to judge the performance and feasibility of the algorithm. In the process of judging the algorithm, 10 environmental seeds are first randomly generated and fixed. It should be noted that these ten environmental seeds do not participate in the training of the model. In the process of training the model, each iteration, which contains five episodes, compares the difference between the solution of the current wireless sensor strategy and the optimal solution in these 10 environmental seeds, and plots it into a curve according to the number of iterations. At the same time, the greedy strategy and the proximal policy optimization algorithm (PPO) are used as the baseline of the algorithm performance for reference, such as Figure 3 This scenario also represents another practical application scenario, namely, when the amount of data to be sent is infinite. Another purpose of simulating this scenario is to verify the feasibility and effectiveness of the algorithm. With limited data, the problem becomes non-convex, and since the scenarios are continuous, it is impossible to find an offline optimal solution. However, with infinite data, the problem becomes convex, and a numerically optimal solution can be found. Therefore, we first assume that the scenario data is infinite.

[0136] Then consider the second more realistic scenario, that is, the data-limited scenario, in which data and energy are randomly received. The greedy strategy and PPO algorithm are used as baselines to verify the effectiveness of the algorithm. The experimental results are as follows: Figure 4As shown in the figure, the SAC algorithm initialization method makes the initial parameters of the algorithm larger. Since the environment accepts illegal actions and cuts them, and what is shown here is the actual data throughput rather than the real reward of the agent, the effect is similar to the greedy strategy.

[0137] Comparing the curves in the two scenarios, the results show that the trained agent is able to achieve the goal of maximizing data throughput without knowing future information, and its performance in scenarios with unlimited and limited data volume is better than that of the greedy algorithm, and is closer to the optimal solution for wireless sensors for data transmission and energy allocation.

[0138] During the implementation process, the problem is formulated as a Markov decision process, and then solved using the flexible actor-critic algorithm (SAC) method. Through some reinforcement learning techniques and careful design of the reward function, a strategy model is obtained. The final strategy can learn and optimize the energy usage strategy by tracking the specific energy, data size, and channel state information received in each time slot when the future information of energy, data, and channel state is unknown, so as to achieve better results.

[0139] A wireless sensor data transmission power control system, comprising:

[0140] A data receiving node and a wireless sensor, wherein the wireless sensor includes a memory for storing data and a battery for storing electricity;

[0141] A collection module, used to collect the remaining energy and data volume in the wireless sensor;

[0142] A training module is used to obtain wireless sensor data and train models. The training module is used to establish an optimization problem. The goal of the optimization problem is to maximize the data transmission volume of wireless sensors within a set time while maintaining energy sustainability. A Markov decision process is established for the optimization problem, and then the flexible actor-critic algorithm is used to solve the process.

[0143] The embodiments of the present application can be provided as methods or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and translated scripting language JavaScript, etc.

[0144] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0145] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0147] The above description is merely a specific example of the present invention and does not constitute any limitation thereto. It is apparent to those skilled in the art, after understanding the content and principles of the present invention, that various modifications and alterations in form and detail may be made without departing from the principles and structure of the present invention. However, such modifications and alterations based on the concepts of the present invention remain within the scope of protection of the claims of the present invention.

Claims

1. A wireless sensor data transmission power control method, characterized in that: The steps include: Step 1: Setting up a communication connection between a wireless sensor and a data receiving node, wherein the wireless sensor includes a storage module for data storage and a battery for energy storage; Step 2: The wireless sensor collects data from the outside world and stores it in its own memory. At the same time, the wireless sensor also obtains energy from the outside world to charge the battery. Step 3: Calculate the amount of data and energy stored in the wireless sensor at set intervals; Step 4: Establish an optimization problem of data transmission and energy consumption. The goal is to maximize the data transmission volume of wireless sensors within a set time period under the premise of energy sustainability. Step 5: Formulate the optimization problem in step 4 as a Markov decision process, and then use the flexible actor-critic algorithm to solve the process to obtain the data transmission strategy of the wireless sensor at the current moment; Step 6: The wireless sensor executes the data transmission strategy and transmits data to the data receiving node; The process of expressing the optimization problem as a Markov decision in step 5 includes the following steps: Step 51: Based on the Markov decision process, a model is built for the energy collection and data transmission process of the wireless sensor. The model is expressed as MDP M = {S, A, P, R, γ}, where S represents all possible states s in the state space; A represents all possible actions a in the action space; P represents the state transition probability, which is the state s of the wireless sensor at time t. t Next, perform action a t After that, the state transition occurs and transfers to state s t+1 The probability of; R represents the set reward value; the parameter γ represents the discount factor, which is the importance of future rewards; Step 52: Building the Energy Harvesting Model Data Collection Model and data transmission model R i , and based on the energy harvesting model, data collection model and data transmission model, define the problem statement of wireless sensor data transmission strategy; The energy collection model represents the amount of electricity stored in the battery at the beginning of the current time slot i, the data collection model represents the amount of data stored by the wireless sensor at the beginning of the current time slot i, and the data transmission model represents the data transmission rate in time slot i; Step 53: Set the reward function of the reward value R; Step 54: Setting up a model algorithm based on the flexible actor-critic algorithm; Step 55: Building and training a model based on the model algorithm; Step 56: Run the model to obtain the data transmission strategy of the wireless sensor at the current moment, and control the output power of the wireless sensor's transmission data; The energy harvesting model in step 52 Using the harvest-store-use (HSU) architecture, it is expressed as formula (1): in is the energy harvesting model of the wireless sensor, which represents the amount of electricity at the beginning of the current i-th time slot; τp i-1 is the energy consumed in the previous time slot i-1; p i-1 is the wireless sensor transmission power of the previous time slot i-1; E i-1 is the energy collected in the battery in the previous time slot i-1; E max The upper limit of the battery capacity of the wireless sensor; The data collection model in step 52 The harvest-store-use (HSU) architecture is adopted, which is expressed as formula (2): in represents the amount of data stored in the wireless sensor storage module at the beginning of the current i-th time slot; τR i-1 The amount of data sent in the previous time slot i-1; R i is the transmission rate of the i-th time slot; D i-1 is the amount of data obtained in the i-1th time slot; D max The upper limit of the data cache capacity of the storage module; The data transmission model R in step 52 i , expressed as formula (3): Among them, R i represents the data transmission rate in time slot i; h i represents the channel gain; σ 2 represents the variance of the noise; The problem of wireless sensor data transmission strategy in step 52 can be expressed as formula (4): Among them, R i represents the data transmission rate in time slot i; h i represents the channel gain; σ 2 represents the variance of the noise; τ represents a time interval; N represents a natural number; In step 54, the model algorithm is set based on the flexible actor-critic algorithm, as shown in formula (9): Among them, the content expressed by formula (9) is that the mean value of the constraint entropy is greater than Under the premise of maximizing expected return; r(s t ,a t ) means in state s t Next take action a t Instant rewards; represents the mean of the cumulative returns obtained under strategy π; Medium (s t ,a t )~ρ π Represents a state-action pair (s t , a t ) is taken from the state-action distribution ρ π , which means that we need to calculate the state-action distribution ρ caused by the policy π π The mean of the negative log-probability of the action, where represents the expected value; π t (a t |s t ) indicates that in the strategy π t Next, in state s t When selecting action a t probability; is the set target value.

2. A wireless sensor data transmission power control method according to claim 1, characterized in that: The reward function of the reward value R in step 53 includes the single-step data transmission amount r1, the action energy excess penalty r2, the action data excess penalty r3 and the energy overflow penalty r4; wherein the single-step data transmission amount r1 represents the amount of data transmitted by the wireless sensor in a single-step action; the action energy excess penalty r2 is related to the distance between the action energy and the boundary value of the set energy range, and the action energy represents the energy consumed by the wireless sensor to transmit a set amount of data during the action; the action data excess penalty r3 is related to the amount of data transmitted by the wireless sensor during the action and the amount of data remaining in the cache; the energy overflow penalty r4 is related to the energy overflowed by the wireless sensor and the amount of untransmitted data stored by the wireless sensor when the energy overflows.

3. The wireless sensor data transmission power control method according to claim 1, characterized in that: Formula (9) is simplified by the Lagrange multiplier method, and the loss function of the coefficient α is obtained as follows: Among them, s t ~R represents state s t is sampled from the experience replay pool R; a t ~π(s t ) indicates action a t is from policy π, in state s t Downsampled.

Citation Information

Patent Citations

  • Information age minimization sensor power distribution method and system

    CN117858188A