Energy Transfer and Data Acquisition Method Based on Multi-UAV-Assisted Wireless Sensor Networks
By decomposing the average information age problem into a sub-problem and adopting a multi-agent deep reinforcement learning method, the drone trajectory and scheduling are optimized, the time-domain conflicts between sensor data transmission and energy collection are solved, efficient data acquisition and energy transmission are achieved, communication and computing costs are reduced, and data freshness is ensured.
Patent Information
- Application Number
- CN202311077526.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-25
AI Technical Summary
The data transmission and energy collection requirements of sensors in wireless sensing networks conflict in the time domain, and centralized decision-making leads to high communication overhead and computing costs, making it difficult to ensure data freshness in a non-stationary environment.
The problem of minimizing the average information age is broken down into two sub-problems, and the multi-agent deep reinforcement learning method is adopted. Through the distributed multi-drone cooperative online scheduling algorithm, the associated scheduling variables and drone trajectory are optimized, the Markov decision-making process is designed, and the hybrid network is used to promote drone cooperation.
It realizes efficient data acquisition and energy transmission in multi-UAV-assisted wireless sensing networks, coordinates drone and sensor scheduling, ensures data freshness, and reduces communication and computing costs.
Smart Images

Figure CN117119490B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for energy transmission and data collection of unmanned aerial vehicles (UAVs) and sensors in a multi-UAV-assisted wireless sensor network in the academic field, and particularly to an energy transmission and data collection method based on a multi-UAV-assisted wireless sensor network. Background Art
[0002] Wireless energy transmission technology and UAV communication technology have become two major mainstream technologies for balancing the resources and demands of wireless sensor networks. Wireless energy transmission allows sensors to collect radio frequency energy signals transmitted from access points to achieve sustainable and controllable energy supply for them. Due to their mobility and agility, UAVs can effectively alleviate the problem of limited coverage of traditional ground access points. The wireless energy-driven multi-UAV-assisted wireless sensor platform combines these two technologies, which can not only achieve efficient sensor data collection but also alleviate the problem of limited sensor energy. For the wireless energy-driven multi-UAV-assisted wireless sensor platform, there are two key problems to be solved. One key problem is that due to the half-duplex hardware limitation of sensors and the time constraints of UAV movement, charging, and data collection, conflicts occur between the data transmission and energy collection requirements of sensors in the time domain, as well as complex UAV trajectory design and time allocation strategies. Therefore, it is necessary to carefully study the problem of reasonable time, space, and energy resource allocation for sensor scheduling, energy collection, and UAV trajectories. Another key problem is that many related studies make centralized decisions based on global state information, which will bring a large amount of communication overhead and computational cost. Therefore, it is necessary to further study distributed UAV scheduling optimization algorithms that can ensure data freshness and adapt to non-stationary environments. Summary of the Invention
[0003] The main object of the present invention is to propose an energy transmission and data collection method based on a multi-UAV-assisted wireless sensor network in view of some deficiencies in existing research.
[0004] The technical solution adopted by the present invention is: an energy transmission and data collection method based on a multi-UAV-assisted wireless sensor network, including the following steps:
[0005] Step 1): Construct a system model, determine the time allocation, energy consumption, and age-of-information models, and construct a problem of minimizing the average age of information;
[0006] Step 2): Decompose the problem of minimizing the average age of information in Step 1 into two sub-problems. Sub-problem 1 is to optimize the association scheduling variables and UAV trajectories, and sub-problem 2 is to optimize the time allocation variables and update the scheduling variables under the UAV association range;
[0007] Step 3): Use the multi-agent deep reinforcement learning theory to establish Markov models for the two sub-problems in Step 2;
[0008] Step 4): Use the multi-agent deep reinforcement learning method based on value function decomposition to decouple the two related decentralized partially observable Markov decision processes in Step 3).
[0009] The beneficial effects of the present invention are as follows:
[0010] The present invention constructs a distributed dynamic scheduling framework for realizing efficient data collection and energy transmission in a multi-UAV assisted wireless sensor network. To ensure data freshness, the present invention first decomposes the problem of minimizing the average age of information into two sub-problems, namely optimizing the association scheduling variables and UAV trajectory design, and optimizing the time allocation and updating scheduling variables under the UAV association range. In addition, the present invention models Sub-problem 1 and Sub-problem 2 as two coupled multi-agent stochastic cooperative games and formulates them as corresponding decentralized partially observable Markov decision processes. To obtain the equilibrium of the above two games, the present invention designs a distributed multi-UAV cooperative online scheduling algorithm based on multi-agent deep reinforcement learning, sets two actor networks for each UAV to learn the strategies of Sub-problem 1 and Sub-problem 2, and uses a hybrid network to fit the relationship between the local action value and the global action value to promote cooperation among UAVs. The experimental results show that the present invention can effectively coordinate the scheduling of UAVs and sensors and its efficiency in terms of the age of information. The present invention provides a new method for distributed energy transmission and data collection applied to a multi-UAV assisted wireless sensor network. Description of the Drawings
[0011] Figure 1 It is a system model diagram of a multi-UAV assisted wireless sensor network;
[0012] Figure 2 It is a time allocation model diagram;
[0013] Figure 3 It is a schematic diagram of algorithm training based on multi-agent deep reinforcement learning;
[0014] Figure 4 and Figure 5 They are the changes of the average estimated energy consumption and the average age of information with time during the training of the GEMINI algorithm designed by the present invention respectively. Except for certain fluctuations in the training curves in the first few rounds, both the average estimated energy consumption and the average age of information converge in less than 100 rounds;
[0015] Figure 6 and Figure 7The performance of the GEMINI algorithm designed for the present invention and the other five baseline algorithms in terms of average estimated energy consumption and average age of information at different transmission times. At a transmission time of 0.2 seconds, the effects of all algorithms are the best because this setting can achieve a balance between the number of sensors that meet the update conditions and the number of updatable sensors for all drones.
[0016] Figure 8 and Figure 9 The performance of the GEMINI algorithm designed for the present invention and the other five baseline algorithms in terms of average estimated energy consumption and average age of information at different numbers of sensors. The GEMINI algorithm designed for the present invention can also maintain the best performance when the number of sensors is large.
[0017] Figure 10 and Figure 11 The performance of the GEMINI algorithm designed for the present invention and the other five baseline algorithms in terms of average estimated energy consumption and average age of information at different numbers of drones. As the number of drones increases, the average estimated energy consumption and average age of information of all algorithms decrease because a sufficient number of drones can provide full coverage and ensure good channel conditions with candidate sensors. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0019] An energy transmission and data collection method based on a multi-UAV-assisted wireless sensor network, the steps are as follows:
[0020] Step 1: Build a system model and determine the time allocation, energy consumption, and age of information models.
[0021] As Figure 1 shown, the present invention constructs a system model, which includes U UAVs and S ground sensors, respectively represented as and To achieve long-term real-time collection of sensor data, in addition to collecting sensor data, the UAVs also need to provide charging services to the sensors. For the convenience of analysis, we divide the time into a series of time slots with a length of T, and assume that the on-board energy of the UAVs can support a total of N time slots for charging and data collection work, that is The position of sensor s is fixed at q s =(x s , y s ), where x s and y s respectively represent the abscissa and ordinate of sensor s. The coordinates of UAV u at time slot t are Among them, and respectively represent the horizontal and vertical coordinates and flight altitude of the unmanned aerial vehicle u. The distance between the sensor s and the unmanned aerial vehicle u is denoted as The channel gain between the sensor s and the unmanned aerial vehicle u is:
[0022]
[0023] Among them, ξ0 is the channel power gain at a reference distance of 1 meter. At time slot t, the data upload rate between the sensor s and the unmanned aerial vehicle u is:
[0024]
[0025] Among them, the variable is the data transmission power of the sensor s at time slot t, and σ 2 are the communication bandwidth and noise power respectively. In addition, the sensor has a maximum transmission power p max . The present invention assumes that the unmanned aerial vehicle u has a coverage area centered on its ground projection, with a radius of θ * is the elevation angle with the best constant value. In addition, the present invention uses the association scheduling variable to represent the association situation between the unmanned aerial vehicle u and the sensor s at time slot t. The association scheduling variable indicates that the sensor is within the coverage area of the unmanned aerial vehicle, that is, Among them is the projection distance between the sensor and the unmanned aerial vehicle. Therefore, the sensor s and the unmanned aerial vehicle u are associated during the time period t. On the contrary, when the association scheduling variable is right, there is
[0026] To satisfy the half-duplex constraint of the sensor and the unmanned aerial vehicle, the present invention designs a time allocation model for the unmanned aerial vehicle u as shown in Figure 2 . Specifically, in time slot t with a time interval length of T seconds, the unmanned aerial vehicle u spends and parts of the time for movement, charging, and sensor data collection respectively, where and respectively represent the time ratios spent by the unmanned aerial vehicle u on movement, charging, and sensor data collection during time slot t. During the charging stage, the unmanned aerial vehicle u transmits radio frequency signals to charge the associated sensors within its coverage area. In addition, the sensors adopt a time division multiple access mode during the data collection stage, and each sensor is allocated the same time φ for data transmission, where Therefore, only A sensor can transmit data, where is a rounding operation. In addition, in time slot t, the UAV u needs to decide which sensors within its coverage area can be used to transmit data. The present invention utilizes the transmission scheduling variable to represent the transmission scheduling vector of the UAV u. The transmission scheduling variable indicates that the sensor s transmits data to the UAV u in time slot t, while indicates that the sensor s does not transmit data to the UAV u.
[0027] The remaining battery energy of the sensor s at the end of time slot t is denoted as and the maximum battery capacity is Emax. When the sensor s is under the coverage of the UAV u, i.e., and the energy that the sensor s can collect from the UAV u is where and are the energy conversion efficiency of the sensor and the energy transmission power of the UAV, respectively. In addition, in time slot t, the data transmission energy consumption of the sensor s is Therefore, the remaining energy of the sensor s in time slot t is In addition, the present invention assumes that all UAVs have an initial on-board battery power and the energy consumption of the UAV mainly includes flight energy consumption, charging energy consumption, and hovering energy consumption. The UAV starts to fly at a speed from the position in the previous time slot to the position at a uniform speed in the current time slot. During this movement stage, its flight energy consumption is expressed as:
[0028]
[0029] where, is the propulsion power consumption of the UAV, i.e.:
[0030]
[0031] where, P0 = δρs0AΩ 3 R 3 / 8 and represent the blade profile power and the derived power when the UAV hovers, respectively. v0 is the average induced velocity of the rotor when hovering, d0 and ρ are the fuselage drag radio and air density, respectively, s0 and A are the rotor stiffness and rotor disk area, respectively. The symbols Ω, W, and k represent the blade angular velocity, the UAV weight, and the incremental correction factor for the induced power, respectively; δ and R are the airfoil drag coefficient and rotor radius, respectively.
[0032] In addition, the hovering and charging energy consumptions of the UAV u are expressed as and Among them, is the hovering power of the UAV, represents the transmission power of the UAV. In time slot t, the total energy consumed by UAV u is expressed as:
[0033]
[0034] This invention assumes that the arrival of sensor data follows a Bernoulli process, that is, when the sensor does not have the latest data packet in time slot t or transmitted the data packet to the UAV in time slot t - 1, it will randomly generate a new data packet size In addition, we define the generation time of the sensor's latest data packet as Then, in time slot t, the age of information of the sensor is the difference between the generation time of its latest data packet and the current time slot, that is:
[0035]
[0036] Specifically, the age of information of the sensor increases linearly with the increase of time slots until the sensor transmits the data to the UAV. In time slot t, with the transmission scheduling variable the age of information of the sensor in time slot t + 1 drops to 1. On the contrary, using the transmission scheduling variable the age of information of the sensor in time slot t + 1 increases to Therefore, the of the sensor in time slot t + 1 is expressed as:
[0037]
[0038] The optimization objective of this invention is to minimize the average age of information, and the problem description is as follows:
[0039]
[0040] s.t.
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048] Constraint 1 requires that the sensors associated with the UAV u must be within its coverage area. Constraint 2 ensures that the transmission scheduling variable is only valid when associated with the scheduling variable . Constraint 3 states that the total duration of UAV movement, charging, and sensor data collection cannot be greater than the time slot length T. Constraint 4 ensures that the energy consumed by the sensor at time slot t must be less than the energy collected at time slot t plus the remaining energy at time slot t - 1. Constraint 5 ensures that the initial on-board energy of the UAV u can support its operation between N time slots. Constraints 6 and 7 are the ranges of the time allocation variable transmission scheduling variable and the associated scheduling variable respectively.
[0049] Step 2: Analyze and transform the average age of information minimization problem in Step 1).
[0050] Since the constraints on the total duration of UAV movement, charging, and sensor data collection pose significant challenges to UAV scheduling, the UAV cannot spend too much time on movement, otherwise there will not be enough time for sensor charging and data collection. At the same time, the charging time allocation and transmission scheduling of sensors under the UAV association range should be further optimized, while ensuring the rationality of the associated scheduling and UAV trajectory design. Therefore, considering the complexity of the average age of information minimization problem in Step 1), the present invention tends to decompose it into two sub-problems to obtain an approximate solution. Sub-problem 1 is to optimize the associated scheduling variable and UAV position, and Sub-problem 2 is to optimize the time allocation variable and update the scheduling variable under the UAV association range.
[0051] The optimization objective of Sub-problem 1 is expressed as minimizing the average estimated energy consumption of the system, that is:
[0052]
[0053] s.t.
[0054]
[0055]
[0056]
[0057] where represents the estimated energy consumption of the UAV. reflects the data arrival of sensor s at time slot t, characterizes whether the sensor s associated with the UAV u meets the data transmission condition.
[0058] In addition, Sub - problem 2 is to optimize the time - allocation variables and update the scheduling variables to minimize the average age of information, with the solution of Sub - problem 1 (associating scheduling variables and UAV positions) as the input, i.e.:
[0059]
[0060] s.t.
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067] Step 3: Utilize the multi - agent deep reinforcement learning theory to establish Markov models for the two sub - problems in Step 2).
[0068] Considering that Sub - problems 1 and 2 are slot - related and need to be solved sequentially, the present invention converts Sub - problems 1 and 2 into two related decentralized partially observable Markov decision processes, denoted as tuples and γ represents the discount factor.
[0069] Among them, the meanings of the elements of the Markov model of Sub - problem 1 are as follows:
[0070] State: represents the state set of the Markov model of Sub - problem 1, where s 1,t represents the state of Sub - problem 1 at time slot t, and represent the UAV coordinates and time - allocation variables at time slot t - 1, respectively. and represent the remaining energies of UAV u and sensor s at time slot t - 1, respectively.
[0071] Observation: The joint observation of Sub - problem 1 is The individual partial observation of UAV u is represents the independent observation of UAV u based on Sub - problem 1 at time slot t.
[0072] Action: The joint action of Sub - problem 1 is where is the independent action of the UAV u, indicating the action of the UAV u at time slot t, represented as the coordinates of the UAV u at time slot t.
[0073] Transition function: is the state transition function from state s 1,t to the next state s 1,t+1 .
[0074] Reward: s1×a1→r 1 indicating the immediate reward obtained by taking the joint action a1. To ensure the overall optimization of sub - problem 1, we define the total immediate reward of all UAVs at time slot t as:
[0075]
[0076] where the first two terms of r 1,t come from the optimization objective of sub - problem 1. In addition, the variables and are the penalties for not satisfying constraint conditions 1 and 3 of sub - problem 1 respectively.
[0077] In addition, the meanings of the elements of the Markov model of sub - problem 2 are as follows:
[0078] State: defines the state set of the Markov model of sub - problem 2, where s 1,t represents the state of sub - problem 1 at time slot t, represents the remaining battery power of the UAV u at time slot t - 1. and represent the remaining battery power and the age of information of the sensor in time slot t - 1 respectively. and represent the coordinates and the associated scheduling variable of the UAV u in time slot t respectively.
[0079] Observation: The joint observation of sub - problem 2 is The partial observation of the UAV u is representing the independent observation of the UAV u based on sub - problem 1 at time slot t.
[0080] Action: The joint action of sub - problem 2 is where represents the action of the UAV u, indicating the action of the UAV u at time slot t. The variables and represent the time - allocation variable and the sensor - transmission scheduling of the UAV u in time slot t respectively.
[0081] Transition function: is a state transition function from state s 2,t to the next state s 2,t+1 .
[0082] Reward: s 2 ×a 2 →r 2 represents the immediate reward obtained by taking the joint action a 2 . The total immediate reward of the UAV in sub - problem 2 at time slot t is:
[0083]
[0084] where the variables and are the penalties for violating constraints 2, 3, and 4 of sub - problem 2 respectively.
[0085] It should be noted that due to the high - dimensional search space problem caused by the associated scheduling variables when modeled as discrete actions, it cannot be solved by the deep reinforcement learning method. Therefore, the present invention studies an association scheduling matching algorithm between sensors and UAVs, and the relevant pseudo - code is given in Table 1.
[0086]
[0087]
[0088] Step 4: Design a multi - agent deep reinforcement learning method based on value - function decomposition to decouple the two related decentralized partially observable Markov decision processes in step 3).
[0089] As Figure 3 shown, the present invention designs two actor networks for each UAV agent to achieve efficient learning of the strategies (π i , i ∈ {1, 2}) of sub - problems 1 and 2, namely actor network 1 and actor network 2. In addition, a critic network is designed for each UAV agent to simultaneously estimate the local action value functions of actor network 1 and actor network 2, and a hybrid network is used to fit the local action value functions of all UAV agents to generate a global action value function. All agents share a global action value which can be decomposed into:
[0090]
[0091] where τ i represents the historical trajectories of sub - problems 1 and 2. The variables and represent the network parameters of the critic network and the local action value function of agent u respectively. g ψIt is a hybrid network, which is a non - linear monotonic function represented by a neural network with parameters ψ. Represents the local action - value function of agent u. And Represent the historical trajectories and policies of agent u for sub - problems 1 and 2 respectively.
[0092] In addition, to evaluate the policies (π i , i ∈ {1, 2}) of sub - problems 1 and 2, the critic network and the hybrid network are trained simultaneously according to the following loss:
[0093]
[0094] Where D is the experience buffer, ζ1 and ζ2 are the influence - adjustment parameters of actor network 1 and actor network 2 on the critic network respectively. The variables and ψ - are the parameters of the target critic network and the target hybrid network respectively. The variable θ i - is the parameter of the target actor network for sub - problems 1 and 2. The variables s i′ and τ i′ are the observed state and the updated action - observation history after taking action a i in state s i respectively.
[0095] The present invention sets that all agents share two policy networks with parameters θ 1 and θ 2 to learn the policies (π i , i ∈ {0, 1}), and performs centralized updates on the joint action space. When evaluating , the actions of all agents are sampled from the current policies of all agents to calculate the policy gradient. Therefore, the centralized policy gradient can be estimated as:
[0096]
[0097] The present invention implements the GEMINI algorithm proposed by the present invention in a way of centralized training and distributed execution. In the centralized training stage, agents share information and utilize the available additional global agent information. However, during the execution process, each agent can only access its historical action - observation records.
[0098] The centralized training process of the GEMINI algorithm is shown in Table 2.
[0099]
[0100]
[0101] The test process of the designed GEMINI algorithm is shown in Table 3.
[0102]
[0103] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated therefrom do not exceed the spirit covered by the specification and the drawings, they should still fall within the protection scope of the present invention.
Claims
1. An energy transfer and data acquisition method based on a multi-UAV-assisted wireless sensor network, characterized in that Including the following steps: Step 1): Construct a system model, determine the time allocation, energy consumption, and age-of-information model, and construct an average age-of-information minimization problem, which is: Among them, and represent the sets of sensors, UAVs, and time slots respectively, and N is the number of time slots; represents the association between UAV u and sensor s in time slot t, represents the position of the current time slot, represents the transmission scheduling vector of UAV u, represents the time ratio that UAV u spends on charging data collection within time slot t, S represents the number of ground sensors, represents the age of information of the sensor, θ * is the elevation angle with the best constant value, is the projection distance between the sensor and the UAV, represents the flight altitude of UAV u, T represents the length of the time interval, and φ represents the time allocated to each sensor, represents the energy collected by sensor s from UAV u, represents the data transmission energy consumption of sensor s in time slot t, represents the remaining battery energy of sensor s at the end of time slot t - 1, represents the initial onboard battery power of the UAV, represents the total energy consumed by UAV u; Step 2): Decompose the average age-of-information minimization problem in Step 1) into two sub-problems. Sub-problem 1 is to optimize the association scheduling variables and the UAV trajectory, and Sub-problem 2 is to optimize the time allocation variables and the update scheduling variables under the UAV association range; Sub-problem 2 minimizes the average age of information by optimizing the time allocation variables and the update scheduling variables with the solution of Sub-problem 1 as the input. The said Sub-problem 1 is: Among them, represents the estimated energy consumption of the UAV, reflects the data arrival situation of sensor s at time slot t, characterizes whether the sensor s associated with the UAV u meets the data transmission condition; The said Sub-problem 2 is: Step 3): Utilize the multi-agent deep reinforcement learning theory to establish a Markov model for the two sub-problems in Step 2). Step 4): Use the multi-agent deep reinforcement learning method based on value function decomposition to decouple the two related partially observable Markov decision processes in Step 3). Two actor networks are designed for each UAV agent to achieve efficient learning of the sub-problem 1 and 2 policies, namely Actor Network 1 and Actor Network 2. Each UAV agent has a critic network to simultaneously estimate the local action value functions of Actor Network 1 and Actor Network 2, and a hybrid network is used to fit the local action value functions of all UAV agents to generate the global action value function. All agents share a global action value It can be decomposed into: where τ i represents the historical trajectories of sub - problems 1 and 2, the variables and represent the network parameters of the local action - value functions of the critic network and the agent u respectively, g ψ is a hybrid network, which is a non - linear monotonic function represented by a neural network with parameters ψ, represents the local action - value function of the agent u, and represent the historical trajectories and policies of the agent u for sub - problems 1 and 2 respectively; s i represents the state; Use a hybrid network to fit the relationship between the local action value and the global action value, promote cooperation among UAVs, and achieve effective coordination of the scheduling of UAVs and sensors.
2. The energy transfer and data acquisition method based on a multi-UAV assisted wireless sensor network according to claim 1, wherein: The time allocation is expressed as the time spent by the drone u and parts of the time are respectively used for movement, charging, and sensor data collection, where and respectively represent the time ratios spent by the drone u on movement, charging, and sensor data collection within the time slot t. During the charging phase, the drone u transmits radio frequency signals to charge the associated sensors within its coverage area; during the sensor data collection phase, a time division multiple access mode is adopted, and each sensor is allocated the same time φ for data transmission; The total energy consumed by UAV u is expressed as: Indicates flight energy consumption, Indicates the propulsion power consumption of the UAV, Indicates the hovering energy consumption of the UAV u, Indicates the charging energy consumption of the UAV u, which are respectively expressed as and Indicates the hovering power of the UAV, Indicates the transmission power of the UAV; The information age model is as follows: represents the generation time of the latest data packet of the sensor, and t represents the current time slot.
3. The energy transfer and data acquisition method based on a multi-UAV-assisted wireless sensor network according to claim 1, wherein: The elements of the Markov model of Sub-problem 1 are: Status: Denote the state set of the Markov model for sub - problem 1, where s 1,t Denote the state of sub - problem 1 at time slot t, and Represent the UAV coordinates and time allocation variables at time slot t - 1 respectively, and Represent the remaining energy of UAV u and sensor s at time slot t - 1 respectively; Observation: The joint observation of sub-problem 1 is The individual partial observation of drone u is yes Denote the independent observation of drone u based on sub-problem 1 at time slot t; Action: The combined action of sub-problem 1 is where is the independent action of the drone u, denotes the action of the drone u at time slot t, is represented as the coordinates of the drone u at time slot t; Transition function: is the state transition function from state s 1,t to the next state s 1,t+1 ; Reward: s 1 × a 1 → r 1 Indicates taking the joint action a 1 The immediate reward obtained.
4. The energy transfer and data acquisition method based on a multi-UAV-assisted wireless sensor network according to claim 3, wherein: The overall immediate reward of the UAV in Sub-problem 1 at time slot t is: Among them, the variables and are the penalties for not satisfying the constraint conditions 1 and 3 of sub-problem 1, respectively.
5. The energy transfer and data acquisition method based on a multi-UAV assisted wireless sensor network according to claim 1, characterized in that: The elements of the Markov model of Sub-problem 2 are: Status: Represents the state set of the Markov model for sub - problem 2, where s 1,t Represents the state of sub - problem 1 at time slot t, Represents the remaining battery power of drone u at time slot t - 1, and Represent the remaining battery power and the age of information of the sensor in time slot t - 1 respectively, and Represent the coordinates of drone u and the associated scheduling variable in time slot t respectively; Observation: The joint observation of sub-problem 2 is The partial observation of drone u is Denote the independent observation of drone u based on sub-problem 1 at time slot t; Action: The joint action of sub-problem 2 is where represents the action of UAV u, represents the action of UAV u at time slot t, and the variable and represent the time allocation variable and sensor transmission schedule of UAV u in time slot t, respectively; Transfer function: is a state transition function from state s 2,t to the next state s 2,t+1 ; Reward: s 2 × a 2 → r 2 Indicates taking the joint action a 2 The immediate reward obtained.
6. The energy transmission and data acquisition method based on a multi-UAV assisted wireless sensor network according to claim 5, wherein: The total immediate reward of the UAV in Sub-problem 2 at time slot t is: Among them, the variables and are the penalties for violating the constraints 2, 3, and 4 of sub-problem 2 respectively.
Citation Information
Patent Citations
Track and resource allocation method based on unmanned aerial vehicle assisted wireless energy supply network
CN116567546A
Multi-unmanned aerial vehicle data acquisition position optimization method based on multi-agent reinforcement learning
CN116627162A