A Trajectory and Resource Allocation Method Based on UAV-Assisted Wireless Energy Transfer Network
By decomposing the problems of drone trajectory control and equipment association allocation, the UAV trajectory is optimized using multi-agent deep reinforcement learning and Lagrangian punishment reward function, solving the problems of uneven access and security of equipment in the drone-assisted wireless energy supply network, and achieving the maximum system energy efficiency.
Patent Information
- Application Number
- CN202310750511.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-25
AI Technical Summary
The prior art has failed to effectively solve the problem of uneven access to user equipment due to differences in equipment geographical distribution in drone-assisted wireless energy supply networks, as well as the problem of maximizing system energy efficiency while ensuring drone safety, especially how to improve system throughput, task completion time and energy efficiency under limited resources.
The multi-objective optimization problem is constructed, and it is decomposed into the optimization problem of drone trajectory control and the optimization problem of drone-related allocation with intelligent sensing devices. The constrained multi-agent deep reinforcement learning and the drone-assisted intelligent sensing device association allocation method is used to optimize the UAV trajectory control strategy by constructing the Lagrangian punishment reward function and attention mechanism.
Under limited resources, the trade-off between drone flight safety and energy consumption is achieved, the equipment service quality and energy consumption constraints are ensured, the system's energy efficiency and resource utilization are improved, and the drone trajectory control strategy is optimized.
Smart Images

Figure CN116567546B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of wireless power supply communication transmission, and specifically to a trajectory and resource allocation method based on an unmanned aerial vehicle (UAV)-assisted edge wireless power supply network. Background Art
[0002] The statements in this part only provide background technical information related to the present disclosure, and these statements may constitute the prior art. In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art.
[0003] In recent years, the sixth-generation mobile communication technology network is expected to enhance the availability of network data, mobile data rate, and provide ubiquitous connections for a large number of Internet of Things (IoT) terminal devices. In addition, in order to meet the energy supply needs for data transmission of intelligent sensing devices, the emergence of energy harvesting technology has brought a new mode to solve this problem. To mitigate the performance impact brought by energy harvesting technology, we use wireless power transfer technology to provide a low-maintenance-cost and highly flexible energy supply method for devices in the IoT network. For the platform of the multi-UAV-assisted wireless power supply network, we use UAVs as wireless power transfer platforms to provide energy transfer services for intelligent sensing devices and assist intelligent sensing devices in communication, and formulate an energy consumption problem according to the constraints of UAV-assisted intelligent sensing device energy transfer and effective data collection.
[0004] Regarding the problem of maximizing the energy efficiency of UAVs, the existing solution is to solve it by the number of bits of transmitted data and the total energy consumption of UAVs. Another example is the patent with the application number 202310173207.8 and the patent name "A Method for UAV Wireless Power Supply and Air Computing-Assisted IoT Data Collection", which alternately solves the constructed resource allocation optimization model P1, UAV trajectory optimization model P2, and the three-dimensional coordinate IoT transmission system of the UAV base station to obtain the flight trajectory of the UAV, wireless power supply time slot allocation, and sensor transmission power when the uplink data collection rate of the three-dimensional coordinate IoT transmission system is maximized.
[0005] However, the prior art including the above patent does not consider the problem that some user devices are rarely accessed or never sensed by UAVs due to the difference in device geographical distribution, and the flight safety factors during the process of multiple UAVs adjusting their moving positions. How to maximize the system energy efficiency while ensuring the safety of UAVs, and consider the mobility of UAVs and the quality of service constraints of intelligent sensing devices, that is, make a trade-off between energy consumption and safety, there is still no good solution at present. Especially under limited resources, how to maximize the performance of the system network, including system throughput, task completion time, and energy efficiency, is worthy of further exploration by researchers. Summary of the Invention
[0006] In view of the above problems, the object of the present invention is to solve a part of the problems in the prior art, or at least alleviate these problems.
[0007] A trajectory and resource allocation method based on an unmanned aerial vehicle (UAV)-assisted wireless power supply network, including an Internet of Things network with U UAVs and I intelligent sensing devices for wireless power supply, for UAV-assisted energy transmission and data collection; the method includes the following steps:
[0008] Construct a system model of a multi-UAV-assisted wireless power supply network, and determine a movement model, a power transmission model, and a data collection model;
[0009] According to the constructed movement model, power transmission model, and data collection model, propose and model a multi-objective optimization problem for maximizing the system energy efficiency:
[0010]
[0011] s.t. Constraint conditions
[0012]
[0013] C6: α i,u (t)l i,u (t)≤l z
[0014]
[0015] where, Ω i,u (t) is the expected size of the collected data, Λ i (t) is the service fairness of the UAV; e u (t) is the total energy consumption for completing UAV movement, energy transmission, and data collection, t is the time slot, T is the time period, is the movement position strategy of UAV u; χ u (t) is the flight action of UAV u, and represent the side length of the target area; h u (t) is the flight altitude of the UAV, and represent the minimum and maximum flight altitudes of the UAV; q u (t) is the position coordinate of the u-th UAV in the t-th time period; α i,u (t) is the user association decision between the u-th UAV and the i-th intelligent sensing device in the t-th time period, l i,u (t) is the three-dimensional distance between the UAV and the sensing device, l z is the sensing distance of the UAV; ω i,u(t) is the data rate of the uplink between the intelligent sensing device and the UAV, p n is the transmission power of the intelligent sensing device, is within the sensing range of the UAV the received power of the intelligent sensing device, is the time spent by the UAV to transmit energy to the intelligent sensing device within the coverage area; E u (0) is the initial battery capacity of the UAV; v max represents the maximum deadline of the task, represents the time spent by the UAV to move to adjust its own position and coverage area, is the time spent by the UAV to collect the data uploaded by the device; x u (t) is the x-axis position coordinate of the u-th UAV in the t-th time period; y u (t) is the y-axis position coordinate of the u-th UAV in the t-th time period; is the minimum safe flight distance of the u-th UAV; q u′ (t) is the position coordinate of the u'-th UAV in the t-th time period; The multi-objective optimization problem of maximizing the system energy efficiency of the proposed system is decomposed into two sub-problems. The first sub-problem is the UAV trajectory control optimization problem, and the second sub-problem is the association allocation optimization problem between the UAV and the intelligent sensing device;
[0016] For the first sub-problem, a constrained multi-agent deep reinforcement learning algorithm is used to solve the UAV trajectory control optimization problem;
[0017] For the second sub-problem, a UAV-assisted intelligent sensing device association allocation algorithm is used to solve the association allocation optimization problem between the UAV and the intelligent sensing device.
[0018] The process of decomposing the multi-objective optimization problem of maximizing the system energy efficiency of the proposed system into two sub-problems includes the following steps:
[0019] Allow all intelligent sensing devices within the sensing range of the UAV to upload data simultaneously;
[0020] Maximizing the system energy efficiency over the entire system time is transformed into maximizing the corresponding energy efficiency for each time slot;
[0021] When the movement position strategy of UAV u is given the upper limit of the energy transmission time of constraint C9 is transformed into:
[0022]
[0023] where, δ c is the sampling frequency of the sensor equipped on the UAV;
[0024] Build a model for the first sub - problem of the UAV trajectory control optimization problem
[0025]
[0026] s.t. Constraint conditions
[0027]
[0028] Build a model for the association allocation optimization problem of UAVs and intelligent sensing devices
[0029]
[0030] s.t. Constraint conditions
[0031]
[0032] α i,u (t)l i,u (t) ≤ l z
[0033]
[0034] The constrained multi - agent deep reinforcement learning algorithm is a constrained Markov decision process model, which combines the deep reinforcement learning algorithm and the Lagrangian primal pairing method. By combining the actor - critic network with a personalized attention mechanism to train the agents, and by adjusting the Lagrange multipliers to assign different attention weights to the optimization objective and penalty constraints, the optimal trajectory control strategy can be obtained.
[0035] Furthermore, the constrained Markov decision process model includes the following four key elements:
[0036] State: Set is the state space, including where the symbol S1 contains the position coordinates of the UAV The symbol S2 contains the relative position between the UAV and the intelligent sensing device, that is g i (t) is the position coordinates of the intelligent sensing device; the symbol S3 contains the relative position between the UAV and other UAVs, that is The symbol S4 contains the remaining battery energy of the UAV, that is The symbol S5 includes the remaining transmission data volume of the intelligent sensing device, that is ε i (t);
[0037] Action: Set is the action space, including Variable a u (t) represents the movement of the UAV, including the flight movement direction χ of the UAV u within the time period t u and the flight movement distance d u (t);
[0038] Reward: After the UAV takes the action a(t) in the state s(t), it obtains an immediate reward
[0039] Penalty: The penalty constraint is represented by j(t) for the collision of the UAV with other UAVs during the movement, that is, the penalty for not satisfying the constraint penalty.
[0040] Furthermore, by adjusting the Lagrange multiplier to allocate different attention weights to the optimization objective and penalty constraints, including constructing the Lagrange penalty reward function as follows:
[0041]
[0042] where is the Lagrange multiplier corresponding to the penalty constraint.
[0043] Furthermore, the steps to obtain the optimal trajectory control strategy are as follows:
[0044] Deploy a constrained deep reinforcement learning agent on each UAV, including multiple deep neural networks, namely the policy actor network, policy critic network, penalty actor network, penalty critic network and their corresponding target networks;
[0045] Train the agent; after training is completed, the agent uses the locally deployed policy network for decision-making;
[0046] Initialize the network parameters of the deep neural network, Lagrange multiplier and the experience replay buffer D; at the beginning of each epoch, all UAVs can serve user equipment from different location points;
[0047] In the time slot t, each UAV agent extracts an action from the policy given by the current policy actor network according to its observation and executes it, that is, determines the joint action set of the movement direction and distance of the UAV;
[0048] The environment updates the state according to the joint action set collected from the UAVs to obtain a new state; if the UAV flies out of the target area, its movement will be terminated; accordingly, the rewards of all UAVs are obtained and the corresponding penalties and construct the Lagrange penalty reward function;
[0049] Construct state transition and store it in the replay experience buffer D; when the buffer space reaches the upper limit, each UAV agent will sample a small batch of experiences from the experiences in the replay experience buffer.
[0050] Solve the UAV trajectory control optimization problem using the constrained multi-agent deep reinforcement learning algorithm for the first sub-problem, which also includes further optimization and further updates of its network parameters, including the following steps:
[0051] Update the network parameters of the policy critic and the penalty critic by minimizing the loss function;
[0052] Update the parameters of the policy actor network through policy gradient;
[0053] After completing the parameter updates of the policy critic network and the policy actor network on the fast time scale, the penalty actor network on the slower time scale is updated using the Lagrange multiplier.
[0054] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the trajectory and resource allocation method based on the UAV-assisted wireless energy supply network.
[0055] The present invention has the following beneficial effects:
[0056] 1. On the premise of limited resources, the present invention makes a trade-off between energy consumption and safety. Under the conditions of considering the flight safety and energy limitation of the UAV, it ensures the quality of service of the device, and while obtaining the optimal trajectory control strategy, it also meets the quality of service and energy consumption constraints of the intelligent sensing device;
[0057] 2. To formulate an optimization problem for maximizing the system energy efficiency while ensuring the safety of the UAV, the present invention proposes a multi-objective optimization problem for maximizing the system energy efficiency; however, through experiments and in-depth research, it is found that this problem is difficult to solve due to the deep coupling of the UAV variables α i,u (t) and q u (t). Therefore, this problem is transformed into a UAV trajectory control sub-problem and a UAV and intelligent sensing device association sub-problem, and the UAV trajectory control and the association allocation of the UAV and the intelligent sensing device are respectively completed by using the constrained multi-agent deep reinforcement learning and the UAV-assisted intelligent sensing device association allocation method, which improves the resource utilization efficiency while obtaining the optimal trajectory control strategy;
[0058] 3. The present invention constructs a Lagrangian penalty reward function, and by adjusting the Lagrange multipliers to assign different attention weights to the optimization objective and penalty constraints, it solves the problem of the uncertainty of the reward-related weight values in the linear combination of the original reward and penalty constraints used in traditional constrained deep reinforcement learning;
[0059] 4. For the problem of optimizing the UAV trajectory control, the present invention designs a constrained multi-agent deep reinforcement learning algorithm and combines it with the attention mechanism to obtain the optimal trajectory control strategy, and theoretically proves that the complexity of this method is polynomial time;
[0060] 5. For the problem of optimizing the association between UAVs and intelligent sensing devices, the present invention designs a UAV-assisted intelligent sensing device association allocation algorithm to determine the association of intelligent sensing devices within the sensing range of the UAV and meet the quality of service and energy consumption constraints of the intelligent sensing devices;
[0061] 6. Through further updating of the deep neural network, the critic network is responsible for evaluating the policy and penalty to guide the decision-making and action selection of the UAV; the actor network is responsible for improving the policy and penalty and generating the actions of the UAV, and optimizing the policy by maximizing the expected reward; updating the parameters of the penalty actor network by adjusting the Lagrange multiplier helps to improve the constraint satisfaction and overall performance of the policy; combined with network parameter updating to obtain the optimal trajectory control strategy of the UAV and ensure the safety of the UAV. Description of the Drawings
[0062] Figure 1 is the system model of the multi-UAV-assisted wireless power supply network, and multiple UAVs transmit energy to intelligent sensing devices within their coverage area and collect the data uploaded by the devices.
[0063] Figure 2 is the schematic diagram of time allocation, and each time slot t is divided into the time spent on UAV movement, UAV transmitting energy to intelligent sensing devices within its coverage area, and collecting the data uploaded by the devices.
[0064] Figure 3 is the schematic diagram of the training of the constrained multi-agent deep reinforcement learning algorithm.
[0065] Figure 4 is the comparison schematic diagram of the total training time of the MCDRL algorithm designed by the present invention and three other algorithms under different update cycles and the number of devices.
[0066] Figure 5 shows the performance of the MCDRL algorithm designed by the present invention and three other algorithms in terms of loss, total reward, and penalty under different training rounds.
[0067] Figure 6 The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is analyzed under different numbers of user devices.
[0068] Figure 7 The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is compared under different numbers of unmanned aerial vehicles (UAVs).
[0069] Figure 8 The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is demonstrated under different aperture angles. Detailed implementation manners
[0070] The following further describes the present invention in conjunction with the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention rather than limit it. Without departing from the technical idea of the present invention, various substitutions and changes made according to the common general knowledge and conventional means in the art shall all be included within the scope of the present invention.
[0071] The present invention considers a wireless power - supplied Internet of Things (IoT) network for UAV - assisted energy transfer and data collection. Specifically, the process is divided into three stages, namely, the UAV moves to adjust its position, then associates with intelligent sensing devices to provide energy within its sensing range, and finally collects data from these devices. The present invention designs a stable energy transfer and reliable data collection scheme for energy - constrained IoT networks with various intelligent sensing devices.
[0072] A trajectory and resource allocation method based on a UAV - assisted wireless power - supplied network, for a wireless power - supplied IoT network including U UAVs and I intelligent sensing devices, for UAV - assisted energy transfer and data collection; the method includes the following steps:
[0073] Construct a system model of a multi - UAV - assisted wireless power - supplied network, and determine the movement model, energy transfer model, and data collection model;
[0074] According to the constructed movement model, energy transfer model, and data collection model, propose and model a multi - objective optimization problem for maximizing the system energy efficiency:
[0075]
[0076] s.t. Constraint conditions
[0077]
[0078] C6: α i,u (t)l i,u (t)≤lz
[0079]
[0080] where, Ω i,u (t) is the expected size of the collected data, Λ i (t) is the service fairness of the UAV; e u (t) is the total energy consumption for UAV movement, energy transfer, and data collection. t is the time slot, and T is the time period. is the movement position strategy of UAV u; χ u (t) is the flight action of UAV u, and denote the side length of the target area; h u (t) is the flight altitude of the UAV, and denote the minimum and maximum flight altitudes of the UAV; q u (t) is the position coordinate of the u-th UAV in the t-th time period; α i,u (t) is the user association decision between the u-th UAV and the i-th intelligent sensing device in the t-th time period, l i,u (t) is the three-dimensional distance between the UAV and the sensing device, l z is the sensing distance of the UAV; ω i,u (t) is the data rate of the upload link between the intelligent sensing device and the UAV, p n is the transmission power of the intelligent sensing device, is the received power of the intelligent sensing device within the sensing range of the UAV, is the time spent by the UAV to transfer energy to the intelligent sensing device within the coverage area; E u (0) is the initial battery capacity of the UAV; v max denotes the maximum deadline of the task, denotes the time spent by the UAV to move and adjust its own position and coverage area, is the time spent by the UAV to collect the data uploaded by the device; x u (t) is the x-axis position coordinate of the u-th UAV in the t-th time period; y u (t) is the y-axis position coordinate of the u-th UAV in the t-th time period; is the minimum safe flight distance of the u-th UAV; q u′ (t) is the position coordinate of the u'-th UAV in the t-th time period;
[0081] The multi-objective optimization problem of maximizing the system energy efficiency is decomposed into two sub-problems. The first sub-problem is the UAV trajectory control optimization problem, and the second sub-problem is the association allocation optimization problem between the UAV and intelligent sensing devices;
[0082] For the first sub-problem, a constrained multi-agent deep reinforcement learning algorithm is used to solve the UAV trajectory control optimization problem;
[0083] For the second sub-problem, a UAV-assisted intelligent sensing device association allocation algorithm is used to solve the association allocation optimization problem between the UAV and intelligent sensing devices.
[0084] To achieve the effective utilization of system resources, the present invention constructs a multi-objective optimization model for maximizing the system energy efficiency under the conditions of considering the flight safety and energy limitation of the UAV and ensuring the quality of service of the devices. Constraints 1, 2, and 3 represent the constraint range of the UAV's moving position, which limits the three-dimensional position range of the UAV trajectory control. Each UAV can only move within the target area. and represent the side length of the target area. and represent the minimum and maximum flight altitudes of the UAV; Constraint 4 ensures the safe distance between UAVs during flight; Constraint 5 means that each intelligent sensing device can only be served by one UAV within a time period; Constraint 6 means that the intelligent sensing device can only be served within the sensing range of the UAV to meet the communication requirements; Constraint 7 ensures that the energy provided by the UAV to the intelligent sensing device can meet the energy consumption for uploading data; Constraint 8 means that the total energy consumption of the UAV should be less than the initial battery capacity E u (0); Constraint 9 limits the total delay time of the task not to exceed the maximum service deadline.
[0085] However, experiments found that the problem is difficult to solve. After in-depth research and repeated experiments by the applicant, it is mainly due to the deep coupling of the UAV variables α i,u (t) and q u (t). To decouple the decision variables and effectively solve the problem, the present invention relaxes Constraints 6 and 7, allowing all intelligent sensing devices within the sensing range of the UAV to upload data simultaneously, that is, deleting α from the variables constrained by i,u and the upload rate ω i,uIn addition, we transform the maximization of the system energy efficiency over the entire system time into the maximization of the corresponding energy efficiency for each time slot. Considering the case of the UAV trajectory, the problem can be transformed into the problem and Thus, the process of decomposing the multi-objective optimization problem of maximizing the proposed system energy efficiency into two sub-problems includes the following steps:
[0086] Allow all intelligent sensing devices within the sensing range of the UAV to upload data simultaneously;
[0087] Transform the maximization of the system energy efficiency over the entire system time into the maximization of the corresponding energy efficiency for each time slot;
[0088] When the movement position strategy of UAV u is given, the upper limit of the energy transmission time allocation of constraint C9 is transformed into:
[0089]
[0090] where δ c is the sampling frequency of the sensors equipped on the UAV;
[0091] Build a model for the first sub-problem of the UAV trajectory control optimization problem
[0092]
[0093] s.t. Constraints
[0094]
[0095] Build a model for the association allocation optimization problem between the UAV and intelligent sensing devices
[0096]
[0097] s.t. Constraints
[0098]
[0099] α i,u (t)l i,u (t)≤l z
[0100]
[0101] After transforming the multi-objective optimization problem of maximizing the system energy efficiency into a UAV trajectory control sub-problem and a sub-problem associated with intelligent sensing devices, the decision variables of UAV position adjustment and intelligent sensing device association are decoupled. Then, by using constrained multi-agent deep reinforcement learning and a UAV-assisted intelligent sensing device association allocation method, the UAV trajectory control and the association allocation between the UAV and intelligent sensing devices are completed respectively, obtaining the optimal trajectory control strategy and satisfying the quality of service and energy consumption constraints of the intelligent sensing devices.
[0102] The experimental results are as Figure 4 、 5 shown in 6, 7, and 8, which are the experimental results of the MCDRL algorithm (constrained multi-agent deep reinforcement learning) designed by the present invention compared with MADDPG (multi-agent deep deterministic policy gradient), MAAC (multi-agent attention actor-critic algorithm), and PPO (proximal policy optimization algorithm).
[0103] Figure 4 It is a comparison schematic diagram of the MCDRL algorithm designed by the present invention and the other three algorithms in terms of the total training time under different update cycles and the number of devices. It can be seen that since the UAV agent is updated at the beginning of each cycle to replace old data in a timely manner and avoid performance degradation, the change trend of the total training time tends to be stable within two update cycles. Secondly, the total training time increases with the increase in the number of devices because the UAV agent needs more time to explore the optimal trajectory control and data collection strategies.
[0104] Figure 5 It shows the performance of the MCDRL algorithm designed by the present invention and the other three algorithms in terms of loss, total reward, and penalty under different numbers of training rounds. It can be seen that MCDRL, MADDPG, MAAC, and PPO gradually converge as the training process extends, and MCDRL has a better convergence speed. Secondly, the total rewards of the MCDRL and MADDPG algorithms increase with the increase in the training process, then decrease and tend to be stable. Since the MAAC algorithm does not have a constraint penalty constraint and the UAV does not consider the collision constraint of flight safety, its total energy efficiency is the highest. While the MCDRL and MADDPG algorithms are subject to the penalty constraint of the safe flight distance, their performance trends decrease partially until the penalty constraint is limited within the agreed range. Due to the adjustment effect of the Lagrange multiplier in the constrained environment, the penalty is limited within the constraint range, and the penalty constraints of the MCDRL and MADDPG algorithms increase and then decrease. However, the MAAC and PPO algorithms have no constraint conditions, and their performance trends almost remain unchanged.
[0105] Figure 6The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is analyzed under different numbers of user devices. As the number of user devices increases, due to the differences in the data uploaded by user devices, the total amount of data increases, so the effective data collection amount of the UAV continuously decreases. As the number of user devices increases, the service fairness monotonically decreases. This is because when the number of UAVs remains unchanged while the amount of data of user devices increases, the service capacity of the UAV is limited and it cannot complete the effective data collection within the same time. In addition, the UAV needs to consume more energy to meet the increasing service demands of user devices. When the number of user devices increases, the energy utilization rate of the UAV monotonically increases. The UAV needs to continuously fly and move to sense the positions of user devices. Therefore, it needs to consume more transmission and data collection energy to adapt to the rich data of user devices.
[0106] Figure 7 The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is compared under different numbers of UAVs. It can be seen that the effective data collection rate and service fairness increase as the number of UAVs increases. This is because the sensing range and battery capacity resources of a single UAV are limited. Facing the same number of user devices, deploying more UAVs can improve the service capacity of the target area to meet the energy transmission and data collection tasks of user devices. However, as the number of UAVs increases, the energy utilization rate shows a downward trend because a sufficient number of UAVs can provide a wider sensing range, sufficient energy transmission resources, and effective data collection services.
[0107] Figure 8 The performance of the MCDRL algorithm designed in the present invention and three other algorithms in terms of data collection, service fairness, and energy utilization rate is shown under different aperture angles. Since the aperture angle affects the size of the sensing range of the UAV, it further affects the service performance of the UAV.
[0108] In summary, it can be seen that the present invention is more efficient than multiple existing algorithms in terms of data collection, service fairness, and energy utilization.
[0109] Construct a system model of a multi-UAV assisted wireless power supply network, and determine the movement model, power transmission model, and data collection model, including the following content:
[0110] The present invention constructs a system model, such as Figure 1As shown, it includes U drones and I intelligent sensing devices. Each intelligent sensing device can only be served by one drone within a time period, and the intelligent sensing device can only be served when it is within the sensing range of the drone. At the beginning, the drones obtain sufficient energy from the charging station and move within a fixed boundary target area. Within this area, the drones move to adjust their positions, associate with the intelligent sensing devices, and provide energy within their sensing range, enabling the intelligent sensing devices to have sufficient ability to upload data. Then, the drones collect the data uploaded by the devices.
[0111] The present invention divides the time range of the system into T time periods, that is As Figure 2 shown. At the beginning of each time period, the drones adjust their positions, and the corresponding positions are approximately considered unchanged during this period. At time slot t, the time spent by the drone to move to adjust its own position and coverage range is and the time spent to transmit energy to the intelligent sensing devices within the coverage range is while the time spent by the drone to collect the data uploaded by the devices is Assume that each intelligent sensing device has a delay-sensitive task, which can be expressed as: F i (t) = {f i (t), v max} where f i (t) represents the size of the original data uploaded by the device, and the variable v max represents the maximum deadline of the task. Then, the data uploaded by the device beyond this deadline is invalid. Since the device needs to consume energy to upload data, the energy obtained by the device from the drone is closely related to f i (t).
[0112] Within the t-th time period, each drone can move in a certain direction χ u (t), where the variables represent horizontal left, horizontal right, horizontal forward, horizontal backward, vertical upward, vertical downward, and stationary in turn. The change in the position of the drone between adjacent time slots can be expressed as d u (t) = ||q u (t + 1) - q u (t)||, where q u (t) is the position coordinate of the u-th drone within the t-th time period. Assume that the drone moves at a constant speed η, then the time cost of the drone's movement can be calculated as follows:
[0113]
[0114] The flight propulsion power of the UAV moving at η can be expressed as:
[0115]
[0116] Among them, p o and p s are the blade section power and the induced power respectively. The variable U p is the tip speed of the rotor, V h is the average induced speed of the rotor. The symbol d0 is the fuselage drag ratio, ρ a is the air density. The variable r s is the rotor solidity, J r is the rotor disk area. Therefore, the energy consumption of the UAV moving is
[0117] After the UAV adjusts its position, assuming that most of the transmission power is concentrated below the UAV, its horizontal coverage range depends on the azimuth angle and flight altitude of the antenna. Then the sensing distance of the UAV is expressed as:
[0118]
[0119] Thus, the sensing range of the UAV can be expressed as:
[0120]
[0121] Among them, τ max is the aperture angle, h u (t) is the flight altitude of the UAV. The three-dimensional distance between the UAV and the sensing device can be expressed as:
[0122] Each sensing device can choose whether to be served by the UAV, that is α i,u (t)=1 means that the u-th UAV provides service to the i-th intelligent sensing device within the t time period, while α i,u (t)=0 means that the u-th UAV does not provide service to the i-th intelligent sensing device within the t time period.
[0123] Considering the path loss channel model in free space, then the channel gain of the UAV is:
[0124]
[0125] Among them, β0 represents the channel gain when the reference distance is 1 meter. Based on the generalized radio frequency power transmission model, the received power of the intelligent sensing device within the sensing range of the UAV is:
[0126]
[0127] Among them, p t is the transmission power of the drone. The symbol μ = c / f c is the wavelength, and c and f c are the speed of light and the carrier frequency of the radio frequency wave transmitted by the drone respectively, and θ is the path loss exponent of the air-to-ground channel. Therefore, the energy obtained by the intelligent sensing device from the drone is The energy consumed by the drone for transmission is
[0128] After the drone provides energy to the intelligent sensing device, it starts to collect the uploaded data of the served device. Based on the orthogonal frequency division multiple access technology, the interference between devices can be ignored when the intelligent sensing device uploads data. The data rate of the upload link between the device and the drone is:
[0129]
[0130] Among them, B is the bandwidth of the wireless channel, and p n is the transmission power of the intelligent sensing device, and σ 2 is the noise power. Considering that the original data sizes of each user device are different, given the sampling frequency δ c of the sensor equipped on the drone, the expected size of the collected data can be expressed as: Among them, ε i (t) is the size of the remaining original data.
[0131] The energy consumption of the drone for data collection is The symbol p c is the data collection power. In addition, the total amount of data collected by the drone is equal to the amount of data effectively uploaded by the device, so the energy consumed by the device for uploading data is:
[0132]
[0133] Based on the above, the total energy consumption for the drone's energy transmission and data collection is:
[0134]
[0135] The effective data collection rate of the drone can be measured by the ratio of the amount of data already collected to the size of the original data, and can be expressed as:
[0136]
[0137] Among them, represents the total amount of data collected from device i when the drone position adjustment strategy is adopted. In addition, the service fairness of the drone can be expressed as:
[0138]
[0139] When Λ i (t) is close to 1, the UAV will collect data from all intelligent sensing devices in a more balanced manner to ensure the fairness of its services. The energy saving κ u (t) of the UAV is defined by the ratio of the energy consumption of the UAV to the initial energy consumption E u (0), that is:
[0140]
[0141] Thus, the optimization objective of the present invention is confirmed as maximizing the energy efficiency of the UAV (i.e., a multi-objective optimization problem of maximizing the system energy efficiency).
[0142] For the optimization problem of UAV trajectory control, the present invention designs a constrained multi-agent deep reinforcement learning algorithm and combines it with an attention mechanism to obtain the optimal trajectory control strategy, and theoretically proves that the complexity of this method is polynomial time. This part first considers the optimal trajectory control and associated allocation strategy that satisfy the flight safety of the UAV, and formulates it as a constrained Markov decision process model. The constrained multi-agent deep reinforcement learning algorithm is a constrained Markov decision process model, which combines the deep reinforcement learning algorithm and the Lagrangian primal pairing method. By combining the actor-critic network with a personalized attention mechanism to train the agent, and by adjusting the Lagrange multipliers to assign different attention weights to the optimization objective and penalty constraints, the optimal trajectory control strategy is obtained.
[0143] To solve the UAV trajectory control sub-problem and find the optimal trajectory control and associated allocation strategy that satisfy the flight safety of the UAV, the present invention constructs a constrained Markov decision process model to solve the problem of the constrained deep reinforcement learning scenario, where the actions of the agent are constrained and in turn affect the interaction with the environment. The constrained Markov decision process model includes the following four key elements:
[0144] State: Set is the state space, including where the symbol S1 contains the position coordinates of the UAV the symbol S2 contains the relative position between the UAV and the intelligent sensing device, that is g i (t) is the position coordinates of the intelligent sensing device; the symbol S3 contains the relative position between the UAV and other UAVs, that is the symbol S4 contains the remaining battery energy of the UAV, that is The symbol S5 includes the remaining transmission data volume of the intelligent sensing device, i.e., ε i (t);
[0145] Action: Set is the action space, including variable a u (t) represents the movement of the drone, including the flight movement direction χ of the drone u within the time period t u (t) and the flight movement distance d u (t);
[0146] Reward: After the drone takes the action a(t) in the state s(t), it obtains an immediate reward
[0147] Penalty: The penalty constraint is represented by j(t) for the collision of the drone with other drones during the movement, i.e., the penalty for not satisfying the constraint of the penalty
[0148] In the traditional constrained deep reinforcement learning setting, the linear combination of the original reward and penalty constraints is used to construct the final reward. However, the weight value related to the reward is uncertain and needs to be found through trial and error. In addition, when there are multiple constraint conditions, the problem becomes quite difficult to solve. Therefore, the present invention constructs a Lagrangian penalty reward function, and assigns different attention weights to the optimization objective and penalty constraints by adjusting the Lagrangian multiplier, including constructing the Lagrangian penalty reward function as:
[0149]
[0150] wherein, is the Lagrangian multiplier corresponding to the penalty constraint
[0151] As Figure 3 shown, the steps to obtain the optimal trajectory control strategy are as follows:
[0152] Deploy a constrained deep reinforcement learning agent on each drone, which includes multiple deep neural networks, namely the policy actor network, the policy critic network, the penalty actor network, the penalty critic network and their corresponding target networks;
[0153] Train the agent; after the training is completed, the agent uses the locally deployed policy network for decision-making;
[0154] Initialize the network parameters of the deep neural network, the Lagrangian multiplier and the experience replay buffer D; at the beginning of each epoch, all drones can serve the user equipment from different location points;
[0155] In time slot t, each UAV agent extracts and executes an action from the policy given by the current policy actor-critic network according to its observation, that is, determines the joint action set of the movement direction and distance of the UAV;
[0156] The environment updates the state according to the joint action set collected from the UAVs to obtain a new state; if the UAV flies out of the target area, its movement will be terminated; accordingly, the rewards of all UAVs and corresponding penalties are obtained and the Lagrangian penalty reward function is constructed;
[0157] The state transition is constructed and stored in the replay experience buffer D; when the buffer space reaches the upper limit, each UAV agent will sample a small batch of experiences in the replay experience buffer.
[0158] Among them, the actor-critic network is responsible for improving the policy and penalty, while the critic network is responsible for evaluating the policy and penalty. During the training process, the agent uploads the observed state to the value network deployed on the central controller and returns the time difference error to the agent. After the training is completed, the agent uses the locally deployed policy network for decision-making.
[0159] The present invention constructs each UAV as an agent and combines an attention mechanism in the network architecture to obtain the best UAV trajectory optimization decision and maximize the safety of UAV flight.
[0160] The execution process pseudocode of the designed energy transmission and data acquisition algorithm for UAV-assisted intelligent sensing devices is shown in Table 1.
[0161]
[0162] Table 1
[0163] Using the constrained multi-agent deep reinforcement learning algorithm to solve the UAV trajectory control optimization problem for the first sub-problem, further optimization is also included, and its network parameters are further updated, including the following steps:
[0164] Update the network parameters of the policy critic and penalty critic by minimizing the loss function;
[0165] Update the parameters of the policy actor network through policy gradient;
[0166] After the parameters of the policy critic network and policy actor-critic network are updated on the fast time scale, the penalty actor-critic network on the slower time scale is updated using the Lagrangian multiplier.
[0167] Specifically, by minimizing the loss function Δ(ζ Q) to update the network parameters ζ of the policy critic network Q , which is defined as follows:
[0168]
[0169] Similarly, the network parameters φ of the penalty critic Q are also updated as follows:
[0170]
[0171] where an entropy term that favors the stochastic policy is introduced, representing the coefficient used to control the randomness of the policy.
[0172] Then, the policy actor network of each agent can update the parameters through policy gradient, which is defined as follows:
[0173]
[0174] where b is a baseline function independent of the action of the UAV agent u, and a u′ (t) represents the actions of agents other than agent u.
[0175] After completing the parameter updates of the policy critic network and the policy actor network on the fast time scale, the penalty actor network on the slower time scale can update the Lagrange multiplier using the following formula:
[0176]
[0177] where ι represents the learning rate for updating the Lagrange multiplier, represents the threshold of the penalty constraint.
[0178] Finally, for their respective target networks, the target network parameters can be updated through the update rate, and the specific formula is as follows:
[0179]
[0180] where represents the update rate for updating the network parameters of the target network. Similarly, the parameter update formulas for the penalty actor target network and the penalty critic target network are also similar. E[·] represents the corresponding expected value; Q u | is the output of the policy critic network, obtained based on the current state, the actions taken, and the network parameters; ξ represents the discount factor for future rewards; represents the UAV reward after being adjusted by the penalty term, considering the trade-off between rewards and penalties, and is used to measure the comprehensive performance of the UAV; represents the policy of the policy actor network parameters; is the network parameter of the policy actor network for the drone u; The policy representing the policy actor network parameter of the drone u; The output value representing the policy critic network parameter of the drone u; The output value representing the penalty critic network parameter; ζ Q′ is the target network parameter corresponding to the policy critic network.
[0181] The pseudo-code of the execution process of the designed network parameter update algorithm is shown in Table 2.
[0182]
[0183] Table 2
[0184] In the steps of using the drone-assisted intelligent sensing device association and allocation algorithm to solve the association and allocation sub-problem of drones and intelligent sensing devices, after the drones move and adjust their positions, each sensing device can choose whether to be served by the drones. In addition, the intelligent sensing device can only be associated with the drones within the sensing range of the drones to be served. Therefore, the present invention uses a set of size I to record the association between the intelligent sensing devices and the drones. After that, by constructing the preference list of the u-th drone to record the user devices that can improve energy efficiency. In addition, the present invention also designs a drone-assisted intelligent sensing device association and allocation algorithm to determine the association and allocation of intelligent sensing devices within the sensing range of the drones and meet the quality of service and energy consumption constraints of the intelligent sensing devices to improve the utilization efficiency of resources.
[0185] The pseudo-code of the execution process of the designed drone and intelligent sensing device association algorithm is shown in Table 2.
[0186]
[0187]
[0188] Table 3
[0189] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the trajectory and resource allocation method based on the drone-assisted wireless power supply network according to any one of claims 1 to 7 are implemented.
[0190] The present invention discloses a method for trajectory and resource allocation based on an unmanned aerial vehicle (UAV)-assisted wireless power supply network. While taking the urban sensing and monitoring scenario as the research background, it studies the problems of difficult data collection and energy supply for intelligent sensing devices in a timely manner. By leveraging the advantage of flexible deployment of UAVs, on the premise of ensuring flight safety, it optimizes UAV data collection and trajectory control. The present invention is a method based on constrained multi-agent deep reinforcement learning that optimizes the trajectory control strategy of UAVs while considering UAV flight safety and energy limitations. Under limited resources, it can obtain the optimal trajectory control strategy while determining the association of intelligent sensing devices within the sensing range of the UAV and satisfying the quality of service and energy consumption constraints of the intelligent sensing devices, making the present invention more effective in terms of data collection, service fairness, and energy utilization.
Claims
1. A trajectory and resource allocation method based on an unmanned aerial vehicle (UAV)-assisted wireless power supply network, including an Internet of Things (IoT) network with U UAVs and I intelligent sensing devices for wireless power supply, for UAV-assisted energy transmission and data collection; characterized in that, The method includes the following steps: Construct a system model of a multi-UAV assisted wireless power supply network, and determine the movement model, power transfer model, and data collection model; According to the constructed movement model, power transfer model, and data collection model, propose a multi-objective optimization problem for maximizing the system energy efficiency and model it: s.t. Constraint conditions C1: C2: C3: C4: C5: C6: α i,u (t)l i,u (t) ≤ l z C7: C8: C9: Among them, Ω i,u (t) is the expected size of the collected data, Λ i (t) is the service fairness of the UAV; e u (t) is the total energy consumption for completing UAV movement, energy transfer, and data collection. t is the time slot, and T is the time period. is the movement position strategy of UAV u; χ u (t) is the flight action of UAV u. and represent the side length of the target area; h u (t) is the flight altitude of the UAV. and represent the minimum and maximum flight altitudes of the UAV; q u (t) is the position coordinate of the u-th UAV in the t-th time period; α i,u (t) is the user association decision between the u-th UAV and the i-th intelligent sensing device in the t-th time period, l i,u (t) is the three-dimensional distance between the UAV and the sensing device, l z is the sensing distance of the UAV; w i,u (t) is the data rate of the upload link between the intelligent sensing device and the UAV, p n is the transmission power of the intelligent sensing device. is within the sensing range of the UAV The received power of the intelligent sensing device. is the time spent by the UAV transmitting energy to the intelligent sensing device within the coverage range; E u (0) is the initial battery capacity of the UAV; v max represents the maximum deadline of the task. represents the time spent by the UAV moving to adjust its own position and coverage range. is the time spent by the UAV collecting data uploaded by the device; x u (t) is the x-axis position coordinate of the u-th UAV in the t-th time period; y u (t) is the y-axis position coordinate of the u-th UAV in the t-th time period. is the minimum safe flight distance of the u-th UAV; q u′ (t) is the position coordinate of the u'-th UAV in the t-th time period. Decompose the proposed multi-objective optimization problem for maximizing the system energy efficiency into two sub-problems. The first sub-problem is the UAV trajectory control optimization problem, and the second sub-problem is the association allocation optimization problem between the UAV and intelligent sensing devices; Use a constrained multi-agent deep reinforcement learning algorithm to solve the UAV trajectory control optimization problem for the first sub-problem; Use a UAV-assisted intelligent sensing device association allocation algorithm to solve the association allocation optimization problem between the UAV and intelligent sensing devices for the second sub-problem.
2. The trajectory and resource allocation method for the UAV-assisted wireless energy supply network according to claim 1, wherein The process of decomposing the proposed multi-objective optimization problem for maximizing the system energy efficiency into two sub-problems includes the following steps: Allow all intelligent sensing devices within the sensing range of the UAV to upload data simultaneously; Convert the maximization of the system energy efficiency over the entire system time into the maximization of the corresponding energy efficiency for each time slot; When the movement position strategy of the UAV u is given the upper limit of the energy transfer time of the constraint condition C9 is converted to: where δ c is the sampling frequency of the sensors equipped on the UAV; Establish a model for the first sub-problem of the UAV trajectory control optimization problem s.t. Constraint conditions Establish a model for the optimization problem of the association allocation between drones and intelligent sensing devices s.t. Constraint conditions 3. The trajectory and resource allocation method for an unmanned aerial vehicle (UAV)-assisted wireless power transfer network according to claim 2, wherein The constrained multi-agent deep reinforcement learning algorithm is a constrained Markov decision process model, which combines the deep reinforcement learning algorithm and the Lagrangian primal pairing method. By combining the actor-critic network with a personalized attention mechanism to train the agent, and by adjusting the Lagrange multiplier to assign different attention weights to the optimization objective and penalty constraints, the best trajectory control strategy is obtained.
4. The trajectory and resource allocation method based on the UAV-assisted wireless power supply network according to claim 3, wherein The constrained Markov decision process model includes the following four key elements: Status: Set is the state space, including where the symbol S1 contains the position coordinates q u (t) The symbol S2 contains the relative position between the UAV and the intelligent sensing device, i.e., ||q u (t)-g i (t)||, g i (t) is the position coordinates of the intelligent sensing device; the symbol S3 contains the relative position between the UAV and other UAVs, i.e., ||q u (t)-q u′ (t)||, The symbol S4 contains the remaining battery energy of the drone, i.e., The symbol S5 includes the remaining amount of transmitted data of the intelligent sensing device, i.e., ε i (t); Action: Aggregation is the action space, including variable a u (t) represents the movement of the UAV, including the flight movement direction χ of the UAV u within the time period t u (t) and the flight movement distance d u (t); Reward: When the drone takes action \(a(t)\) in state \(s(t)\), it receives an immediate reward Penalty: The penalty constraint is represented by j(t) for the collision of the UAV with other UAVs during the movement, that is, the penalty for not satisfying the constraint penalty.
5. The trajectory and resource allocation method based on an unmanned aerial vehicle-assisted wireless power supply network according to claim 4, wherein Assign different attention weights to the optimization objective and penalty constraints by adjusting the Lagrange multiplier, including constructing the Lagrangian penalty reward function as: where θ is the Lagrange multiplier corresponding to the penalty constraint.
6. The trajectory and resource allocation method for an unmanned aerial vehicle-assisted wireless power supply network according to claim 4, wherein The steps to obtain the best trajectory control strategy are as follows: Deploy a constrained deep reinforcement learning agent on each UAV, which contains multiple deep neural networks, namely the policy actor network, policy critic network, penalty actor network, penalty critic network, and their corresponding target networks; Train the agent; after training is completed, the agent uses the locally deployed policy network for decision-making; Initialize the network parameters of the deep neural network, the Lagrange multipliers θ1, ..., θ m and the experience replay buffer D; at the beginning of each epoch, all drones can serve user equipment from different location points; In time slot t, each UAV agent extracts an action from the policy given by the current policy actor network according to its observation and executes it, that is, determines the joint action set of the movement direction and distance of the UAV; The environment updates its state according to the combined action set collected from the drones to obtain a new state; if a drone flies out of the target area, its movement will be terminated; based on this, the rewards of all drones are obtained and corresponding penalties and a Lagrangian penalty reward function is constructed; Construct state transition and store it in the replay experience buffer D; when the buffer space reaches the upper limit, each UAV agent will sample a small batch of experiences from the experiences in the replay experience buffer.
7. The trajectory and resource allocation method for the UAV-assisted wireless power supply network according to claim 6, wherein Using the constrained multi-agent deep reinforcement learning algorithm to solve the UAV trajectory control optimization problem for the first sub-problem also includes further optimization and further updating of its network parameters, including the following steps: Update the network parameters of the policy critic and penalty critic by minimizing the loss function; Update the parameters of the policy actor network through policy gradients; After completing the parameter update of the policy critic network and policy actor network on the fast time scale, the penalty actor network on the slower time scale is updated using the Lagrange multiplier.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the trajectory and resource allocation method for the drone-assisted wireless power supply network according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned aerial vehicle wireless energy supply air calculation assisted Internet of Things data acquisition method
CN116170776A
Multi-unmanned aerial vehicle autonomous navigation and task allocation algorithm for wireless self-powered communication network
CN113776531A
Fairness-based sensor energy consumption and service life balancing method and unmanned aerial vehicle Internet of Things system
CN113985917A