A Deployment and Resource Optimization Method for a Wireless Energy Supply Network Based on Multi-UAV Assistance
Through clustering algorithms and multi-agent collaborative decision-making algorithms, drone task allocation and service scheduling are optimized, combined with the approximate strategy of Lagrangian dual optimization, the balance between drone battery energy and equipment demand is solved, efficient wireless energy supply network deployment and resource optimization are achieved, and the throughput and energy-saving performance of communication services are improved.
Patent Information
- Application Number
- CN202311077525.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-08-25
AI Technical Summary
The prior art is difficult to effectively balance the energy demand of the limited battery energy of the drone with the wearable device, and it is difficult to optimize the tight coupling problem of the drone's altitude position and equipment scheduling, resulting in high cost and low efficiency of communication services.
The clustering algorithm and multi-agent collaborative decision-making algorithm are used to optimize task allocation and drone service scheduling, combined with the approximate strategy optimization algorithm of Lagrangian original dual optimization, adjust the drone altitude and user time slot scheduling by constructing a constrained Markov decision-making process to achieve high maneuverability and low cost of drone services.
On the premise of ensuring that the energy of the drone is non-negative, the network throughput is improved, energy consumption is reduced, service completion rate and service fairness of the equipment are improved, and good energy-saving performance is shown.
Smart Images

Figure CN117119489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for deploying and optimizing resources of a multi-UAV-assisted wireless power supply network in the academic field, and particularly to an approximate strategy optimization method based on Lagrangian primal-dual optimization that can simultaneously optimize the UAV altitude and user scheduling strategy. Background Art
[0002] In recent years, both wireless energy transfer and UAV communication technologies have been regarded as key technologies to improve device performance. Wireless energy transfer can achieve wireless charging by transmitting energy through the wireless spectrum. Trajectory-controllable UAVs, as aerial edge servers, can provide dynamically controllable services for users. The multi-UAV-assisted wireless power supply platform combines these two technologies, which can not only achieve highly mobile communication services but also reduce the energy supply cost. For the multi-UAV-assisted wireless power supply platform, there are two key issues to be solved. One key issue is how to balance the limited battery energy of UAVs and the energy requirements of wearable devices. Another key issue is how to solve the problem of the tight coupling between the altitude position of UAVs and device scheduling. Therefore, simultaneously achieving the balanced load of UAVs and the task allocation of devices and reasonably optimizing UAV altitude adjustment and device scheduling remains to be further explored by researchers. Summary of the Invention
[0003] The main purpose of the present invention is to propose a method for deploying and optimizing resources of a multi-UAV-assisted wireless power supply network in view of some deficiencies in existing research, and use a clustering algorithm, a multi-agent collaborative trajectory optimization algorithm, and an approximate strategy optimization algorithm based on Lagrangian primal-dual optimization to respectively complete task allocation, UAV service scheduling, and user time slot scheduling and UAV altitude adjustment.
[0004] The technical solution adopted by the present invention is as follows: 1. A method for deploying and optimizing resources of a multi-UAV-assisted wireless power supply network, characterized by including the following steps:
[0005] 1) Construct a system model, determine the communication model and energy model, and construct an optimization problem with the optimization objectives of maximizing throughput and minimizing task completion time;
[0006] 2) Divide the optimization problem in step 1) into three sub-problems, namely a task allocation sub-problem, a UAV service scheduling sub-problem, and a user time slot scheduling and UAV altitude adjustment sub-problem;
[0007] 3) Use the clustering algorithm and the multi-agent collaborative decision-making algorithm to solve the task allocation sub-problem and the UAV service scheduling sub-problem in step 2) respectively;
[0008] 4) Analyze the user time slot scheduling and UAV altitude adjustment sub - problems in step 2) using the theory of deep reinforcement learning, and construct a constrained Markov problem;
[0009] 5) Establish an agent model for the problem in step 4) and train this model.
[0010] Advantages of the present invention:
[0011] The present invention constructs a dynamic service deployment framework for realizing highly mobile communication services and low - cost energy supply in a multi - UAV - assisted wireless power supply network. On the premise of ensuring non - negative energy of UAVs, considering the limited coverage of multiple UAVs and the service requirements of wearable devices, the present invention first proposes a clustering algorithm for balancing task loads to determine candidate hovering point positions, then designs a multi - agent collaborative decision - making algorithm to obtain service scheduling decisions of UAVs among user clusters, and then the UAVs in the user clusters perform user time slot scheduling and UAV altitude position adjustment, aiming to maximize network throughput. To solve the time slot scheduling of wearable devices and altitude control of UAVs in an energy - constrained UAV - assisted network, the present invention transforms this problem into a constrained Markov decision process, and on the basis of Lagrangian primal - dual policy optimization, proposes a nested constrained approximate policy optimization algorithm, where the UAV decides the moving distance for altitude adjustment and the wearable devices to be served within the current coverage range. Simulation results show that compared with two existing deep reinforcement learning algorithms, the solution of this study has good service completion rate and energy - saving performance. Description of the Drawings
[0012] Figure 1 It is a downlink network system for wireless power supply of a UAV - assisted intelligent wearable device network.
[0013] Figure 2 It is a network architecture diagram of a multi - agent collaborative decision - making algorithm.
[0014] Figure 3 and Figure 4 shows the performance of the algorithm designed by the present invention and two other baseline algorithms in terms of total energy consumption. The experimental data results show that the method of using deep reinforcement learning by the algorithm designed by the present invention to simultaneously learn load and altitude adjustment strategies among multiple UAVs is effective. Compared with the other two algorithms, the present invention can obtain the lowest total energy consumption.
[0015] Figure 5 and Figure 6 shows the performance of the algorithm designed by the present invention and two other baseline algorithms in terms of service fairness. The experimental data results show that the present invention can still obtain high service fairness when the number of wearable devices and the number of time slots are large. Detailed implementation manners
[0016] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe in detail the specific implementation manners of the present invention.
[0017] An embodiment of the present invention provides a method for deploying and resource optimizing a wireless power supply network assisted by multiple unmanned aerial vehicles. The method includes:
[0018] Step 1: Construct a system model and determine a communication model and an energy model.
[0019] As Figure 1 shown, the present invention constructs a system model, which includes I intelligent wearable devices, denoted as and U unmanned aerial vehicles, denoted as Before the task starts, the unmanned aerial vehicles load sufficient power resources at the base station. When the intelligent wearable devices send out energy requests, the unmanned aerial vehicles fly to a given hovering position (for example, directly above the center of the user cluster) according to a continuous hovering flight strategy to provide energy supply to the user devices. Considering the limited coverage range of the unmanned aerial vehicles, it can adjust its height to serve the intelligent wearable devices in the hot spot area. The energy requirements of the intelligent wearable devices in the user cluster are represented by w i . The present invention assumes that most of the transmission power of the unmanned aerial vehicle is concentrated within the aperture angle of α directly below the unmanned aerial vehicle.
[0020] Adopt a time slot-frame structure, and divide the time domain into a set of L max equal-length time frames. l represents the l-th time frame. Each frame consists of K time periods, representing the set of time periods, and the duration of each time period is δ. As the height position of the unmanned aerial vehicle changes, the set of wearable devices currently covered by the unmanned aerial vehicle is In addition, the intelligent wearable devices arranged within this time period are defined as a service device group, and the optional candidate service device groups are denoted as z represents the z-th candidate service device group. The maximum number is On the k-th time period of the l-th frame, the number and set of the currently served device group z are represented by C k,z,l and .
[0021] The present invention uses 3D Euclidean coordinates to model the positions of the unmanned aerial vehicles and the intelligent wearable devices. The coordinates of the intelligent wearable devices are (x i , y i , 0) and the coordinates of the unmanned aerial vehicle are (x u,k,l , y u,k,l , h u,k,l), where h u,k,l is the altitude of the UAV. Therefore, the distance between the UAV u and the smart wearable device is expressed as:
[0022]
[0023] In the UAV communication model of the present invention, the influence of the line-of-sight (LoS for short) link and the non-line-of-sight (NLoS for short) link on the communication model modeling is considered. The LoS link probability between the UAV u and the smart wearable device i can be expressed as:
[0024]
[0025] Where and are constants related to environmental conditions. The symbol θ i,u represents the elevation angle between the UAV and the user equipment, and follows θ i,u =(180 / π)arctan(h u,k,l / χ i,u ). The symbol represents the horizontal distance between the UAV and the device.
[0026] Correspondingly, the probability of the NLoS link is Then, the path loss models of the LoS link and the NLoS link between the UAV and the smart wearable device are expressed as:
[0027]
[0028]
[0029] Among them, the symbols f c and c represent the carrier frequency and the speed of light respectively. The symbol κ represents the path loss exponent. and describe the additional path losses of the LoS link and the NLoS link respectively. Further, the average path loss is expressed as:
[0030]
[0031] In addition, G i,u =1 / L i,u is defined as the average channel gain. Then, the signal-to-noise ratio (SNR for short) from the UAV to the smart wearable device can be expressed as:
[0032]
[0033] Among them, P t represents the transmission power, and σ 2 is the variance of the additive white Gaussian noise (AWGN for short). Further, the total throughput of the wearable device network in the present invention is defined as:
[0034]
[0035] where B is the system bandwidth.
[0036] Each drone is equipped with a battery with an initial energy of E init to provide power. In each time period k of the l-th frame, the drone moves at a certain speed η, and then interacts with the environment at the new hovering position. The propulsion power of the drone's flight motion is calculated as follows:
[0037]
[0038] where the symbols P0 and P s represent the blade section power and the induced power in the hovering state of the drone respectively. The symbols u tip and v0 represent the tip speed of the drone's rotor blade and the average rotor induced speed in the hovering state respectively. The symbol d0 represents the fuselage drag ratio, ρ represents the air density, s represents the rotor solidity, and J represents the rotor disk area. Therefore, the flight energy consumption of the drone can be defined as: where l f is the flight time for the drone to adjust its altitude on the l-th frame, expressed as l f = ||h u (k + 1)- h u (k)|| / η, and the symbols h u (k + 1) and h u (k) represent the altitudes of the drone at the (k + 1)-th and k-th time periods of the l-th frame respectively.
[0039] The distance that the drone u flies from the starting point to the first user cluster it serves is denoted as d u , and d n,n+1 represents the distance between the user cluster μ n and the user cluster μ n+1 . Therefore, the flight time l u of the drone u can be expressed as:
[0040]
[0041] where M u represents the total number of user clusters traversed by the drone u. Since the total number of user clusters is N, then 1 ≤ M u≤N and In addition, the hovering energy consumption of the UAV is:
[0042]
[0043] Where the symbol θ represents the drag coefficient, and the symbols V and R represent the angular velocity of the rotor blade and the rotor radius respectively. The symbol W is the weight of the UAV, W = mg, where m is the mass of the UAV and g is the acceleration due to gravity. The variable l represents the incremental correction coefficient.
[0044] In the energy supply phase of the downlink, the UAV u transmits energy to the smart wearable device in the hovering state, where the hovering time l of the UAV h is equal to the energy supply time l c , therefore, the total energy consumption of the UAV for energy supply and hovering can be expressed as:
[0045]
[0046] Where P c is the transmission power of the UAV during energy transmission to the smart wearable device. Therefore, the battery energy of the UAV in the time frame l is Furthermore, the time for the UAV u to complete the task is defined as:
[0047]
[0048] In addition, the energy collected by the smart wearable device i from the UAV u is given by:
[0049]
[0050] Where, is the radio frequency to DC conversion rate, and its range is
[0051] The optimization objective of the present invention is to maximize the throughput and minimize the task completion time. The multi-objective optimization problem is described as follows:
[0052]
[0053]
[0054] s.t.
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066] Among them, the variable Υ i,n allocates metrics for wearable devices. Υ i,n = 1 indicates that the user device i is allocated to the user cluster μ n . The variable Ψ u,n is the association metric between the drone and the user cluster. Ψ u,n = 1 indicates that the drone u selects to serve the user cluster μ n . Let the decision variable p u,z,k,l be the scheduling metric for wearable devices. p u,z,kl = 1 indicates that the wearable device group is selected by the drone u for service at time slot k. Another decision variable represents the distance at which the drone adjusts its altitude at frame l in time slot k.
[0067] Constraints 1 and 4 respectively determine the validity of the decision variables Υ i,n and Ψ u,n . Constraint 2 indicates that the user device i can only be allocated to one user cluster. In Constraint 3, by setting the variable Π max the service threshold of a user cluster is indicated. Constraint 5 ensures that each user cluster can only be served by one drone. Constraint 6 indicates that the number of user clusters selected by the drone u for service is within the range [1, N]. Constraints 7 and 8 limit the ranges of the decision variables p u,z,k,l and h u,kl , where h max represents the maximum flight altitude of the drone. Constraint 9 ensures that the energy transfer requirements of all wearable devices are met within L max time frames. Constraint 10 indicates that each drone u can only serve one group of wearable devices within a period of time. Constraint 11 sets the SNR threshold Λ th , and if the SNR is greater than the threshold, the transmission is considered successful.
[0068] Step 2: Use the clustering algorithm and the multi-agent collaborative decision-making algorithm to solve the task allocation sub-problem and the UAV service scheduling sub-problem in step 1) respectively.
[0069] In the optimization problem described in step 1), the altitude position of the UAV and the scheduling of the wearable device are tightly coupled. To solve the optimization problem described in step 1), the present invention divides it into three sub-problems, namely the task allocation sub-problem, the UAV service scheduling sub-problem, and the user time slot scheduling and UAV altitude adjustment sub-problem. In this part, an improved clustering algorithm is first considered to ensure the task balance of each user cluster, and then a multi-agent collaborative decision-making algorithm is used to complete the UAV service scheduling.
[0070] Decompose the problem in step 1) into a task allocation sub-problem, a UAV service scheduling sub-problem, and a user time slot scheduling and UAV altitude adjustment sub-problem. Among them, the task allocation sub-problem is expressed as follows:
[0071]
[0072] s.t.
[0073]
[0074]
[0075]
[0076] Among them, and (x i , y i ) are the two-dimensional position of the cluster center of the user equipment cluster and the position of the smart wearable device respectively, and μ n is the vector mean of the equipment cluster of user cluster n .
[0077] The present invention designs an improved clustering algorithm to balance the load of each UAV. First, according to the distance between user devices, the users are divided into N user clusters (μ1, μ2,..., μ N ), and the user devices in the cluster are divided into the range of the nearest cluster center. In the process of dividing the user devices into the corresponding user clusters, if the total demand of the current user cluster is less than the service threshold of the user cluster, then the user device is divided into the nearest user cluster, otherwise the user is divided into the second-nearest user cluster.
[0078] The specific details of the improved clustering algorithm are shown in Table 1.
[0079]
[0080] In addition, the task allocation sub-problem is solved through the decision variable Υi,n Partition the user equipment into user clusters to obtain candidate hovering point positions and balance the task assignment of the user clusters, which is used as the input of the UAV service scheduling sub-problem. The UAV service scheduling sub-problem determines the UAV service scheduling by selecting candidate hovering points to minimize the task completion time. This sub-problem is expressed as follows:
[0081]
[0082] s.t.
[0083]
[0084]
[0085] After obtaining the candidate hovering point positions by solving the task assignment sub-problem, in order to minimize the flight time of the UAVs between different user clusters as much as possible u , the optimization objective of the UAV service scheduling sub-problem is transformed into minimizing the distance of the longest sub-path. Then, a multi-agent collaborative decision-making algorithm is proposed to solve the UAV service scheduling problem. The network architecture of the multi-agent collaborative decision-making algorithm is as Figure 2 shown. The present invention constructs a network architecture including a shared graph network and a distributed policy network. Each UAV is regarded as an agent, and an allocation hovering position policy network and a determining service sequence policy network are deployed on the agent to obtain an optimal UAV solution. Considering the limited computing power of the UAVs, this research places this network architecture on the edge cloud.
[0086] The pseudo-code of the multi-agent collaborative decision-making algorithm is shown in Table 2. The input of this algorithm is the candidate hovering point position information, the number of UAVs, and the position of the take-off point. By allocating the candidate hovering points to the task sets of different UAV agents, the initial solution of each agent is generated. Then, an improved neighborhood search algorithm is used to determine the sequence of hovering points visited by an agent. Finally, the output solution consists of the service hovering point trajectories of multiple agents.
[0087]
[0088] Specifically, the present invention constructs a Markov decision model for the allocation hovering position policy network, where the state S t represents the candidate hovering point position state at round t, which contains multiple solutions, that is, the set of candidate hovering points selected by each agent for service. The action A t represents which agent's service task set each candidate hovering node is assigned to in the current S t state. The reward R t represents minimizing the distance of the longest sub-path.
[0089] In addition, the agents can utilize the observed hovering point position information x(μ1), x(μ2),..., x( N ), to construct the graph structure of the hovering points through the shared graph network. In this graph structure, the positions of the candidate hovering points are regarded as nodes, and the distances between the candidate hovering points are regarded as edges, forming a complete graph, which shows the connection relationship of the candidate hovering points. Between adjacent hovering points, the feature vectors of the hovering nodes can be obtained by means of message passing and message aggregation and the graph embedding information as the input information for each hovering position allocation policy network
[0090] Since each UAV agent departs from the starting point h1 and flies to each candidate hovering point, except for the starting point which will be visited by multiple UAV agents, the rest of the candidate hovering points will only be visited once. Let h c = [g h , h1] be used as the context embedding. In addition, in the hovering position allocation policy network, the attention mechanism is used to first construct the respective agent embedding information for each agent through the global hovering node feature information, as shown in the following formula
[0091]
[0092] where and represent the query matrix, the key-value matrix, and the value matrix respectively and represent their corresponding matrix parameters respectively. d k and d v represent the dimensions of the key-value matrix and the value matrix respectively. Additionally represent the dimensions of the hovering node embedding information and the agent embedding information respectively. Then, combining the previous hovering point graph embedding information g h , the scoring coefficient obtained by using the similar attention mechanism mentioned above is used to measure the importance probability of each hovering point position for each agent, as shown in the following formula
[0093]
[0094] where the matrix result is restricted to [-C, C] by using the tanh function, C and They respectively represent the adjustment parameter, the query transpose matrix, and the key-value transpose matrix. Then, each agent determines the allocation decision of which hovering points each agent serves respectively according to the importance probability. After obtaining the allocation candidate hovering point strategy of the UAV agents, the candidate hovering point sequence visited by each agent is determined through an improved neighborhood search algorithm. The output of each agent of the multi-agent collaborative decision-making algorithm is used as the input initial solution of each agent of the improved neighborhood search algorithm, and a deep learning framework is used to learn a specific pairwise local operator selection strategy to improve the initial solution. The process of the improved neighborhood search algorithm is shown in Table 3.
[0095]
[0096] Step 3: Analyze the user time slot scheduling and UAV altitude adjustment sub-problems in step 2) using the theory of deep reinforcement learning.
[0097] The user time slot scheduling and UAV altitude adjustment sub-problems in step 2) are still difficult to solve, mainly because it is difficult to balance the limited coverage and battery energy of the UAV and the energy requirements of the wearable devices. In this part, the user time slot scheduling and UAV altitude adjustment sub-problems are constructed as a constrained Markov decision process, and an optimization problem based on deep reinforcement learning is established.
[0098] After solving the task allocation sub-problem and the UAV service scheduling sub-problem, each UAV maximizes the throughput of each user cluster by adjusting its altitude position and selecting user devices:
[0099]
[0100] s.t.
[0101]
[0102]
[0103]
[0104]
[0105]
[0106] After balancing the task load and determining the candidate hovering positions, at this time, the UAV provides services to the intelligent wearable devices in each user cluster. Considering the coverage range of the UAV, the limited battery energy constraint, and the energy requirements of the wearable devices, the present invention first constructs the user time slot scheduling and UAV altitude adjustment sub-problems as a constrained Markov decision process model, and then proposes a constrained approximate policy optimization algorithm.
[0107] First, the state definition of the system consists of three parts: the current altitude h of the drone u,k,l ∈[H min H max , the remaining battery capacity of the drone and the energy demand D not obtained by the intelligent wearable device i i,l . The state S l can be defined as:
[0108] Secondly, the actions of the drone consist of two parts. One is the altitude movement distance in the l time frame and the other is the user time slot allocation p of the drone u in the time period frame l u,k,l ={p u,1,l , …, p u,K,l}, where p u,k,l ∈{1, …, z, …Z u,k,l}, p u,k,l =z means that the z user equipment candidate group in the k time period of the l time frame is selected by the drone u for service. Therefore, the action space is defined as
[0109] Then, the definitions of rewards and penalties will be introduced below. The immediate reward is defined as the throughput of the wearable device network, that is Considering the maximization of the long-term discounted network throughput, a discount factor β is introduced. Therefore, the constrained Markov problem can be formulated as follows:
[0110]
[0111] s.t.
[0112]
[0113]
[0114]
[0115]
[0116]
[0117] where, π ζ represents the user time slot scheduling and drone altitude adjustment strategy. denotes taking the expectation according to the policy π ζ .
[0118] Step 4: Establish an agent model for the problem in step 3) and train this model.
[0119] The present invention regards the drone as an intelligent agent interacting with the wearable device, and learns the optimal altitude adjustment decision and the wearable device scheduling strategy. The present invention proposes a nested constrained approximate policy optimization algorithm to solve this problem by constructing a Lagrangian-based reward and punishment function to relax the constrained optimization problem. In this algorithm, an actor-critic scheme is used to improve the policy, updating the value network and policy network parameters on a faster time scale and updating the Lagrangian parameters on a slower time scale.
[0120] In the proposed constrained deep reinforcement learning algorithm, the policy network updates the network parameters by maximizing the objective function L CLIP (ζ),
[0121]
[0122] where denotes the expectation of the l-th frame, A l and S l represent the action and state at the l-th frame respectively. ζ is the network parameter of the policy network, and the symbols π ζ and π ζol represent the old and new policies respectively. The symbol ∈ represents the hyperparameter for adjustment and clipping, and clip represents the clipping function, which can be regarded as a regularizer. In this way, local improvement can be made to the current policy in each iteration by setting the ratio of updating the old and new policies within a specific range boundary to improve the policy and solve the problems of update instability and low data efficiency. In addition, the symbol denotes the generalized advantage estimation, which is the key to balancing the influence of bias and variance.
[0123] Secondly, the mean square error of the minimized policy is used to learn the state value function
[0124]
[0125] where denotes the estimated value of the expectation of the l-th frame, and represent the state value function and the Lagrangian penalty reward function from the l-th frame to the l + k-th frame respectively, and β k represents the discount factor at the k-th time period. At the same time, the penalty critic optimizes the penalty value function parameter w q , and the calculation of
[0126]
[0127] where Represents the long-term action value function from frame l to frame l + k. Represents the value estimation function that penalizes the critic.
[0128] By performing the stochastic gradient ascent method, update the Lagrangian parameter γl+1 at frame l + 1 with the penalized actor:
[0129]
[0130] where ε represents the learning rate of the Lagrangian multiplier, and γ l represents the Lagrangian parameter at frame l. and respectively represent the battery levels of the UAV u at frames l and l + 1.
[0131] The pseudocode for the nested constrained approximate policy optimization algorithm is shown in Table 4.
[0132]
[0133] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated still do not exceed the spirit covered by the specification and the drawings, they should still fall within the protection scope of the present invention.
Claims
1. A deployment and resource optimization method for a wireless power supply network based on multi-UAV assistance, characterized in that Including the following steps: 1) Construct a system model, determine the communication model and energy model, and construct an optimization problem with the optimization objectives of maximizing throughput and minimizing task completion time. The optimization problem is: s.t Among them, the variable γ i,n is an allocation metric for wearable devices. γ i,n = 1 indicates that the user device i is assigned to the user cluster μ n Among them, the variable Ψ u,n is an association metric between the UAV and the user cluster. Ψ u,n = 1 indicates that the UAV u selects to serve the user cluster μ n , and the decision variable p u,z,k,l is a scheduling metric for wearable devices. p u,z,k,l = 1 indicates that the wearable device group is selected by the UAV u for service in time slot k. Another decision variable represents the distance at which the UAV adjusts its altitude on frame l at time slot k, are the sets of intelligent wearable devices and UAVs respectively, specifically represented as and represents the set of device groups currently being served in time slot k of frame l, represents the set of candidate service device groups in time slot k of frame l, represents a set in terms of frames; w i represents the energy requirement of the intelligent wearable device in the user cluster. Π max represents the service threshold of a user cluster. N represents the total number of user clusters, and n represents the nth user cluster; h max represents the maximum flight altitude of the UAV; Λ i,u represents the signal-to-noise ratio between the UAV and the intelligent wearable device; T l represents the total throughput of the wearable device network; F u represents the time for the UAV u to complete the task; represents the set of all time slots in frame l; represents the energy collected by the intelligent wearable device i from the UAV u; Λ th represents the SNR threshold; 2) Divide the optimization problem in step 1) into three sub-problems, namely the task allocation sub-problem, the UAV service scheduling sub-problem, and the user time slot scheduling and UAV altitude adjustment sub-problem; 3) Use the clustering algorithm and the multi-agent collaborative decision-making algorithm to solve the task allocation sub-problem and the UAV service scheduling sub-problem in step 2) respectively; 4) Use the deep reinforcement learning theory to analyze the user time slot scheduling and UAV altitude adjustment sub-problem in step 2), and construct a constrained Markov problem; 5) Adopt a nested constrained approximate policy optimization algorithm to solve the problem in step 4) by constructing a Lagrangian-based reward and punishment function, and use an actor-critic scheme to improve the policy and update the penalty value network parameter ω q , the policy network parameter ζ, and the Lagrangian parameter γ, and learn and obtain the optimal height adjustment decision and wearable device scheduling strategy; The policy network updates the network parameters by maximizing the objective function L CLIP (ζ). Among them, A l and S l respectively represent the actions and states of the l-th frame, ζ are the policy network parameters, π ζ and π ζold respectively represent the old and new policies, ∈ represents the hyperparameter for adjustment and clipping, clip represents the clipping function, represents the generalized advantage estimation, represents taking the expectation of the l-th frame; Learn the state value function using the mean squared error with a minimization strategy Among them, and respectively represent the state value function and the Lagrangian penalty reward function from the l-th frame to the (l + k)-th frame. β k represents the discount factor for the time period k, represents the estimated value of the expectation of the l-th frame. At the same time, the penalty critic optimizes the penalty value network parameter ω q and is calculated in a similar way to the state value function; Among them, represents the long-term action value function from the l-th frame to the (l + k)-th frame, represents the value estimation function for punishing the critic; Update the Lagrangian parameter γ at frame l + 1 by penalizing the actor through performing the stochastic gradient ascent method l+1 : where ε represents the learning rate of the Lagrange multiplier, and γ l represents the Lagrange parameter of frame l, and represent the power levels of UAV u at frames l and l + 1, respectively.
2. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 1, characterized in that: The path loss models of the Los link and the NLos link between the UAV and the smart wearable device in the communication model are expressed as: Among them, is the distance between the drone u and the smart wearable device i, (x u,k,l , y u,k,l , h u,k,l ) and (x i , y i ) are the positions of the drone and the smart wearable device respectively; f c and c represent the carrier frequency and the speed of light respectively, κ represents the path loss exponent, and represent the additional path losses of the Los link and the NLos link respectively; and represent the Los link probability and the NLos link probability respectively, and the average path loss is expressed as: Let G i,u = 1 / L i,u be defined as the average channel gain. The signal-to-noise ratio between the UAV and the smart wearable device is expressed as: where P t represents the transmission power, and σ 2 is the variance of the additive white Gaussian noise; the total throughput of the wearable device network is: B is the system bandwidth, represents the set of all time periods in frame l, and k represents the k-th time period; The time for the UAV u to complete the task in the energy model is: Among them, is the set of wearable devices currently covered by the drone, l f is the flight time for the drone to adjust its altitude, l h and c are the hovering time and the energy supply time of the drone respectively. The energy collected by the intelligent wearable device i from the drone u is: Among them, is the radio frequency to direct current conversion rate, P c is the transmission power when the drone transmits energy to the smart wearable device, z represents the z-th candidate service device group, and l represents the l-th frame.
3. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 1, characterized in that: The task allocation sub-problem is: s.t. in, and (x i ,y i ) are the two-dimensional position of the center of the user device cluster and the position of the smart wearable device, μ n is the device cluster of user cluster n The vector mean of ; the drone service scheduling sub-problem is: s.t.
4. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 3, characterized in that: Step 3) The clustering algorithm includes: dividing users into N user clusters (μ1, μ2,..., μ N ), dividing the user devices in the cluster into the range of the nearest cluster center. In the process of dividing the user devices into the corresponding user clusters, if the total demand of the current user cluster is less than the service threshold of the user cluster, then divide the user device into the nearest user cluster, otherwise divide the user into the second nearest user cluster.
5. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 3, characterized in that: For the multi-agent collaborative decision-making algorithm in step 3), each UAV is regarded as an agent, and an allocation hover position policy network and a determine service sequence policy network are deployed on the agent. The input is the candidate hover point position information, the number of UAVs, and the position of the take-off point. An initial solution for each agent is generated by allocating the candidate hover points to the task sets of different UAV agents, and then an improved neighborhood search algorithm is used to determine the hover point sequence visited by an agent; finally, the output solution consists of the service hover point trajectories of multiple agents.
6. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 1, characterized in that: The user time slot scheduling and UAV altitude adjustment sub-problem is: s.t.
7. The deployment and resource optimization method of a wireless power supply network based on multi-UAV assistance according to claim 6, characterized in that: The constrained Markov problem is: s.t. Among them, β l and π ζ represent the discount factor of frame l and the user time slot scheduling and UAV altitude adjustment strategy respectively, denotes taking the expectation according to the strategy π ζ to find the expectation.
Citation Information
Patent Citations
Unmanned aerial vehicle relay system resource allocation method based on wireless energy-carrying communication network
CN110166107A
Multi-link new radio (NR)-physical downlink control channel (PDCCH) design
CN110226295A