A path planning method based on multi-unmanned aerial vehicle assisted data collection

By optimizing drone trajectories through multi-drone collaboration and the Dueling-DDQN algorithm, the problem of insufficient coverage by a single drone is solved, and dynamic planning and maximum coverage are achieved in a distributed user environment.

CN114879726BActive Publication Date: 2025-12-16GUANGDONG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210468940.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-12-16
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In situations where users are dispersed and move freely, existing technologies cannot provide sufficient coverage for a single drone, and traditional optimization decision-making methods cannot solve model-free dynamic programming problems.

Method used

A path planning method for multi-UAV-assisted data collection is adopted, which combines the Dueling-DDQN algorithm to optimize UAV trajectories. Through model-free dynamic programming and deep reinforcement learning, user coverage is maximized.

Benefits of technology

It achieves more comprehensive user coverage in situations where users are dispersed, and optimizes the shortest path to the destination, making it suitable for drone path planning in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114879726B_ABST
    Figure CN114879726B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of path planning methods based on multi-unmanned aerial vehicle assisted data collection, comprising the following steps: in target area, target is divided into several clusters, user coordinates are randomly generated in cluster, there are several users in cluster, and user randomly moves but does not exceed the boundary of region;The communication channel of unmanned aerial vehicle and user is dominated by time limit link, dynamic programming is carried out without model using multiple unmanned aerial vehicles;Using Dueling-DDQN algorithm optimizes the trajectory of unmanned aerial vehicle to maximize user coverage.When the distribution of user is dispersed and can move freely in the whole target area, in order to make up the problem of insufficient coverage of single unmanned aerial vehicle in the case of more dispersed users, more user coverage is realized using multiple unmanned aerial vehicles to assist data collection, and a shortest path to the end point can be optimized, so as to realize maximum user coverage;Also propose Dueling-DDQN algorithm, can accurately estimate neural network output value, plan the action of each step movement of unmanned aerial vehicle, applicable to other different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of path planning, and more particularly, to a path planning method based on multi-unmanned aerial vehicle (UAV) assisted data collection. BACKGROUND

[0002] To improve the network performance and coverage of wireless communication, unmanned aerial vehicles (UAVs) have been deployed in various communication fields, such as wireless sensor networks, cache, heterogeneous cellular networks, massive multiple-input multiple-output (MIMO), disaster communication, and device-to-device communication (D2D). For example, in L. D. Nguyen, A. Kortun, and T. Q. Duong, “An introduction of real-time embedded optimisation programming for UAV systems under disaster communication,” EAI Endorsed Transactions on Industrial Networks and Intelligent Systems, vol. 5, no. 17, pp. 1-8, Dec. 2018, UAVs are deployed to provide network coverage for people in remote areas and disaster areas. In J. Gong, T.-H. Chang, C. Shen, and X. Chen, “Flight time minimization of UAV for data collection over wireless sensor networks,” IEEE J. Select. Areas Commun., vol. 36, no. 9, pp. 1942-1954, Sept. 2018, UAVs are also used to collect data in wireless sensor networks.

[0003] However, the above research either ignores the strict flight time limit in real applications or usually sets the environment as a static environment or the users are too concentrated, but in general, the users or targets to be covered are actually free to move and are generally more dispersed. Due to the limitation of the on-board power and flight time of the UAV, how to obtain the optimal resource allocation scheme under the premise of the fastest arrival at the destination and achieve the maximization of user coverage is a challenge.

[0004] In the prior art, a Chinese invention patent discloses a method for path planning of a UAV group for information collection. The method models an environment monitoring point that needs to be monitored, then one-to-one correspondence between each region and a UAV base is established for task allocation, and finally path planning is performed on a single UAV performing a monitoring task. A K-means task grouping algorithm based on simulated annealing is used to solve the minimum cost UAV flight path under the evaluation model, so as to obtain a path planning method for multiple UAVs to cooperate. The method improves the K-means clustering algorithm by using the simulated annealing algorithm, so that the grouping result is more balanced, the task grouping effect can be more effectively improved, and the path planning distance is shortened. However, the method does not consider that users in a region can move freely, and cannot solve the problem of dynamic planning without a model. SUMMARY

[0005] The present application provides a path planning method based on multi-UAV assisted data collection to solve the technical defects of the existing single UAV in the case of more scattered users, insufficient coverage, and the inability of traditional optimization decision methods to solve the problem of dynamic planning without a model.

[0006] To achieve the above application purposes, the technical solutions adopted are as follows:

[0007] A path planning method based on multi-UAV assisted data collection, comprising the following steps:

[0008] S1: In a target region, the target is divided into several clusters, and user coordinates are randomly generated in the clusters. There are several users in the cluster, and the users move randomly but do not exceed the boundary of the region;

[0009] S2: The communication channel between the UAV and the user is dominated by time limit link, and multiple UAVs are used for dynamic planning without a model;

[0010] S3: The Dueling-DDQN algorithm is used to optimize the UAV trajectory to maximize user coverage.

[0011] In the above scheme, when the distribution of users is scattered and can move freely in the entire target region, in order to compensate for the problem of insufficient coverage of a single UAV in the case of more scattered users, multiple UAVs are used to achieve more user coverage and can optimize a shortest path to the end point, thereby maximizing user coverage. The Dueling-DDQN algorithm based on deep reinforcement learning can accurately estimate the output value of the neural network, make accurate strategies, and plan the action of each step of the UAV movement, which is suitable for other different scenarios.

[0012] Preferably, in step S1, the users in the target area are divided into M clusters, each cluster corresponds to a circle with radius R, the coordinates of the users in these circles are randomly generated, there are K users in each cluster, the position of the kth user in the mth cluster at time step t is Meanwhile, the users move randomly at a speed lower than the maximum speed v, but will not exceed the boundary of the target area, that is and

[0013] In the above scheme, the flight height of the UAV is H, and all clusters are accessed by a single antenna to maximize the coverage of users, the three-dimensional coordinates of the UAV at time step t are defined as The starting point and the ending point of the two UAVs are the same, the maximum coverage range is determined by the flight height H of the UAV and the antenna emission angle θ, that is R max = H tan(θ), at the same time, the UAV can only fly in the specified area, that is 0≤X(t)≤X max and 0≤Y(t)≤Y max , where X max and Y max are the length and width of the area.

[0014] Preferably, in step S2, the communication channel between the UAV and the user is dominated by the line-of-sight link, the distance from the kth user in the mth cluster to the first UAV at time step t is: At time step t, the channel between the first UAV and the kth user in the mth cluster follows the free space path loss model, which is expressed as where β0 represents the power gain of the channel at the reference distance d = 1 m.

[0015] Preferably, in step S3, when the user satisfies the distance constraint and is within the coverage of the UAV, the achievable throughput of the kth user in the mth cluster to the UAV at time t is defined as: If it is simultaneously within the overlapping coverage of multiple UAVs, the throughput of this user at time t is the sum of the throughputs generated by communicating with two UAVs respectively, where B and α 2 are the bandwidth and noise power respectively, the total throughput of the kth user in the mth cluster to the UAV and at time step T is:

[0016] Preferably, in the multi-UAV data collection system, since both UAVs will be subject to end position, UAV coverage, throughput threshold, step penalty and boundary constraint, the maximum coverage of users is achieved by optimizing the trajectories of the two UAVs, and the target problem is as follows

[0017]

[0018] s.tR final1 = X target ,

[0019] d m,k ≤ d cons ,

[0020] R m,k ≥ r min ,

[0021] P(m, k) = {0, 1},

[0022] 0 ≤ X(t) ≤ X max ,

[0023] 0 ≤ Y(t) ≤ Y max ,

[0024] Distance constraint d cons represents the straight-line distance between the served user and the UAV, X now , X now1 , X now2 and X target represent the current position of the UAV under the single-UAV data collection system, the current position of the two UAVs under the multi-UAV data collection system, and the key position, respectively; if the UAV touches the boundary, it will be punished by the boundary punishment R bp = -100, and R sp = -1000 is defined as the step punishment, and the UAV will receive a negative reward for each additional step, and the UAV can only fly in the specified area, i.e., 0 ≤ X(t) ≤ X max and 0 ≤ Y(t) ≤ Y max , where X max and Y max are the length and width of the area, and X(t), Y(t) represent the horizontal and vertical coordinates of the current position of the UAV, respectively.

[0025] In the above scheme, when there is only one UAV in the system, if the UAV reaches the end point, it will directly obtain the end point reward R final1 , and only when both UAVs reach the end point under the multi-UAV data collection system can the end point reward R final2 be obtained, i.e.

[0026]

[0027]

[0028] When there is only one drone in the system, the drone is subject to constraints such as destination location, drone coverage area, throughput threshold, step penalty, and boundary constraints. By optimizing the drone trajectory to maximize user coverage, we can obtain the following objective problem:

[0029]

[0030] s.tR final1 X final =X target ,

[0031] d m,k ≤d cons ,

[0032] R m,k ≥r min ,

[0033] P(m,k)={0,1},

[0034] 0≤X(t)≤X max ,

[0035] 0≤Y(t)≤Y max .

[0036] The above scheme aims to maximize user coverage and provide communication services, while ensuring that the drone takes off from the starting point and reaches the destination in the shortest possible time. Therefore, we define P(m,k) = {0,1}, where P(m,k) is the total throughput R of the k-th user in the m-th cluster. m,k Greater than the threshold r min When the time is right, it means that the user has made contact with the drone and will no longer communicate with the drone in this round of missions. At this time, it is marked as P(m,k)=1, otherwise P(m,k)=0.

[0037] Preferably, in step S3, the Dueling-DDQN algorithm is used so that each drone starts from the starting point and ends at the destination.

[0038] During the training phase, before each scene begins, the starting and ending positions of the drone are initialized, and the positions of M*K users are randomly initialized; at each time step t, the drone adjusts its position based on the observed state information s. t The output action a(t) represents the drone's flight direction. If the user is within the drone's coverage area, the agent will calculate the throughput of communication with each user separately, accumulating this process up to one step, until R... m,k ≥r min If the drone's next location exceeds the designated area, the flight maneuver is cancelled; a corresponding reward is obtained based on the maneuver. t and the state information s at the next moment t+1 ,Will The network parameters are updated by randomly sampling N sets of experience data from the experience buffer at the end of each time step.

[0039] Preferably, the Dueling-DDQN algorithm is a model-free reinforcement learning algorithm for iteratively solving the Bellman equation, and its state-action value function is: in Let π(·) represent the probability that the agent will transition to state s' after taking action a in state s, and let π(·) represent the agent's choice of policy.

[0040] Preferably, the Dueling-DDQN algorithm is configured with a parameter θ. - The target network Q2(s',a) max ;θ - The network consists of an estimation network Q1(s',a;θ) with parameter θ, and a target network that is a copy of the estimation network, with parameter θ of the target network. - The update frequency is slower than the estimated network;

[0041] The Dueling-DDQN algorithm also includes an experience buffer, which tracks the current state, action, reward, and next state. Stored in an experience buffer for later random access to update weights.

[0042] In the above scheme, this paper proposes a Dueling-DDQN algorithm based on reinforcement learning to calculate the optimal trajectory, achieving the goal of reaching the destination in the shortest time while maximizing user coverage. Whether using a single UAV or multiple UAVs, a single agent learns the mapping from state space to action space by continuously interacting with the environment and learning based on environmental feedback. Each step the UAV takes observes its current state s(t) from the environment, inputs s(t) into a deep neural network to obtain the corresponding action a(t), interacts with the environment through action a(t), and receives a reward r(t) and a new state s(t+1) from the environment. The experience (s(t), a(t), r(t), s(t+1)) obtained from the above process is then stored in an experience buffer for training the deep neural network.

[0043] When only one drone is performing a task in the system, the drone is an intelligent agent that interacts with the environment and seeks the peak of reward; in the case of multiple drones, the two drones belong to the same intelligent agent.

[0044] This patent defines the position of the drone as a state space, i.e., S = {x, y, H}, or S = {x1, y1, H1, x2, y2, H2} in a multi-drone data collection system. At time step t, the state of the drone in the above scenarios is defined as s.t = {x t ,y t ,H t} and

[0045] When there is only one UAV in the system performing the task, at time step t, the UAV in state s t can choose an action a t from the action space A according to the policy, by dividing the area into a grid

[0046] A = {left, right, forward, backward}

[0047] In the multi-UAV data collection system, two UAVs belong to one agent, and each action controls the movement of both UAVs at the same time, for example, a t = {forward, right}, which means the first UAV moves forward (up) while the second UAV moves right. When a user is within the coverage of the UAV, the UAV moves in the environment and starts collecting information from the user , however, when enough information is collected, i.e., R m,k ≥ r min , the user will be marked as collected, i.e., P(m, k) = 1, and the UAV can not visit the user again.

[0048] In deep reinforcement learning, the reward is used to evaluate the goodness of the action taken by the agent in the current state. In joint trajectory and data collection optimization, the reward function is designed to depend on both the ratio of user coverage and the reward collected by the UAV along the whole path. The optimization goal is to maximize the user coverage while the UAV flies from the start point to the end point in the shortest time. If one more user is covered at each step, the reward brought by the average throughput will be greater. At the same time, in the single-UAV scenario, the faster the UAV reaches the end point, the greater the end point reward R final will be, and the less the total step penalty brought by R sp will be; for multi-UAV, both UAVs reach the end point (not required to reach at the same time) to get the reward R final , and any extra step taken by either UAV will result in more total R sp . When the constraints are not met, a series of penalties are set, i.e., the penalty for flying out of the specified area is R bp . Therefore, the reward expression is:

[0049]

[0050] Compared with the prior art, the present application has the advantages of:

[0051] This invention provides a path planning method based on multi-UAV assisted data collection. When users are distributed and can move freely throughout the target area, in order to compensate for the insufficient coverage of a single UAV when users are more dispersed, multiple UAVs are used to achieve greater user coverage and optimize a shortest path to the destination, thereby maximizing user coverage. The proposed Dueling-DDQN algorithm based on deep reinforcement learning can accurately estimate the output value of the neural network, make accurate strategies, and plan the actions of each step of the UAV, and is applicable to other different scenarios. Attached Figure Description

[0052] Figure 1 This is a flowchart of the method of the present invention;

[0053] Figure 2 The diagram shows the structure of the neural networks in the DQN (left) and Dueling-DDQN (right) of this invention.

[0054] Figure 3 Trajectory graph for methods that do not use deep reinforcement learning;

[0055] Figure 4 Trajectory graph for using deep reinforcement learning methods;

[0056] Figure 5 To and Figure 4 Compare the waveforms after they converge;

[0057] Figure 6 A comparison chart showing the use of the Dueling-DDQN algorithm and the conventional algorithm;

[0058] Figure 7 Multiple drone trajectory maps;

[0059] Figure 8 This is a comparison chart showing the coverage per step between multiple drones and a single drone. Detailed Implementation

[0060] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0061] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0062] Example 1

[0063] like Figure 1 As shown, a path planning method based on multi-UAV assisted data collection includes the following steps:

[0064] S1: The target is divided into several clusters within the target area. User coordinates are randomly generated within the cluster. There are several users in the cluster, and the users move randomly but will not exceed the area boundary.

[0065] S2: The communication channel between the UAV and the user is dominated by the time-limited link, and multiple UAVs are used to dynamically plan without a model;

[0066] S3: The UAV trajectory is optimized using the Dueling-DDQN algorithm to maximize user coverage.

[0067] In the above scheme, when the user distribution is scattered and can move freely throughout the target area, in order to compensate for the insufficient coverage of a single UAV in the case of more scattered users, multiple UAVs are used to achieve more user coverage and can optimize a shortest path to the end point, thereby maximizing user coverage; The proposed Dueling-DDQN algorithm based on deep reinforcement learning can accurately estimate the output value of the neural network and make accurate strategies to plan the action of each step of the UAV, which is suitable for other different scenarios.

[0068] Preferably, in step S1, in the target area, the users are divided into M clusters, each cluster is equivalent to a circle with a radius R, and the coordinates of the users are randomly generated within these circles, there are K users in each cluster, and the position of the kth user in the mth cluster at time step t is Meanwhile, the user moves randomly at a speed lower than the maximum speed v, but will not exceed the boundary of the target area, i.e. and

[0069] In the above scheme, the flight height of the UAV is H, and a single antenna is used to access all clusters to maximize user coverage, and the three-dimensional coordinates of the UAV at time step t are defined as The starting point and the end point of the two UAVs are the same, and the maximum coverage range is determined by the flight height H of the UAV and the antenna emission angle θ, i.e. R max = H tan(θ), and the UAV can only fly in a specified area, i.e. 0 ≤ X(t) ≤ X max and 0 ≤ Y(t) ≤ Y max , where X max and Y max are the length and width of the area.

[0070] Preferably, in step S2, the communication channel between the UAV and the user is dominated by the line-of-sight link, and the distance from the kth user in the mth cluster to the first UAV at time step t is: At time step t, the channel between the first UAV and the kth user in the mth cluster follows the free space path loss model, which is represented as where β0 represents the power gain of the channel at a reference distance d = 1 m.

[0071] Preferably, in step S3, when the user satisfies the distance constraint and is within the coverage of the UAV, the achieved throughput of the kth user in the mth cluster to the UAV at time t is defined as follows: If the user is within the overlapping coverage of multiple UAVs at the same time, the throughput of the user at time t is the sum of the throughput produced by communicating with two UAVs respectively, where B and a 2 are the bandwidth and noise power respectively, the total throughput of the kth user in the mth cluster to the UAV and at time step T is:

[0072] Preferably, in the multi-UAV data collection system, since both UAVs are subject to the end position, UAV coverage, throughput threshold, step penalty and boundary constraint, the user coverage maximization is achieved by optimizing the trajectories of the two UAVs, and the target problem is as follows

[0073]

[0074] s.t R final1 = X target ,

[0075] d m,k ≤ d cons ,

[0076] R m,k ≥ r min ,

[0077] P(m, k) = {0, 1},

[0078] 0 ≤ X(t) ≤ X max ,

[0079] 0 ≤ Y(t) ≤ Y max ,

[0080] The distance constraint d cons represents the straight-line distance between the served user and the UAV, X now , X now1 , X now2 and X target represent the current position of the UAV in the single-UAV data collection system, the current positions of the two UAVs in the multi-UAV data collection system and the key position respectively; if the UAV touches the boundary, it will be subject to a boundary penalty R bp = -100, and R sp = -1000 is defined as the step penalty, the UAV will receive a negative reward for each additional step, and the UAV can only fly in the specified area, i.e. 0 ≤ X(t) ≤ X max and 0 ≤ Y(t) ≤ Y maxwhere X max and Y max are the length and width of the area, X(t), Y(t) represent the horizontal and vertical coordinates of the current position of the UAV respectively.

[0081] In the above scheme, when there is only one UAV in the system, if the UAV reaches the end point, it will directly obtain the end point reward R final1 , and in the multi-UAV data collection system, only when both UAVs reach the end point can the end point reward R final2 be obtained, that is

[0082]

[0083]

[0084] When there is only one UAV in the system, the UAV is subject to the end point position, the UAV coverage range, the throughput threshold, the step penalty and the boundary constraint, and by optimizing the UAV trajectory, the user coverage maximization is realized, and the following target problem can be obtained, that is

[0085]

[0086] s.tR final1 X final =X target ,

[0087] d m,k ≤d cons ,

[0088] R m,k ≥r min ,

[0089] P(m,k)={0,1},

[0090] 0≤X(t)≤X max ,

[0091] 0≤Y(t)≤Y max .

[0092] In the above scheme, the purpose is to maximize the coverage of users to provide communication services for them, while the UAV needs to take off from the starting point to the end point in the shortest time. Therefore, we define P(m,k)={0,1}, when the total throughput R m,k of the kth user in the mth cluster is greater than the threshold r min , it means that this user has contacted the UAV and this round of task no longer communicates with the UAV, and is marked as P(m,k)=1, otherwise P(m,k)=0.

[0093] Preferably, the Dueling-DDQN algorithm in step S3, each episode of the UAV starts from the starting point and ends at the destination;

[0094] In the training phase, the starting point and the end point of the UAV are initialized before each episode, and the positions of the M*K users are randomly initialized; at each time step t, the UAV calculates the state information s t The output action a(t) is the flight direction of the UAV. At this time, if the user is within the coverage range of the UAV, the agent will calculate the throughput of communication with each user and accumulate it until R m,k ≥r min If the next position of the UAV exceeds the specified area, the action is cancelled; according to the action, the corresponding reward r t and the state information s t+1 at the next time are obtained, and are stored in the experience buffer buffer, and N groups of experiences are randomly sampled from the experience buffer at the end of each time to update the network parameters.

[0095] Preferably, the Dueling-DDQN algorithm is a model-free reinforcement learning algorithm for iteratively solving Bellman equations, and the state-action value function is: Where represents the probability of the agent moving from state s to state s' after taking action a, and π(·) represents the selection strategy of the agent.

[0096] Preferably, the Dueling-DDQN algorithm is provided with a target network Q2(s',a - ; θ max ) with parameters θ - and an estimation network Q1(s',a; θ) with parameters θ, the target network is a copy of the estimation network, and the parameter θ - of the target network is updated at a slower frequency than the estimation network.

[0097] At the same time, the Dueling-DDQN algorithm also sets an experience buffer, and the current state-action-reward-next state is stored in the experience buffer, and is randomly accessed later for weight update.

[0098] Embodiment 2

[0099] As Figure 2As shown, the Dueling-DDQN algorithm based on reinforcement learning is proposed herein to calculate the optimal trajectory to maximize user coverage while reaching the destination in the shortest time. Whether it is a single UAV or multiple UAVs, a single agent is used to learn the mapping from state space to action space by constantly interacting with the environment, and learning according to the feedback information of the environment. The UAV will observe the current state s(t) from the environment every step, input the state s(t) into the deep neural network to obtain the corresponding action a(t), interact with the environment through the action a(t), and the agent will get the current reward r(t) and the new state s(t+1) from the environment. Then the above process is stored in the experience buffer (s(t), a(t), r(t), s(t+1)), and the deep neural network is trained.

[0100] As Figure 2 , in the neural network, V π (s) and A π (s, a) are between the output layer and the last hidden layer, and the dimensions of V π (s) and A π (s, a) are the same as the output layer. Compared with DQN, Dueling-DDQN has been greatly improved, which not only reduces overestimation, but also speeds up convergence.

[0101] When there is only one UAV in the system to perform the task, the UAV is an agent that interacts with the environment to find the peak of the reward; in the case of multiple UAVs, the two UAVs belong to the same agent.

[0102] The position of the UAV is defined as the state space, i.e. S = {x, y, H}, and S = {x1, y1, H1, x2, y2, H2} under the multi-UAV data collection system. At time step t, the state of the UAV in the above scenarios is defined as s t = {x t , y t , H t} and

[0103] When there is only one UAV in the system to perform the task, at time step t, the UAV in state s t can choose an action a t belonging to the action space A according to the policy, which is divided into a grid

[0104] A = {left, right, forward, backward}

[0105] In the multi-UAV data collection system, the two UAVs belong to the same agent, and each action controls the movement of the two UAVs at the same time, for example a t={forward,right} means the first drone moves forward (up), while the second drone moves to the right. When the user is within the drone's coverage area, the drone moves through the environment and begins to move away from the user. Information is collected in the middle, however, when enough information is collected, i.e., R m,k ≥r min The user will be marked as collected, i.e., P(m,k)=1, and the drone may no longer visit the user.

[0106] In deep reinforcement learning, rewards are used to evaluate the quality of an agent's actions in the current state. In joint trajectory and data collection optimization, the designed reward function depends on both the user coverage ratio and the rewards collected by the drone along the entire path. The optimization objective is to maximize user coverage while minimizing the time it takes for the drone to fly from the starting point to the destination. Covering one more user per step increases the average throughput reward. Furthermore, in a single-drone scenario, reaching the destination faster not only yields a larger endpoint reward Rfinal but also reduces the total step penalty Rsp. For multiple drones, both drones reaching the destination (not necessarily simultaneously) receive the reward Rfinal; any drone taking an extra step increases the total Rsp. When constraints are not met, a series of penalties are set, with Rbp as the penalty for a drone flying out of the designated area. Therefore, the reward expression is:

[0107]

[0108] Example 3

[0109] like Figures 3-8 As shown, this patent will use the coverage rate per step of the drone as the benchmark. To measure performance, Users c This indicates the number of users covered in each scene, and Steps indicates the total number of steps taken by the drone in each scene.

[0110] Due to the step penalty, Steps will be large at the beginning of training, but will gradually decrease until the endpoint is reached with the fewest steps. At the start of training, because Steps are large, the drone's flight path is long, and the number of users covered per screen is limited. c It will also be relatively large (Users) c The maximum number of users is 50, while the maximum total number of steps for a drone is much greater than 50, resulting in a relatively small coverage rate C per step. As training progresses, the drone will balance its trajectory and the users' positions, aiming to reach the destination with fewer flight steps while simultaneously optimizing the target—average throughput. The larger the value, the more efficient the optimization process becomes, until convergence is achieved, at which point a shortest path to the destination is found, maximizing user coverage. represents the number of users that the current UAV has covered, when the UAV reaches the end point is equivalent to Users c .

[0111] This paper takes 50 users as an example, that is, 10 users are randomly generated in each cluster, a total of 5 clusters. At the same time, X max = 1000 and Y max = 1000 are the length and width of the target area. Whether it is a single UAV or a multi-UAV scenario, the starting point of each UAV is set to (0, 0, 200), and the end point is (1000, 10000, 200), and the moving distance of each UAV is 40.

[0112] When there is only one UAV in the data collection system to perform the task, we first compare it with the method without using deep reinforcement learning to highlight the advantages of reinforcement learning, Figure 3 is the trajectory diagram of the method without using deep reinforcement learning, Figure 4 is the trajectory diagram of the method using deep reinforcement learning. Figure 4 and Figure 5 The experimental results show that when using the method of deep reinforcement learning, after the waveform converges, the UAV can cover more users with the least number of steps, and the coverage rate of each step of the UAV is larger, that is, larger, so the effect is better. Secondly, this paper compares the Dueling-DDQN algorithm with the traditional DQN algorithm, as shown in Figure 6 , when using the Dueling-DDQN algorithm, the coverage rate of each step of the UAV rises faster and stabilizes faster, that is, the UAV can find the shortest path to the end point faster, and better balance the UAV trajectory and user location, so that the number of covered users is maximized. Therefore, we conclude that the Dueling-DDQN algorithm has better performance and faster convergence.

[0113] Finally, under the multi-UAV data collection system, this paper compares Figure 6 with the single-UAV scenario also using the Dueling-DDQN algorithm, after training convergence, the coverage rate of each step of the UAV is close to 1.0, that is, Users c = Steps = 50, since the UAV needs at least 50 steps to reach the end point, therefore, compared with a single UAV, a multi-UAV can not only optimize a path to reach the end point in the shortest time, but also can cover more users, even can achieve full user coverage, the performance is improved obviously.

[0114] Obviously, the above embodiments of the present application are merely exemplary but not intended to limit the embodiments of the present application. Based on the above description, any other variations or changes can be made by those skilled in the art without departing from the spirit and principles of the present application. It is not necessary to list all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall fall within the scope of the claims of the present application.

Claims

1. A path planning method based on multi-UAV assisted data collection, characterized in that, Includes the following steps: S1: Within the target area, users are divided into several clusters. User coordinates are randomly generated within the cluster. There are several users in each cluster, and users move randomly but will not exceed the area boundary. S2: The communication channel between the drone and the user is dominated by time-limited links, and dynamic planning is performed using multiple drones without a model. S3: Use the Dueling-DDQN algorithm to optimize drone trajectories to maximize user coverage; In step S1, within the target area, users are divided into M clusters, each cluster corresponding to a circle with radius R. The coordinates of the users are randomly generated within these circles. Each cluster contains K users, and at time step t, the position of the k-th user in the m-th cluster is... Meanwhile, the user moves randomly at a speed lower than the maximum speed v, but will not exceed the boundary of the target area, i.e. and In step S2, the communication channel between the drone and the user is dominated by a time-limited link. At time step t, the distance from the k-th user in the m-th cluster to the first drone is: At time step t, the three-dimensional coordinates of the UAV are defined as follows: H represents the drone's flight altitude. The channel between the first drone and the k-th user in the m-th cluster follows a free-space path loss model, denoted as... Where β0 represents the power gain of the channel at a reference distance d = 1m; In a multi-drone data collection system, both drones are subject to constraints such as destination location, drone coverage area, throughput threshold, step penalty, and boundary constraints. By optimizing the trajectories of the two drones, the goal is to maximize user coverage, resulting in the following objective problem: s.tR final1 =X target , d m,k ≤d cons , R m,k ≥r min , P(m,k)={0,1}, 0≤X(t)≤X max , 0≤Y(t)≤Y max , Distance constraint d cons Indicates the user being served The straight-line distance to the drone, X now X now1 X now2 and X target These represent the current position of the drone in a single-drone data collection system, the current positions of two drones in a multi-drone data collection system, and the endpoint position, respectively. If a drone touches the boundary, it will be subject to a boundary penalty R. bp = -100, and define R at the same time sp =-1000 is the step penalty; the drone will receive a negative reward for each extra step it takes. The drone can only fly within a designated area, i.e., 0≤X(t)≤X. max and 0≤Y(t)≤Y max , where X max and Y max Let X(t) represent the length and width of the region, and Y(t) represent the x-coordinate and y-coordinate of the drone's current position, respectively. In a multi-drone data collection system, the finish line reward R is only obtained when both drones reach the finish line. final2 R m,k This represents the total throughput to the drone from the k-th user in cluster m at time step T.

2. The path planning method based on multi-UAV assisted data collection according to claim 1, characterized in that, In step S3, when a user meets the distance constraint and is within the drone's coverage area, the throughput from the k-th user in the m-th cluster to the drone at time t is defined as follows: If a user is simultaneously within the overlapping coverage area of ​​multiple drones, then the throughput of this user at time t is the sum of the throughput generated by communicating with two drones separately, where B and α 2 These represent bandwidth and noise power, respectively. The total throughput to the drone from the k-th user in cluster m at time step T is:

3. The path planning method based on multi-UAV assisted data collection according to claim 1, characterized in that, In step S3, the Dueling-DDQN algorithm starts from the starting point and ends at the destination for each scene of the drone.

4. The path planning method based on multi-UAV assisted data collection according to claim 3, characterized in that, During the training phase, before each scene begins, the starting and ending positions of the drone are initialized, and the positions of M*K users are randomly initialized; at each time step t, the drone adjusts its position based on the observed state information s. t The output action a(t) represents the drone's flight direction. If the user is within the drone's coverage area, the agent will calculate the throughput of communication with each user separately, accumulating this process up to one step, until R... m,k ≥r min If the drone's next location exceeds the designated area, the flight maneuver is cancelled; a corresponding reward is obtained based on the maneuver. t and the state information s of the next moment t+1 , will [s t ,a t ,r t ,s t+1 The network parameters are updated by randomly sampling N sets of experience data from the experience buffer at the end of each time step.

5. The path planning method based on multi-UAV assisted data collection according to claim 4, characterized in that, The Dueling-DDQN algorithm is a model-free reinforcement learning algorithm for iteratively solving the Bellman equation, and its state-action value function is: in Let π(·) represent the probability that the agent will transition to state s' after taking action a in state s, and let π(·) represent the agent's policy selection.

6. The path planning method based on multi-UAV assisted data collection according to claim 5, characterized in that, The Dueling-DDQN algorithm is configured with a parameter θ. - The target network Q2(s',a) max ;θ - ) and the estimation network Q1(s',a) with parameter θ; θ), where the target network is a copy of the estimation network, and the parameters θ of the target network are... - The update frequency is slower than that of the estimated network.

7. The path planning method based on multi-UAV assisted data collection according to claim 6, characterized in that, The Dueling-DDQN algorithm also sets up an experience buffer, which is the current state-action-reward-next state. Stored in an experience buffer for later random access to update weights.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster trajectory optimization and task unloading method based on digital twinning

    CN114125708A