Unmanned aerial vehicle path planning method

By designing the distributed architecture of the multi-drone cooperative data acquisition system and the multi-agent DQN path planning, the robustness and insufficient emergency response of the centralized system are solved, and low-complexity and high-rootment data acquisition without repeated acquisition and collision in drone cooperation is achieved.

CN120335472APending Publication Date: 2025-07-18BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474686.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

Smart Images

  • Figure CN120335472A_ABST
    Figure CN120335472A_ABST
Patent Text Reader

Abstract

An unmanned aerial vehicle path planning method belongs to the field of unmanned aerial vehicles, and comprises the following steps: designing a distributed architecture of an MUC-DCS, taking each unmanned aerial vehicle as an independent intelligent agent, communicating with a neighbor unmanned aerial vehicle to perform information interaction, and performing flight path decision based on the interacted information to realize non-repetitive and non-collision multi-unmanned aerial vehicle cooperative data acquisition. In a system with multiple unmanned aerial vehicles and users, the distributed architecture of the MUC-DCS has relatively low calculation complexity, excellent expandability and relatively high robustness. The DPPS-MDQN provided by the invention comprises a network structure, a state space, an action space and a reward function of the unmanned aerial vehicle and an emergency response strategy aiming at the failure condition of the unmanned aerial vehicle, fully considers the complexity and uncertainty of an unmanned aerial vehicle data acquisition system, and is suitable for the normal working and failure conditions of the unmanned aerial vehicle; the data acquisition time is shortened; and the emergency response capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicles, and particularly relates to a method for path planning of an unmanned aerial vehicle. Background Art

[0002] In an Internet of Things (IoT) system, unmanned aerial vehicles are considered for collecting data of IoT devices. The paper (J. Zhong, Y. Hu, Y. Li, Y. Xu, R. Gao and J. Wang, "UAV Data Collection With Deep Reinforcement Learning for Grant-Free IoT," 2024 IEEE Wireless Communications and Networking Conference (WCNC), Dubai, United Arab Emirates, 2024, pp. 1-6, doi: 10.1109 / WCNC57260.2024.10571061.) discloses a UAV-assisted data collection algorithm based on a deep Q network (DQN). This UAV-assisted data collection algorithm maximizes the average throughput of data transmission by optimizing the flight path of the UAV on the premise that the device data is successfully collected by the UAV. In this paper, the network architecture of the UAV-assisted data collection algorithm mainly includes an environment module and an agent. The agent consists of an experience replay sub-module, a Q network, and a target Q network. Specifically, after the Q network uses the ε-greedy strategy for action selection, the agent obtains the state and reward from the environment module and stores them together in the experience replay sub-module. Subsequently, samples are randomly selected from the experience replay sub-module and input into the Q network to generate Q values. Finally, the Q values and the target Q values obtained from the target Q network are input into the optimizer to update the weights of the Q network and the target Q network.

[0003] The UAV-assisted data collection algorithm based on a deep Q network (DQN) disclosed in this paper mainly has the following technical problems:

[0004] (1) First, this paper uses a centralized UAV data collection system, which has poor robustness, is not easy to expand, and the computational complexity increases exponentially with the increase in the number of users.

[0005] (2) Second, in the centralized UAV data collection system, there is no cooperation and information interaction among UAVs, and problems such as UAV collisions and duplicate collection of user data are likely to occur during the data collection process.

[0006] (3)Finally, the paper only considered the case where the UAVs were working properly. However, due to emergencies such as energy depletion and accidental damage, the UAVs may fail and cannot work properly. Considering only the case where the UAVs are working properly will lead to insufficient emergency response capabilities of the system. SUMMARY OF THE INVENTION

[0007] To solve the problems existing in the prior art, the present invention provides a UAV path planning method. The present invention considers a multi-UAV data acquisition system. However, insufficient cooperation among multiple UAVs will result in a longer data acquisition time. Moreover, in the case of failure where the UAVs cannot work properly due to emergencies, insufficient emergency response capabilities will also extend the data acquisition time. To solve these problems, the present invention proposes a distributed architecture and a distributed path planning method for a multi-UAV data acquisition system, which can effectively reduce the data acquisition time and improve the robustness of the system.

[0008] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0009] A UAV path planning method provided by the present invention includes the following steps:

[0010] Step S1: Design a distributed architecture for a multi-UAV collaborative data acquisition system;

[0011] In the distributed architecture of the multi-UAV collaborative data acquisition system, each UAV serves as an independent agent, communicates with neighboring UAVs for information interaction, and makes flight path decisions based on the interacted information to achieve non-redundant and collision-free multi-UAV collaborative data acquisition;

[0012] Step S2: Design a distributed path planning scheme based on multi-agent DQN;

[0013] The distributed path planning scheme based on multi-agent DQN includes an algorithm applicable to the case where the UAVs are working properly and an algorithm applicable to the case where the UAVs fail;

[0014] The algorithm applicable to the case where the UAVs are working properly includes an environment module and N UAVs. The N UAVs have the same network structure, state space, action space, and reward function; the network architecture of each UAV includes an information interaction sub-module, an experience replay sub-module, a loss function sub-module, a Q network, and a target Q network; the information interaction sub-module is used to generate the current time slot state s n (k) and the next time slot state s n (k + 1); the Q network is used to output an estimated Q value according to the current time slot state s n (k) and select the corresponding action a n (k); the experience replay sub-module is used to store the tuple (s n(k), a n (k), r n (k), s n (k + 1)), r n (k) represents the reward; the target Q-network is used to output the maximum Q-value according to a partial tuple randomly sampled from the experience replay sub-module; the loss function sub-module is used to minimize the loss function based on the estimated Q-value and the maximum Q-value to update the Q-network weights;

[0015] The algorithm applicable to the UAV failure situation includes an environment module and N + N d UAVs, and N d UAVs are used as supplementary UAVs. Each UAV has the same network structure, state space, action space, and reward function as the UAVs in the algorithm applicable to the normal operation situation of UAVs; the supplementary UAVs cannot fly and perform data collection tasks when in the dormant state, and when in the active state, they are the same as the normally operating UAVs and can cooperate with the normally operating UAVs to complete data collection tasks.

[0016] Further, in the distributed architecture of the multi-UAV collaborative data collection system, N UAVs complete the data collection tasks of M stationary ground users in the target area; the total time for completing the data collection tasks is divided into K time slots, and the length of each time slot k ∈ {1,..., K} is τ(k); the UAVs fly over the target area, and at time slot k, the position of UAV n ∈ {1,..., N} is q n (k) = (x n (k), y n (k), z n (k)), x n (k),, y n (k), z n (k) respectively represent the three-dimensional coordinates of the UAV.

[0017] Further, the specific implementation process of the distributed architecture of the multi-UAV collaborative data collection system is as follows:

[0018] S1.1: Multiple UAVs take off from different starting positions and start executing data collection tasks simultaneously; each UAV completes part of the data collection tasks, and all UAVs cooperate to complete the data collection tasks of all users;

[0019] S1.2: At time slot k, each UAV communicates with its neighbor UAVs for information interaction;

[0020] S1.3: At time slot k, each UAV autonomously decides the flight path for executing the data collection task based on its own observation information and the information interacted with the neighbor UAVs;

[0021] S1.4: In time slot k, if there are no users within the ground coverage of the UAV or the user data has been collected, the UAV will move according to the determined path; conversely, if the users within the ground coverage of the UAV have not been collected, the UAV will first stay in place, collect the users' data, and after the data collection is completed, the UAV will then move according to the determined path.

[0022] S1.5: When the UAV moves to the target location according to the determined path, this time slot k ends and the next time slot k + 1 begins.

[0023] S1.6: Continuously repeat S1.2 - S1.5 until multiple UAVs cooperate to complete all users' data collection tasks.

[0024] Furthermore, in step S1.6, the data collection task completion time T depends on one or more UAVs that take the longest time to complete a partial data collection task, that is, the data collection task completion time T is equal to the sum of the time slot lengths τ(k) of K time slots.

[0025] Furthermore, the specific implementation process of the algorithm applicable to the normal operation of the UAV is as follows:

[0026] Define the state, action, and reward function of the UAV. After each module is initialized, the information interaction sub-module generates the state and inputs it into the Q-network. The Q-network outputs the estimated Q value, selects an action according to the ε-greedy policy, the UAV moves to the corresponding position according to the selected action, calculates the reward and the length of each time slot, and the information interaction sub-module generates the state of the next time slot; store the tuple into the experience replay sub-module; randomly extract some tuples from the experience replay sub-module and input them into the target Q-network. The target Q-network outputs the maximum Q value. Input the estimated Q value output by the Q-network and the maximum Q value output by the target Q-network into the loss function sub-module, and minimize the loss function based on the gradient descent method to update the Q-network weights. Through the neural network backpropagation, copy the weights of the Q-network to the target Q-network; repeat the above steps until all data collection tasks are completed.

[0027] Furthermore, in time slot k, the state of UAV n is defined as s n (k) = {q n (k), F n (k), Z n (k), ζ n (k)}, where q n (k) represents the position of UAV n; F n (k) represents the data collection state after the information interaction between UAV n and its neighboring UAVs, ensuring that the users' data is not collected repeatedly; Z n(k) represents the target position of the neighboring UAV in the next time slot k + 1 to avoid UAV collisions; ζ n (k) represents the pheromone of UAV n.

[0028] Furthermore, in time slot k, the action of UAV n is defined as a n (k) = {1, 2, 3, 4, 5, 6}, where {1, 2, 3, 4, 5, 6} respectively represent that UAV n flies forward, backward, left, right, up, and down from the current position.

[0029] Furthermore, in time slot k, the reward function r n (k) is defined as:

[0030]

[0031] where r tanh (ζ n (k)) represents the reward shaping function of the pheromone ζ n (k) of UAV n; K re = K max -K represents the remaining time slot reward for encouraging UAVs to complete the data collection task as soon as possible, that is, the maximum time slot number K max minus the time slot number K required to complete the task; β cov represents the pheromone reward factor.

[0032] Furthermore, the specific implementation process of the algorithm applicable to the UAV failure situation is as follows:

[0033] After each module is initialized, when a UAV failure occurs, it is judged whether additional UAVs need to be activated based on the emergency response strategy; if not, the failed UAV is removed from the system, and the remaining normally operating UAVs execute the algorithm process applicable to the normal operating situation of UAVs until all data collection tasks are completed; if so, in the time slot when the UAV failure occurs, additional UAVs are activated, the additional UAVs fly to the initial position calculated by the emergency response strategy and replace the failed UAV, and the remaining normally operating UAVs execute the algorithm process applicable to the normal operating situation of UAVs; in the subsequent time slots, the remaining normally operating UAVs and the additional UAVs repeat the algorithm process applicable to the normal operating situation of UAVs until all data collection tasks are completed.

[0034] Furthermore, the emergency response strategy includes:

[0035] 1) Judge whether additional UAVs need to be activated;

[0036] Define the task completion degree χ(k) as:

[0037]

[0038] Among them, and respectively represent the average pheromone values at time slot k and task completion time slot K; the task completion degree when the UAV fails is denoted as χ(k lose ), k lose represents the corresponding UAV failure time slot;

[0039] By comparing the task completion degree χ(k lose ) and the weighted time slot ratio ρ·k lose / K to determine whether it is necessary to activate the supplementary UAV at the failure time slot k lose , ρ∈[0,1] represents the supplementary weight factor; if the task completion degree χ(k lose ) is greater than the weighted time slot ratio ρ·k lose / K, it indicates that the system task completion degree is relatively high, and there is no need to activate the supplementary UAV. That is, on the premise of ensuring the timeliness of the task, the flight paths of the remaining normal working UAVs are adjusted to complete the data collection task; otherwise, the task completion degree is relatively low, and it is necessary to activate the supplementary UAV to replace the failed UAV and cooperate with the remaining normal working UAVs to complete the data collection task on time.

[0040] 2) Calculate the initial position of the supplementary UAV after activation;

[0041] The center point (x center , y center ) of the uncollected user cluster in the target area is obtained through the K-means clustering algorithm; when the system needs to be supplemented, the supplementary UAV is activated, and within one time slot, the supplementary UAV flies to the initial position (x center , y center , H), and then cooperates with the remaining normal working UAVs to complete the data collection task.

[0042] The beneficial effects of the present invention are:

[0043] (1) In the distributed architecture of the MUC-DCS designed by the present invention, each UAV acts as an independent intelligent agent, can communicate with neighboring UAVs for information interaction, and make flight path decisions based on the information after interaction, realizing multi-UAV cooperative data collection without repetition and collision. In addition, in a large-scale system with a large number of UAVs and users, the designed distributed architecture of the MUC-DCS has low computational complexity, excellent scalability, and strong robustness.

[0044] (2) The present invention proposes a DPPS-MDQN for optimizing the flight paths of multiple UAVs. The DPPS-MDQN designs the network structure, state space, action space, reward function, and emergency response strategy for UAV failures of UAVs, fully considering the complexity and uncertainty of the UAV data acquisition system, applicable to the normal operation and failure of UAVs, reducing the data acquisition time and improving the emergency response ability, and having great application prospects in real environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a distributed architecture diagram of MUC-DCS. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following further describes the present invention in detail with reference to the drawings.

[0047] A UAV path planning method provided by the present invention has the following specific implementation process:

[0048] I. Distributed architecture of a multiple UAVs collaboration data collection system (MUC-DCS);

[0049] The present invention realizes the data collection task through the collaboration of multiple UAVs by designing a distributed architecture of a multiple UAVs collaboration data collection system (MUC-DCS). Through the information interaction between multiple UAVs, duplicate data collection and UAV collisions can be effectively avoided.

[0050] As Figure 1 shown, the present invention considers a MUC-DCS, where N UAVs complete the data collection tasks of M stationary ground users in the target area. The total time for completing the data collection task is divided into K time slots, and the length of each time slot k∈{1,...,K} is τ(k). The UAVs fly over the target area. At time slot k, the position of UAV n∈{1,...,N} is q n (k)=(x n (k),y n (k),z n (k)), where x n (k), y n (k), z n(k) represents the three-dimensional coordinates of the UAV. When the user is within the ground coverage range of the UAV, the communication and data collection between the UAV and the user can be carried out. In addition, each UAV can only communicate with other UAVs within the communication range, and other UAVs within the communication range are called neighbor UAVs.

[0051] In order to fully realize the collaborative collection of multiple UAVs, a distributed architecture of MUC-DCS is designed. In the distributed architecture of MUC-DCS, there is no central controller to uniformly schedule and assign tasks to multiple UAVs, and each UAV operates as an independent agent. Multiple UAVs collaborate to complete the data collection tasks of all users and achieve global optimization of the system. The specific implementation process of the distributed architecture of MUC-DCS is as follows:

[0052] Step S1.1: Multiple UAVs take off from different starting positions and start executing data collection tasks simultaneously. Each UAV completes part of the data collection tasks, and all UAVs collaborate to complete the data collection tasks of all users.

[0053] Step S1.2: In time slot k, each UAV communicates with its neighbor UAVs, that is, broadcasts its own observation information and receives the information of neighbor UAVs for information interaction. These information include the data collection status of users and the positions of UAVs to avoid duplicate data collection and UAV collisions.

[0054] Step S1.3: In time slot k, each UAV autonomously decides the flight path for executing the data collection task according to its own observation information and the information interacted with neighbor UAVs. The decision-making times of different UAVs may be different, but no matter which UAV makes a decision, its neighbor UAVs will provide it with the latest information.

[0055] Step S1.4: In time slot k, if there is no user within the ground coverage range of the UAV or the user data has been collected, the UAV will move according to the decision-making path. On the contrary, if the user within the ground coverage range of the UAV has not been collected, the UAV will first keep its position unchanged, collect the user's data, and after the data collection is completed, the UAV will move according to the decision-making path.

[0056] Step S1.5: When the UAV moves to the target position according to the decision-making path, this time slot k ends and enters the next time slot k + 1.

[0057] Step S1.6: Continuously repeat Step S1.2 to Step S1.5 until multiple UAVs collaborate to complete the data collection tasks of all users. The data collection task completion time T depends on one or more UAVs that take the longest time to complete part of the data collection tasks, that is, the data collection task completion time T is equal to the sum of the time slot lengths τ(k) of K time slots.

[0058] II. Distributed Path Planning Scheme Based on Multi-Agent DQN (DPPS-MDQN);

[0059] The present invention proposes a distributed path planning scheme based on multi-agent deep Q network (DPPS-MDQN), which is applicable to the normal operation and failure of unmanned aerial vehicles, reduces the completion time of data collection tasks, and improves the emergency response ability of the system.

[0060] To minimize the completion time T of the data collection task, it is necessary to optimize the positions {q1(k),..., q n (k),..., q N (k)} of multiple unmanned aerial vehicles in each time slot k. The present invention proposes DPPS-MDQN to solve this path optimization problem, which specifically includes Algorithm 1 applicable to the normal operation of unmanned aerial vehicles and Algorithm 2 applicable to the failure of unmanned aerial vehicles.

[0061] On the one hand, in the case of the normal operation of unmanned aerial vehicles, Algorithm 1 applicable to the normal operation of unmanned aerial vehicles is proposed to optimize the flight paths of multiple unmanned aerial vehicles and minimize the completion time T of the data collection task. The specific implementation process is as follows:

[0062] Step S2.1.1: Algorithm 1 includes an environment module and N unmanned aerial vehicles. The N unmanned aerial vehicles have consistent network structures, state spaces, action spaces, and reward functions. The network architecture of each unmanned aerial vehicle includes an information interaction sub-module, an experience replay sub-module, a loss function sub-module, a Q network, and a target Q network.

[0063] Step S2.1.2: In time slot k, according to its own observation information and the information interacted with neighboring unmanned aerial vehicles, the state of unmanned aerial vehicle n is defined as s n (k) = {q n (k), F n (k), Z n (k), ζ n (k)}, where q n (k) represents the position of unmanned aerial vehicle n; F n (k) represents the data collection state after the information interaction between unmanned aerial vehicle n and neighboring unmanned aerial vehicles, ensuring that the user's data is not collected repeatedly; Z n (k) represents the target position of neighboring unmanned aerial vehicles in the next time slot k + 1, avoiding collisions between unmanned aerial vehicles; ζn (k) represents the pheromone of UAV n.

[0064] Step S2.1.3: At time slot k, the action of UAV n is defined as a n (k) = {1, 2, 3, 4, 5, 6}, where {1, 2, 3, 4, 5, 6} respectively represent that UAV n flies forward, backward, left, right, up, and down from the current position.

[0065] Step S2.1.4: In time slot k, the reward function r n (k) has the following mathematical expression:

[0066]

[0067] where r tanh (ζ n (k)) represents the reward shaping function of the pheromone ζ n (k) of UAV n, ensuring the continuous and smooth change of the reward value; K re = K max - K represents the remaining time slot reward for encouraging UAVs to complete the data collection task as soon as possible, that is, the maximum time slot number K max minus the time slot number K required to complete the task, where the maximum time slot number K max is a relatively large constant; β cov represents the pheromone reward factor.

[0068] Step S2.1.5: The specific implementation process of Algorithm 1 for optimizing the flight paths of multiple UAVs is as follows:

[0069] 1) Input: UAV and user positions, time slot k = 0.

[0070] 2) Output: The data collection task completion time T and the positions of multiple UAVs in each time slot k {q1(k),..., q n (k),..., q N (k)}.

[0071] 3) Initialize the information interaction sub-module, experience replay sub-module, loss function sub-module, Q network, and target Q network.

[0072] 4) Based on its own observation information from the environment module and the information interacted with neighboring UAVs, the information interaction sub-module generates the state s n (k), and inputs it into the Q network.

[0073] 5) The Q network outputs the estimated Q value Q(s n (k), a n (k); θ n ), θ nRepresent the weights of the Q-network and select the action a according to the ε-greedy policy n (k).

[0074] 6) According to the selected action a n (k), the UAV moves to the corresponding position, calculates the reward r n (k) and the length τ(k) of each time slot, and the information interaction sub-module generates the state s n (k + 1) of the next time slot.

[0075] 7) Store the tuple (s n (k), a n (k), r n (k), s n (k + 1)) into the experience replay sub-module.

[0076] 8) Randomly extract some tuples (s n (k'), a n (k'), r n (k'), s n (k' + 1)) from the experience replay sub-module and input them into the target Q-network. The target Q-network outputs the maximum Q value representing the weights of the target Q-network.

[0077] 9) Input the estimated Q value Q(s n (k), a n (k); θ n ) output by the Q-network and the maximum Q value output by the target Q-network into the loss function sub-module, and minimize the loss function based on the gradient descent method to update the Q-network weights.

[0078] 10) Through the backpropagation of the neural network, copy the weights of the Q-network to the target Q-network.

[0079] 11) Repeat steps 4) - 10) for each UAV, and update the time slot k = k + 1 until all data collection tasks are completed.

[0080] On the other hand, due to the complexity and uncertainty of the actual environment, emergencies often occur during the data collection process, resulting in the failure of UAVs and their inability to work properly. In the case of UAV failure, an algorithm 2 applicable to the UAV failure situation is proposed to optimize the flight paths of multiple UAVs and minimize the completion time T of the data collection task. The specific implementation process is as follows:

[0081] Step S2.2.1: In the case of UAV failure, in order to avoid the problem that the data collection task times out or even cannot be completed, deploy N in MUC-DCS dA drone is used as a supplementary drone, and their states are all set to the sleep state. The supplementary drones in the sleep state cannot fly and perform data collection tasks. After the supplementary drones are activated, they can cooperate with the normally working drones to complete data collection tasks, just like the normally working drones.

[0082] Step S2.2.2: To ensure the robustness of the system and the timeliness of tasks, it is necessary to adjust the deployment of multiple drones in a timely manner in the event of drone failure. Therefore, the present invention proposes an emergency response strategy, which mainly includes two key parts: one is to determine whether to activate the supplementary drones according to the task completion degree, and the other is to calculate the initial positions of the activated supplementary drones. The specific implementation process is as follows:

[0083] 1) Determine whether it is necessary to activate the supplementary drones;

[0084] Specifically, the task completion degree χ(k) is defined as:

[0085]

[0086] Where, and respectively represent the average pheromone values at time slot k and the task completion time slot K. The task completion degree at the time of drone failure is denoted as χ(k lose ), k lose represents the corresponding drone failure time slot.

[0087] To determine whether it is necessary to activate the supplementary drones at the failure time slot k lose , it is necessary to compare the task completion degree χ(k lose ) and the weighted time slot ratio ρ·k lose / K, where ρ ∈ [0, 1] represents the supplementary weight factor. If the task completion degree χ(k lose ) is greater than the weighted time slot ratio ρ·k lose / K, it means that the system task completion degree is relatively high, and there is no need to activate the supplementary drones. That is, on the premise of ensuring the timeliness of tasks, the flight paths of the remaining normally working drones can be adjusted in a timely manner to complete the data collection task. Otherwise, the task completion degree is relatively low, and it is necessary to activate the supplementary drones to replace the failed drones and cooperate with the remaining normally working drones to complete the data collection task on time.

[0088] 2) Calculate the initial positions of the activated supplementary drones;

[0089] The center point (x center , y center ) of the uncollected user clusters in the target area can be obtained through the K-means clustering algorithm. When the system needs to be supplemented, the supplementary drones are activated, and within one time slot, the supplementary drones fly to the initial position (xcenter , y center , H), and then cooperate with the remaining normally operating drones to complete the data collection task. The height H of the supplementary drone's initial position can be determined according to the actual system environment and maintained between 0 and the maximum flight height H max in between.

[0090] Step S2.2.3: Algorithm 2 includes an environment module and N + N d drones. Each drone has the same network structure, state space, action space, and reward function as the drones in Algorithm 1.

[0091] Step S2.2.4: Based on Algorithm 1 and the emergency response strategy, the specific implementation process for Algorithm 2 to optimize the flight paths of multiple drones is as follows:

[0092] 1) Input: Drone and user positions, time slot k = 0.

[0093] 2) Output: The data collection task completion time T and the positions of multiple drones in each time slot k {q1(k),..., q n (k),..., q N (k)}.

[0094] 3) Initialize the information interaction sub-module, experience replay sub-module, loss function sub-module, Q-network, and target Q-network.

[0095] 4) When a drone failure occurs, based on the emergency response strategy, determine whether it is necessary to activate the supplementary drone.

[0096] 5) If the judgment result of the emergency response strategy is that it is not necessary to activate the supplementary drone, remove the failed drone from the system. Repeat steps 4) - 10) of step S2.1.5 of Algorithm 1 for the remaining normally operating drones, and update the time slot k = k + 1 until all data collection tasks are completed.

[0097] 6) If the judgment result of the emergency response strategy is that it is necessary to activate the supplementary drone, then in the time slot when the drone failure occurs, activate the supplementary drone. The supplementary drone flies to the initial position calculated by the emergency response strategy, replaces the failed drone, and moreover, the remaining normally operating drones execute steps 4) - 10) of step S2.1.5 of Algorithm 1. In subsequent time slots, repeat steps 4) - 10) of step S2.1.5 of Algorithm 1 for the remaining normally operating drones and the supplementary drone, and update the time slot k = k + 1 until all data collection tasks are completed.

[0098] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for path planning of an unmanned aerial vehicle, characterized in that, It includes the following steps: Step S1: Design the distributed architecture of the multi-UAV collaborative data collection system; In the distributed architecture of the multi-UAV collaborative data collection system, each UAV acts as an independent agent, communicates with neighboring UAVs for information interaction, and makes flight path decisions based on the interacted information to achieve non-redundant and collision-free multi-UAV collaborative data collection; Step S2: Design a distributed path planning scheme based on multi-agent DQN; The distributed path planning scheme based on multi-agent DQN includes an algorithm applicable to the normal operation of UAVs and an algorithm applicable to the failure of UAVs; The algorithm applicable to the normal operation of the drone includes an environment module and N drones. The N drones have the same network structure, state space, action space, and reward function. The network architecture of each drone includes an information interaction sub-module, an experience replay sub-module, a loss function sub-module, a Q-network, and a target Q-network. The information interaction sub-module is used to generate the current time slot state s n (k) and the next time slot state s n (k + 1). The Q-network is used to output an estimated Q value based on the current time slot state s n (k) and select the corresponding action a n (k). The experience replay sub-module is used to store the tuple (s n (k), a n (k), r n (k), s n (k + 1)), where r n (k) represents the reward. The target Q-network is used to output the maximum Q value based on a partial tuple randomly selected from the experience replay sub-module. The loss function sub-module is used to minimize the loss function based on the estimated Q value and the maximum Q value to update the Q-network weights. The algorithm applicable to the UAV failure situation includes an environment module and N+N d UAVs, where N d UAVs are supplementary UAVs. Each UAV has the same network structure, state space, action space, and reward function as the UAVs in the algorithm applicable to the normal operation of UAVs. The supplementary UAVs cannot fly and perform data collection tasks when in the dormant state. When the supplementary UAVs are in the active state, they are the same as the UAVs operating normally and can cooperate with the UAVs operating normally to complete data collection tasks.

2. The method for path planning of an unmanned aerial vehicle according to claim 1, wherein, In the distributed architecture of the multi-UAV collaborative data acquisition system, N UAVs complete the data acquisition tasks of M stationary ground users in the target area. The total time for completing the data acquisition tasks is divided into K time slots, and the length of each time slot k∈{1,...,K} is τ(k). The UAVs fly over the target area. At time slot k, the position of UAV n∈{1,...,N} is q n (k)=(x n (k),y n (k),z n (k)), where x n (k), y n (k), and z n (k) respectively represent the three-dimensional coordinates of the UAV.

3. A method for path planning of an unmanned aerial vehicle according to claim 1, characterized in that, The specific implementation process of the distributed architecture of the multi-UAV collaborative data collection system is as follows: S1.1: Multiple UAVs take off from different starting positions and start executing data collection tasks simultaneously; each UAV completes part of the data collection task, and all UAVs collaborate to complete the data collection tasks of all users; S1.2: In time slot k, each UAV communicates with neighboring UAVs for information interaction; S1.3: In time slot k, each UAV autonomously decides the flight path for executing the data collection task based on its own observation information and the information interacted with neighboring UAVs; S1.4: In time slot k, if there are no users within the ground coverage range of the UAV or the user data has been collected, the UAV will move according to the decision path; otherwise, if the users within the ground coverage range of the UAV have not been collected, the UAV will first stay in place, collect the user data, and after the data collection is completed, the UAV will then move according to the decision path; S1.5: When the UAV moves to the target position according to the decision path, this time slot k ends and enters the next time slot k + 1; S1.6: Continuously repeat S1.2 to S1.5 until multiple UAVs collaborate to complete the data collection tasks of all users.

4. A method for path planning of an unmanned aerial vehicle according to claim 3, characterized in that, In step S1.6, the data collection task completion time T depends on one or more UAVs that take the longest time to complete part of the data collection task, that is, the data collection task completion time T is equal to the sum of the time slot lengths τ(k) of K time slots.

5. A method for path planning of an unmanned aerial vehicle according to claim 1, characterized in that, The specific implementation process of the algorithm applicable to the normal operation of UAVs is as follows: Define the state, action, and reward function of the UAV. After each module is initialized, the information interaction sub-module generates a state and inputs it into the Q network. The Q network outputs the estimated Q value, selects an action according to the ε-greedy strategy, the UAV moves to the corresponding position according to the selected action, calculates the reward and the length of each time slot, and the information interaction sub-module generates the state of the next time slot; Store the tuple into the experience replay sub-module; Randomly extract some tuples from the experience replay sub-module and input them into the target Q network. The target Q network outputs the maximum Q value. Input the estimated Q value output by the Q network and the maximum Q value output by the target Q network into the loss function sub-module, and minimize the loss function based on the gradient descent method to update the Q network weights. Through the backpropagation of the neural network, copy the weights of the Q network to the target Q network; Repeat the above steps until all data collection tasks are completed.

6. A method for path planning of an unmanned aerial vehicle according to claim 5, characterized in that, At time slot k, the state of UAV n is defined as s n (k) = {q n (k), F n (k), Z n (k), ζ n (k)}, where q n (k) represents the position of UAV n; F n (k) represents the data collection state of UAV n after information interaction with neighboring UAVs, ensuring that the user's data is not collected repeatedly; Z n (k) represents the target position of neighboring UAVs in the next time slot k + 1 to avoid collisions between UAVs; ζ n (k) represents the pheromone of UAV n.

7. A method for path planning of an unmanned aerial vehicle according to claim 5, characterized in that At time slot k, the action of UAV n is defined as a n (k) = {1, 2, 3, 4, 5, 6}, where {1, 2, 3, 4, 5, 6} respectively represent that UAV n flies forward, backward, left, right, up, and down from its current position.

8. A method for path planning of an unmanned aerial vehicle according to claim 5, characterized in that In time slot k, the reward function r n (k) of UAV n is defined as: Among them, r tanh (ζ n (k)) represents the reward shaping function of the pheromone ζ n (k); K re = K max - K represents the remaining time slot reward for encouraging the UAV to complete the data collection task as soon as possible, that is, the maximum number of time slots K max minus the number of time slots K required to complete the task; β cov represents the pheromone reward factor.

9. A method for path planning of an unmanned aerial vehicle according to claim 1, characterized in that, The specific implementation process of the algorithm applicable to the UAV failure situation is as follows: After the initialization of each module, when a UAV failure occurs, it is judged whether it is necessary to activate a supplementary UAV based on the emergency response strategy; if not, the failed UAV is removed from the system, and the algorithm process applicable to the normal working situation of the UAV is executed for the remaining normally working UAVs until all data collection tasks are completed; if so, at the time slot when the UAV failure occurs, the supplementary UAV is activated, and the supplementary UAV flies to the initial position calculated by the emergency response strategy and replaces the failed UAV, and the remaining normally working UAVs execute the algorithm process applicable to the normal working situation of the UAV; in the subsequent time slots, the algorithm process applicable to the normal working situation of the UAV is repeatedly executed for the remaining normally working UAVs and the supplementary UAVs until all data collection tasks are completed.

10. A method for path planning of an unmanned aerial vehicle according to claim 9, characterized in that, The emergency response strategy includes: 1) Judging whether it is necessary to activate a supplementary UAV; The task completion degree χ(k) is defined as: Among them, and respectively represent the average pheromone values at time slot k and task completion time slot K; the task completion degree when the UAV fails is denoted as χ(k lose ), k lose represents the corresponding UAV failure time slot; By comparing the task completion degree χ(k lose ) and the weighted time slot ratio ρ·k lose / K to determine whether it is necessary to activate the supplementary drone in the failure time slot k lose , ρ∈[0,1] represents the supplementary weight factor; if the task completion degree χ(k lose ) is greater than the weighted time slot ratio ρ·k lose / K, it indicates that the system has a high task completion degree and there is no need to activate the supplementary drone, that is, on the premise of ensuring the timeliness of the task, adjust the flight path of the remaining normally working drones to complete the data collection task; otherwise, the task completion degree is low, and it is necessary to activate the supplementary drone to replace the failed drone and cooperate with the remaining normally working drones to complete the data collection task on time. 2) Calculating the initial position of the supplementary UAV after activation; The center point (x center , y center ) of the uncollected user clusters in the target area is obtained through the K-means clustering algorithm; when the system needs to be replenished, the replenishment drone is activated, and within one time slot, the replenishment drone flies to the initial position (x center , y center , H), and then cooperates with the remaining normally operating drones to complete the data collection task.