A multi-mission point path planning method for unmanned aerial vehicles

By optimizing the flight trajectory of the CUAV through the HT-D3QN algorithm, the communication security and mission efficiency issues of the CUAV multi-task point inspection in complex urban environments are solved, and efficient path planning is achieved in a dynamic environment.

CN120524837BActive Publication Date: 2025-10-03NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511021183.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-10-03
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

In complex urban environments, the communication security of connected unmanned aerial vehicles (CUAVs) faces the risk of eavesdropping. Existing path planning algorithms are unable to cope with unknown or dynamically changing environments, and it is difficult to balance communication confidentiality and mission efficiency during multi-task point inspections.

Method used

The hierarchical reinforcement learning framework HT-D3QN algorithm, combined with a deep neural network, optimizes the flight trajectory of the CUAV to minimize the total mission completion time and communication interruption time. It decomposes complex problems through high-level task allocation and low-level path planning, and realizes three-dimensional trajectory and cellular association.

Benefits of technology

It improves the task execution efficiency and communication security of CUAV in complex dynamic environments, adapts to the uncertainty of eavesdropper's position, and optimizes the path planning of multi-task point inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524837B_ABST
    Figure CN120524837B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task path planning method for unmanned aerial vehicles (UAVs), specifically comprising the following steps: method initiation; establishing an air-ground system channel model and an uncertainty model for the eavesdropper's location; constructing a weighted sum minimization model for the total CUAV mission completion time and communication interruption time based on the air-ground system channel model and the uncertainty model for the eavesdropper's location; discretizing the optimization problem corresponding to the weighted sum minimization model for the total CUAV mission completion time and communication interruption time; modeling the discretized optimization problem using a Markov decision process, and proposing an HT-D3QN algorithm for solving it. The proposed path planning method, combining HRL and T-D3QN, can effectively improve the mission execution efficiency and communication security of CUAVs in complex dynamic environments, providing reliable technical support for the application of CUAVs in smart cities, emergency rescue, and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) path planning, and in particular to a multi-mission point path planning method for an UAV. Background Art

[0002] In recent years, connected unmanned aerial vehicles (CUAV) technology has rapidly developed, showing broad application prospects in areas such as agricultural monitoring, logistics distribution, and public safety. CUAVs communicate with ground base stations in real time via cellular networks, offering advantages such as wide coverage, high speed, and strong interference resistance. However, due to the open nature of line-of-sight (LoS) wireless channels, which CUAVs rely on for data transmission, the high probability of LoS channels makes communication signals between CUAVs and ground base stations more susceptible to interception or eavesdropping by malicious nodes. Therefore, enhancing CUAV communication security without compromising mission efficiency has become a pressing technical challenge.

[0003] The airspace freedom introduced by UAV trajectory planning offers a new solution to this challenge. By dynamically adjusting the CUAV's flight trajectory in real time, channel characteristics can be optimized, reducing the risk of eavesdropping on information transmission. Therefore, efficiently optimizing the CUAV's flight trajectory is crucial for enhancing CUAV communication security. Traditional path planning algorithms, such as fast randomized tree and artificial potential field methods, rely heavily on prior knowledge of the environment, making them difficult to adapt to unknown or dynamically changing environments, limiting their application in complex scenarios.

[0004] To address these challenges, deep reinforcement learning (DRL) has been gradually introduced into CUAV path planning, becoming a powerful tool for addressing these challenges. DRL autonomously learns through interaction with the environment and dynamically adjusts its strategy, providing a flexible solution for path planning. Its advantages include adapting to complex and dynamic environments without the need for precise modeling; efficiently processing high-dimensional data using deep neural networks; optimizing strategies in real time and rapidly adjusting decisions; and balancing multiple optimization objectives through a multi-objective reward function. Therefore, DRL has shown significant potential for CUAV path planning. However, research on secure CUAV path planning in complex urban environments with eavesdroppers remains relatively scarce. Furthermore, most existing research assumes that the eavesdropper's location is fully known. However, in practice, due to the eavesdropper's passivity, its location information is often inaccurate, and the CUAV can only obtain an approximate range of its location, significantly increasing the complexity of CUAV path planning. Furthermore, existing literature primarily focuses on single-task point inspection scenarios, with limited research on the three-dimensional multi-task point inspection problem in the presence of eavesdroppers. In such missions, CUAVs not only need to address traditional challenges such as altitude control, obstacle avoidance, and managing the positional relationships of mission points, but also must effectively prevent eavesdroppers from intercepting and eavesdropping on communication links to ensure the confidentiality of communications. This multi-constraint, information security-focused nature of these missions places even higher demands on the design and optimization of CUAV path planning algorithms. Therefore, it is imperative to develop a path planning method that balances communication confidentiality, mission efficiency, and adaptability to complex environments to address the diverse challenges faced in real-world scenarios. Summary of the Invention

[0005] The purpose of the present invention is to improve and innovate the shortcomings and problems existing in the background technology and provide a multi-mission point path planning method for unmanned aerial vehicles.

[0006] According to a first aspect of the present invention, a method for multi-mission point path planning for an unmanned aerial vehicle is provided, which specifically comprises the following steps:

[0007] Step S1: method starts;

[0008] Step S2: establishing an air-ground system channel model and an uncertainty model of the eavesdropper's location, and constructing a weighted sum minimization model of the CUAV mission total completion time and communication interruption time based on the air-ground system channel model and the uncertainty model of the eavesdropper's location;

[0009] Step S3: Discretize the optimization problem corresponding to the weighted sum minimization model of the total CUAV task completion time and communication interruption time;

[0010] Step S4: Model the Markov decision process for the discretized optimization problem and propose the HT-D3QN algorithm to solve it;

[0011] Step S5: Modeling the communication system environment and initializing the parameters of the neural network corresponding to the HT-D3QN algorithm, and iteratively optimizing the parameters of the neural network corresponding to the HT-D3QN algorithm;

[0012] Step S6: Determine whether the neural network corresponding to the trained HT-D3QN algorithm has converged. If so, output the optimal parameters of the neural network corresponding to the HT-D3QN algorithm. Otherwise, return to step S5 to continue iterative optimization.

[0013] Step S7: After training is completed, the optimal path planning strategy is output.

[0014] A further solution is that the establishment of the air-ground system channel model and the uncertainty model of the eavesdropper's position in step S2 specifically includes:

[0015] Assume that the eavesdropper is located in a radius of The distance between CUAV and the eavesdropper can be modeled as:

[0016] ;

[0017] Where, Indicates that CUAV is The coordinates of the moment, Indicates that CUAV is The distance between its projection on the ground and the eavesdropper at that moment, Indicates that CUAV is The height of the moment, the location of the eavesdropper ; represents the Euclidean distance between CUAV and the eavesdropper.

[0018] A further solution is that the step S2 of establishing the air-ground system channel model and the uncertainty model of the eavesdropper's position further includes:

[0019] Get Moment from CUAV to Path loss per cell

[0020] ;

[0021] in and Represents CUAV and The probability of LoS and Non-LoS links occurring when the cells communicate;

[0022] in ; Where and is a constant;

[0023] Indicates the The elevation angle from the cell to the CUAV, Where, and Indicates the The coordinates of the cells, Indicates that CUAV is coordinates of the moment;

[0024] CUAV to Path loss of the LoS link of a cell

[0025] ;

[0026] CUAV to Path loss of the NLoS link for each cell

[0027]

[0028] in Indicates CUAV to The Euclidean distance of cells, is the carrier frequency of the base station communicating with the CUAV, For CUAV Flight altitude at the moment;

[0029] Get CUAV to Information rate per cell ;

[0030] Where, represents the transmit power of CUAV, Indicates CUAV to The channel gain of each cell, 、 、 Represent base station antenna gain, path loss, and small-scale fading, respectively. It is related to the cellular unit gain and the antenna array gain, and its calculation formula is: , where element and Array represent unit gain and array gain respectively; is a random variable; represents the interference power of other cells that are not communicating with CUAV to CUAV at this time, where , Indicates the The transmit power of a cell.

[0031] A further solution is that the step S2 of establishing the air-ground system channel model and the uncertainty model of the eavesdropper's position further includes:

[0032] Get Moment from CUAV to Path loss of an eavesdropper

[0033] ;

[0034] Get from CUAV to Path loss of LoS and NLoS links for each eavesdropper

[0035] ;

[0036] ;

[0037] in Indicates reference distance Path loss at is the path loss exponent; is a product with a mean of 0 and a standard deviation of Gaussian random variable; represents the Euclidean distance between CUAV and the eavesdropper;

[0038] in and Represents CUAV and The probability of LoS and Non-LoS links occurring when two eavesdroppers communicate;

[0039] in ; ; is the elevation angle from e eavesdroppers to the CUAV; , and represents the coordinates of the e-th eavesdropper;

[0040] Get CUAV to eavesdropper's eavesdropping rate ;in, is natural Gaussian noise.

[0041] A further solution is that the weighted sum minimization model of the total CUAV mission completion time and the communication interruption time is constructed based on the air-ground system channel model and the uncertainty model of the eavesdropper's position in step S2, specifically including:

[0042] Get the worst confidentiality rate between CUAV and base station ;

[0043] Determine CUAV in Is the confidentiality rate at the moment lower than the interruption threshold? ;

[0044] If yes, it means that the secure communication between CUAV and the ground base station has been interrupted, and the probability of communication interruption is ;

[0045] Get the CUAV communication interruption time within the total task completion time ,in Indicates the total completion time of the CUAV task; Indicates that CUAV is t The cell you are connected to at any moment;

[0046] According to the communication interruption time of CUAV within the total task completion time, an optimization problem is constructed to minimize the weighted sum of the total task completion time and the communication interruption time:

[0047]

[0048] in is the weighted coefficient for balancing the total task completion time and the communication interruption time; middle Indicates the maximum flight speed of CUAV, Indicates the CUAV flight direction; middle This means that the flight speed of the CUAV is not affected by direction; Specifies the starting point of CUAV; The airspace scope for CUAV mission execution is specified; It is stipulated that CUAV cannot collide with obstacles when performing tasks; Indicates which cell the CUAV is communicating with at a certain moment; and It is stipulated that CUAV must complete the inspection of all task points in sequence; S represents the flight path of the CUAV; Indicates the i A towering obstacle, Indicates the j The coordinates of the mission points, express t Cellular communication at all times, represents the total number of tall obstacles, represents the total number of cells, Indicates the total number of mission points, Indicates the flight airspace of CUAV.

[0049] A further solution is that step S3 specifically includes:

[0050] The total completion time of the CUAV task Discretized into ,in Indicates the total number of steps of CUAV flight, Indicates the time required for each step of CUAV;

[0051] The flight trajectory of CUAV is discretized into ;

[0052] The communication interruption probability of CUAV Discretized into ;

[0053] In the unit time step, the soft handover mechanism is used to measure the base station cells in the airspace. Second-rate , and put the The CUAV confidentiality rate obtained by measuring Recorded as ,in represents small-scale fading, then the new communication interruption probability is:

[0054] ;

[0055] ;

[0056] Prioritize establishing a connection with the base station cell with the lowest probability of communication interruption. The probability of communication interruption during the CUAV mission execution is expressed as:

[0057] ;

[0058] Then the optimization problem P0 can be discretized as:

[0059] ;

[0060] Where, Indicates the displacement of CUAV per unit time step.

[0061] A further solution is that step S4 specifically includes:

[0062] Each task point that the CUAV needs to inspect is defined as a subtask, which is managed by the high-level agent and executed by the low-level agent;

[0063] MDP quadruple for low-level agents Mapping; state space ; Action Space ; State transition probability The value of is governed by the conditional probability of CUAV selecting an action in the current state; They represent the six directions of CUAV in three-dimensional space: front, back, left, right, up, and down;

[0064] No. k The reward function corresponding to the time step

[0065] ;

[0066] in Represents the location of the current subtask, Indicates the position of the current time step CUAV, Indicates the location of CUAV at the previous time step; Represents low-level agent subtasks Total number of steps completed;

[0067] MDP quadruple for high-level agents Mapping

[0068] State Space ; Action Space ;

[0069] Reward Function Where, is the discount factor; , Indicates the first k Step reward, Represents low-level agent subtasks Total number of steps completed;

[0070] State Transfer It is determined by the final position of CUAV when it completes or fails to complete the subtask, i.e. .

[0071] A further solution is that step S4 further includes:

[0072] Initialize CUAV position and the reward function ;

[0073] The state of the high-level agent and the reward function Input into its corresponding two groups of Q networks so that the high-level agent outputs the subtask to be performed next ;

[0074] The low-level agent calculates the Q value of each action based on the current position of the CUAV and the subtask of the high-level agent, and selects the action with the largest Q value. Feedback to CUAV;

[0075] CUAV performs the action corresponding to the maximum Q value, and the CUAV position changes from Transfer to , and generate a new reward function ; At the same time, the state sequence of the low-level agent and high-level agent state sequence Stored in the experience pool for subsequent network updates;

[0076] Judgment subtask Is it completed? If so, a new subtask is generated. Give it to the lower-level agent to execute; if not, return to the lower-level agent to calculate the Q value of each action based on the current position of CUAV and the subtask of the higher-level agent, and select the action corresponding to the maximum Q value Feedback to CUAV until the number of CUAV steps reaches the total number of steps specified by the subtask;

[0077] Among them, the judgment subtask Completion specifically includes:

[0078] Determine the new position of CUAV With subtasks Whether the distance between the coordinates satisfies a preset distance threshold; if so, it indicates that the subtask has been completed; if not, it indicates that the subtask has not been completed.

[0079] According to a second aspect of the present invention, there is provided an electronic device, comprising: a memory and a processor;

[0080] The memory is used to store programs;

[0081] The processor is used to call the program stored in the memory to execute the UAV multi-mission point path planning method as described above.

[0082] According to a third aspect of the present invention, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements a multi-mission point path planning method for a drone as described above.

[0083] Compared with the existing technology, the present invention has the following advantages: the present invention addresses the need for CUAVs to perform inspections of multiple mission points in complex urban environments. Considering that CUAVs need to simultaneously meet complex constraints such as flight altitude control, obstacle avoidance, and multi-mission point inspections in three-dimensional space, the present invention proposes a hybrid algorithm, HT-D3QN, that integrates HRL and T-D3QN. The algorithm aims to minimize the weighted sum of the total CUAV task completion time and communication interruption time, thereby achieving efficient and reliable path planning. Specifically, to address the uncertainty of the eavesdropper's position and the sparse reward problem of multi-mission point inspections in three-dimensional space, the present invention introduces a hierarchical reinforcement learning framework to decompose the complex multi-mission point inspection problem into two sub-problems: high-level task allocation and low-level path planning. The high-level task allocation module is responsible for determining the optimal order in which the CUAV visits multiple mission points, while the low-level path planning module focuses on generating a three-dimensional trajectory that meets flight altitude, obstacle avoidance, and communication constraints. Combining the advantages of DRL, the present invention uses deep neural networks to extract environmental features, achieving efficient perception and response to complex environments, thereby optimizing the CUAV's three-dimensional trajectory and cellular association strategy. In summary, the HT-D3QN path planning method combining HRL and T-D3QN proposed in this paper can effectively improve the task execution efficiency and communication security of CUAV in complex dynamic environments, and provide reliable technical support for the application of CUAV in smart cities, emergency rescue and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0085] Figure 1 This is an air-to-ground communication model for CUAV to perform multi-task inspections in a complex urban environment provided by the first embodiment of the present invention;

[0086] Figure 2 This is a flow chart of a multi-mission point path planning method for a UAV provided by the first embodiment of the present invention;

[0087] Figure 3 is the uncertainty model under the imperfect eavesdropper position provided by the first embodiment of the present invention;

[0088] Figure 4 is a schematic diagram of the interaction between CUAV and HRL provided by the first embodiment of the present invention;

[0089] Figure 5 The network architecture of the high-level agent and the low-level agent provided by the first embodiment of the present invention;

[0090] Figure 6 is a schematic diagram of an urban environment provided by the first embodiment of the present invention;

[0091] Figure 7 1. A comparison diagram of the probability distribution of CUAV communication interruption under perfect and imperfect eavesdropper positions provided by the first embodiment of the present invention;

[0092] Figure 8 This is a comparison diagram of the two-dimensional and three-dimensional flight paths of the CUAV multi-task point inspection provided by the first embodiment of the present invention;

[0093] Figure 9 This is the average reward graph of the HT-D3QN algorithm provided by the first embodiment of the present invention under different multi-task inspection orders and different numbers of task points;

[0094] Figure 10 This is a drone trajectory diagram under different obstacles, eavesdropper numbers, and eavesdropper positions provided by the first embodiment of the present invention;

[0095] Figure 11 The weighted sum comparison results of task completion time and communication interruption time for two-dimensional and three-dimensional paths planned by the HT-D3QN algorithm provided in the first embodiment of the present invention with different numbers of task points are shown;

[0096] Figure 12 This is a comparison chart of average rewards of different algorithms provided by the first embodiment of the present invention. DETAILED DESCRIPTION

[0097] In order to make the objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0098] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0099] Example 1

[0100] See also Figure 1 The present invention provides a method for planning a multi-task path of an unmanned aerial vehicle, which specifically includes the following steps:

[0101] like Figure 2As shown in the figure, the flowchart of the three-dimensional path planning method for CUAV to perform multi-task point inspection when the eavesdropper's position is imperfect consists of 7 steps, namely, step S1: method start; step S2: establish the air-ground system channel model and the uncertainty model of the eavesdropper's position, and construct the weighted sum minimization model of the CUAV task total completion time and communication interruption time based on the air-ground system channel model and the uncertainty model of the eavesdropper's position; step S3: discretize the optimization problem corresponding to the weighted sum minimization model of the CUAV task total completion time and communication interruption time; step S4: model the discretized optimization problem with a Markov decision process, and propose the HT-D3QN algorithm to solve it; step S5: model the communication system environment and initialize the parameters of the neural network corresponding to the HT-D3QN algorithm, and iteratively optimize the parameters of the neural network corresponding to the HT-D3QN algorithm; step S6: determine whether the trained neural network corresponding to the HT-D3QN algorithm has converged. If so, output the optimal parameters of the neural network corresponding to the HT-D3QN algorithm. Otherwise, return to step S5 to continue iterative optimization; step S7: training is completed and the optimal path planning strategy is output. The following is a detailed introduction to these 7 steps:

[0102] Step S1: method starts;

[0103] Step S2: establishing an air-ground system channel model and an uncertainty model of the eavesdropper's location, and constructing a weighted sum minimization model of the CUAV mission total completion time and communication interruption time based on the air-ground system channel model and the uncertainty model of the eavesdropper's location;

[0104] like Figure 1 As shown in Figure 1, in a complex urban environment with multiple eavesdroppers, a CUAV performs an air-to-ground communication system for multiple inspection tasks. The CUAV needs to fly from a designated starting point A to an end point B to perform the task, and during the task, it needs to pass through multiple task points in sequence to complete the fixed-point inspection. Assume that the flight airspace of the CUAV is , that is, the flight airspace of CUAV is set as a cuboid, where and Respectively represent the minimum and maximum lengths of the cuboid boundary; and Respectively represent the minimum and maximum width of the cuboid boundary, and Respectively represent the minimum and maximum heights of the cuboid boundary. Base stations, each ground base station Located in a fixed position , and equipped with 3 fan-shaped honeycomb structures, the number of honeycombs can be calculated get. The eavesdropper's position is And it remains unchanged during the CUAV mission execution, where Assume that the starting position A of CUAV is expressed as , the task end position B is expressed as ,in and The CUAV starts from the starting point and visits multiple predetermined mission points in sequence. Carry out fixed-point inspection and finally reach the destination. Consider the existence of N OBS towering obstacles, whose set is recorded as ,in .

[0105] The establishment of the uncertainty model of the eavesdropper's location specifically includes:

[0106] like Figure 3 As shown, it is assumed that the eavesdropper is located in a radius of The distance between CUAV and the eavesdropper can be modeled as:

[0107] ;

[0108] Where, Indicates that CUAV is The coordinates of the moment, Indicates that CUAV is The distance between its projection on the ground and the eavesdropper at that moment, Indicates that CUAV is The height of the moment, the location of the eavesdropper ; represents the Euclidean distance between CUAV and the eavesdropper.

[0109] Among them, establishing the air-ground system channel model specifically includes:

[0110] (1) Obtain the path loss from CUAV to each cellular unit at each moment;

[0111] (2) Obtain the path loss from CUAV to each eavesdropper at each moment;

[0112] Specifically, the present invention adopts the probabilistic line-of-sight LoS channel model. Moment from CUAV to The path loss of a cellular unit can be expressed as: ,in and Represent the probability of LoS and Non-LoS (NLoS) link occurrence respectively. The probability of a LoS channel when a cell is communicating is given by It is concluded that and is a constant, which can be taken as a constant with different values ​​according to rural and urban environments. It is The elevation angle from the cell to the CUAV is given by Calculated, and Indicates the The coordinates of the cells, represents the coordinates of CUAV at time t. The NLoS probability is .

[0113] CUAV to The path loss of the LoS link of a cell is expressed as , the corresponding NLoS link path loss can be expressed as

[0114] ;

[0115] in Indicates CUAV to The distance between cells, is the carrier frequency of the base station communicating with the CUAV, is the height of CUAV.

[0116] Finally, CUAV to the The information rate of a cell can be written as:

[0117] ;

[0118] Where, represents the transmit power of CUAV, Indicates CUAV to The channel gain of each cell, 、 、 Represent base station antenna gain, path loss, and small-scale fading, respectively. It is mainly related to the cellular unit gain and the antenna array gain, and its calculation formula is: , where element and Array represent unit gain and array gain respectively; is a random variable; represents the interference power of other cells that are not communicating with CUAV to CUAV at this time, where , Indicates the The transmit power of a cell.

[0119] Similarly, according to the probabilistic line-of-sight LoS channel model, the UAV and the The path loss between eavesdroppers is expressed as:

[0120] ;

[0121] From drones to The path loss of the LoS and NLoS links of an eavesdropper can be expressed as:

[0122] ;

[0123] ;

[0124] in is the reference distance The path loss at is the path loss exponent, is a product with a mean of 0 and a standard deviation of is a Gaussian random variable used to represent the effect of shadow fading. As mentioned above, represents the Euclidean distance between CUAV and the eavesdropper.

[0125] It should be noted that the reference distance and path loss index can be determined by those skilled in the art according to actual conditions, and are not subject to specific limitations. = 1m; the path loss exponent n ranges from 2 to 6, depending on the propagation environment. For example, the free-space path loss exponent is 2. However, the path loss exponent increases in the presence of obstacles. For example, the path loss exponent n for urban macrocells ranges from 3.7 to 6.5, and the path loss exponent n for urban microcells ranges from 2.7 to 3.5.

[0126] The probability that the CUAV and the e-th eavesdropper communicate in a LoS channel is given by It is concluded that and is a constant, which can be taken as a constant with different values ​​according to rural and urban environments. is the elevation angle from the e-th eavesdropper to the CUAV, which is given by Calculated, and represents the coordinates of the e-th eavesdropper. The NLoS probability is .

[0127] Finally, CUAV to the The eavesdropping rate of an eavesdropper can be expressed as:

[0128] ;

[0129] in, is natural Gaussian noise. Once the CUAV position is determined, the eavesdropping rate of the CUAV to each eavesdropper can be calculated.

[0130] Furthermore, based on the air-ground system channel model and the uncertainty model of the eavesdropper's location, a weighted sum minimization model of the CUAV mission total completion time and communication interruption time is constructed. Specifically, the model includes:

[0131] The worst confidentiality rate between CUAV and base station is obtained as The present invention defines the communication interruption threshold as , if CUAV is If the confidentiality rate at the moment is lower than the threshold, it means that the secure communication between the CUAV and the ground base station is interrupted. Therefore, the probability of communication interruption can be expressed as The corresponding total task completion time, CUAV communication interruption time can be written as ,in Indicates the total completion time of the CUAV task.

[0132] Therefore, the optimization problem of minimizing the weighted sum of total task completion time and communication interruption time for CUAVs performing multi-task inspections in complex urban environments can be expressed as:

[0133]

[0134] in is the weighted coefficient for balancing the total task completion time and the communication interruption time; middle Indicates the maximum flight speed of CUAV, Indicates the CUAV flight direction; middle This means that the flight speed of the CUAV is not affected by direction; Specifies the starting point of CUAV; The airspace scope for CUAV mission execution is specified; It is stipulated that CUAV cannot collide with obstacles when performing tasks; Indicates which cell the CUAV is communicating with at a certain moment; and It is stipulated that CUAV must complete the inspection of all task points in sequence; S represents the flight path of the CUAV; Indicates the i A towering obstacle, Indicates the j The coordinates of the mission points, express t Cellular communication at all times, represents the total number of tall obstacles, represents the total number of cells, Indicates the total number of mission points.

[0135] Step S3: Discretize the optimization problem corresponding to the weighted sum minimization model of the total CUAV task completion time and communication interruption time;

[0136] It should be noted that the optimization problem corresponding to the weighted sum minimization model of the discretized CUAV task total completion time and communication interruption time is to reconstruct the continuous state space and action space into a form suitable for deep reinforcement learning.

[0137] First, the total completion time of the CUAV task Discretized into ,in Indicates the total number of steps of CUAV flight, represents the time required for each step of CUAV, so the flight trajectory of CUAV can be discretized as , in problem P0 and Rewritten as: and ,in represents the displacement of CUAV in unit time step, Indicates the flight direction of CUAV per unit time. Therefore, the probability of communication interruption of CUAV is Then it is discretized into: .

[0138] It is worth noting that because The closed-form expression of is related to the CUAV location, the associated base station cell, and the small-scale fading Related, but is a random variable, which results in each measurement due to the random variable existence, the resulting However, the expected method can be used to calculate the communication interruption time of CUAV in a unit time step, that is, in a given time step k , the CUAV position and the associated base station cell are fixed, and the CUAV confidentiality rate during the entire mission execution process is Discrete ,now that Including the coefficients of random small-scale fading in all cells, the discretized expression of communication interruption probability can be obtained, that is,

[0139] ;

[0140] in, Expressing hope, is a discriminant function and can be written as:

[0141] ;

[0142] In practical applications, it is usually difficult to obtain the CUAV confidentiality rate The complete probability distribution of , so the time average value is used instead of the statistical average value (ie expected value) to estimate its statistical characteristics. Specifically, within the unit time step, the soft handover mechanism is used to measure the base station cells in the airspace. Second-rate , and put the The measurement results of Recorded as ,in represents small-scale fading. It can be expressed as , further communication interruption probability Can be rewritten as:

[0143] ;

[0144] According to the law of large numbers, as the number of trials increases, the sample mean of a random event will approach its theoretical mean (i.e., expected value). Therefore, when the CUAV confidentiality rate is measured sufficiently many times within a unit time step, When the value is large enough, the sample mean can be used to approximate the expected value, so that In order to minimize the probability of communication interruption during CUAV flight, it is necessary to prioritize establishing a connection with the base station cell with the lowest probability of communication interruption. Therefore, the probability of communication interruption during CUAV mission execution can be expressed as:

[0145] ;

[0146] In summary, the optimization problem P0 can be discretized as:

[0147] .

[0148] Step S4: Model the Markov decision process for the discretized optimization problem and propose the HT-D3QN algorithm to solve it;

[0149] The CUAV task execution process is converted to a Markov Decision Process (MDP) because MDP provides a formal framework for describing sequential decision problems. By defining states, actions, transition probabilities, and reward functions, it makes the problem description clearer and more structured. It also facilitates the algorithm's prediction of future states and rewards based on the current state and action, enabling more efficient learning of optimal policies. This paper adopts the HRL framework, first defining each task point that the CUAV needs to inspect as a subtask, which is managed by a high-level agent in the HRL. Once the CUAV meets specific conditions to complete a subtask, the low-level agents are responsible for executing it.

[0150] First, the MDP quadruple of the low-level agent Mapping: The state space can be expressed as ; represents the coordinates of the current time step CUAV, Represents the coordinates of the subtask executed at the current time step. The action space of CUAV can be formally expressed as ; 、 、 、 、 、 They represent the six directions of the drone in three-dimensional space: front, back, left, right, up, and down. At each time step, the drone selects an action from the action space to perform, thereby realizing path planning in three-dimensional space. The value of is determined by the conditional probability of CUAV selecting an action in the current state.

[0151] Taking into account various factors such as the need for the CUAV to fly from the starting point to the mission destination and the need to avoid areas with a high probability of communication interruption as much as possible, the reward function is designed as follows:

[0152] ;

[0153] Where, Indicates the first k The reward at the time step, Represents the location of the corresponding subtask, Indicates the position of the current time step CUAV, Indicates the location of CUAV at the previous time step; Represents low-level agent subtasks The total number of steps completed.

[0154] After completing the MDP mapping of the low-level agent, map the MDP of the high-level agent and define the MDP quadruple of the high-level agent , since the state space of the high-level agent is the entire airspace of the flight, that is, ; The high-level agent action space is set to The goal of the high-level agent is to guide the low-level agent to complete the subtasks by selecting subtasks reasonably to maximize the reward of the total task. Therefore, the reward of each subtask of the low-level agent affects the overall reward of the high-level agent. The reward at the end is , Indicates the first k Step reward, Represents low-level agent subtasks The total number of steps completed, is the discount factor, so the reward function of the high-level agent is set to: ;in represents the first subtask; where, is the discount factor.

[0155] For example, the time steps of the first subtask are 1, 2, ..., ;The time step of the second subtask is , , ... , ;The time step of the second subtask is , , ... , .

[0156] State Transfer It is determined by the final position of CUAV when it completes or fails to complete the subtask, i.e. .

[0157] It should be noted that for each subtask, the lower-level agent may complete the subtask or may not complete the subtask. During the step, if the distance between the CUAV coordinate and the subtask coordinate is less than or equal to the preset distance threshold, it means that the low-level agent has completed the subtask, and the position of the CUAV when the low-level agent completes the subtask is used as the ; If executing subtask If the distance between the CUAV coordinates and the subtask coordinates is still greater than the preset distance threshold after the step, it means that the lower-level agent has not completed the subtask, and the last position of the CUAV is used as ; Therefore, the position of the low-level agent after completing the subtask is , then the high-level agent needs to reselect the task, and the position of the low-level agent after failing to complete the subtask is , at this time, the high-level agent does not need to reallocate subtasks, and the low-level agent continues to perform the original task; that is, the high-level agent will only be based on Generate new tasks without Generate a new task.

[0158] Specifically, the interaction between CUAV and HRL in the proposed HT-D3QN algorithm is as follows: Figure 4 and Figure 5 The whole task execution steps are divided into: Step 1, initialize CUAV position and the reward function Specifically, the initial position of CUAV is randomly generated by the program. The second step is to convert the state of the high-level intelligent body and the reward function Input to its network, since CUAV has not yet performed actions and moved positions, the state of the high-level agent at this time The coordinates and position of CUAV The coordinates are equal, then the reward function and The default value is 0; the network corresponding to the high-level agent outputs the subtask to be executed next based on the input state ; The third step is that the low-level agent is based on the CUAV position and subtasks of high-level agents , calculate the Q value of each action and select the action with the largest Q value Feedback to CUAV; Step 4, CUAV performs the action, and CUAV position changes from Transfer to , and generate a new reward function At the same time, the state sequence of the low-level agent and high-level agent state sequence It is stored in the experience pool for subsequent network updates; in the fifth step, the high-level agent is based on the new position of CUAV With subtasks The distance between coordinates is used to determine the subtask Completed or not, if CUAV completes the subtask, the high-level agent will generate a new subtask Submit it to the lower-level agent, and then start a new round of cycles. Finally, CUAV gradually learns the optimal strategy through interaction with the environment.

[0159] For example, both the high-level agent and the low-level agent have two sets of Q networks. In this embodiment, the Q network is a feedforward neural network, and each set of networks adopts the D3QN network structure to transform the state of the high-level agent into the state of the low-level agent. and the reward function After input into the network, each group of networks outputs all the actions corresponding to the high-level intelligent agent The maximum value, then select two groups of networks The smaller value of the maximum value, and finally select the corresponding action according to the smaller value, that is, output the optimal subtask; when the low-level agent obtains the subtask, it inputs the current position of the drone into the neural network and the current subtask Then, similarly, the lower-level agents follow the two corresponding networks The smaller value of the maximum value selects the optimal action that the UAV needs to perform; CUAV performs the optimal action, and the CUAV position changes from Transfer to , and generate a new reward function ; Determine the new position of CUAV With subtasks Is the distance between the coordinates less than or equal to the preset distance threshold? If so, it means that the subtask is completed and the new position of CUAV is set. As the state of a high-level agent , to generate the next subtask; if not, return to the neural network to input the current position of the drone and the current subtask until the CUAV moves the total number of steps specified by the subtask; if the CUAV moves the total number of steps specified by the subtask, and the distance between the CUAV position coordinates and the subtask coordinates is still greater than the preset distance threshold, then the final position of the CUAV is used as the state of the high-level agent ; The lower-level agent re-executes the original task.

[0160] Step S5: Modeling the communication system environment and initializing the parameters of the neural network corresponding to the HT-D3QN algorithm, and iteratively optimizing the parameters of the neural network corresponding to the HT-D3QN algorithm;

[0161] In order to make the simulation results more consistent with the real environment, the present invention first uses a building statistical model published by the International Telecommunication Union to generate an urban environment, such as Figure 6 As shown in the 2D building distribution map, the CUAV flight airspace is limited to 2000m In a 2000m environment, there are 4 base stations in the area, and their locations are m、 m、 m、 m, represented by black hexagons, while rectangles represent buildings. The ratio of the area occupied by buildings to the total area in this area is 0.3, and the average height of the buildings is , the average number of buildings per square kilometer is 300. The 3D image of buildings in the environment generated according to these values ​​is as follows Figure 6As shown in the middle right picture.

[0162] In this area, the base station height is designed to be 25 meters, with a total of 12 cells. The base station antenna design strictly adheres to 3GPP standards and is configured as a vertically arranged 8-element uniform linear array (ULA). Each element in the array is directional, with a half-power beamwidth set to specific values ​​in both the vertical and horizontal directions. This vertical antenna array also incorporates a phase offset, tilting the main beam downward at a specific angle, thus forming a directional antenna array with fixed three-dimensional radiation characteristics.

[0163] In the HT-D3QN algorithm system constructed by the present invention, the Q network involved is built on a fully connected feedforward artificial neural network, and its structure includes 5 hidden layers. In the first 4 hidden layers, the number of neurons is configured as 512, 256, 128, and 128 respectively. The last hidden layer of the network adopts a dueling layer design, and the number of neurons in the dueling layer is , one neuron specifically outputs the evaluation data of the state-value function, while the remaining neurons are responsible for outputting the evaluation data of the advantage function. Regarding the network's operating mechanism, the ReLu function is used as the activation function for each hidden layer to promote effective neuron activation. Furthermore, the Adam optimizer is used to reduce mean squared error and optimize network training results.

[0164] Step S6: Determine whether the neural network corresponding to the trained HT-D3QN algorithm has converged. If so, output the optimal parameters of the neural network. Otherwise, return to step S5 to continue iterative optimization.

[0165] Specifically, when it is judged that the CUAV reward curve converges, it indicates that the neural network corresponding to the trained HT-D3QN algorithm has converged.

[0166] It should be noted that the horizontal axis of the CUAV reward curve corresponds to the number of iterations, and the vertical axis of the CUAV reward curve corresponds to the average reward value corresponding to multiple iterations, where the average reward value is obtained by dividing the reward function obtained by each step when CUAV performs the subtask in each iteration. The reward value for each iteration is accumulated, and then the reward values ​​corresponding to the preset number of iterations are accumulated, and finally the average value of the preset number of iterations is calculated.

[0167] Whether a curve converges can refer to whether points on the curve gradually approach a certain value or line as the independent variable changes. Whether a curve converges can also refer to whether the residual value of the curve gradually decreases with the increase in the number of iterations until the residual value is less than a preset value. Those skilled in the art can choose according to actual circumstances. All existing methods for determining curve convergence are within the scope of protection of the present invention.

[0168] Step S7: After training is completed, the optimal path planning strategy is output;

[0169] Specifically, after the training is completed, the optimal path planning strategy is output based on the optimal parameters of the neural network.

[0170] To verify the effectiveness of the proposed CUAV multi-task point 3D path planning method under imperfect eavesdropper locations, we conducted simulations using Anaconda, Python 3.7.8, and TensorFlow 2.11.0. The specific simulation parameters are shown in Table 1.

[0171] Table 1 Simulation parameters

[0172]

[0173] Figure 7 Figure 1 shows the distribution of communication interruption probabilities for perfect and imperfect eavesdropper positions at an altitude of 100 meters. Black diamonds represent eavesdroppers, black dots represent obstacles, and blue triangles represent the mission endpoint. Comparing the areas enclosed by dashed circles in Figures (a) and (b), it can be seen that the probability of communication interruption under the imperfect eavesdropper position is significantly higher than that under the perfect eavesdropper position in some areas. Figure (c) further presents the three-dimensional distribution of communication interruption probabilities for an altitude of 80 to 120 meters. The figure shows significant differences in the probability of communication interruption at different altitudes, and the area with higher interruption probability expands as altitude decreases.

[0174] Figure 8 The comparison of the 2D and 3D inspection paths of CUAV under different numbers of task points using the HT-D3QN algorithm proposed in this paper is shown. Figure (a) shows a scenario with 3 task points, and Figure (b) shows a scenario with 4 task points. As can be seen from Figure (a), in the initial flight phase of CUAV, m range, since the communication interruption probability at the 80m and 85m altitude planes is similar, their two-dimensional and three-dimensional flight paths basically overlap. Starting from the m area, the probability of communication interruption at the 80-meter altitude plane increases significantly, while the probability of communication interruption at the 85-meter altitude plane is lower. Therefore, the CUAV chooses to climb to avoid the high-risk area. Furthermore, when the CUAV encounters an area with low communication interruption probability, it flexibly adjusts its flight altitude, effectively overcoming the limitations of two-dimensional path planning and making the flight strategy more realistic. Furthermore, the CUAV prioritizes inspecting green task points because the red and blue task points are far apart and there is a large area of ​​high communication interruption probability between them. As can be seen in Figure (b), the CUAV optimizes the inspection sequence based on the distribution of interruption probability areas and the distance between task points. These results demonstrate that the HRL combined with the T-D3QN algorithm demonstrates good adaptability, flexibility, and robustness when handling the three-dimensional multi-task point inspection problem under imperfect eavesdropper positions.

[0175] Figure 9 The average reward performance of the HT-D3QN algorithm proposed in this invention for two-dimensional and three-dimensional path planning in CUAV multi-task point inspection tasks is compared. Comparing Figures (a) and (b), it can be seen that under the same traversal order, the average reward obtained by CUAV under three-dimensional path planning is significantly higher than that under two-dimensional path planning. This is because under three-dimensional path planning, CUAV can avoid areas with a high probability of communication interruption by adjusting the flight altitude, thereby obtaining higher rewards. Comparing Figures (c) and (d), it can be seen that under the same number of task points, three-dimensional path planning significantly improves the task reward of the UAV by introducing vertical degrees of freedom, and its performance is always better than two-dimensional path planning. And as the number of task points increases, the algorithm can still maintain a high reward level and achieve stable convergence, which reflects strong robustness.

[0176] Figure 10 The proposed HT-D3QN algorithm demonstrates the flight trajectory of a CUAV under varying conditions of obstacles, eavesdropper counts, and eavesdropper locations. As the figure shows, while the complexity of the flight environment increases significantly with the number of obstacles and eavesdroppers, the HT-D3QN algorithm still manages to effectively avoid areas with a high probability of communication interruption by dynamically adjusting the CUAV's flight altitude and path, ensuring successful mission completion. In particular, the proposed HT-D3QN algorithm demonstrates enhanced environmental adaptability and robustness in scenarios with a high density of obstacles and eavesdroppers, fully demonstrating its superiority.

[0177] Figure 11The weighted sum of task completion time and communication interruption time for two-dimensional and three-dimensional paths planned by the HT-D3QN algorithm proposed in this invention with different numbers of task points is compared. As can be seen from Figure (a), the weighted sum of the three-dimensional path planning is significantly lower than that of the two-dimensional path planning. As can be seen from Figure (b), the time-saving advantage of three-dimensional path planning becomes more significant as the number of task points increases. This is mainly due to the introduction of vertical degrees of freedom in three-dimensional path planning, which enables the CUAV to avoid areas with high probability of communication interruption by adjusting its altitude, thereby optimizing the flight path and improving mission execution efficiency.

[0178] Figure 12 The average reward curves of different algorithms are compared. As can be seen from the figure, the proposed HT-D3QN algorithm outperforms the baseline algorithm in terms of convergence time and operational stability. Specifically, the HT-D3QN algorithm reaches stable convergence after 3000 training rounds, and its final average reward value is significantly higher than that of the comparison algorithm. This superior performance is mainly attributed to two aspects: first, the layered training mechanism effectively improves the algorithm's learning efficiency; second, T-D3QN enhances algorithm stability and avoids the overestimation problem in traditional Q-Learning.

[0179] Example 2

[0180] The present invention provides an electronic device, comprising: a memory and a processor;

[0181] The memory is used to store programs;

[0182] The processor is used to call the program stored in the memory to execute the drone multi-mission point path planning method as described in Example 1.

[0183] Example 3

[0184] The present invention provides a readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a multi-mission point path planning method for a drone as described in Example 1 is implemented.

[0185] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.

Claims

1. A multi-task point path planning method for an unmanned aerial vehicle, characterized in that: The specific steps include: Step S1: method starts; Step S2: establishing an air-ground system channel model and an uncertainty model of the eavesdropper's location, and constructing a weighted sum minimization model of the CUAV mission total completion time and communication interruption time based on the air-ground system channel model and the uncertainty model of the eavesdropper's location; Step S3: Discretize the optimization problem corresponding to the weighted sum minimization model of the total CUAV task completion time and communication interruption time; Step S4: Model the Markov decision process for the discretized optimization problem and propose the HT-D3QN algorithm to solve it; Step S5: Modeling the communication system environment and initializing the parameters of the neural network corresponding to the HT-D3QN algorithm, and iteratively optimizing the parameters of the neural network corresponding to the HT-D3QN algorithm; Step S6: Determine whether the neural network corresponding to the trained HT-D3QN algorithm has converged. If so, output the optimal parameters of the neural network corresponding to the HT-D3QN algorithm. Otherwise, return to step S5 to continue iterative optimization. Step S7: After training is completed, the optimal path planning strategy is output; The step S2 of establishing the air-ground system channel model and the uncertainty model of the eavesdropper's position specifically includes: Assume that the eavesdropper is located in a radius of The distance between CUAV and the eavesdropper can be modeled as: ; Where, Indicates that CUAV is The coordinates of the moment, Indicates that CUAV is The distance between its projection on the ground and the eavesdropper at that moment, Indicates that CUAV is The height of the moment, the location of the eavesdropper ; represents the Euclidean distance between CUAV and the eavesdropper; The step S2 of establishing the air-ground system channel model and the uncertainty model of the eavesdropper's position also includes: Get Moment from CUAV to Path loss per cell ; in and Represents CUAV and The probability of LoS and Non-LoS links occurring when the cells communicate; in ; Where and is a constant; Indicates the The elevation angle from the cell to the CUAV, Where, and Indicates the The coordinates of the cells, Indicates that CUAV is coordinates of the moment; CUAV to Path loss of the LoS link of a cell ; CUAV to Path loss of the NLoS link for each cell in Indicates CUAV to The Euclidean distance of cells, is the carrier frequency of the base station communicating with the CUAV, For CUAV Flight altitude at the moment; Get CUAV to Information rate per cell ; Where, represents the transmit power of CUAV, Indicates CUAV to The channel gain of each cell, 、 、 Represent base station antenna gain, path loss, and small-scale fading, respectively. It is related to the cellular unit gain and the antenna array gain, and its calculation formula is: , where element and Array represent unit gain and array gain respectively; is a random variable; represents the interference power of other cells that are not communicating with CUAV to CUAV at this time, where , Indicates the The transmit power of each cell; The step S2 of establishing the air-ground system channel model and the uncertainty model of the eavesdropper's position also includes: Get Moment from CUAV to Path loss of an eavesdropper ; Get from CUAV to Path loss of LoS and NLoS links for each eavesdropper ; ; in Indicates reference distance Path loss at is the path loss exponent; is a product with a mean of 0 and a standard deviation of Gaussian random variable; represents the Euclidean distance between CUAV and the eavesdropper; in and Represents CUAV and The probability of LoS and Non-LoS links occurring when two eavesdroppers communicate; in ; ; is the elevation angle from e eavesdroppers to the CUAV; , and represents the coordinates of the e-th eavesdropper; Get CUAV to eavesdropper's eavesdropping rate ;in, is natural Gaussian noise; The step S2 in which the weighted sum minimization model of the total completion time of the CUAV mission and the communication interruption time is constructed based on the air-ground system channel model and the uncertainty model of the eavesdropper's position specifically includes: Get the worst confidentiality rate between CUAV and base station ; Determine CUAV in Is the confidentiality rate at the moment lower than the interruption threshold? ; If yes, it means that the secure communication between CUAV and the ground base station has been interrupted, and the probability of communication interruption is ; Get the CUAV communication interruption time within the total task completion time ,in Indicates the total completion time of the CUAV task; Indicates that CUAV is t The cell you are connected to at any moment; According to the communication interruption time of CUAV within the total task completion time, an optimization problem is constructed to minimize the weighted sum of the total task completion time and the communication interruption time: in is the weighted coefficient for balancing the total task completion time and the communication interruption time; middle Indicates the maximum flight speed of CUAV, Indicates the CUAV flight direction; middle This means that the flight speed of the CUAV is not affected by direction; Specifies the starting point of CUAV; The airspace scope for CUAV mission execution is specified; It is stipulated that CUAV cannot collide with obstacles when performing tasks; Indicates which cell the CUAV is communicating with at a certain moment; and It is stipulated that CUAV must complete the inspection of all task points in sequence; S represents the flight path of the CUAV; Indicates the i A towering obstacle, Indicates the j The coordinates of the mission points, express t Cellular communication at all times, represents the total number of tall obstacles, represents the total number of cells, Indicates the total number of mission points, Indicates the flight airspace of CUAV.

2. A multi-task point path planning method for an unmanned aerial vehicle according to claim 1, characterized in that: The step S3 specifically includes: The total completion time of the CUAV task Discretized into ,in Indicates the total number of steps of CUAV flight, Indicates the time required for each step of CUAV; The flight trajectory of CUAV is discretized into ; The communication interruption probability of CUAV Discretized into ; In the unit time step, the soft handover mechanism is used to measure the base station cells in the airspace. Second-rate , and put the The CUAV confidentiality rate obtained by measuring Recorded as ,in represents small-scale fading, then the new communication interruption probability is: ; ; Prioritize establishing a connection with the base station cell with the lowest probability of communication interruption. The probability of communication interruption during the CUAV mission execution is expressed as: ; Then the optimization problem P0 can be discretized as: ; Where, Indicates the displacement of CUAV per unit time step.

3. A multi-task point path planning method for an unmanned aerial vehicle according to claim 2, characterized in that: The step S4 specifically includes: Each task point that the CUAV needs to inspect is defined as a subtask, which is managed by the high-level agent and executed by the low-level agent; MDP quadruple for low-level agents Mapping; state space ; Action Space ; State transition probability The value of is governed by the conditional probability of CUAV selecting an action in the current state; They represent the six directions of CUAV in three-dimensional space: front, back, left, right, up, and down; No. k The reward function corresponding to the time step ; in Represents the location of the current subtask, Indicates the position of the current time step CUAV, Indicates the location of CUAV at the previous time step; Represents low-level agent subtasks Total number of steps completed; MDP quadruple for high-level agents Mapping State Space ; Action Space ; Reward Function Where, is the discount factor; , Indicates the first k Step reward, Represents low-level agent subtasks Total number of steps completed; State Transfer It is determined by the final position of CUAV when it completes or fails to complete the subtask, i.e. .

4. A multi-task point path planning method for an unmanned aerial vehicle according to claim 3, characterized in that: The step S4 further includes: Initialize CUAV position and the reward function ; The state of the high-level agent and the reward function Input into its corresponding two groups of Q networks so that the high-level agent outputs the subtask to be performed next ; The low-level agent calculates the Q value of each action based on the current position of the CUAV and the subtask of the high-level agent, and selects the action with the largest Q value. Feedback to CUAV; CUAV performs the action corresponding to the maximum Q value, and the CUAV position changes from Transfer to , and generate a new reward function ; At the same time, the state sequence of the low-level agent and high-level agent state sequence Stored in the experience pool for subsequent network updates; Judgment subtask Is it completed? If so, a new subtask is generated. Give it to the lower-level agent to execute; if not, return to the lower-level agent to calculate the Q value of each action based on the current position of CUAV and the subtask of the higher-level agent, and select the action corresponding to the maximum Q value Feedback to CUAV until the number of CUAV steps reaches the total number of steps specified by the subtask; Among them, the judgment subtask Completion specifically includes: Determine the new position of CUAV With subtasks Whether the distance between the coordinates satisfies a preset distance threshold; if so, it indicates that the subtask has been completed; if not, it indicates that the subtask has not been completed.

5. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the drone multi-mission point path planning method according to any one of claims 1 to 4.

6. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the multi-mission point path planning method for a drone as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Unmanned aerial vehicle flight route off-line and on-line hybrid optimization method for secure communication

    CN113765579A

  • Marine area safety communication unmanned aerial vehicle track real-time planning method based on reinforcement learning

    CN115407794A