Multi-unmanned aerial vehicle cooperative coverage path planning method based on reinforcement learning

Through the multi-UAV collaborative coverage path planning method based on reinforcement learning, the problem of adaptive learning and limited computing resources of multiple drones in dynamic environments is solved, and efficient path planning is achieved, suitable for post-disaster search and rescue fields.

CN120295339APending Publication Date: 2025-07-11KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510459024.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

现有技术难以有效解决多无人机协同覆盖路径规划问题,尤其是在动态环境中无人机的自适应性学习和有限计算资源下的高效路径规划。

Method used

The multi-unmanned aerial vehicle collaborative coverage path planning method based on reinforcement learning is adopted. By establishing a UAV energy consumption model, a multi-agent area coverage path model and an improved multi-agent reinforcement learning algorithm framework, combining the back-and-forth path algorithm, the regional access sequence and internal path planning are optimized, and cross-layer connection networks are used to reduce computing complexity.

Benefits of technology

It realizes the adaptive environment of drones under limited computing power, efficiently completes search tasks, improves the real-time and autonomy of path planning, and is suitable for important areas such as post-disaster search and rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295339A_ABST
    Figure CN120295339A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-unmanned aerial vehicle cooperative coverage path planning method based on reinforcement learning, and belongs to the field of unmanned aerial vehicle autonomous navigation and path planning. The method is suitable for scenes requiring a multi-agent area coverage path planning (MACPP) problem, such as post-disaster search and rescue, agricultural irrigation and the like. According to the method, the MACPP is decomposed into two sub-problems of inter-region access sequence planning (MTSP) and sub-region coverage (CPP), and a lightweight cross-layer connection multi-agent reinforcement learning framework is designed. According to the framework, a complete collaborative learning strategy is adopted to optimize the MTSP, meanwhile, an exploration mechanism dynamically adjusted along with the interaction frequency is designed and used so as to optimize an entrance and exit combination strategy in the training process, and the CPP problem is efficiently solved in combination with a back-and-forth path algorithm. Through the method, multiple unmanned aerial vehicles can efficiently complete a multi-region search task in the shortest time to meet the requirements of actual application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous navigation and path planning of unmanned aerial vehicles, and more particularly, relates to a multi-UAV cooperative coverage path planning method based on reinforcement learning. Background Art

[0002] In the past few years, the application of unmanned aerial vehicles (UAVs) has experienced unprecedented growth and has been widely used in various scenarios, such as post-disaster search and rescue, land surveying, agricultural irrigation, and disaster assessment. These application scenarios usually involve the multi-agent area coverage path planning problem (MACPP). Specifically, MACPP can be further divided into two types of tasks: point-to-point search path planning and coverage path planning (CPP).

[0003] The point-to-point search path planning problem is often modeled as an extended form of the traveling salesman problem (TSP) or vehicle routing problem (VRP). Solutions to such problems can be classified into three categories: classical algorithms, heuristic methods, and deep learning-based methods. Compared with the first two methods, path planning based on deep reinforcement learning provides a new perspective for the point-to-point search path planning of UAVs due to its high flexibility and environmental adaptability. However, in practical applications, this method faces problems such as a large scale of training samples and high computational iteration costs, and the complex neural network structure poses challenges to the limited computational resources of UAVs.

[0004] The core of coverage path planning (CPP) is to select the optimal path to achieve full coverage of the area at the lowest cost. Given the need for multi-aircraft cooperation in actual search tasks, the present invention defines MACPP as a special type of multi-traveling salesman problem (MTSP). Although there are many existing algorithms for solving MTSP and CPP problems, it is difficult to effectively solve the MACPP problem by directly extending the existing methods because the unique feature of this problem is that the entry and exit positions of each sub-region are dynamically changing, which directly affects the access order and the selection of the navigation path within the region. In addition, the MACPP problem also needs to meet a series of constraints, including the maximum flight distance limit of UAVs, the optimal allocation of target sub-regions, the optimization of the cross-region access sequence, and the planning of the coverage path within the region.

[0005] Facing the above challenges, two core problems need to be solved urgently in current research: one is to enable UAVs to learn adaptively to the environment; the other is to develop a lightweight algorithm architecture suitable for limited computational resources. The breakthrough of these problems will push the UAV path planning technology a big step forward towards real-time, autonomous, and efficient directions, thus serving important fields such as post-disaster search and rescue more effectively. Summary of the Invention

[0006] The object of the present invention is to provide a multi-UAV cooperative coverage path planning method based on reinforcement learning, enabling the UAVs to adaptively learn the post-disaster environment under limited computing power and efficiently complete the search task. The technical solution adopted by the present invention is: a multi-UAV cooperative coverage path planning method based on reinforcement learning, comprising the following steps:

[0007] Step S100: Establish a UAV energy consumption model;

[0008] Step S200: Establish a multi-agent area coverage path model;

[0009] Step S300: Establish a multi-agent reinforcement learning algorithm framework;

[0010] Step S301: Plan the area access order of the UAV swarm. The UAV selects an area to work. If the energy of the UAV is not exhausted in the current selection, go to step S302. If the energy of the UAV is exhausted in the current selection, re-plan the area access order of the UAV swarm, and the UAV re-selects an area to work;

[0011] Step S302: Based on the round-trip path algorithm, plan the internal path of the area currently selected by the UAV. If all areas have not been searched, go to step S303; if all areas have been searched, go to step S400;

[0012] Step S303: The entrance and exit exploration factor adjusts the exploration probability, optimizes the path, and then returns to step S301;

[0013] Step S400: The UAV returns to the base station to complete the current task.

[0014] Specifically, the step S100 includes:

[0015] At the beginning of the task, multiple UAVs with sufficient power are deployed to the search area, and it is necessary to ensure that the search task is completed before the UAVs run out of power. Assume that the UAVs fly at a constant flight speed V and can turn with any radius of curvature at speed V. It should be noted that the energy consumption of wireless communication for transmitting information back to the command post is relatively small. Therefore, in the present invention, only the flight energy consumption of the UAVs is concerned, which is expressed as:

[0016]

[0017] Among them, represents the blade profile power of the UAV, is the induced power in the hover state, V t is the tip speed of the rotor blade; represents the induced power of the UAV, is the blade profile power in the hover state, ι0 is the average induced velocity of the rotor; represents the parasitic power of the UAV, Ψ is the ratio of the blade area to the disk area, γ represents the fuselage drag ratio, Ξ represents the air density, and Ω represents the rotor disk area.

[0018] Let the farthest flight distance of the UAV be D max , then the remaining energy of the UAV at flight duration t is expressed as:

[0019] E t (t) = E max - t×P(V)

[0020] where represents the maximum energy limit of the UAV:

[0021]

[0022] Specifically, the step S200 includes:

[0023] Establish a UAV swarm composed of multi-rotor UAVs, and establish non-overlapping convex polygon regions distributed on the map divided by a command post or base station. Assume that there are no obstacles higher than the minimum flight altitude of the UAVs inside the regions. The positions, shapes, and sizes of each search region i, i ∈ [M] are represented by P i = {p ij | j = 1, 2,..., v i}, where p ij represents the j-th vertex of the i-th region, is the total number of vertices of region i. The mission objective is: all UAVs start from the same temporary command post or base station, fly to each convex polygon region for search, and each region can only be searched once. The real-time data is transmitted back to the command post or base station, and the UAV swarm returns to the command post or base station after all regions have been searched.

[0024] Model MACPP as a high-dimensional discrete optimization problem and decompose it into two sub-problems: the inter-region access order MTSP and the sub-region coverage CPP.

[0025] Specifically, the step S300 includes:

[0026] Construct a multi-agent reinforcement learning algorithm to solve the MTSP problem. MTSP mainly focuses on the region access order of the UAVs and abstracts it into an observable Markov model under multi-agents: the environment of the multi-agent reinforcement learning algorithm can be composed of a five-tuple where the state of the agent is represented as where is the private observation of the UAV. represents the action set of the agent at time t, is the action taken by the nth UAV. At each time slot t, each UAV obtains its own executes its own action and receives its own reward and enters the next observation If the UAV runs out of power, then done = True, and a new action is reselected. The specific definition is as follows:

[0027] State space: The state of each UAV includes:

[0028] 1) E n [t], n ∈ N: The current remaining energy of the UAV.

[0029] 2) includes the states q0 of M search areas to be searched and the command post or base station, designed as a one-hot code. The area is set to 0 after being selected, otherwise it is 1.

[0030] Action space: The action space of each UAV n is defined as A t = {a t , stop}, where a t = {M + 1}, M + 1 represents the set of all search areas to be searched and a command post or base station. Set a t in the form of a one-hot encoding, indicating that the agent can choose the area to visit, including the command post or base station. Each area corresponds to an index. When a certain area is selected, the corresponding index position is set to 1, and the rest of the positions are set to 0. stop means that at the current time step, the UAV is still performing the search task within the area.

[0031] Reward function: Introduce the binary variable x ij , i, j ∈ [M] ∪ {0} indicates that if a certain UAV visits the next area j immediately after visiting area i, then x ij = 1, otherwise it is 0. Introduce the variable indicating the coverage path of area i. The UAV can completely cover area i along f i , while F i represents the set of all internal paths of area i. Where k represents the internal path node of the coverage path, is the total number of internal nodes obtained by calculating area i. Connect the solved x ij with the entrance f i and the exit in sequence, and then connect back to the base P0 to obtain a complete UAV path. Therefore, a method for calculating the reward function is proposed:

[0032] R = -(αr1 + r2 + r3)

[0033] wherein represents the total distance of the current movement of the UAV, and dis(a, b) represents the Euclidean distance from area a to area b, x ij , i, j ∈ [M] ∪ {0} means that if a certain UAV visits location j (which can be the starting base or a certain area) immediately after visiting location i (which can also be another area or return to the base), then x ij = 1, otherwise it is 0, and α is an adjustment factor. And calculates the internal coverage path distance of each area, and f ki represents the exit k of area i.

[0034] r2 represents the penalty for the UAV's energy. When the UAV falls due to exceeding the energy limit and cannot fly, a penalty is imposed. r2 is calculated as follows:

[0035]

[0036] r3 represents the preemption penalty caused by unreasonable task allocation during the UAV training, specifically as follows:

[0037]

[0038] Based on the above observable Markov framework, a multi-agent reinforcement learning algorithm framework is established to solve the MTSP. In the present invention, this method is named the multi-agent reinforcement learning framework with cross-layer connection under energy constraint, hereinafter referred to as CLMPO for short. Because the MACPP problem has the characteristic of unknown regional distance and needs to optimize both the internal path coverage and the regional access order at the same time, traditional multi-agent reinforcement learning frameworks including the MAPPO algorithm cannot be directly used and need to be improved. In the present invention, CLMPO is improved based on the MAPPO algorithm framework.

[0039] Specifically, the improvement of CLMPO based on the MAPPO algorithm framework includes:

[0040] CLMPO has N agents deployed in the same temporary command post or base station, representing the UAVs about to execute tasks. These agents can make decisions by themselves and complete tasks. Each agent interacts with its corresponding UAV through actions selected from a series of discrete action sets A = {a t , t ∈ T}. In addition, the agent interacts with the environment represented by the state set S = {s t , t ∈ T}, and calculates the current action through the cross-layer connection network. Where the state is represented as The action is defined as wherein is the target area selected by the UAV. At each time slot t, each agent n obtains Select to go to a target area Then, after the agent enters the area, it executes the back-and-forth path algorithm (step S301) and calculates the reward Finally, the UAV environment updates the current state s t and transitions to the new state s t+1 . In the entire CLMPO, each agent n, n ∈ N independently maintains an experience pool B n , a critic network V(s t |θ V ) and an actor network π(a t |s t , θ π ).

[0041] Specifically, the calculation via the neural network includes:

[0042] To reduce the computational complexity of the algorithm to adapt to the limited computing power of the UAV, it is proposed to use a cross-layer connection network to replace the original fully connected network of the MAPPO algorithm, and the critic network V(s t |θ V ) and the actor network π(a t |s t , θ π ) of each agent have the same network structure and both use a cross-layer connection network. The calculation method of the cross-layer connection network is as follows:

[0043]

[0044] Among them, θ is the neural network parameter matrix, represents the nth parameter of the i-th layer, θ i+1 is obtained from θ i by calculating the weight matrix and the bias matrix ζ i .

[0045] Specifically, the execution of the back-and-forth path algorithm includes:

[0046] To obtain the internal path of each area i, the back-and-forth path algorithm (BFP) is used to solve the CPP of the target area. BFP selects the longest side within the area i as the reference and determines the direction perpendicular to the longest side as the direction of the scan line. After the direction of the scan line is given, the length W of the vertex with the farthest distance from the longest side in the direction of the scan line is calculated. Based on W, the number of internal flight paths required to completely cover the area can be obtained from It is calculated that w represents the scanning width of the drone. To ensure the stability and non - repetition of the data collected by the drone, the distance between each flight path needs to be consistent. Therefore, the width of the first and last cells is set to w, and the width of the middle part is set to d to maintain the constant distance between the flight paths, where d is expressed as:

[0047]

[0048] Then, find the shortest line segment that can completely cover each sub - region for each sub - region respectively. Finally, connect the line segments of adjacent regions in sequence to obtain the internal flight path of the drone, and calculate the

[0049] Different from traditional path - planning methods and heuristic algorithms, in the reinforcement learning framework, each action of the agent will significantly affect the overall quality of the path in each round of training. Therefore, it is also necessary to improve the interaction process of the agent (step S303).

[0050] Specifically, improving the interaction process of the agent includes:

[0051] For the longest side of each region, there are four possible BFP routes, that is, the total length of the internal path remains unchanged, but the combination of entrances and exits is different. Therefore, only need to calculate the internal path length L once, and introduce a variable E xy representing the combination of entrances and exits, where x ∈ {1, 2, 3, 4}, y ∈ {0, 1}, x represents the four possible BFP routes in region i, y = 0 represents the entrance under this BFP route, and y = 1 represents the exit under this BFP route. Based on the above, introduce an exploration probability ε that changes with the interaction frequency for selecting entrances and exits:

[0052]

[0053] Among them, let the current position of drone i, i ∈ [N] be U i , then D i = dis(U i , E x0 ) + L represents the distance generated when drone i selects a specific region, and dis(U i , E x0 ) represents the distance for the drone to reach the target region.

[0054] The exploration probability ε is expressed as:

[0055] ε(t)= ε min +(εmax - ε min )·e -Δepoch

[0056] where ε minis the minimum value of the final iteration; ε max is the initial exploration value; epoch is the global step counter, which increases with the number of training times; Δ is the decay factor.

[0057] A larger ∈ value can encourage the agent to try different combinations of entrances and exits, thus potentially discovering better strategies and planning more effective paths. The exploration probability ∈ can be expressed as:

[0058] ε(t) = ε min +(ε max -ε min )·e -kn

[0059] where ε min is the minimum value of the final iteration; ε max is the initial exploration value; n is the global step counter, which increases with the number of training times; k is the decay factor.

[0060] In the initial stage of training, due to the environmental location, a larger ε encourages the UAV to explore and master skills through continuous trial and error. On the contrary, in the later stage, after the UAV has learned certain strategies and become familiar with the environment, ε decreases continuously with the number of rounds to determine the selected strategy. Description of the Drawings

[0061] Figure 1 is the schematic diagram of MACPP of the present invention;

[0062] Figure 2 is the framework diagram of the round-trip path algorithm of the present invention;

[0063] Figure 3 is the framework diagram of the CLMPO algorithm of the present invention;

[0064] Figure 4 is the flowchart of executing MACPP of the present invention. Detailed Embodiments

[0065] For a more detailed description of the present invention and for the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the description of the drawings and embodiments. The embodiments in this part are used to explain and illustrate the present invention for the purpose of understanding, and do not limit the present invention.

[0066] Embodiment 1: A multi-UAV cooperative coverage path planning method based on reinforcement learning, comprising the following steps:

[0067] As Figure 1 , establish a MACPP model, including determining the UAV energy consumption model and the mathematical definition of the area to be searched.

[0068] The UAV energy consumption model is set as follows:

[0069] At the start of the mission, multiple drones with sufficient power are deployed to the search area, and it is necessary to ensure that the search mission is completed before the drones run out of power. Assume that the drones fly at a constant speed V and can turn with any radius of curvature at speed V. It should be noted that the energy consumption of the wireless communication for transmitting information back to the command post is relatively small. Therefore, in this invention, only the flight energy consumption of the drones is concerned, which is expressed as:

[0070]

[0071] Among them, represents the blade profile power of the drone, is the induced power in the hover state, V t is the tip speed of the rotor blade; represents the induced power of the drone, is the blade profile power in the hover state, ι0 is the average induced velocity of the rotor; represents the parasite power of the drone, Ψ is the ratio of the blade area to the disk area, Υ represents the fuselage drag ratio, Ξ represents the air density, and Ω represents the rotor disk area.

[0072] Let the farthest flight distance of the drone be D max , then the remaining energy of the drone at flight duration t is expressed as:

[0073] E t (t) = E max - t × P(V)

[0074] Among them, represents the maximum energy limit of the drone:

[0075]

[0076] The specific parameters used in the drone energy consumption model are as follows:

[0077]

[0078]

[0079] The mathematical definition of the area to be searched is as follows:

[0080] Establish a drone swarm consisting of multi-rotor drones, and establish non-overlapping convex polygon areas (areas to be searched) distributed on the map by the command post or base station. Assume that there are no obstacles higher than the minimum flight altitude of the drones inside the areas. The position, shape, and size of each search area i, i ∈ [M] are determined by P i = {p ij |j = 1, 2,..., vi} indicates, where p ij represents the j-th vertex of the i-th region, is the total number of vertices of region i.

[0081] Figure 1 The complete MACPP process is shown. That is, all drones start from the same temporary command post or base station, fly to each convex polygon region for search, and each region can only be searched once. Then, the real-time data is transmitted back to the command post or base station, and after all regions are searched by the drone swarm, they return to the command post or base station.

[0082] Furthermore, MACPP is an NP-hard problem and cannot be directly solved by other linear planners, etc. The problem needs to be decomposed. In the present invention, MACPP is decomposed into two sub-problems: the inter-region access order (MTSP) and the sub-region coverage (CPP).

[0083] Furthermore, the multi-agent reinforcement learning algorithm and the round-trip path algorithm are respectively used to solve the MTSP and CPP problems.

[0084] A multi-agent reinforcement learning algorithm is constructed to solve the MTSP problem. MTSP mainly focuses on the region access order of drones and abstracts it into an observable Markov model under multiple agents: the environment of the multi-agent reinforcement learning algorithm can be composed of a five-tuple where the state of the agent is represented as where is the private observation of the drone. represents the action set of the agent at time t, is the action taken by the n-th drone. At each time slot t, each drone obtains its own executes its own action and receives its own reward and enters the next observation If the drone runs out of power, then done = True, and a new action is reselected. The specific definition is as follows:

[0085] State space: The state of each drone includes:

[0086] 1) E n [t], n ∈ N: The current remaining energy of the drone.

[0087] 2) includes the states q0 of M regions to be searched and the command post or base station, which is designed as a one-hot code. After a region is selected, it is set to 0, otherwise it is 1.

[0088] Action space: The action space of each drone n is defined as At = {a t , stop}, where a t = {M + 1}, and M + 1 represents the set of all areas to be searched and a command post or base station. Set a t in the form of a one - hot encoding, indicating the areas that the agent can choose to visit, including the command post or base station. Each area corresponds to an index. When a certain area is selected, the corresponding index position is set to 1, and the rest are set to 0. stop means that at the current time step, the UAV is still in the area and performing a search task.

[0089] Reward function: Introduce a binary variable x ij , where i, j ∈ [M] ∪ {0} means that if a certain UAV visits area j immediately after visiting area i, then x ij = 1, otherwise it is 0. Introduce a variable representing the coverage path of area i. The UAV can completely cover area i along f i , while F i represents the set of all internal paths of area i. Where k represents the internal path node of the coverage path, is the total number of internal nodes obtained by calculating area i. Connect the solved x ij with the entrance f i and the exit in order, and then connect back to the base P0, then a complete UAV path can be obtained. Therefore, a reward function calculation method is proposed:

[0090] R = -(αr1 + r2 + r3)

[0091] where represents the total distance that the UAV moves in this action, dis(a, b) represents the Euclidean distance from area a to area b, x ij , where i, j ∈ [M] ∪ {0} means that if a certain UAV visits position j (which can be the starting base or a certain area) immediately after visiting position i (which can also be another area or return to the base), then x ij = 1, otherwise it is 0, and α is a regulation factor. And calculates the internal coverage path distance of each area, and f ki represents the exit k of area i.

[0092] r2 represents the penalty for the UAV's energy. When the UAV falls due to exceeding the energy limit and cannot fly, a penalty is imposed. r2 is calculated as follows:

[0093]

[0094] r3 represents the preemption penalty caused by unreasonable task allocation during UAV training, which is specifically as follows:

[0095]

[0096] However, considering that in practical applications, when the number of agents and the number of regions both increase, the architecture of the multi-agent reinforcement learning method has a large number of parameters, which will greatly reduce the model operation speed. Therefore, a feasible solution is to modify the fully connected network of each agent, improve the fully connected network used in traditional MAPPO, and design a cross-layer connection structure.

[0097] Furthermore, the cross-layer connection structure is as follows:

[0098] The network structures of the critic network V(s t |θ V ) and the actor network π(a t |s t , θ π ) of each agent are the same, and both use a cross-layer connection network. The calculation method of the cross-layer connection network is as follows:

[0099]

[0100] Among them, θ is the neural network parameter matrix, represents the nth parameter of the i-th layer, and θ i+1 is obtained from θ i by calculating the weight matrix and the bias matrix ζ i .

[0101] Specifically, the advantage of the cross-layer connection structure is that it minimizes the neuron parameters required by the network without losing important features, reduces the computational amount, and at the same time introduces new information into the last layer of the network, preventing overfitting and enhancing the robustness to a certain extent.

[0102] Secondly, the CPP problem mainly focuses on the optimal path selection for achieving full coverage of the region at the minimum cost. As Figure 2 shown, it is proposed to use the Back and Forth Path algorithm (BFP) to solve the CPP problem. To obtain the internal path of each region i, BFP selects the longest side within region i as a reference and determines the direction perpendicular to the longest side as the direction of the scan line. After the direction of the scan line is given, the length W of the vertex with the farthest distance from the longest side in the direction of the scan line is calculated. Based on W, the number of internal flight paths required to fully cover the region can be obtained from It is calculated that w represents the scanning width of the drone. To ensure the stability and non - repetition of the data collected by the drone, the distance between each flight path needs to be kept consistent. Therefore, the width of the first and last cells is set to w, and the width of the middle part is set to d to maintain the constant distance between flight paths, where d is expressed as:

[0103]

[0104] Then, find the shortest line segment that can completely cover each sub - region for each sub - region. Finally, connect the line segments of adjacent regions in sequence to obtain the internal flight path of the drone, and calculate the

[0105] Different from traditional path - planning methods and heuristic algorithms, in the reinforcement learning framework, each action of the agent will significantly affect the overall quality of the path in each round of training. Therefore, it is also necessary to improve the interaction process of the agent.

[0106] Furthermore, to improve the interaction process of the agent, for the longest side of each region, there are four possible BFP routes, that is, the total length of the internal path remains unchanged, but the combination of entrances and exits is different. Therefore, only need to calculate the internal path length L once, and introduce a variable E xy representing the combination of entrances and exits, where x ∈ {1, 2, 3, 4}, y ∈ {0, 1}, x represents the four possible BFP routes in region i, y = 0 represents the entrance under this BFP route, and y = 1 represents the exit under this BFP route. Based on the above, introduce an exploration probability ε that changes with the interaction frequency for selecting entrances and exits:

[0107]

[0108] Among them, let the current position of drone i, i ∈ [N] be U i , then D i = dis(U i , E x0 ) + L represents the distance generated when drone i selects a specific region, and dis(U i , E x0 ) represents the distance for the drone to reach the target region.

[0109] The exploration probability ε is expressed as:

[0110] ε(t)= ε min +(ε max - ε min )·e -Δepoch

[0111] where ε min is the minimum value of the final iteration; ε maxis the initial exploration value; epoch is a global step counter that increases with the number of training times; Δ is the decay factor.

[0112] In the initial stage of training, due to the environmental location, a larger ε encourages the drone to explore and master skills through continuous trial and error. On the contrary, in the later stage, after the drone has learned certain strategies and is familiar with the environment, ε decreases continuously with the rounds to determine the selected strategy.

[0113] Furthermore, as Figure 3 , the multi-agent reinforcement learning algorithm and the round-trip path algorithm are combined and named the multi-agent reinforcement learning framework with cross-layer connection under energy constraint (CLMPO). The specific process includes:

[0114] In the entire CLMPO, each agent n, n ∈ N independently maintains an experience pool B n , a critic network V(s t |θ V ) and an actor network π(a t |s t ,θ π ). There are N agents in the CLMPO deployed in the same temporary command post or base station, representing the drones about to perform tasks. These agents can make decisions on their own and complete tasks. Each agent interacts with its corresponding drone through actions selected from a series of discrete action sets A = {a t , t ∈ T}. In addition, the agent interacts with the environment represented by the state set S = {s t , t ∈ T} and calculates the current action through the cross-layer connection network. The state is represented as The action is defined as where is the target area selected by the drone. At each time slot t, each agent n obtains selects to go to a target area Then, after the agent enters the area, it executes the round-trip path algorithm (step S301) and calculates the reward Finally, the drone environment updates the current state s t and transitions to the new state s t+1 . Until all the areas to be searched are completed, the drone returns to the base station to complete this task.

[0115] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A multi-UAV collaborative coverage path planning method based on reinforcement learning, characterized in that It includes the following steps: Step S100: Establish a UAV energy consumption model; Step S200: Establish a multi-agent area coverage path model; Step S300: Establish a multi-agent reinforcement learning algorithm framework; Step S301: Plan the area access order of the UAV swarm. The UAV selects an area to work. If the UAV's energy is not exhausted in the current selection, go to Step S302. If the UAV's energy is exhausted in the current selection, re-plan the area access order of the UAV swarm, and the UAV re-selects an area to work; Step S302: Based on the round-trip path algorithm, plan the internal path of the area currently selected by the UAV. If all areas have not been searched, go to Step S303; if all areas have been searched, go to Step S400; Step S303: The entrance and exit exploration factor adjusts the exploration probability, optimizes the path, and then returns to Step S301; Step S400: The UAV returns to the base station to complete this task.

2. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 1, wherein The said Step S100 includes: At the start of the task, multiple UAVs with sufficient power are deployed to the search area. Assume the UAVs fly at a constant flight speed V and can turn at any curvature radius at speed V. The flight energy consumption of the UAV is expressed as: Among them, represents the blade profile power of the UAV, is the induced power in the hover state, V t is the tip speed of the rotor blade; represents the induced power of the UAV, is the blade profile power in the hover state, ι0 is the average induced velocity of the rotor; represents the parasite power of the UAV, Ψ is the ratio of the blade area to the disk area, Υ represents the fuselage drag ratio, Ξ represents the air density, and Ω represents the rotor disk area; Let the maximum flight distance of the drone be D max , then the remaining energy of the drone under the flight duration TH is expressed as: E t E(t) = max -TH × P(V) Among them, E max represents the maximum energy limit of the UAV:

3. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 1, wherein, The said Step S200 includes: Establish a drone swarm composed of N multi-rotor drones, and establish M non-overlapping convex polygon regions on the map divided by a command post or base station, that is, the areas to be searched. Assume that there are no obstacles higher than the minimum flight altitude of the drones inside the regions. The positions, shapes, and sizes of each search area i, i ∈ [M] are represented by P i ={p ij |j = 1, 2,..., v i}, where p ij represents the j-th vertex of the i-th region. is the total number of vertices of region i. The mission objective is: all drones start from the same temporary command post or base station, fly to each convex polygon region for search, and each region can only be searched once, and the real-time data is transmitted back to the command post or base station. After the drone swarm has searched all regions, it returns to the command post or base station; Model MACPP as a high-dimensional discrete optimization problem and decompose it into two sub-problems: the inter-area access order MTSP and the sub-area coverage CPP.

4. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 3, characterized in that The said Step S300 includes: Construct a multi-agent reinforcement learning algorithm to solve the MTSP problem. MTSP focuses on the regional access order of drones and abstracts it into an observable Markov model under multi-agent. Let the set of execution times for the entire task be T, and N agents are deployed on the command post or base station, representing the drones about to depart. The environment of the multi-agent reinforcement learning algorithm consists of a five-tuple which, at the current time t ∈ T, the state of the agent is represented as where is the private observation of the drone, represents the set of actions of all agents at the current time t, is the action taken by the nth drone, where n ∈ N. At the current time t, each drone obtains its own executes its own action and receives its own reward and enters the next observation If the drone runs out of power, then done = True, and a new action is reselected. The specific definition is as follows: State space: State of each UAV n including: 1) E n [t], n ∈ N: The current remaining energy of the drone; 2) Including M areas to be searched and the state q0 of the command post or base station, which is designed as a one-hot code. After an area is selected, it is set to 0, otherwise it is 1; Action space: The action space of each drone \(n\) is defined as \(\mathcal{A}\) t =\{a t , stop\}, where \(a t = \{M + 1\}\), \(M + 1\) represents the set of all areas to be searched and a command post or base station. Set \(a t in a one - hot encoding form, indicating the areas that the agent can choose to visit, including the command post or base station. Each area corresponds to an index. When choosing a certain area, set the corresponding index position to 1 and the rest to 0. stop means that at the current time step, the drone is still in the area and performing the search task; Reward function: Introduce a binary variable x ij , where i, j ∈ [M] ∪ {0} indicates that if a drone visits the next area j immediately after visiting area i, then x ij = 1, otherwise it is 0. Introduce a variable to represent the coverage path of area i. The drone can completely cover area i along f i , while F i represents the set of all internal paths of area i, where k represents the internal path nodes of the coverage path, is the total number of internal nodes obtained by calculating area i. Connect the solved x ij with the entrance f i and the exit in sequence, and then connect back to the base P0 to obtain a complete drone path. Therefore, a reward function calculation method is proposed: R=-(αr1+r2+r3) Among them represents the total distance of the UAV's movement in this action, dis(a, b) represents the Euclidean distance from area a to area b, and calculates the internal coverage path distance of each area, and α is an adjustment factor; r2 represents the penalty for the UAV's energy. When the UAV falls due to exceeding the energy limit and cannot fly, a penalty is imposed. r2 is calculated as follows: r3 represents the preemption penalty caused by unreasonable task allocation during the UAV training, specifically as follows: Based on the above observable Markov framework, establish a multi-agent reinforcement learning algorithm framework to solve MTSP, and name this method the multi-agent reinforcement learning framework with cross-layer connection under energy constraint, abbreviated as CLMPO.

5. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 4, wherein, The said CLMPO is improved based on the MAPPO algorithm framework: There are N agents in the CLMPO deployed in the same temporary command post or base station, representing the drones that are about to execute tasks. These agents can make decisions on their own and complete tasks. Each agent interacts with its corresponding drone through actions selected from a series of discrete action sets A = {a t , t ∈ T}. In addition, the agent interacts with the environment represented by the state set S = {s t , t ∈ T}, and calculates the current action through the cross-layer connection network. At the current moment t, each agent n obtains selects to go to a target area Then, after the agent enters the area, it executes the round-trip path algorithm and calculates the reward Finally, the drone environment updates the current state s t and transitions to the new state s t+1 . In the entire CLMPO, each agent n, n ∈ N independently maintains an experience pool B n , a critic network V(s t |θ V ) and an actor network π(a t |s t , θ π ).

6. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 5, wherein The calculation of the cross-layer connection network includes: Replace the original fully connected network of the MAPPO algorithm with a cross-layer connection network, and the critic network V(s t |θ V ) and the actor network π(a t |s t , θ π ) of each agent have the same network structure, both using a cross-layer connection network. The calculation method of the cross-layer connection network is as follows: Among them, is the neural network parameter matrix, represents the nth parameter of the ith layer, which is obtained by calculating the weight matrix and the bias matrix ζ i .

7. The multi-UAV cooperative coverage path planning method based on reinforcement learning according to claim 5, characterized in that, When the agent enters the area and executes the round-trip path algorithm, it includes: To obtain the internal path of each region i, the back-and-forth path algorithm BFP is used to solve the CPP of the target region. BFP selects the longest side within region i as a reference and determines the direction perpendicular to the longest side as the direction of the scanning line. After the direction of the scanning line is given, the length W of the vertex with the farthest distance from the longest side in the direction of the scanning line is calculated. Based on W, the number of internal flight lines required to completely cover the region can be obtained by calculated, where w represents the scanning width of the UAV. The widths of the first and last cells are set to w, while the width of the middle part is set to d to maintain a constant distance between flight lines. Here, d is expressed as: Then, find the shortest line segment that can completely cover each sub-region for each sub-region respectively, that is, calculate the internal nodes of each sub-region Then make connections. Finally, connect the line segments of adjacent regions in sequence to obtain the internal flight path of the UAV, and calculate To improve the quality of the path selected in each round of training, the interaction process of the agent is improved.

8. The multi-UAV collaborative coverage path planning method based on reinforcement learning according to claim 7, characterized in that The improvement of the interaction process of the agent includes: For the longest side of each region, there are four possible BFP routes, that is, the total length of the internal path remains unchanged, but the combination of the entrance and exit is different. Therefore, it is only necessary to calculate the internal path length L once and introduce a variable E xy It represents the combination of the entrance and exit, where x ∈ {1, 2, 3, 4}, y ∈ {0, 1}, x represents the four possible BFP routes in region i, y = 0 represents the entrance under this BFP route, and y = 1 represents the exit under this BFP route. Based on the above, an exploration probability ε that changes with the interaction frequency for selecting the entrance and exit is introduced: Among them, let the current position of the UAV i, where i ∈ [N], be U i , then D i = dis(U i , E x0 ) + L represents the distance generated by the UAV selecting a specific area, and dis(U i , E x0 ) represents the distance of the UAV reaching the target area; The exploration probability ε is expressed as: ε(t) = ε min +(ε max -ε min )·e -△epoch where ε min is the minimum value of the final iteration; ε max is the initial exploration value; epoch is a global step counter that increases as the number of training times increases; Δ is the decay factor.