A Coverage Path Planning Method for Multiple Unmanned Aerial Vehicles in the Inspection of Abnormal Environments of Dams

Through the DRL-based multi-UAV coverage path planning model, combined with TP and MADDPG structures, the communication interference and energy consumption problems of drones in abnormal dams are solved, and drone inspection tasks with larger coverage and lower energy consumption are achieved.

CN119440050BActive Publication Date: 2025-07-22HOHAI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411558437.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-07-22
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

In the abnormal environment of the dam, the existing drone coverage path planning model cannot effectively solve the problems of wireless communication interference and energy consumption, resulting in insufficient coverage and excessive energy consumption, and the inability to achieve efficient inspection tasks.

Method used

The multi-UAV coverage path planning model based on DRL is adopted, combined with TP and MADDPG structures, and the communication media between drones is designed through ant colony pheromone-inspired, and a reward function is introduced to optimize coverage and energy consumption, and path planning is used to use MADDPG.

Benefits of technology

A drone inspection with greater coverage and lower energy consumption is achieved in abnormal dam environments, improving the reliability and efficiency of coverage path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119440050B_ABST
    Figure CN119440050B_ABST
Patent Text Reader

Abstract

The present invention discloses a coverage path planning method for multiple unmanned aerial vehicles (UAVs) in the inspection of abnormal environments of dams, and formally defines the coverage path planning problem for multiple UAVs in the inspection of abnormal environments of dams: Firstly, the abnormal environment is defined, that is, the dam area after a disaster occurs. Secondly, the coverage path planning problem for multiple UAVs in the inspection of abnormal environments of dams is defined as an optimization problem, and the optimization objectives are respectively the coverage rate of multiple UAVs in the abnormal environment and the energy consumption; Model design: For the optimization problem defined in the previous step, a coverage path planning model for multiple UAVs based on deep reinforcement learning (DRL) is constructed; Update the parameters of the coverage path planning model for multiple UAVs and evaluate the performance of the model. Compared with the prior art, the present invention has the advantages of good practicability, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a coverage path planning method for multiple unmanned aerial vehicles (UAVs) in the inspection of abnormal environments of dams, which is used for the coverage path planning when multiple UAVs perform inspection tasks in the dam area after natural disasters, and belongs to the field of intelligent water conservancy technology. Background Technique

[0002] Inspection is of great significance for the safe operation of dams. Comprehensive and effective dam inspection can timely detect potential structural problems such as cracks and leaks, and at the same time, monitor the impact of natural disasters on the reservoir ecological environment in real time. Recently, the rapid development of UAV technology has provided new ideas for dam inspection. The cooperation of multiple UAVs has the advantages of low cost, small size, and strong mobility, and has the potential to achieve comprehensive and efficient coverage in dam inspection. The key technology for the cooperation of multiple UAVs is the coverage path planning of UAVs.

[0003] In recent years, a large number of scholars have studied the coverage path planning of UAVs. The coverage path planning models of UAVs are mainly divided into two categories. The first category is the non-self-organizing control model, including cooperative control and stochastic control. Cooperative control is a strategy that enables multiple UAVs to work together to achieve a common goal (such as covering an area). For example, to solve the cooperation mechanism and distributed strategy problems in UAV cooperative search, some researchers have developed a distributed online heuristic method. However, the robustness of cooperative control is limited. If a small number of UAVs in the UAV cluster are damaged, the entire UAV network needs to be redeployed. Stochastic control attempts to cover the target area by randomly selecting path points, thereby reducing the computational overhead and time delay. For example, some scholars have designed a strategy that integrates random walk, artificial intelligence, and coverage control to deploy multiple helicopter-type UAVs in the target area. Compared with cooperative control, stochastic control further improves the robustness, but it exposes the problems of too long coverage time and overlapping coverage.

[0004] To address these issues, a second type of UAV coverage path planning model has been proposed, known as the self-organizing control model. This model can achieve area coverage tasks based on the decisions of autonomous UAVs. Specifically, researchers designed a multi-agent DRL (Deep Reinforcement Learning) model to monitor industrial production. This model can cover a larger area and minimize overlapping coverage as much as possible. To meet the long-term knowledge and seamless coverage requirements of IoT nodes, researchers developed a cooperative scheme for multi-rechargeable UAVs, achieving system energy consumption optimization. In engineering inspections, the self-organizing control model has also been widely applied. For example, researchers used a heat equation-driven area coverage method to achieve trajectory planning for UAVs to inspect three-dimensional infrastructures such as wind turbines and bridges. This method uses a potential field and a distance field to generate the flight trajectory of UAVs and prevent collisions between UAVs respectively.

[0005] However, during dam inspections, due to the existence of natural disasters, the dam area inevitably faces an abnormal environment. The abnormal environment of the dam gives rise to two problems: 1) Natural disasters such as thunderstorms, strong winds, and earthquakes can interfere with wireless communication signals, leading to communication obstacles between UAVs; 2) In the abnormal environment of the dam, the coverage range and energy consumption of UAVs are non-negligible factors. For example, a larger UAV coverage range supports obtaining more disaster information. Additionally, natural disasters can damage the energy supply devices of UAVs. The low energy consumption of UAVs ensures the execution of long-term inspection tasks with limited energy supply. Therefore, it is very necessary to achieve the joint optimization of coverage range and energy consumption. Currently, most self-organizing control models aim to explore dam inspections in normal environments. When natural disasters occur in the dam area, these models do not take into account the above two problems. This is the motivation for the research of this invention. Summary of the Invention

[0006] Object of the Invention: Aiming at the problems and deficiencies of the existing technology, the present invention provides a coverage path planning method for multi-UAVs to inspect in the abnormal environment of dams, which has good performance and strong practicability.

[0007] Technical Solution: A coverage path planning method for multi-UAVs to inspect in the abnormal environment of dams includes the following steps:

[0008] Step 1: Formally define the coverage path planning problem for multi-UAVs to inspect in the abnormal environment of dams;

[0009] Step 2: Construct a multi-UAV coverage path planning model based on DRL;

[0010] Step 3: Update the model parameters.

[0011] Preferably, the specific content of the above Step 1 is:

[0012] The formal definition of the coverage path planning problem for multiple UAVs in the inspection of the abnormal environment of a dam can be divided into two steps, namely, defining the abnormal environment and defining the optimization problem.

[0013] More preferably, the definition of the abnormal environment is specifically as follows:

[0014] Dam inspection is of great significance for the safe operation of the dam. Especially after natural disasters such as thunderstorms, strong winds, and earthquakes occur in the dam area, comprehensive and effective dam inspection can timely detect potential structural problems (such as cracks and leaks) and take corrective measures, while monitoring the impact of natural disasters on the reservoir ecological environment in real time. In the present invention, the dam area after a natural disaster is called the abnormal environment of the dam. Specifically, the abnormal environment of the dam is jointly defined by the following components, namely, the dam area including boundaries and obstacles, the disaster points in the dam area, the severity of each disaster point, static sensors (deployed before the disaster occurs), multiple UAVs, and the coverage area of each UAV.

[0015] More preferably, the definition of the optimization problem is specifically as follows:

[0016] After defining the abnormal environment of the dam, the present invention defines the coverage path planning problem for multiple UAVs in the inspection of the abnormal environment of the dam as an optimization problem. The optimization objectives include the coverage rate of multiple UAVs and the energy consumption of multiple UAVs, aiming to minimize the energy consumption of multiple UAVs and maximize the coverage area of multiple UAVs.

[0017] Preferably, step 2 is specifically as follows:

[0018] According to the principles of deep learning and reinforcement learning, the present invention constructs a multi-UAV coverage path planning model based on DRL. This model can realize the joint optimization of the coverage rate and energy consumption of multiple UAVs during the inspection of the abnormal environment of the dam.

[0019] More preferably, the multi-UAV coverage path planning model based on DRL mainly consists of two parts, namely, TP (Trace Pheromone, tracking pheromone, abbreviated as TP) and MADDPG (Multi-Agent Deep Deterministic Policy Gradient, multi-agent deep deterministic policy gradient, abbreviated as MADDPG). TP establishes a communication medium between UAVs and realizes the indirect communication of UAVs in the abnormal environment of the dam. At the same time, TP can effectively manage the path information of UAVs and simplify the complex interaction between UAVs into a low-dimensional representation. Then MADDPG is used to realize the coverage path planning of multiple UAVs. In MADDPG, the reward function is used to simultaneously achieve the goals of maximizing the coverage rate of multiple UAVs and minimizing the energy consumption of multiple UAVs.

[0020] More preferably, the MADDPG is specifically as follows:

[0021] Each drone has an independent MADDPG, and each MADDPG includes two networks, namely an actor network and a critic network. The actor network consists of two FC layers (Fully Connected Layers, simply referred to as FC layers) and a softmax layer. This network aims to generate the probability of an action (such as flight direction and flight angle) that a drone can execute. Based on these probabilities, the drone selects the next action to be executed. The critic network consists of three FC layers and is mainly used to generate Q values.

[0022] Preferably, step 3 is specifically as follows:

[0023] MADDPG updates the parameters by separately updating the actor network and the critic network for each drone. Each drone has an independent actor-critic network pair, where the actor network is updated according to the gradient of the expected return, and the critic network updates the Q value according to the Bellman equation. During the training process, a centralized training method is used, that is, the critic network of each drone will consider the actions of all drones. However, the execution is decentralized, that is, each drone will act independently according to its own strategy.

[0024] Finally, by comparing the performance of the model proposed in the present invention and some advanced DRL models in the simulated abnormal dam environment, it can be found that the model proposed in the present invention has a greater drone coverage rate and lower drone energy consumption in the multi-drone coverage path planning.

[0025] A computer device, characterized in that: the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, the steps of the above-mentioned multi-drone coverage path planning method in the abnormal dam environment inspection are implemented.

[0026] A computer-readable storage medium, characterized in that: the computer-readable storage medium stores a computer program for executing the above-mentioned multi-drone coverage path planning method in the abnormal dam environment inspection.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] 1. A novel model based on DRL is proposed in the present invention. This model inherits the TP and MADDPG structures and is used for the dam inspection by multiple unmanned aerial vehicles (UAVs) in abnormal environments. Specifically, inspired by the ant colony pheromone, the present invention develops a communication scheme based on TP to control the movement of UAVs. This scheme uses TP to establish a communication medium between UAVs, realizing the indirect communication of UAVs in the abnormal dam environment. At the same time, this scheme can effectively manage the path information of UAVs, simplifying the complex interaction between UAVs into a low-dimensional representation.

[0029] 2. The present invention introduces MADDPG to perform the coverage path planning for multiple UAVs. In MADDPG, a reward function is designed to optimize the coverage rate and energy consumption of UAVs simultaneously. Moreover, the present invention introduces the penalties for overlapping coverage and the boundary constraints of the task area into the reward function, thus greatly improving the reliability of the continuous motion control decision of UAVs. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is the flow chart of the coverage path planning for multiple UAVs in the abnormal dam environment inspection in the embodiment of the present invention;

[0031] Figure 2 It is the overall structure diagram of the DRL-based multi-UAV coverage path planning model in the embodiment of the present invention. DETAILED IMPLEMENTATION MANNER

[0032] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0033] The present invention utilizes the learning ability of deep learning for big data and the decision-making ability of reinforcement learning to establish a DRL-based multi-UAV coverage path planning model. The method for the coverage path planning of DRL-based multi-UAVs in the abnormal dam environment inspection, as Figure 1 shown, includes the following steps:

[0034] Step 1: Formalize the coverage path planning problem for multiple UAVs in the abnormal dam environment inspection;

[0035] The formal definition of the coverage path planning problem for multi-UAVs in the inspection of the abnormal environment of the dam can be divided into two steps, namely, defining the abnormal environment and defining the optimization problem. First, the abnormal environment of the dam is defined. Specifically, the entire dam area is divided into a two-dimensional grid of L×W, which includes DPs (Disaster Points, abbreviated as DPs), boundaries, and obstacles. After a natural disaster occurs, the entire dam area is regarded as the mission area of the UAVs. Before the disaster occurs, some static sensors have been deployed in the dam area. Normal static sensors can collect data and then transmit this data to the base station for observation and analysis. Given P UAVs and Q static sensors, the sets of UAVs and static sensors are respectively represented as and where x p , p ∈ [1, P] and e q , q ∈ [1, Q] are the p-th UAV and the q-th static sensor respectively. Since disasters may damage the static sensors in the mission area, abnormal static sensors cannot collect reliable data or transmit data to the base station. At the same time, due to natural disasters, the communication between UAVs may be interrupted. The set of DPs is represented as where a d , d ∈ [1, D] represents the d-th DP. D represents the total number of DPs. The severity of each DP is different and is quantified as a value, called the disaster value. For each DP, the set of disaster values is represented as In addition, it is ensured that each DP is covered by at least one UAV within a certain time slot. More specifically, the coverage area of each UAV is If a DP is located within the of at least one UAV, it means that the DP has been covered by the UAV. Before performing the inspection task, all UAVs dock at the same starting position.

[0036] After defining the abnormal environment of the dam, the coverage path planning problem for multi-UAVs in the inspection of the abnormal environment of the dam is defined as an optimization problem. Specifically, the coverage rate refers to the proportion of DPs that can be covered by UAVs within a certain period of time. If a DP is within the coverage area of the UAV, it means that the DP is covered. Each DP is represented by a corresponding disaster value, and the coverage rate of the P-th UAV in the t-th generation is calculated as:

[0037]

[0038] where, represents the sum of the disaster values covered by the p-th UAV. T p represents the number of DPs within the coverage area of the p-th UAV. Represents the sum of the disaster values of all DPs within the mission area.

[0039] Achieve the highest coverage rate on the premise of the energy consumption of the drones. Therefore, it is necessary to pay attention to the energy consumption of each drone in each generation, which includes the energy requirements for data collection and continuous horizontal flight. Specifically, in the t-th generation, the energy consumption of the p-th drone is calculated as follows:

[0040]

[0041] Among them, Represents the moving distance of the p-th drone. λ d Is the energy consumption coefficient for the moving distance of the drone. λ v Represents the energy consumption coefficient when the drone collects data.

[0042] In order to balance the coverage rate of the drones and the energy consumption of the drones, considering the energy consumption of the drones, the coverage energy efficiency is designed. It represents the number of newly covered DPs by the drones within one generation. Specifically, in the t-th generation, the coverage energy efficiency Is expressed as:

[0043]

[0044] Among them, ω represents whether each DP is fairly covered. Is the incremental coverage of the p-th drone. Represents the effective incremental coverage. This is because the drones must ensure that the DPs are fairly covered. Represents the incremental energy consumption of the p-th drone. In the coverage energy efficiency, the numerator is understood as the coverage benefit of the drones, and the denominator is understood as the coverage cost of the drones. Maximizing the energy coverage efficiency can ensure the joint optimization of the coverage rate and energy consumption of the drones.

[0045] Step 2: Construct a multi-drone coverage path planning model based on DRL;

[0046] According to the principles of deep learning and reinforcement learning, construct a multi-drone coverage path planning model based on DRL. As Figure 2As shown, the model mainly consists of two parts, namely TP and MADDPG. In nature, ants release chemical pheromones during foraging. These chemical pheromones generate a pheromone map, showing the location information of food and the ant nest. Ants use the pheromone map to exchange information with each other to obtain the specific locations of food and the ant nest. Inspired by this, a TP for drones is designed. The TP can store global state information such as the flight time and location of drones. These TPs generate a trajectory tracking map of drones within the mission area. The trajectory tracking map serves as a communication medium for drones in the abnormal environment of the dam, solving the problem of communication interruption between drones caused by natural disasters. In addition, the concentration of TP shows dynamic changes. For example, the concentration of TP gradually decreases over time. At the same time, the concentration within the grid can accumulate. Within a time period, the higher the concentration in a grid, it means that more drones are patrolling in that grid. In view of this, TP can be used to reduce the overlapping coverage of drones within the mission area.

[0047] Initialization of TP: Ants use chemical pheromones to guide the foraging routes of their companions. Driven by the chemical pheromones, this invention designs TP to record the flight routes of drones, aiming to achieve two goals: 1) serving as a communication medium between drones in the abnormal environment of the dam, 2) reducing the repeated visits of drones to a grid within a time period. Specifically, in the 0th generation, the TP concentration at the starting position of the drone is initialized as:

[0048]

[0049] where (x,y) represents the starting coordinates of the drone. All drones have the same starting position. Through the trajectory tracking map, each drone can sense the concentration of TP. The grids with high concentration are regarded as the negative impacts of adjacent drones, that is, drones tend to fly towards the grids with low concentration.

[0050] Evaporation of TP: Evaporation refers to the update of the TP concentration, indicating that the concentration of the previous TP will dissipate over time. In the tth generation, the evaporated TP concentration in the ith grid is expressed as:

[0051]

[0052] where ξ, ξ∈[0,1] represents the evaporation coefficient in each generation. In the tth generation, represents the concentration of TP in the ith grid.

[0053] Aggregation of TP: When the drone arrives at the ith grid in the next generation (the (t + 1)th generation), due to the deposition of pheromones, the concentration of TP in this grid increases. At the same time, as described above, TP evaporates over time, resulting in a decrease in concentration. Therefore, in the (t + 1)th generation, the concentration of TP in the ith grid is calculated as:

[0054]

[0055] Among them, represents the concentration of TP left by the UAV in the i-th grid and is calculated as:

[0056]

[0057] where represents the concentration of TP left when the h-th UAV arrives at the i-th grid. P i is the total number of UAVs in the i-th grid.

[0058] Finally, in order to avoid excessive fluctuations or negative values in the concentration of TP, the concentration of TP is limited to the interval Therefore, is updated to:

[0059]

[0060] All in all, the present invention uses TP to mark the flight trajectory of the UAV in the mission area. The UAV selects actions according to the concentration of TP (such as selecting the next target grid).

[0061] In data collection, each UAV interacts with the environment, thus accumulating rich experience in the mission area. This process is represented as an MDP (Markov Decision Process). In addition, the UAV can choose any flight distance and flight direction. Therefore, how to obtain a multi-UAV coverage path planning that meets the above optimization objectives is an MDP problem with a continuous action space. The MADDPG based on TP has strong decision-making and perception capabilities and is suitable for solving this problem. Specifically, according to the current state in the state space S, each UAV adopts a policy to select a suitable action from the action space A, and then obtains the corresponding reward. Furthermore, in order to find actions with high rewards, random noise is introduced when performing actions, so as to ensure that the policy is more divergent. The details of the MADDPG based on TP are introduced as follows.

[0062] Design of state, action, and reward:

[0063] I. Environmental state space: The state of each UAV mainly includes the TP concentration of the grid, the coverage status of DP, and the current energy consumption of the UAV. Specifically, for the p-th UAV, the environmental state s p (t) of the t-th generation includes the following three parts:

[0064] 1) The TP concentration of the i-th grid;

[0065] 2) The coverage status of the d-th DP. If the DP is covered, the value is 1. Otherwise, the value is 0.

[0066] 3) The energy consumption of the p-th UAV.

[0067] For the p-th UAV, these three parts reflect the environmental information of the mission area. At the t-th generation, the environmental state space of all P UAVs can be expressed as:

[0068]

[0069] At each generation, each UAV obtains the local environmental state space information within its coverage area According to this information, the UAV makes a decision, that is, selects the next target grid.

[0070] II. Action Space: Using the policy Each UAV selects the action to be executed according to its current state space. The action consists of two components, namely the flight distance and flight direction of the UAV. Specifically, at the t-th generation, the action of the p-th UAV is expressed as:

[0071] 1) The flight distance of the p-th UAV. Normalize using the maximum flight distance. 0 represents that the UAV is in a hovering state, indicating that the UAV reaches the maximum flight distance.

[0072] 2) The flight angle of the p-th UAV.

[0073] The action space a(t) is the flight distance and flight direction of all P UAVs at the t-th generation, and can be expressed as:

[0074]

[0075] Since the decision variables of the UAV are continuous, the action space is defined as a continuous control task.

[0076] III. Reward: To solve the practical problems in DRL, using a suitable reward function can enable the UAV to learn an excellent policy. Therefore, to solve the multi-UAV coverage path planning problem, a new reward function is designed. Specifically, for the p-th UAV at the t-th generation, a reward is given after the UAV executes a task This reward is regarded as a state - action function, aiming to guide the p - th drone to select the optimal strategy. It consists of the following parts:

[0077] 1) If DP is included in the coverage area then the reward is obtained by the p - th drone. is equal to

[0078] 2) If the p - th drone collides with an obstacle, then a penalty whose value is related to the number of collisions is obtained.

[0079] 3) If the energy consumption of the p - th drone reaches the threshold and the drone has returned to the base, then this drone gets a reward

[0080] 4) If the p - th drone selects the wrong flight direction, it means that the drone does not choose the direction with the lowest TP concentration. At this time, the p - th drone gets a penalty

[0081] 5) The static sensors pre - deployed in the dam area can quickly collect disaster information and directly transmit it to the base station. However, after a natural disaster occurs, the static sensors may be damaged. Abnormal static sensors can cause serious consequences such as misjudgment of the disaster situation. Suppose the abnormal probability of the static sensor is δ, and δ < 0.05 indicates that the static sensor is functioning normally. At this time, if the drone collects disaster information around the normal sensor, it will waste the energy of the drone. This is because the normal static sensor can directly transmit valid data back to the base station, and the drone does not need to collect this data again. In view of this, the drone needs to be penalized. On the contrary, 0.05 <= δ < 1 means that the static sensor is abnormal.

[0082] In the t - th generation, the total reward of the p - th drone is expressed as:

[0083]

[0084] where ω represents whether each DP is fairly covered.

[0085] Actor Network: MADDPG contains two networks, namely the actor network and the critic network. In an abnormal environment, each drone uses these two networks for coverage path planning in the abnormal environment of the dam. The structure of the actor network includes two FC layers and a softmax layer. The two FC layers select ReLU as the activation function. Specifically, for the p-th drone, its state s p (t) is the input of the actor network. The two FC layers are used to extract the features in s p (t). Then, the softmax layer is used to generate the probabilities of the corresponding actions. Finally, the actor network selects the next action to be executed according to the probabilities of each action.

[0086] Critic Network: The critic network contains 3 FC layers. Specifically, for the p-th drone, the actions a(t) and states s(t) of all P drones are used as the input of the p-th critic network. The first two FC layers use ReLU as the activation function to capture the hidden features in the input. The last FC layer generates the Q value, denoted as Q(s,a).

[0087] Step 3: Update the model parameters of MADDPG in Step 2;

[0088] For the actor network, the policy set is represented as ψ = ψ1, ψ2, …, ψ P , where φ = φ1, φ2, …, φ P is the set of policy parameters of all P drones. At the beginning of each generation, each drone randomly selects samples from the experience replay buffer as a batch. Specifically, when this batch is input into the p-th actor network, the gradient of the expected reward is represented as:

[0089]

[0090] where m represents the index of the sample. Based on the observed state of policy ψ, is the expected discounted reward when performing actions (a1, a2, …, a P ). represents the gradient of the policy parameters of the p-th actor network.

[0091] After that, the weight parameters of the critic network are updated by the temporal difference error as:

[0092]

[0093] where, is the target Q value of the p-th critic network and is defined as:

[0094]

[0095] Among them, represents an approximate function. is the target policy. represents the reward value of the p-th critic network on the m-th sample. γ represents the weight coefficient. represents the policy function. Furthermore, the loss function l(φ p ) of the p-th critic network is expressed as:

[0096]

[0097] Finally, based on the actor network and the critic network, the parameters of the target network are updated. Soft update is a method for updating the target network. This update method can be expressed as:

[0098]

[0099] where τ represents the moving distance from the target network to the evaluation network during each update. φ Q are the parameters of the trained critic network. represents the policy parameters of the target network. represents the parameters of the critic network in the target network. Finally, by comparing the performance of the model proposed in the present invention and some advanced DRL models in the simulated abnormal dam environment, it can be found that the model proposed in the present invention has a higher UAV coverage rate and lower UAV energy consumption in multi-UAV coverage path planning.

[0100] This embodiment also relates to a storage medium in which the above-mentioned coverage path planning method is stored.

[0101] Obviously, those skilled in the art should understand that each step of the above-mentioned coverage path planning method for multi-UAVs in the abnormal dam environment inspection of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in the storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A coverage path planning method for multiple unmanned aerial vehicles in the inspection of abnormal environments of dams, characterized in that, It includes the following steps: Step 1: Formalize the coverage path planning problem of multiple UAVs in the inspection of the dam abnormal environment; Step 2: Construct a coverage path planning model of multiple UAVs based on DRL; Step 3: Update the model parameters; The formalization of the coverage path planning problem of multiple UAVs in the inspection of the dam abnormal environment is divided into two steps, namely, defining the abnormal environment and defining the optimization problem; The specific definition of the abnormal environment is as follows: The dam area after suffering natural disasters is called the dam abnormal environment; The entire dam area is divided into a two-dimensional grid of L×W. The dam abnormal environment is jointly defined by the following components, namely, the dam area including boundaries and obstacles, the disaster points DP in the dam area, the severity of each disaster point, static sensors, multiple UAVs, and the coverage area of each UAV; After natural disasters occur, the entire dam area is regarded as the task area of the UAVs. The static sensors can collect data and then transmit these data to the base station for observation and analysis. The severity of each DP is different and is quantified as a value, called the disaster value; Ensure that each DP is covered by at least one UAV within a certain time slot; Before performing the inspection task, all UAVs dock at the same starting position; The specific definition of the optimization problem is as follows: After defining the dam abnormal environment, the coverage path planning problem of multiple UAVs in the inspection of the dam abnormal environment is defined as an optimization problem. The optimization objectives include the coverage rate of multiple UAVs and the energy consumption of multiple UAVs, aiming to minimize the energy consumption of multiple UAVs and maximize the coverage area of multiple UAVs; Construct a coverage path planning model of multiple UAVs based on DRL, which can realize the joint optimization of the coverage rate and energy consumption of multiple UAVs during the inspection of the dam abnormal environment; The coverage path planning model of multiple UAVs based on DRL consists of two parts, namely, TP and MADDPG; TP establishes the communication medium between UAVs, realizes the indirect communication of UAVs in the dam abnormal environment, manages the path information of UAVs on the ground, and simplifies the complex interaction between UAVs into a low-dimensional representation; MADDPG is used to realize the coverage path planning of multiple UAVs; In MADDPG, the reward function is used to simultaneously achieve the goals of maximizing the coverage rate of multiple UAVs and minimizing the energy consumption of multiple UAVs; To solve the multi-UAV coverage path planning problem, a reward function is designed; for the p-th UAV in the t-th generation, a reward is given after the UAV executes a task This reward is regarded as a state-action function, aiming to guide the p-th UAV to select the optimal strategy; It consists of the following parts: 1) If DP is included in the coverage area then the reward is obtained by the p-th drone, equal to denotes the sum of the disaster values covered by the p-th drone; 2) If the p-th drone collides with an obstacle, a penalty is incurred The value is related to the number of collisions; 3) If the energy consumption of the p-th drone reaches the threshold and the drone has returned to the base, then this drone gets a reward 4) If the p-th drone selects the wrong flight direction, it means that the drone does not select the direction with the lowest TP concentration. At this time, the p-th drone receives a penalty 5) Assume that the abnormal probability of the static sensor is δ. If δ < the set value, it means that the static sensor functions normally. At this time, if the UAV collects disaster information around the normal sensor, it will waste the energy of the UAV and the UAV needs to be punished. On the contrary, 0.05 <= the set value < 1 means that the static sensor is abnormal; In the t-th generation, the total reward of the p-th UAV is expressed as:

2. The coverage path planning method for multiple unmanned aerial vehicles in dam abnormal environment inspection according to claim 1, wherein, The specific MADDPG is as follows: Each UAV has an independent MADDPG, and each MADDPG contains two networks, namely, the actor network and the critic network; The actor network consists of two FC layers and a softmax layer. This network aims to generate the probability of an action that the UAV can execute. Based on this probability, the UAV selects the next action to be executed; The critic network consists of three FC layers and is used to generate Q values; MADDPG updates parameters by separately updating the actor network and the critic network for each drone; each drone has an independent actor-critic network pair, where the actor network is updated according to the gradient of the expected return, and the critic network updates the Q-value according to the Bellman equation; during training, a centralized training method is used, that is, the critic network of each drone considers the actions of all drones; but the execution is decentralized, that is, each drone acts independently according to its own strategy.

3. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the coverage path planning method for multiple drones in the inspection of the dam abnormal environment as described in any one of claims 1-2.

4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that executes the coverage path planning method for multiple drones in the inspection of the dam abnormal environment as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle joint task allocation and flight path planning method under limited communication

    CN115981369A

  • Multi-unmanned aerial vehicle autonomous coverage method based on pheromone inspiration

    CN116643587A