Coverage Path Planning Method for Daily Inspection of Dams Based on DRL and Multiple UAVs
By applying the DRL-based EA-MATD3 model in daily inspection of dams, the task allocation and flight path of multiple UAVs are optimized, and the problem of difficulties in the existing technology is solved, and the inspection effect of efficient and low-energy consumption is achieved.
Patent Information
- Application Number
- CN202411460833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-18
AI Technical Summary
The existing multi-UAV coverage path planning algorithm fails to take into account both coverage and energy consumption during daily inspections of dams, resulting in inefficiency.
A method for daily inspection coverage path planning for multi-UAV dams based on DRL is proposed. Through the EA-MATD3 model, combined with MATD3 and stacked LSTM network, the task allocation and flight path planning of UAV groups are optimized, the coverage area weight is dynamically adjusted, and energy consumption is reduced.
It improves the coverage and efficiency of dam inspection, reduces the energy consumption of UAV groups, and ensures the safe operation of the dam.
Smart Images

Figure CN119416995B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a coverage path planning method for daily inspection of dams based on DRL multi-UAVs, belonging to the research fields of artificial intelligence and intelligent water conservancy. Background Art
[0002] In recent years, as an efficient autonomous mobile platform, UAV (Unmanned Aerial Vehicle, also known as drone, abbreviated as UAV) has played an increasingly important role in many fields such as agricultural monitoring, urban planning, disaster response, and environmental protection. Especially when performing the dam area coverage task, the UAV swarm can complete the large-scale coverage task faster and more efficiently than a single UAV due to its unique collaborative ability and flexibility. The dam area coverage task usually requires the UAV swarm to conduct comprehensive search or monitoring within a specific area, ensuring that every area is effectively covered and avoiding repeated coverage as much as possible to save resources. During this process, the challenges mainly include how to efficiently allocate the task areas of UAVs, how to plan the flight paths of each UAV, and how to coordinate the actions of UAVs to avoid resource waste. In addition, the energy limitation of UAVs and the dynamic changes of the environment also add additional complexity to the dam area coverage path planning problem.
[0003] In recent years, artificial intelligence technology has developed rapidly and been widely applied. Deep learning and reinforcement learning have achieved great success in many fields. Deep learning has powerful perception capabilities, while reinforcement learning has decision-making capabilities. DRL (Deep Reinforcement Learning) combines the advantages of deep learning and reinforcement learning, providing solutions for the perception and decision-making problems of complex systems. DRL can effectively solve problems in continuous state spaces and continuous action spaces. At present, DRL has been widely applied in the field of multi-UAV coverage path planning. For example: Some studies use Actor-Critic neural networks with multi-resolution observations to train multi-UAV trajectory planning strategies, but all UAVs are considered to belong to the same entity and execute the same strategy. Some scholars have studied the trajectories of multi-UAV uplink data collection tasks based on multi-agent DRL algorithms. The researchers provided an algorithm based on MADDPG (Multi-Agent Deep Deterministic Policy Gradient) to independently manage the trajectories of each UAV. Some studies have applied MADDPG to trajectory planning in UAV motion control. Some scholars have designed an incentive mechanism to alleviate the problem of sparse rewards. To make better navigation and obstacle avoidance decisions, the researchers proposed a trajectory planning method based on the RNN (Recurrent Neural Network) architecture and temporal attention, ignoring multi-UAV formation obstacle avoidance. Some work has formulated multi-UAV trajectory planning as a dynamic non-cooperative game and used a DRL algorithm based on ESN (Echo State Network) units to solve this game. Some studies have proposed a multi-UAV collision avoidance method based on two-stage reinforcement learning, but have not considered the policy differences between individual UAVs and the constraint conditions of multi-UAV cooperative formation trajectory planning. Some work has used a decentralized DRL method in a two-dimensional grid-based map while ignoring the complex environment, considering multi-UAV trajectory planning in non-cooperative multi-UAV scenarios. To avoid ineffective action exploration in DRL and improve convergence, the researchers used Bayesian optimization to estimate the flight decisions of UAVs. For multi-UAV coverage trajectory planning, some work has proposed the DCDQN (Distributed Cooperative Deep Q-Learning) algorithm to obtain the UAV coverage trajectory online, ignoring the cooperative obstacle avoidance of multi-UAV formations.Some studies have proposed an effective reward mechanism that comprehensively considers the orbital changes of spacecraft based on DQN (Deep Q-Network), which improves the success rate of orbital interception in uncertain environments. Some work uses artificial potential fields to guide the anti-v multi-UAV formation navigation and obstacle avoidance, ignoring continuous actions and complex environments. However, the current multi-UAV coverage path planning algorithms encounter a key problem during the daily inspection of dams, that is, these algorithms do not take into account the coverage rate and energy consumption of multi-UAVs during the daily inspection of dams, resulting in low efficiency of the multi-UAV coverage path planning algorithms. Summary of the Invention
[0004] Object of the Invention: Aiming at the problems and deficiencies of the existing technology, the present invention provides a coverage path planning method for daily inspection of dams based on DRL multi-UAVs. Using a group of UAVs for dam inspection not only greatly improves the accuracy and efficiency of inspection, but also reduces labor costs, which is of great significance for ensuring the safe operation of dams.
[0005] Technical Solution: A coverage path planning method for daily inspection of dams based on DRL multi-UAVs includes the following steps:
[0006] (1) Problem Formalization: Define the multi-UAV coverage path planning problem as an optimization problem, which contains two optimization objectives, namely the energy consumption of multi-UAVs and the coverage rate. The specific optimization problem is:
[0007]
[0008] where, represents the total coverage rate, that is, the ratio of the covered grid area to the entire area, is the repeated coverage rate, that is, the ratio of the area of the same area covered by different UAVs to the entire area, θ c represents the threshold of the repeated coverage rate, E represents the total energy consumption of the UAV group, and β1 and β2 represent weight parameters, T t represents the total number of time steps.
[0009] (2) Model establishment: The multi-UAV coverage path planning problem can be modeled as a Markov random process. The Markov random process describes the transition of the system state, and its characteristic is that the probability of the next state depends only on the current state. Based on the characteristics of the Markov random process, the present invention proposes an EA-MATD3 (Energy Aware Multi-Agent Twin Delayed Deep Deterministic Policy Gradient) model to solve the multi-UAV coverage path planning problem. Specifically, EA-MATD3 adopts a centralized training and decentralized execution mechanism based on MATD3 (Multi-Agent Twin Delayed Deep Deterministic Policy Gradient), while fully considering the interaction between multiple UAVs, maintaining flexibility and scalability during execution; by introducing a dual Q-network and a policy smoothing technique, the stability of learning is significantly improved; using stacked LSTM (Long Short-Term Memory) to process time series data enables it to better predict and adapt to the dynamic changes of the environment, and reduces the repeated path by analyzing time series data, thereby reducing energy consumption.
[0010] The EA-MATD3 model mainly consists of two parts, namely the MATD3 network and the stacked LSTM network. The specific content is as follows:
[0011] In the MATD3 network, in order to improve the exploration performance and accelerate the convergence speed, Gaussian noise is introduced for action exploration. Specifically, a noise vector is generated through a deterministic policy, and then it is multiplied by a learnable parameter and added to the selected action. The Gaussian noise can be limited within a set range using a clipping function.
[0012] In order to better capture the experience obtained by the UAV during the interaction with the environment during inspection, a two-layer LSTM network has a better effect. The position, speed, and direction of the UAV are used as the input of the first LSTM layer. Then, the output of the first layer, as well as the reward and action obtained by the UAV at the previous time step, are used as the input of the second LSTM layer. In addition, the memory function of LSTM enables it to store and utilize the previous state information, which is beneficial to solving the path planning problem under partially observable conditions. Finally, the output of LSTM is combined with the observation vector in MATD3 to obtain the state representation of the UAV.
[0013] The Q-function update formula in the MATD3 algorithm is where, Represents the loss function of the UAVi critic network, represents the expected value, and the tuple (s, a, r i , s′) represents the state, action, reward of the UAV, and the state at the next moment, represents the experience replay buffer, and y i is the target value for updating the critic network, representing the TD target of the MATD3 algorithm, and Q(s, a; θ - ) represents the Q value.
[0014] The MATD3 algorithm uses the target critic network to calculate the TD target y t = r t + γQ′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) where r t represents the reward value at time step t, γ represents the discount factor, and Q′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) represents the target Q value. To improve the exploration performance and accelerate the convergence speed, Gaussian noise is introduced for action exploration. Specifically, a noise vector ∈ is generated through the deterministic policy μ(s|θ μ ), and then it is multiplied by the learnable parameter σ and added to the selected action, i.e.: a = μ(s|θ μ ) + ∈, ∈ ∼ N(0, σ 2 ). Therefore, the noisy Q value can be expressed as Q(s t , a t |θ Q ) + ∈, where θ Q represents the parameters of the critic network, s t and a t represent the state and action at time step t, respectively. Therefore, the noisy target y t can be expressed as: y t = r t + γQ′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) + clip(∈, -c, c). Where clip(∈, -c, c) represents the clipping function, ∈ represents the Gaussian noise, and [-c, c] represents the range of the Gaussian noise.
[0015] The update rule of the actor network in the MATD3 is where f i represents the actor network of the i-th UAV, oj denotes the observation vector of the j-th UAV, and N represents the number of UAVs. denotes the gradient of the actor network objective function of UAVi with respect to its parameter θ i . The update is approximated using samples from the experience replay buffer, and s j is the state observation of the j-th sample. denotes the gradient of the Q function of the actor network of UAVi with respect to its parameter θ i .
[0016] To better capture the experience obtained by the UAV during the inspection process in interaction with the environment, a two-layer LSTM network is more effective. The position, speed, and direction of the UAV are used as the input to the first LSTM layer. Then, the output of the first layer, along with the reward and action obtained by the UAV at the previous time step, are used as the input to the second LSTM layer. Additionally, the memory function of the LSTM enables it to store and utilize previous state information, which is beneficial for solving the path planning problem under partially observable conditions. The output of the LSTM is denoted as h t = LSTM(s t , h t-1 ). Where s t represents the state at time t, and h t represents the output state of the LSTM. Finally, the output of the LSTM is combined with the observation vector in MATD3 to obtain the state representation of the UAV as where represents the observation vector of UAVi at time t, represents the LSTM output state of UAVi at time t.
[0017] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the above-mentioned coverage path planning method for daily inspection of dams based on DRL multi-UAV.
[0018] A computer-readable storage medium, which stores a computer program for executing the above-mentioned coverage path planning method for daily inspection of dams based on DRL multi-UAV.
[0019] Beneficial effects: The present invention has the following advantages compared with the prior art:
[0020] In the system modeling part, a regional dynamic weight model is constructed, which can dynamically adjust the weight of the coverage area according to the actual scenario of dam inspection.
[0021] In the problem construction section, the UAV swarm action selection problem can be modeled as a decentralized partially observable Markov decision process, which is in line with the actual application scenario.
[0022] In the network model establishment section, the stacked LSTM network is integrated into the MATD3 algorithm to propose the EA-MATD3 method, which can maintain a high coverage rate of the target area while keeping the energy consumption of the UAV swarm low. Description of the Drawings
[0023] Figure 1 It is a realization diagram of the UAV swarm coverage path planning framework based on EA-MATD3 in a specific embodiment. Detailed Embodiments
[0024] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention fall within the scope defined by the appended claims of this application.
[0025] During the daily inspection of the dam, the coverage path planning method for multi-UAV daily inspection of the dam proposed by the present invention can be divided into two steps, namely, formal definition of the problem and model establishment.
[0026] First, the present invention defines the multi-UAV coverage path planning problem as an optimization problem, which includes two optimization objectives, namely, the energy consumption of multi-UAV and the coverage rate. The specific optimization problem is as follows:
[0027]
[0028] The meaning of each parameter is the same as above. Secondly, after formalizing the multi-UAV coverage path planning problem, the present invention finds that this problem is a Markov random process. The Markov random process describes the transition of the system state, and its characteristic is that the probability of the next state depends only on the current state. Based on the characteristics of the Markov random process, the present invention proposes the EA-MATD3 model to solve the multi-UAV coverage path planning problem in the daily inspection of the dam. Specifically, EA-MATD3 adopts a centralized training and decentralized execution mechanism based on MATD3, which can maintain flexibility and scalability during execution while fully considering the interaction between multi-UAV. By introducing the double Q network and policy smoothing technology, the stability of learning is significantly improved. Then, EA-MATD3 uses a two-layer LSTM to process time series data, enabling it to better predict and adapt to the dynamic changes of the environment, and reducing the repeated path by analyzing the time series data, thereby reducing the energy consumption.
[0029] Such as Figure 1As shown, the EA-MATD3 model mainly consists of two parts, namely the MATD3 network and the two-layer LSTM network.
[0030] In the MATD3 network, to improve the exploration performance and accelerate the convergence speed, Gaussian noise is introduced for action exploration. Specifically, a noise vector is generated through a deterministic policy, then multiplied by a learnable parameter and added to the selected action. The Gaussian noise can be limited within a set range using a clipping function.
[0031] To better capture the experience obtained by the UAV during the inspection process in interaction with the environment, using a two-layer LSTM network has better effects. The position, speed, and direction of the UAV are used as the input to the first LSTM layer. Then, the output of the first layer, along with the reward and action obtained by the UAV at the previous time step, are used as the input to the second LSTM layer. Additionally, the memory function of the LSTM enables it to store and utilize previous state information, which is beneficial for solving the path planning problem under partially observable conditions. Finally, the output of the LSTM is combined with the observation vector in MATD3 to obtain the state representation of the UAV.
[0032] The update formula for the Q function in the MATD3 algorithm is where represents the loss function of the UAVi critic network, represents the expected value, and the tuple (s, a, r i , s′) represents the state, action, reward, and the state at the next moment of the UAV, represents the experience replay buffer, y i is the target value for updating the critic network, representing the TD target of the MATD3 algorithm, and Q(s, a; θ - ) represents the Q value.
[0033] The MATD3 algorithm uses the target critic network to calculate the TD target y t = r t + γQ′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ). To improve the exploration performance and accelerate the convergence speed, Gaussian noise is introduced for action exploration. Specifically, a noise vector ∈ is generated through the deterministic policy μ(s|θ μ ), then multiplied by the learnable parameter σ and added to the selected action, that is: a = μ(s|θ μ ) + ∈, ∈ ~ N(0, σ 2 ). Therefore, the noisy Q value can be expressed as Q(s t , a t |θ Q) + ∈, the noisy target y t can be expressed as: y t = r t + γQ′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) + clip(∈, -c, c).
[0034] The update rule of the actor network in MATD3 is where f i represents the actor network of the i-th UAV, o j represents the observation vector of the j-th UAV, and N represents the number of UAVs. represents the gradient of the actor network objective function of UAVi with respect to its parameter θ i . The update is approximated using samples from the experience replay buffer, s j is the state observation of the j-th sample. represents the gradient of the Q function of the actor network of UAVi with respect to its parameter θ i .
[0035] To better capture the experience obtained by the UAV during the inspection process in interaction with the environment, a two-layer LSTM network has a better effect. The position, speed, and direction of the UAV are used as the input to the first LSTM layer. Then, the output of the first layer, along with the reward and action obtained by the UAV at the previous time step, are used as the input to the second LSTM layer. Additionally, the memory function of the LSTM enables it to store and utilize previous state information, which is beneficial for solving the path planning problem under partially observable conditions. The output of the LSTM is represented as h t = LSTM(s t , h t-1 ). Where s t represents the state at time t, and h t represents the output state of the LSTM. Finally, the output of the LSTM is combined with the observation vector in MATD3 to obtain the state representation of the UAV as
[0036] Obviously, those skilled in the art should understand that each step of the above-mentioned coverage path planning method for daily inspection of dams based on DRL multi-UAVs of the embodiments of the present invention can be implemented by a general computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0037] Matters not covered by the present invention are well-known technologies.
[0038] The above embodiments are only for illustrating the technical concept and features of the present invention. The purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A coverage path planning method for daily inspection of dams based on DRL multi-UAV, characterized in that: The following steps are involved: (1) Problem formalization: The multi-UAV coverage path planning problem is defined as an optimization problem, which includes two optimization objectives: the energy consumption of multiple UAVs and the coverage rate; (2) Model building: The multi-UAV coverage path planning problem is modeled as a Markov random process. Based on the characteristics of the Markov random process, the EA-MATD3 model is proposed to solve the multi-UAV coverage path planning problem. EA-MATD3 adopts the centralized training decentralized execution mechanism based on MATD3, introduces the dual Q network and policy smoothing technology, uses stacked LSTM to process time series data, and reduces repeated paths by analyzing time series data. The EA-MATD3 model contains a MATD3 network and a stacked LSTM network: In the MATD3 network, Gaussian noise is introduced for action exploration. A noise vector is generated by a deterministic strategy, which is then multiplied by a learnable parameter and added to the selected action. The Gaussian noise is limited to a set range using a clipping function. The stacked LSTM network is a two-layer LSTM network that captures the experience gained by the UAV in interacting with the environment during the inspection process. The position, speed, and direction of the UAV are used as the input of the first LSTM layer. Then, the output of the first layer and the reward and action obtained by the UAV in the previous time step are used as the input of the second LSTM layer. Finally, the output of the stacked LSTM network is combined with the observation vector in the MATD3 network to obtain the state representation of the UAV.
2. The coverage path planning method based on DRL multi-UAV daily dam inspection according to claim 1 is characterized in that: The Q function update formula in the MATD3 network is: in, represents the loss function of the UAVi critic network, Represents the expected value, tuple (s, a, r i , s′) represents the state, action, reward and state of the UAV at the next moment, represents the experience replay buffer, y i To update the target value of the critic network, Q(s,a;θ) represents the TD target of the MATD3 algorithm. - ) represents the Q value; The MATD3 algorithm uses a target critic network to compute the TD target y t =r t +γQ′(s t+1 ,μ′(s t+1 ∣θ μ′ )|θ Q′ ) Among them, r t represents the reward value at time step t, γ represents the discount factor, Q′(s t+1 ,μ′(s t+1 ∣θ μ′ )|θ Q′ ) represents the target Q value; in order to improve the exploration performance and accelerate the convergence speed, Gaussian noise is introduced for action exploration. Specifically, through the deterministic strategy μ(s|θ μ ) generates a Gaussian noise ∈, which is then multiplied by the learnable parameter σ and added to the selected action, i.e.: a = μ(s|θ μ )+∈,∈~N(0,σ 2 ); therefore, the Q value with noise is expressed as Q(s t ,a t ∣θ Q )+∈, where θ Q represents the parameters of the critic network, s t and a t Represent the state and action at time step t respectively; TD target y t Expressed as: y t =r t +γQ′(s t+1 ,μ′(s t+1 ∣θ μ′ )|θ Q′ )+clip(∈,-c,c); where clip(∈,-c,c) represents the clipping function, ∈ represents Gaussian noise, and [-c,c] represents the range of Gaussian noise.
3. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the coverage path planning method based on DRL multi-UAV daily dam inspection as described in any one of claims 1-2 is implemented.
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the coverage path planning method based on DRL multi-UAV daily dam inspection as described in any one of claims 1-2.