Multi-uav cooperative reconnaissance method based on dcddpg algorithm

CN117762159BActive Publication Date: 2026-09-25NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311775670.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2026-09-25
Estimated Expiration
2043-12-21

AI Technical Summary

Technical Problem

MCRS可以被建模为含有多个智能体的马尔可夫决策过程(MMDP),因此可以采用多智能体深度强化学习(MADRL)的方法来解决,但是目前深度强化学习(MADRL)在集中式训练分布式执行的框架下使用一个集中的评论家克服了环境不稳定的问题,集中式的评论家成为CTDE的一个标准的选择,一个集中的评论家会在策略更新中产生比局部评论家更高的方差,出现值函数估计的偏差与策略更新的方差的平衡问题,使得多无人机协同侦察协调能力差

Benefits of technology

[0042]上述基于DCDDPG算法的多无人机协同侦察方法,本申请将多无人机侦察搜索问题视为部分可观测马尔可夫决策过程,根据改进的深度学习算法-DCDDPG算法来对多无人机侦察搜索问题进行求解,在这个基础上,使用了局部评论家和集中式评论家对动作价值函数估计,通过设置参数控制着集中评论家和局部评论家的权重,可以有效地平衡偏差和方差。局部评论家网络仅输入当前智能体的动作和观测,优化自身动作,集中式评论家将所有智能体的动作和观测作为输入,以评估联合动作的优劣,并反馈给动作网络以学习到合作行为,进而提高了多无人机侦察搜索的协调能力,另外根据经验回放技术对智能体的策略梯度进行优化,减少了训练数据的相关性,提高了策略梯度的准确度,在动作空间设计时根据无人机的动力学模型设计动作空间可以真实的模拟无人机侦察并提供更高的控制精度,最后智能体利用所有智能体的奖励之和计算时间差误差来更新集中式评论家网络,利用智能体各自的奖励计算时间差误差以更新局部评论家网络,并利用优化后的智能体的策略梯度更新动作网络,根据更新后的集中式评论家网络和局部评论家网络以及更新后的动作网络进行多无人机协同侦察还可以大大提高协同侦察的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117762159B_ABST
    Figure CN117762159B_ABST
Patent Text Reader

Abstract

The application relates to a multi-unmanned aerial vehicle cooperative reconnaissance method based on a DCDDPG algorithm. The method comprises the following steps: modeling a multi-unmanned aerial vehicle reconnaissance search problem as a partially observable Markov decision process, solving the partially observable Markov decision process according to the DCDDPG algorithm, estimating a strategy gradient of an intelligent agent, optimizing the strategy gradient of the intelligent agent according to an experience replay technology, designing a state space of the intelligent agent, designing an action space and a reward function of the intelligent agent, calculating a reward obtained by each intelligent agent at each time point, using the rewards of all intelligent agents to calculate a time difference error to update a centralized critic network and a local critic network, using the optimized strategy gradient of the intelligent agent to update an action network, and obtaining an acceleration of each unmanned aerial vehicle at a current time point according to the updated centralized critic network and local critic network and the updated action network. The method can improve the coordination capability of multi-unmanned aerial vehicle reconnaissance search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) reconnaissance and search technology, and in particular to a multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm. Background Technology

[0002] With the continuous development of Unmanned Aircraft Systems (UAS), an increasing number of UAVs are being applied to tasks such as mineral exploration, agricultural monitoring, traffic mapping, and forest search and rescue. UAVs possess the ability to quickly and effectively cover large areas, making them an effective tool in reconnaissance and target search missions. In recent years, the coordinated use of multiple UAVs for reconnaissance and search has become increasingly popular. First, Multi-UAV Cooperative Reconnaissance Search (MCRS) allows for broader target area coverage because they can simultaneously reconnoiter different areas. Furthermore, each UAV can be equipped with specialized sensors for searching various types of targets. However, coordinating multiple UAVs for reconnaissance and search is a challenging problem. First, each UAV needs to know the position, speed, and reconnaissance strategy of other UAVs to coordinate actions and avoid collisions; second, each UAV needs to be able to autonomously plan the optimal search path to maximize search area coverage and avoid trajectory overlap; third, multi-UAV systems need to possess high robustness and flexibility to quickly adapt to different tasks and environments, thereby maintaining system stability and reliability. Furthermore, for reconnaissance and search missions, due to errors such as sensor accuracy, changes in target appearance, and environmental noise, target detection algorithms may not be able to return absolutely reliable results for the search area, thereby increasing the uncertainty of the presence of targets in the area and affecting the accuracy of reconnaissance results.

[0003] However, traditional methods for solving the MCRS problem consider two aspects. Firstly, to address the coordination problem among multiple UAVs, traditional centralized control methods use a centralized ground control station to manage the operation of all UAVs, ensuring centralized decision-making and promoting effective coordination among them. However, this approach has limitations in scalability and robustness, as a single point of failure can disrupt the entire system. Secondly, to solve the reconnaissance search problem, traditional mission planning techniques model the UAV reconnaissance search problem as coverage path planning or mixed-integer nonlinear programming, using coverage as the objective function. This optimization method performs well in simple mission environments with a small number of UAVs. However, as the number of UAVs increases and the scale of the environment expands, this multi-UAV path planning problem proves NP-hard, with significantly increased time and space complexity. Furthermore, reconnaissance search that only considers area coverage ignores potential target detection errors from UAV sensors. In reality, covering the reconnaissance area once does not yield a high-confidence result of target presence, leading to unreliable search results. In recent years, the emergence of machine learning methods such as reinforcement learning (RL) has provided new solutions to the MCRS problem. Reinforcement learning is typically modeled within the framework of Markov Decision Processes (MDPs). To enhance its ability to solve complex problems, researchers have combined deep learning (DL) with reinforcement learning (RL), proposing Deep Reinforcement Learning (DRL). This approach utilizes neural networks to approximate the value function or policy function and optimizes network parameters through agent-environment interactions to maximize cumulative rewards. Deep reinforcement learning can autonomously learn and make decisions in dynamic and complex environments and has achieved significant results in many fields, such as competitive games, intelligent manufacturing, and autonomous driving. MCRS can be modeled as a Markov Decision Process (MMDP) with multiple agents, thus allowing for the application of Multi-Agent Deep Reinforcement Learning (MADRL). However, current MADRL approaches, within a framework of centralized training and distributed execution, use a centralized critic to overcome environmental instability. While a centralized critic has become a standard choice for CTDE (Comprehensive Collaborative Reconnaissance), it can also introduce higher variance in policy updates compared to local critics. This leads to a trade-off between the bias in value function estimation and the variance in policy updates, resulting in poor coordination among multiple UAVs in collaborative reconnaissance. Summary of the Invention

[0004] Therefore, it is necessary to provide a multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm that can improve the coordination capability of multi-UAV reconnaissance and search, in order to address the above-mentioned technical problems.

[0005] A multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm, the method comprising:

[0006] The multi-UAV reconnaissance and search problem is modeled as a partially observable Markov decision process. Local commentator network, centralized commentator network and action network are designed in the partially observable Markov decision process, where the agent represents the UAV.

[0007] The DCDDPG algorithm is used to solve a partially observable Markov decision process. The action value function of the agent is estimated using local and centralized commentator networks. The agent's policy gradient is determined using deterministic policy gradient theory and the action value function, and then optimized using experience replay techniques. The state space of the agent is designed based on its own position and velocity, relative positions between agents, and no-fly zone information. The action space is designed using the UAV's dynamics model. The reward function is designed based on cognitive reward, target discovery reward, behavioral penalty, and boundary avoidance. The reward obtained by each agent at each time step is calculated based on the action space, state space, and reward function. Each agent updates the centralized and local commentator networks using its own reward calculation time difference error, and updates the action network using the optimized agent's policy gradient. The current acceleration of each UAV is obtained based on the updated centralized and local commentator networks and the updated action network. Multi-UAV cooperative reconnaissance is then performed based on the current acceleration of each UAV.

[0008] In one embodiment, estimating the agent's action value function based on a local critic network and a centralized critic network includes:

[0009] Based on the estimation of the action value function using local and centralized critic networks, the action value function of the agent is obtained as follows:

[0010]

[0011] in, The parameter is A local network of critics, The parameter is The centralized critic network, where α represents the weight of the local critic network, and i represents the i-th agent.

[0012] In one embodiment, the loss function of the local commentator network is:

[0013]

[0014] in, Indicates the value of the target action, r i γ represents the reward of agent i, and γ represents the discount factor. Represents the target local commentator network, μ i Let o represent the policy of agent i, where o = {o1, ... o}. N} represents the joint observation of all agents, a = {a1, ... a2} N} indicates a joint action;

[0015] The loss function of a centralized critic network is:

[0016]

[0017] in, Indicates the value of the target action. This indicates a network of commentators focusing on a specific target audience.

[0018] In one embodiment, the agent's policy gradient is determined using deterministic policy gradient theory and action value function, including:

[0019] The agent's policy gradient is determined using deterministic policy gradient theory and the action-value function.

[0020]

[0021] Wherein, J(μ) i ) represents the goal of each agent. It is an action value function. This represents the expected value sampled from the experience replay buffer D. The parameter is θ i The gradient, μ i Let a represent the policy network of agent i. i |o i The observation is o i The agent's action a i , Indicates parameter a i gradient, μ represents the action value function of agent i. i (o i ) indicates that the observation is o i The agent's strategy.

[0022] In one embodiment, the agent's policy gradient is optimized using an experience replay technique to obtain the optimized agent's policy gradient, including:

[0023] The agent's policy gradient is optimized using experience replay techniques, resulting in the optimized policy gradient of the agent.

[0024]

[0025] Where α represents the weight of the overall policy gradient for local critics, CentralizedCritic represents the policy gradient of the centralized critics, and LocalCritic represents the policy gradient of the local critics. This represents a lumped action value function with inputs of joint observation o and joint action a. This indicates that the input is a local observation. i and local action a i The local action value function, where i represents the i-th agent.

[0026] In one embodiment, the state space of the agent is designed based on the agent's own position and velocity, the relative positions between agents, and no-fly zone information, including:

[0027] The state space of the agent is designed based on its own position and velocity, the relative positions between agents, and no-fly zone information.

[0028]

[0029] Where M is the number of no-fly zones detected by the agent within the reconnaissance and search map, (u i,t ,v i,t ) represents the position and velocity of agent i at time t, u i,t v represents the position of agent i at time t. i,t z represents the velocity of agent i at time t. M This indicates the location of the Mth no-fly zone. This indicates that N has not yet been detected. Z -M no-fly zones are assigned a fixed maximum value R.

[0030] In one embodiment, the motion space is designed using the dynamics model of the drone, including:

[0031] The motion space is designed using the dynamic model of the UAV.

[0032]

[0033] Where clip is the truncation function, F i,t The input control force is represented by Δt, which represents the time step for updating the drone's speed and position. min v represents the minimum speed of the drone. max v represents the maximum speed of the drone. i,t Let u represent the velocity of drone i at time t, where i represents the i-th drone. i,tLet m represent the position of drone i at time t, and m be the mass of the drone.

[0034] In one embodiment, the agent's reward function is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries, including:

[0035] The reward function for the agent is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries.

[0036] R t =R 1,t +R 2,t +R 3,t +R 4,t

[0037]

[0038]

[0039] R 3,t =ω3(r entry +r collision )

[0040]

[0041] Among them, R 1,t Let L represent cognitive reward, ω1 represent the cognitive reward weight, and L represent the cognitive reward weight. x L represents the number of cells along the x-axis of the environment map. y This indicates the number of cells along the y-axis of the environment map. Cell C represents time t. x,y Uncertainty of R 2,t ω2 represents the target discovery reward, and ω2 represents the target discovery reward weight. Cell C represents time t. x,y The probability that the target exists is τ, which represents the threshold for believing the target exists, and R. 3,t ω3 represents the behavioral penalty, and r represents the weight of the behavioral penalty. entry This indicates the penalty for a drone entering a no-fly zone. collision R represents the penalty for drone collisions. 4,t This indicates preventing exceeding the boundary, ω4 represents the reward weight for preventing exceeding the boundary, and u i,t Let AoI represent the position of UAV i at time t, and let AoI represent the reconnaissance and search map.

[0042] The aforementioned multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm treats the multi-UAV reconnaissance and search problem as a partially observable Markov decision process. It solves the multi-UAV reconnaissance and search problem using an improved deep learning algorithm—the DCDDPG algorithm. On this basis, it uses local commentators and centralized commentators to estimate the action value function. By setting parameters to control the weights of centralized commentators and local commentators, bias and variance can be effectively balanced. The local critic network only takes into account the current agent's actions and observations to optimize its own actions. The centralized critic takes the actions and observations of all agents as input to evaluate the quality of joint actions and feeds them back to the action network to learn cooperative behaviors, thereby improving the coordination ability of multi-UAV reconnaissance and search. In addition, the policy gradient of the agents is optimized by using experience replay technology, which reduces the correlation of training data and improves the accuracy of policy gradient. The action space is designed according to the dynamic model of UAVs to realistically simulate UAV reconnaissance and provide higher control precision. Finally, the agents use the sum of the rewards of all agents to calculate the time difference error to update the centralized critic network, use the individual rewards of the agents to calculate the time difference error to update the local critic network, and use the optimized policy gradient of the agents to update the action network. Multi-UAV cooperative reconnaissance based on the updated centralized critic network, local critic network, and updated action network can also greatly improve the accuracy of cooperative reconnaissance. Attached Figure Description

[0043] Figure 1 This is a flowchart illustrating a multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm in one embodiment.

[0044] Figure 2 This is a schematic diagram of the action network structure in one embodiment;

[0045] Figure 3 This is a schematic diagram of the structure of the critic network in one embodiment. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] In one embodiment, such as Figure 1 As shown, a multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm is provided, including the following steps:

[0048] Step 102: Model the multi-UAV reconnaissance and search problem as a partially observable Markov decision process. In the partially observable Markov decision process, design a local commentator network, a centralized commentator network, and an action network, where the agent represents the UAV.

[0049] In UAV reconnaissance and search operations, due to the limited detection range of sensors, UAVs cannot directly observe the complete state of the environment and can only infer the environmental state based on partial observation information. Therefore, a partially observable Markov decision process is typically used to model this process. Treating the multi-UAV reconnaissance and search problem as a partially observable Markov decision process, this paper uses an improved deep learning algorithm—DCDDPG—to solve the multi-UAV reconnaissance and search problem. DCDDPG represents a dual-critic deep deterministic policy gradient network. Based on this, a local critic network, a centralized critic network, and an action network are designed. DCDDPG can balance the value function of the learning set with the local value function, effectively balancing bias and variance, thereby improving the coordinated reconnaissance capability of multiple UAVs. Experiments demonstrate that the local critic provides an estimate with more bias and less variance, while the centralized critic provides an estimate with less bias and more variance. DCDDPG uses both centralized and local critics to estimate the true action value function and controls the weights of the centralized and local critics through parameters, thus effectively balancing bias and variance.

[0050] Step 104: Solve the partially observable Markov decision process using the DCDDPG algorithm; estimate the agent's action value function using local and centralized commentator networks; determine the agent's policy gradient using deterministic policy gradient theory and the action value function; optimize the agent's policy gradient using experience replay technology; design the agent's state space based on the agent's own position and velocity, relative positions between agents, and no-fly zone information; design the action space using the UAV's dynamics model; design the agent's reward function based on cognitive reward, target discovery reward, behavioral penalty, and prevention of exceeding boundaries; calculate the reward obtained by each agent at each time step based on the agent's action space, state space, and reward function.

[0051] Step 106: Each agent uses its own reward to calculate the time difference error to update the centralized critic network and the local critic network respectively, and uses the optimized agent's policy gradient to update the action network. Based on the updated centralized critic network, the local critic network and the updated action network, the current acceleration of each UAV is obtained; multi-UAV cooperative reconnaissance is performed based on the current acceleration of each UAV.

[0052] The action value function describes the expected reward after taking a joint action a in state s. The MARL algorithm learns a set of action value functions with the goal of maximizing the overall expected reward G for all agents. However, learning a set of action value functions cannot guarantee the learning of cooperative behavior and instead faces problems such as high variance of policy gradients and scalability issues. To address this problem, this application proposes a novel MADRL algorithm called Double Critic Deep Deterministic Policy Gradient (DCDDPG). The DCDDPG algorithm maximizes the expected reward for each agent. It also provides additional global observations to each agent, in the expectation that a single agent will learn cooperative behavior. i The updates are based on the rewards of each agent. Therefore, each agent calculates the time difference error using its own reward. The time difference error is the loss function described in this application. The centralized critic network and the local critic network are updated using the calculated loss functions of the local critic network and the centralized critic network, respectively. The DCDDPG algorithm framework adopts an Actor-Critic framework with centralized training and distributed execution. During the centralized training phase, each agent is provided with additional global observations, thus ensuring a stable environment. During the execution phase, each agent only uses its own observations to output actions, thus it is distributed and requires no centralized control or communication.

[0053] DCDDPG consists of two parts: an actor network and a critic network. The network structure is as follows: Figure 2 and Figure 3 As shown. The Actor network for each drone (parameter θ) i Input local observation o i Output a definite control force Fi. Because the action a is output at each time step... i =μ i (o i Since π is deterministic, the stochastic policy π becomes a deterministic policy μ. The Actor updates the parameters θ of the deterministic policy μ using stochastic gradient ascent. iBased on deterministic policy gradient theory and the action value function, the policy gradient of the i-th agent can be determined. The biggest innovation of DCDDPG lies in its simultaneous use of local and centralized critics to estimate the action value function. To avoid instability caused by simultaneously acquiring and updating the Q-function, a target actor network and a target critic network are used to calculate the target value. Therefore, each agent has two actor networks and four critic networks. The target network parameters are updated with a delay, preventing the agent from overemphasizing recent actions. Simultaneously, by setting parameters to control the weights of the centralized and local critics, bias and variance can be effectively balanced. The local critic network only takes the current agent's actions and observations as input to optimize its own actions. The centralized critic takes the actions and observations of all agents as input to evaluate the quality of joint actions and feeds this feedback to the action network to learn cooperative behavior, thereby improving the coordination ability of multi-UAV reconnaissance and search. Furthermore, the policy gradient of the agent is optimized using experience replay technology, reducing the correlation of training data and improving the accuracy of the policy gradient. In motion space design, designing the motion space based on the UAV's dynamic model can realistically simulate UAV reconnaissance and provide higher control precision. The UAV's state space represents the observed values ​​and consists of three parts. The first part is the UAV's own position and velocity, used to help the UAV determine its next direction of movement. The second part is the UAV's relative position to other UAVs {u i,t -u 1,t ,…,u i,t -u N,t Relative position, compared to absolute position, helps in better perceiving the location of other drones. The third part is no-fly zone information; drones can perceive the location of no-fly zones through a circular field of view (FOV) to avoid unintended entry. To ensure a fixed dimension for the neural network input, the relative distances of all no-fly zones within the FOV are chosen as input, while the relative distances of undetected no-fly zones outside the FOV are considered a fixed maximum value. Through the design of the state space, observations can be output more accurately, thereby improving the reward calculation accuracy of drones when interacting with the computational environment.

[0054] The process of calculating the reward obtained by each agent at each time step using the determined action space, state space, and reward function after determining the agent's action space, state space, and reward function is the existing technology and will not be elaborated on in this application. Each agent uses its own reward calculation time difference error to update the centralized critic network and the local critic network respectively, and uses the optimized agent's policy gradient to update the action network. The UAV's observations are input into the updated action network to output the UAV's acceleration. Then, the centralized critic network and the local critic network are used to evaluate the output acceleration, and multi-UAV reconnaissance and search planning is performed based on the evaluation results.

[0055] In the aforementioned multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm, this application treats the multi-UAV reconnaissance and search problem as a partially observable Markov decision process. The improved deep learning algorithm, DCDDPG, is used to solve the multi-UAV reconnaissance and search problem. Based on this, local critics and centralized critics are used to estimate the action value function. By setting parameters to control the weights of centralized critics and local critics, bias and variance can be effectively balanced. The local critic network only takes into account the current agent's actions and observations to optimize its own actions. The centralized critic takes the actions and observations of all agents as input to evaluate the quality of joint actions and feeds them back to the action network to learn cooperative behaviors, thereby improving the coordination ability of multi-UAV reconnaissance and search. In addition, the policy gradient of the agents is optimized by using experience replay technology, which reduces the correlation of training data and improves the accuracy of policy gradient. The action space is designed according to the dynamic model of UAVs to realistically simulate UAV reconnaissance and provide higher control precision. Finally, each agent uses its own reward to calculate the time difference error to update the centralized critic network and the local critic network respectively, and uses the optimized agent's policy gradient to update the action network. Multi-UAV cooperative reconnaissance based on the updated centralized critic network, local critic network and the updated action network can also greatly improve the accuracy of cooperative reconnaissance.

[0056] In one embodiment, estimating the agent's action value function based on a local critic network and a centralized critic network includes:

[0057] Based on the estimation of the action value function using local and centralized critic networks, the action value function of the agent is obtained as follows:

[0058]

[0059] in, The parameter is A local network of critics, The parameter is The centralized critic network, where α represents the weight of the local critic network, and i represents the i-th agent.

[0060] In one embodiment, the loss function of the local commentator network is:

[0061]

[0062] in, Indicates the value of the target action, r i γ represents the reward of agent i, and γ represents the discount factor. Represents the target local commentator network, μ i Let o represent the policy of agent i, where o = {o1, ... o}. N} represents the joint observation of all agents, a = {a1, ... a2} N} indicates a joint action;

[0063] The loss function of a centralized critic network is:

[0064]

[0065] in, Indicates the value of the target action. This indicates a network of commentators focusing on a specific target audience.

[0066] In one embodiment, the agent's policy gradient is determined using deterministic policy gradient theory and action value function, including:

[0067] The agent's policy gradient is determined using deterministic policy gradient theory and the action-value function.

[0068]

[0069] Wherein, J(μ) i ) represents the goal of each agent. It is an action value function. This represents the expected value sampled from the experience replay buffer D. The parameter is θ i The gradient, μ i Let a represent the policy network of agent i. i |o i The observation is o i The agent's action a i , Indicates parameter a i gradient, μ represents the action value function of agent i. i (o i ) indicates that the observation is o iThe agent's strategy.

[0070] In one embodiment, the agent's policy gradient is optimized using an experience replay technique to obtain the optimized agent's policy gradient, including:

[0071] The agent's policy gradient is optimized using experience replay techniques, resulting in the optimized policy gradient of the agent.

[0072]

[0073] Where α represents the weight of the overall policy gradient for local critics, CentralizedCritic represents the policy gradient of the centralized critics, and LocalCritic represents the policy gradient of the local critics. This represents a lumped action value function with inputs of joint observation o and joint action a. This indicates that the input is a local observation. i and local action a i The local action value function, where i represents the i-th agent.

[0074] In a specific embodiment, an experience replay technique is used. Previous experiences are stored in an experience pool, and a small batch of experiences is randomly selected for training the network to reduce the correlation between training data. The experience pool D is designed as a tuple (o,o',a1,…a...). N ,r1,...,r N ), and Both are differentiable function approximators. By optimizing the agent's policy gradient using empirical replay techniques, we can obtain the policy gradients of both the centralized critic and the local critic simultaneously.

[0075] In one embodiment, the state space of the agent is designed based on the agent's own position and velocity, the relative positions between agents, and no-fly zone information, including:

[0076] The state space of the agent is designed based on its own position and velocity, the relative positions between agents, and no-fly zone information.

[0077]

[0078] Where M is the number of no-fly zones detected by the agent within the reconnaissance and search map, (u i,t ,v i,t ) represents the position and velocity of agent i at time t, u i,t v represents the position of agent i at time t. i,t z represents the velocity of agent i at time t. MThis indicates the location of the Mth no-fly zone. This indicates that N has not yet been detected. Z -M no-fly zones are assigned a fixed maximum value R.

[0079] In one embodiment, the motion space is designed using the dynamics model of the drone, including:

[0080] The motion space is designed using the dynamic model of the UAV.

[0081]

[0082] Where clip is the truncation function, F i,t The input control force is represented by Δt, which represents the time step for updating the drone's speed and position. min v represents the minimum speed of the drone. max v represents the maximum speed of the drone. i,t Let u represent the velocity of drone i at time t, where i represents the i-th drone. i,t Let m represent the position of drone i at time t, and m be the mass of the drone.

[0083] In one embodiment, the agent's reward function is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries, including:

[0084] The reward function for the agent is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries.

[0085] R t =R 1,t +R 2,t +R 3,t +R 4,t

[0086]

[0087]

[0088] R 3,t =ω3(r entry +r collision )

[0089]

[0090] Among them, R 1,t Let L represent cognitive reward, ω1 represent the cognitive reward weight, and L represent the cognitive reward weight. x L represents the number of cells along the x-axis of the environment map. y This indicates the number of cells along the y-axis of the environment map. Cell C represents time t. x,y Uncertainty of R2,t ω2 represents the target discovery reward, and ω2 represents the target discovery reward weight. Cell C represents time t. x,y The probability that the target exists is τ, which represents the threshold for believing the target exists, and R. 3,t ω3 represents the behavioral penalty, and r represents the weight of the behavioral penalty. entry This indicates the penalty for a drone entering a no-fly zone. collision R represents the penalty for drone collisions. 4,t This indicates preventing exceeding the boundary, ω4 represents the reward weight for preventing exceeding the boundary, and u i,t Let AoI represent the position of UAV i at time t, and let AoI represent the reconnaissance and search map.

[0091] In a specific implementation, cognitive rewards aim to guide the drone's exploration, reducing the uncertainty of the entire reconnaissance and search map, i.e., minimizing the uncertainty of the entire area. Target discovery rewards require multiple reconnaissance attempts for cells with inconsistent reconnaissance results to increase the confidence in the target's presence. Behavioral penalties prevent drones from entering no-fly zones and from colliding with each other, providing negative feedback when a drone enters a no-fly zone or when the distance between drones is less than the safe distance. Both are considered equally important and therefore given equal weight.

[0092]

[0093]

[0094] Preventing Boundary Excursions: To ensure that drones conduct reconnaissance and search within the reconnaissance and search map, a penalty is imposed when a drone enters the Padding Area.

[0095] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this application, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Furthermore, Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0096] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A multi-UAV cooperative reconnaissance method based on the DCDDPG algorithm, characterized in that, The method includes: The multi-UAV reconnaissance and search problem is modeled as a partially observable Markov decision process. In this partially observable Markov decision process, a local commentator network, a centralized commentator network, and an action network are designed, where the agent represents the UAV. The partially observable Markov decision process is solved using the DCDDPG algorithm. The action value function of the agent is estimated based on the local and centralized commentator networks. The agent's policy gradient is determined using deterministic policy gradient theory and the action value function. The policy gradient is then optimized using experience replay technology to obtain the optimized agent's policy gradient. The agent's state space is designed based on its own position and velocity, relative positions between agents, and no-fly zone information. The action space is designed using the UAV's dynamics model. The agent's reward function is designed based on cognitive reward, target discovery reward, behavioral penalty, and boundary avoidance. The reward obtained by each agent at each time step is calculated based on the agent's action space, state space, and reward function. Each agent updates the centralized and local commentator networks using its own reward calculation time difference error, and updates the action network using the optimized agent's policy gradient. The current acceleration of each UAV is obtained based on the updated centralized and local commentator networks and the updated action network. Multi-UAV cooperative reconnaissance is performed based on the current acceleration of each UAV. The action value function of the agent is estimated based on local critic networks and centralized critic networks, including: Based on the estimation of the action value function using local and centralized critic networks, the action value function of the agent is obtained as follows: in, The parameter is A local network of critics, The parameter is A centralized network of critics This indicates the weight of a local commentator network. Let i represent the i-th intelligent agent.

2. The method according to claim 1, characterized in that, The method further includes: The loss function of the local critic network is: in, Indicates the value of the target action. This represents the reward for agent i. Indicates the discount factor. This represents a local network of commentators targeting a specific target. This represents the policy of agent i. This represents the joint observation of all intelligent agents. Indicates a joint action; The loss function of the centralized critic network is: in, Indicates the value of the target action. This indicates a network of commentators focusing on a specific target audience.

3. The method according to claim 1, characterized in that, Determining the agent's policy gradient using deterministic policy gradient theory and action-value functions includes: The agent's policy gradient is determined using deterministic policy gradient theory and the action-value function. in, This represents the goal of each agent. It is an action value function. This represents the expected value sampled from the experience replay buffer D. The parameter is gradient, Represents the policy network of agent i. Indicates observation as Actions of the intelligent agent , The parameter is gradient, The function representing the action value of agent i. Indicates observation as The agent's strategy.

4. The method according to claim 3, characterized in that, The agent's policy gradient is optimized using experience replay technology to obtain the optimized agent's policy gradient, including: The policy gradient of the agent is optimized using empirical replay techniques, resulting in the optimized policy gradient of the agent. in, This represents the weight of the overall policy gradient assigned to a particular commentator. This represents the strategy gradient of the group of commentators. This represents the policy gradient of local commentators. This represents a lumped action value function with inputs of joint observation o and joint action a. This indicates that the input is a local observation. and local actions The local action value function, Let i represent the i-th intelligent agent.

5. The method according to claim 1, characterized in that, The state space of an agent is designed based on its own position and velocity, the relative positions between agents, and no-fly zone information, including: The state space of the agent is designed based on its own position and velocity, the relative positions between agents, and no-fly zone information. Where M is the number of no-fly zones detected by the agent within the reconnaissance and search map. This represents the position and velocity of agent i at time t. This represents the position of agent i at time t. This represents the velocity of agent i at time t. This indicates the location of the Mth no-fly zone. Indicates that no detection has been made yet. Each no-fly zone is assigned a fixed maximum value R.

6. The method according to claim 1, characterized in that, Designing the motion space using the dynamics model of the UAV, including: The motion space is designed using the dynamic model of the UAV. Where clip is the truncation function. Indicates the input control force, The time step indicating the drone's update speed and location. Indicates the minimum speed of the drone. Indicates the maximum speed of the drone. This represents the velocity of drone i at time t. i Indicates the i-th drone. Let m represent the position of drone i at time t, and m be the mass of the drone.

7. The method according to claim 1, characterized in that, The reward function for the agent is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries, including: The reward function for the agent is designed based on cognitive reward, goal discovery reward, behavioral penalty, and prevention of exceeding boundaries. in, Indicates cognitive reward. Indicates the cognitive reward weight, This indicates the number of cells along the x-axis of the environment map. This indicates the number of cells along the y-axis of the environment map. Cell representing time t Uncertainty Indicates a reward for target discovery. Indicates the target discovery reward weight. Cell representing time t The goal is possible. This represents the threshold at which the target is considered to exist. Indicates behavioral punishment. Indicates the weight of the punishment for the behavior. This indicates the penalties for drones entering no-fly zones. This indicates the penalty for drone collisions. This indicates to prevent exceeding the boundary. This indicates that the reward weight should not exceed the boundary. This represents the position of drone i at time t. This indicates a reconnaissance and search map.

Citation Information

Patent Citations

  • Unmanned aerial vehicle network hovering position optimization method based on multi-agent deep reinforcement learning

    CN111786713A

  • Underwater vehicle target area floating control method based on double-commentator reinforcement learning technology

    CN113033119A