Joint optimization method for task allocation and route planning of multiple unmanned aerial vehicles in dynamic environment

By constructing a three-dimensional simulation environment and a heuristic reward model in a dynamic adversarial scenario, and utilizing the MATD3 algorithm and a partially observable Markov decision process, the problem of insufficient decision-making efficiency in UAV task allocation and trajectory planning is solved, and efficient collaborative task planning of UAV swarms in a three-dimensional dynamic environment is realized.

CN120875356APending Publication Date: 2025-10-31CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510969074.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing research has failed to effectively address the problem of insufficient efficiency in UAV task allocation and trajectory planning decisions in dynamic adversarial scenarios. In particular, in three-dimensional dynamic environments, the impact of obstacle dynamic displacement and sudden threat sources on the planning system has not been fully considered, and most algorithm verifications are limited to simplified two-dimensional plane scenarios, lacking adaptive collaborative planning strategies.

Method used

The improved TD3 algorithm is used to construct the MATD3 framework, establish a three-dimensional stochastic dynamic simulation environment, build a heuristic reward value model, and improve the real-time performance and decision stability of UAV mission planning in complex dynamic environments through the multi-agent reinforcement learning algorithm MATD3. Task allocation and trajectory planning are carried out by utilizing partially observable Markov decision processes and multi-agent reinforcement learning neural networks.

Benefits of technology

It improves the real-time performance and decision stability of UAV mission planning in complex and dynamic environments, realizes efficient collaborative mission planning of UAV swarms, reduces the number of iterations of intelligent optimization algorithms, and improves mission completion rate and timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875356A_ABST
    Figure CN120875356A_ABST
Patent Text Reader

Abstract

The invention discloses a joint optimization method for task allocation and route planning of multiple unmanned aerial vehicles in a dynamic environment, and the method comprises the steps: building a three-dimensional scene model of task planning of the multiple unmanned aerial vehicles in the dynamic environment, and building an optimization objective function of a single unmanned aerial vehicle based on the three-dimensional scene model; based on an optimization target of a single unmanned aerial vehicle, converting a multi-unmanned aerial vehicle task allocation and flight path planning joint optimization problem in a dynamic environment into a partially observable Markov decision model; and based on the constructed partially observable Markov decision model, performing task allocation and flight path planning on the multiple unmanned aerial vehicles in a dynamic environment by using the trained multi-agent reinforcement learning neural network to obtain an optimal strategy of task allocation and flight path planning of the multiple unmanned aerial vehicles. According to the invention, unmanned aerial vehicle cluster dynamic obstacle avoidance, low-energy-consumption flight path planning and efficient task allocation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, specifically relating to a joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments. Background Technology

[0002] With continuous breakthroughs in UAV technology, its applications have expanded from basic reconnaissance and transportation in safe airspace to complex tasks such as penetration strikes and electromagnetic suppression in adversarial airspace. In the context of information-based and systemic warfare, facing dynamic battlefields and diverse missions, traditional single-unit UAVs, limited by payload and range, are ill-suited for wide-area surveillance and multi-target strikes. This has made UAV swarm collaborative operations the mainstream development direction. Currently, swarm control mainly relies on swarm intelligence algorithms, such as ant colony, particle swarm, and artificial bee colony algorithms, to achieve autonomous collaboration. While these algorithms endow the system with some intelligent characteristics, they also have significant bottlenecks: first, the control model is complex to construct and highly sensitive to parameters; second, the algorithms have inherent defects such as high computational load and susceptibility to local optima, making it difficult to achieve high-precision real-time trajectory planning in complex dynamic environments.

[0003] In recent years, Multi-Agent Deep Reinforcement Learning (MADRL) has provided an innovative path for this endeavor. MADRL constructs a swarm of agents with situational awareness and information processing capabilities, enabling cooperative autonomous learning through environmental interaction. This not only overcomes the limitations of traditional algorithms in collaborative decision-making but also endows UAV swarms with intelligent collaborative capabilities to cope with complex adversarial tasks, opening up new directions for autonomous decision-making in dynamic battlefield environments. Currently, researchers have conducted exploratory research on UAV swarm task allocation and trajectory planning using deep reinforcement learning methods. Lai Jun and Rao Rui proposed a curiosity-driven deep reinforcement learning method based on spatial location annotation, solving the problems of low efficiency and low accuracy in random target search for indoor UAVs. Wang Tao et al. proposed an autonomous navigation control algorithm under multiple constraints based on reinforcement learning methods, namely fuzzy Q-learning, improving the adaptability and robustness of the autonomous navigation control system for unmanned robots in complex environments.

[0004] However, existing research mostly focuses on cluster task allocation and trajectory planning in static environments, without fully considering the impact of uncertain factors such as the dynamic displacement of obstacles and sudden threat sources on the planning system in actual adversarial scenarios. It also ignores the cluster topology reconstruction and task reallocation problems caused by UAVs being damaged due to obstacle avoidance failure. Most algorithm verifications are limited to simplified two-dimensional plane scenarios, and there is a lack of research on adaptive collaborative planning strategies in three-dimensional dynamic uncertain environments. Summary of the Invention

[0005] To address the bottleneck problem of insufficient efficiency in UAV task allocation and trajectory planning decisions in dynamic adversarial scenarios, this invention proposes a joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments.

[0006] The main objective of this invention is to propose a joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments. By improving the traditional TD3 algorithm, a MATD3 framework is constructed, a three-dimensional stochastic dynamic simulation environment is established, and a heuristic reward value model is built to solve the problem of sparse reward values ​​in dynamic environments, thereby improving the real-time performance and decision stability of multi-UAV collaborative task planning in complex dynamic environments.

[0007] This invention establishes a multi-UAV mission planning scenario model under dynamic environments, models the joint optimization problem of UAV mission allocation and trajectory planning under dynamic environments as a partially observable Markov Decision Process (POMDP), and then employs a multi-agent reinforcement learning algorithm (MATD3) with improved traditional TD3 (TwinDelayed Deep Deterministic Policy Gradient) deep reinforcement learning to construct a three-dimensional complex dynamic environment. Furthermore, it uses a corresponding heuristic reward strategy to address the problem of sparse reward values ​​in dynamic environments, thereby improving the real-time mission planning efficiency of UAVs in complex dynamic scenarios.

[0008] To address the problems of existing technologies, this invention proposes a joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments. This method includes:

[0009] S1: Establish a three-dimensional scene model for multi-UAV mission planning in a dynamic environment, and establish an optimization objective function for a single UAV based on the three-dimensional scene model;

[0010] S2: Based on the optimization objective of a single UAV, the joint optimization problem of multi-UAV task allocation and trajectory planning in a dynamic environment is transformed into a partially observable Markov decision model;

[0011] S3: Based on the constructed partially observable Markov decision model, the trained multi-agent reinforcement learning neural network is used to perform task allocation and trajectory planning for multiple UAVs in a dynamic environment, and the optimal strategy for task allocation and trajectory planning for multiple UAVs is obtained.

[0012] Furthermore, the three-dimensional scene model includes multiple drones, multiple dynamic obstacles, multiple static obstacles, and multiple interception zones. The positions of the dynamic obstacles, static obstacles, and interception zones are randomly generated, and multiple drone teams need to collaborate to complete multiple tasks.

[0013] Furthermore, under the aforementioned three-dimensional scene model, a kinematic model of the UAV and a radar identification and interception zone model of the UAV are established.

[0014] The kinematic model of the UAV is represented as follows:

[0015]

[0016] In the formula, Let v(t) represent the three-dimensional coordinates of the drone at time t, v(t) represent the velocity of the drone at time t, and θ(t) represent the turning angle of the drone at time t. This represents the pitch angle of the UAV at time t.

[0017] The radar identification and interception zone model for UAVs in the interception zone includes an interception zone, an escape zone, and a no-escape zone.

[0018] Furthermore, under the aforementioned three-dimensional scene model, an optimization objective function for a single UAV is established with the objectives of avoiding collisions with dynamic and static obstacles, avoiding interception zones, and minimizing UAV energy consumption.

[0019] The specific objective function for the optimization of a single UAV is expressed as follows:

[0020]

[0021] In the formula, f en This represents the total energy consumption of the m-th drone. This represents the energy consumption of the m-th drone from time t-1 to time t.

[0022] Furthermore, in the partially observable Markov decision model, the joint state space S, joint action space A, and reward function R of the multi-UAV system are defined; where,

[0023] The joint state space S is defined as: S = [s1, s2, ..., s...]. n (i = 1, 2, ..., n)

[0024] In the formula, s1 represents the state of the first drone, s2 represents the state of the second drone, and s n This represents the state of the nth drone, where n represents the number of drones.

[0025] Wherein, the state space s of the i-th drone i Represented as:

[0026] s i ={(p i ,v i ,r i ),(p t ,r t ),(p u ,v u ,r u ),(p o ,v o ,r o ),(p d ,r d ),en i}

[0027] In the formula, s i Let p represent the state space of the i-th drone. i The three-dimensional coordinates (x, y) of the i-th UAV i ,y i ,z i ), that is, p i =(x i ,y i ,z i ), v i Let r represent the speed of the i-th drone. i p represents the radius of the i-th drone; t p represents the relative distance from the i-th drone to the target location. t =(x t -x i ,y t -y i ,z t -z i ), (x t ,y t ,z t ) represents the three-dimensional coordinates of the target location, r t p represents the radius from the i-th UAV to the target point; u p represents the relative distance between the i-th drone and other drones. u =(x u -x i ,y u -y i ,z u -z i ), (x u ,y u ,z u ) represents the three-dimensional coordinates of other drones, v u r represents the speed of other dronesu p represents the radius of other drones. o p represents the relative distance between the dynamic or static obstacle and the i-th drone. o =(x o -x i ,y o -y i ,z o -z i ), (x o ,y o ,z o ) represents the three-dimensional coordinates of a dynamic or static obstacle, v o r represents the speed of a dynamic or static obstacle. o These represent the radii of dynamic or static obstacles, respectively; p d p represents the relative distance between the interception zone and the i-th drone. d =(x d -x i ,y d -y i ,z d -z i ), (x d ,y d ,z d ) represents the three-dimensional coordinates of the interception zone, r d Represents the radius of the interception zone, en i This represents the remaining energy consumption of the i-th drone.

[0028] The joint action space A is defined as: A = [a1, a2, ..., a...]. n (i = 1, 2, ..., n)

[0029] a1 represents the motion space of the first drone, a2 represents the motion space of the second drone, a n Let n represent the action space of the nth drone, where n represents the number of drones.

[0030] Wherein, the motion space a of the i-th drone i Represented as:

[0031] In the formula, Δθ represents the drone's turn angle increment, Δv represents the drone's pitch angle increment, and Δv represents the drone's speed increment.

[0032] The formula for calculating the real-time reward function R is: R = r1 + r2 + r3 + r4

[0033] In the formula, r1 represents the energy reward of the drone, r2 represents the reward obtained by the drone from avoiding dynamic obstacles, static obstacles and collisions between drones, r3 represents the reward obtained by the drone from avoiding the interception zone, and r4 represents the reward obtained by the drone from reaching the target point.

[0034] The energy consumption reward r1 of the drone is calculated using the following formula:

[0035]

[0036] In the formula, P u d represents the energy consumption of the drone per unit time. t-1,s v represents the remaining distance from the UAV to the target point s at time t-1. t-1,t d represents the speed of the drone. init β represents the initial distance between the drone and the target point, and β is the attenuation coefficient.

[0037] The specific formula for calculating the reward r2 is as follows:

[0038]

[0039] In the formula, dist uu Dist represents the distance between the drone and the nearest drone. uo R represents the distance between the drone and moving or stationary obstacles. u R represents the radius of the drone. o D represents the radius of the obstacle. u Represents the width of the pseudo-collision region, -(2R) u +D u -dist uu ) represents the reward when drones collide, -(R u +R o +D u -dist uo This represents the reward between the drone and the obstacle. When the distance between drones and each other, as well as the distance between a drone and an obstacle, is less than the sum of their radii and the width of the pseudo-collision zone, a negative reward will be given.

[0040] The specific formula for calculating the reward r3 obtained by the drone for avoiding the interception zone is as follows:

[0041]

[0042] Where r3 represents the reward the drone receives for avoiding the interception zone, T p T represents the response value of the interception zone to the drone. σ This represents the threat threshold.

[0043] The reward r4 obtained by the drone when it reaches the target point is calculated using the following formula:

[0044]

[0045] In the formula, r4 represents the reward received by the drone when it reaches the target point. r represents the distance from the i-th UAV to the target point at time t. i r represents the radius of the i-th drone. t Represents the radius of the target point.

[0046] Furthermore, in the multi-agent reinforcement learning neural network, each UAV agent has an Actor network, a Target-Actor network, two Critic networks, and two Target-Critic networks.

[0047] Furthermore, the training process of the multi-agent reinforcement learning neural network includes:

[0048] Step a: Initialize the multi-agent reinforcement learning neural network, which specifically includes initializing each network parameter of each UAV agent, initializing the experience pool size D, the number of samples B, and the training episodes;

[0049] Step b: For each training round, obtain the state space s of each UAV agent through the simulation environment. i The action space a of each drone agent is obtained based on the Actor network. i The next state s is calculated based on the kinematic model of the UAV. i ', calculate the reward value r obtained by each drone agent. i .

[0050] Sample <s i ,a i ,s i ′,r i > Store in the experience pool. If the number of samples in the experience pool is greater than the number of samples used for training, proceed to step c; otherwise, continue to step b.

[0051] Step c: Randomly select B samples from the sample pool to form the training batch, and use the loss function... Update both Critic networks separately;

[0052] Step d: Evaluate the current UAV strategy based on the updated two Critic networks, maximize the Q-value predicted by the Critic network, and optimize the Actor network based on the maximum Q-value predicted by the Critic network.

[0053] Furthermore, during the training process of the multi-agent reinforcement learning neural network, the loss function The calculation formula is:

[0054]

[0055] In the formula, E (s,a,r,s′)~D This represents the expectation of sampling (s,a,r,s′) from the priority replay buffer D, where s represents the joint state space of the UAV, a represents the joint action of the UAV, r represents the joint reward of the UAV, and s′ represents the next joint state of the UAV. This represents the Q-value estimate of the j-th Critic network for the i-th UAV agent on joint action s and joint action a; y i Represents the target Q value, r i This indicates that the i-th drone agent performs action a in the current state s. i Then, the instantaneous reward from environmental feedback; γ represents the discount factor. The Q-value represents the target network.

[0056] Furthermore, during the training process of the multi-agent reinforcement learning neural network, the Actor network is optimized based on the maximum Q-value predicted by the Critic network, and its specific update formula is expressed as follows:

[0057] φ′ i ←τφ i +(1-τ)φ′ i

[0058]

[0059] In the formula, φ′ represents the network parameters of the Target-Actor network for the i-th UAV agent, τ is the soft update coefficient used to control the update speed, and φ i This represents the current Actor network parameters of the i-th drone agent. This represents the first Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the first Critic network for the i-th drone agent. This represents the second Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the second Critic network for the i-th drone agent.

[0060] The beneficial effects of this invention are:

[0061] First, this invention models the application scenario of multi-UAV task allocation and trajectory planning in dynamic environments, and transforms the joint optimization problem of multi-UAV task allocation and trajectory planning in dynamic environments into a partially observable Markov decision model.

[0062] Secondly, this invention employs the Multi-Agent Dual-Delay Deep Deterministic Policy Gradient (MATD3) algorithm, using a centralized training and distributed execution approach. This allows UAV agents to interact and learn from each other, enabling them to converge to higher reward values ​​and task completion rates in a shorter time. This improves the timeliness of task planning strategy decisions, reduces the number of iterations in the intelligent optimization algorithm, and enhances the timeliness of the method.

[0063] Furthermore, this invention constructs a reward value function based on a policy set, which improves the training efficiency and stability of reinforcement learning, solves the problem of reward value sparsity in dynamic environments, and enhances the convergence speed of multi-agent reinforcement learning algorithms.

[0064] In summary, this invention can achieve dynamic obstacle avoidance for UAV swarms, low-energy trajectory planning, and efficient task allocation. Attached Figure Description

[0065] Figure 1 This is a flowchart illustrating the steps of an embodiment of the present invention;

[0066] Figure 2 This is a schematic diagram of a three-dimensional scene model for multi-UAV mission planning in a dynamic environment, as shown in an embodiment of the present invention.

[0067] Figure 3 This is a schematic diagram illustrating the training process of a multi-agent reinforcement learning neural network in an embodiment of the present invention.

[0068] Figure 4 This is a comparison chart of the simulation results obtained by this invention and the existing MADDPG and DDPG algorithms. Detailed Implementation

[0069] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.

[0070] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0071] This invention proposes a joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments, referring to... Figure 1 As shown, the method includes:

[0072] S1: Establish a three-dimensional scene model for multi-UAV mission planning in a dynamic environment, and establish an optimization objective function for a single UAV based on the three-dimensional scene model.

[0073] The three-dimensional scene model includes multiple drones, multiple dynamic obstacles, multiple static obstacles, and multiple interception zones. The positions of the dynamic obstacles, static obstacles, and interception zones are randomly generated, and multiple drone teams need to collaborate to complete multiple tasks.

[0074] Specifically, refer to Figure 2 As shown, a 3D scene model for multi-UAV mission planning in a dynamic environment is established in a simulation platform. For example, scene modeling is performed in the PyCharm simulation platform. This scene model includes 4 UAVs, 6 obstacles (2 dynamic obstacles and 4 static obstacles), 2 interception zones, and 4 mission points. The positions of the dynamic obstacles, static obstacles, and interception zones are randomly generated. The mission objective of the multi-UAVs is to reach the designated target point while avoiding all dynamic and static obstacles and interception zones.

[0075] S101: Under the aforementioned three-dimensional scene model, establish the kinematic model of the UAV, specifically as follows:

[0076]

[0077] In the formula, Let v(t) represent the three-dimensional coordinates of the drone at time t, v(t) represent the velocity of the drone at time t, and θ(t) represent the turning angle of the drone at time t. This represents the pitch angle of the UAV at time t.

[0078] Multiple drone teams collaborate to complete several pre-set tasks. During mission execution, the drones must avoid multiple dynamic and static obstacles, as well as the dangers of the interception zone. If a drone collides with a dynamic or static obstacle, the mission fails. The interception zone is typically identified by radar; therefore, each interception zone is limited by the radar's maximum detection range and maximum interception area. Each drone is also limited by its escape zone.

[0079] S102: Under the aforementioned three-dimensional scene model, establish a radar identification and interception zone model for the UAV within the interception zone.

[0080] Specifically, refer to Figure 2 As shown, based on the radar's maximum detection range and the radius of the maximum interception zone, the distance between the UAV and the interception zone is divided into three zones from largest to smallest: no interception zone, escape zone, and no escape zone. The interception zone is a dangerous area for the UAV.

[0081] S103: Establish a response value model for the interception zone to the UAV, specifically as follows:

[0082]

[0083] In the formula, T p R represents the response value of the interception zone to the drone, D represents the distance between the interception zone and the drone, and R represents the response value of the drone to the drone. Rmax R represents the maximum detection range that the interceptor radar can detect; Mmax R represents the maximum radius of the hazardous impact on the drone within the interception zone. Mkmax This indicates the maximum extent of the no-escape zone.

[0084] In the aforementioned 3D scene model, based on the drone's response value to the interception zone, and the moving speed and direction of dynamic or static obstacles, the drone needs to avoid these areas and evade interception. Since the drone's battery energy is limited, it must not only avoid dynamic or static obstacles and interception zones, but also choose the optimal strategy to minimize energy consumption and ultimately complete the mission.

[0085] S104: To avoid collisions with dynamic and static obstacles, avoid interception zones, and minimize UAV energy consumption, an optimization objective function for a single UAV is established, specifically expressed as follows:

[0086]

[0087] In the formula, f en This represents the total energy consumption of the m-th drone. This represents the energy consumption of the m-th drone from time t-1 to time t.

[0088] Step S1 completes the 3D scene modeling. This modeling, through unified quantification of three core constraints—dynamic obstacle avoidance, interception avoidance, and energy consumption optimization—enables the UAV to achieve multi-target collaborative decision-making in complex adversarial environments. When battery power is low, energy-saving paths are prioritized (extending mission endurance), while high-risk maneuvers (such as penetrating interception zones) are permitted when battery power is high.

[0089] S2: Based on the optimization objective of a single UAV, the joint optimization problem of multi-UAV task allocation and trajectory planning in a dynamic environment is transformed into a partially observable Markov decision model.

[0090] S201: Establish a joint state-space model for multiple unmanned aerial vehicle (UAV) systems, specifically:

[0091] Define the joint state space S of a multi-UAV system as: S = [s1, s2, ..., s n (i = 1, 2, ..., n)

[0092] In the formula, s1 represents the state of the first drone, s2 represents the state of the second drone, and s n This represents the state of the nth drone, where n represents the number of drones.

[0093] The state space of the i-th UAV consists of the UAV's internal state and the maximum detection distance d of the onboard sensors. det The environmental information within the space is composed of, and therefore, the state space of the i-th UAV can be defined as:

[0094] s i ={(p i ,v i ,r i ),(p t ,r t ),(p u ,v u ,r u ),(p o ,v o ,r o ),(p d ,r d ),en i}

[0095] In the formula, s i Let p represent the state space of the i-th drone. i The three-dimensional coordinates (x, y) of the i-th UAV i ,y i ,z i ), that is, p i =(x i ,y i ,z i ), v iLet r represent the speed of the i-th drone. i p represents the radius of the i-th drone; t p represents the relative distance from the i-th drone to the target location. t =(x t -x i ,y t -y i ,z t -z i ), (x t ,y t ,z t ) represents the three-dimensional coordinates of the target location, r t p represents the radius from the i-th UAV to the target point; u p represents the relative distance between the i-th drone and other drones. u =(x u -x i ,y u -y i ,z u -z i ), (x u ,y u ,z u ) represents the three-dimensional coordinates of other drones, v u r represents the speed of other drones u p represents the radius of other drones. o p represents the relative distance between the dynamic or static obstacle and the i-th drone. o =(x o -x i ,y o -y i ,z o -z i ), (x o ,y o ,z o ) represents the three-dimensional coordinates of a dynamic or static obstacle, v o r represents the speed of a dynamic or static obstacle. o These represent the radii of dynamic or static obstacles, respectively; p d p represents the relative distance between the interception zone and the i-th drone. d =(x d -x i ,y d -y i ,z d -z i ), (x d ,y d ,z d ) represents the three-dimensional coordinates of the interception zone, rd Represents the radius of the interception zone, en i This represents the remaining energy consumption of the i-th drone.

[0096] S202: Establish a joint action space model for multiple unmanned aerial vehicle systems;

[0097] Define the joint action space A of a multi-UAV system as: A = [a1, a2, ..., a n (i = 1, 2, ..., n)

[0098] a1 represents the motion space of the first drone, a2 represents the motion space of the second drone, a n Let n represent the action space of the nth drone, where n represents the number of drones.

[0099] The action space of the i-th drone is represented as follows:

[0100] In the formula, Δθ represents the drone's turn angle increment, Δv represents the drone's pitch angle increment, and Δv represents the drone's speed increment.

[0101] S203: Construct a reward value function for a single drone based on a set of policies.

[0102] With the optimization objective of avoiding collisions with dynamic and static obstacles, avoiding interception zones, and minimizing drone energy consumption (i.e., the optimization objective function based on a single drone), and under the incentive of promoting the drone to reach the target position as quickly as possible, a reward structure is used to minimize drone energy consumption: traverse all target points, calculate the initial distance between the drones closest to the target point, introduce a target distance decay factor into this distance, weaken the energy consumption constraint when moving away from the target, and incentivize the drone to accelerate to shorten the time to reach the target point; strengthen energy consumption optimization when approaching the target.

[0103] The specific formula for calculating the energy consumption bonus value r1 of the drone is as follows:

[0104]

[0105] In the formula, P u d represents the energy consumption of the drone per unit time. t-1,s v represents the remaining distance from the UAV to the target point s at time t-1. t-1,t d represents the speed of the drone. init β represents the initial distance between the drone and the target point, and β is the attenuation coefficient.

[0106] To optimize collision avoidance, the drones maintain spatial coordination, meaning they avoid collisions simultaneously. When a collision occurs, the agent receives a negative reward. A false collision zone is added to enhance the collision warning mechanism. The reward structure for collision avoidance among drones is trained by iterating through all other drones and calculating the collision avoidance reward value for each drone.

[0107] The reward r2 obtained by a drone for avoiding dynamic obstacles, static obstacles, and collisions between drones is calculated using the following formula:

[0108]

[0109] In the formula, dist uu Dist represents the distance between the drone and the nearest drone. uo R represents the distance between the drone and moving or stationary obstacles. u R represents the radius of the drone. o D represents the radius of the obstacle. u Represents the width of the pseudo-collision region, -(2R) u +D u -dist uu ) represents the reward when drones collide, -(R u +R o +D u -dist uo This represents the reward between the drone and the obstacle. When the distance between drones and each other, as well as the distance between a drone and an obstacle, is less than the sum of their radii and the width of the pseudo-collision zone, a negative reward will be given.

[0110] Based on radar identification of the UAV within the interception zone and the interception zone model, the response value T of the UAV within the interception zone is calculated. p Then according to T p s The specific formula for calculating the reward value r3 for avoiding the interception zone is as follows:

[0111]

[0112] Where r3 represents the reward the drone receives for avoiding the interception zone, T p T represents the response value of the interception zone to the drone. σ This represents the threat threshold.

[0113] If the drone reaches the target point, a reward of r4 is given, the specific calculation formula is as follows:

[0114]

[0115] In the formula, r4 represents the reward received by the drone when it reaches the target point. r represents the distance from the i-th UAV to the target point at time t. i r represents the radius of the i-th drone. t Represents the radius of the target point.

[0116] The real-time reward function R for a single drone consists of the above four parts, namely R = r1 + r2 + r3 + r4.

[0117] Step 2 constructs a multi-UAV joint POMDP model, incorporating dynamic obstacle avoidance, interception avoidance, energy consumption optimization, and mission collaboration into a unified decision-making framework, bringing three core advantages: First, the joint state space (integrating position, speed, energy consumption, obstacle, and interception zone information) enables full-domain situational awareness and partial observable environment modeling; second, a hierarchical reward mechanism dynamically balances multi-target conflicts—adaptively adjusting energy consumption weights through distance attenuation factors (weakening energy-saving incentives for acceleration at long range, and strengthening energy-saving precision maneuvers at close range), combined with false collision zone warnings and radar response values ​​to quantify interception threats, achieving proactive safety protection; finally, the joint action space supports distributed collaborative trajectory planning, optimizing formation energy efficiency while avoiding collisions and interception zones, ultimately achieving a simultaneous leap in safe passage, swarm collaboration, and endurance, providing a verifiable autonomous decision-making kernel for complex adversarial environments.

[0118] S3: Based on the constructed partially observable Markov decision model, the trained multi-agent reinforcement learning neural network is used to perform task allocation and trajectory planning for multiple UAVs in a dynamic environment, and the optimal strategy for task allocation and trajectory planning for multiple UAVs is obtained.

[0119] S301: Construct a multi-agent reinforcement learning neural network.

[0120] The system architecture of the multi-agent reinforcement learning neural network is divided into a target optimization layer, a model training layer, and an action execution layer. The target optimization layer is designed with a reward function to minimize collisions between drones and between drones and the interception area, and to minimize drone energy consumption. The model training layer consists of a training environment and a training algorithm. The environment includes drones, targets, and the interception area, and the training algorithm is the multi-agent reinforcement learning algorithm MATD3.

[0121] The system utilizes a dual Critic network to suppress Q-value overestimation, improves stability by delaying Actor network updates and adding target policy noise, and optimizes the policy by incorporating global state during centralized training. The execution layer deploys the trained Actor network to a real UAV, enabling it to generate actions in real time based solely on local observations, thus achieving decentralized autonomous control and ultimately completing efficient collaborative obstacle avoidance and trajectory planning tasks.

[0122] It's important to note that the Multi-Agent Twin Delayed DeepDeterministic Policy Gradient (MATD3) algorithm employs a centralized training approach. During training, each agent's Critic network can access the states and actions of all agents to maximize the cumulative reward for each agent. However, during execution, agents select actions solely based on their own observations, achieving distributed execution. Each agent in MATD3 has two independent Critic networks (Q-value functions) to reduce bias in Q-value estimation. These Q-value functions are updated similarly to those in TD3, calculating the target Q-value and updating the network parameters using the mean squared error (MSE) loss function. To further reduce estimation bias, MATD3 introduces noise into each agent's actions when calculating the target Q-value. This prevents the policy from overfitting to certain actions, thereby improving the policy's robustness.

[0123] Reference Figure 3 As shown, in the multi-agent reinforcement learning neural network, each UAV agent has a policy network (i.e., Actor network), a target evaluation network (i.e., Target-Actor network), two evaluation networks (i.e., Critic network) and two target evaluation networks (i.e., Target-Critic network). Figure 3 In this context, losses L1 and L2 are based on the following loss function. Calculated.

[0124] S302: Reference Figure 3 As shown, the training process of the multi-agent reinforcement learning neural network includes:

[0125] Step a: Initialize the multi-agent reinforcement learning neural network, specifically:

[0126] Initialize the network parameters for each drone agent, including: network parameters φ of the Actor network, network parameters φ′ of the Target-Actor network, network parameters θ1 of the first Critic network, network parameters θ2 of the second Critic network, network parameters θ1′ of the first Target-Critic network, and network parameters θ2′ of the second Target-Critic network; initialize the experience pool size D, the number of samples B, and the training episodes.

[0127] Step b: For each training round, obtain the state space s of each UAV agent through the simulation environment. i The action space a of each drone agent is obtained based on the Actor network. iThe next state s is calculated based on the kinematic model of the UAV. i ', calculate the reward value r obtained by each drone agent. i .

[0128] Sample <s i ,a i ,s i ′,r i > Store the samples in the experience pool. If the number of samples in the experience pool is greater than the number of samples used for training, proceed to step c; otherwise, continue to step b.

[0129] Step c: Randomly select B samples from the sample pool to form the training batch, and use the loss function... Update both Critic networks separately.

[0130] The loss function The calculation formula is:

[0131]

[0132] In the formula, E (s,a,r,s′)~D This represents the expectation of sampling (s,a,r,s′) from the priority replay buffer D, where s represents the joint state space of the UAV, a represents the joint action of the UAV, r represents the joint reward of the UAV, and s′ represents the next joint state of the UAV. This represents the Q-value estimate of the j-th Critic network for the i-th UAV agent on joint action s and joint action a; y i Represents the target Q value, r i This indicates that the i-th drone agent performs action a in the current state s. i Then, the instantaneous reward from environmental feedback; γ represents the discount factor, which indicates the importance of decaying future rewards, to avoid training divergence caused by infinite accumulation of rewards; The Q value represents the target network. The Q value is estimated using two target Critic networks, and the minimum value is taken to suppress overestimation.

[0133] Step d: Evaluate the current UAV strategy based on the updated two Critic networks, maximize the Q-value predicted by the Critic network, and optimize the Actor network based on the maximum Q-value predicted by the Critic network.

[0134] The Actor network's updates depend on the Critic network's evaluation of the current policy. Specifically, the policy is optimized by maximizing the Q-value predicted by the Critic network.

[0135]

[0136] in, This represents the direction of parameter updates for the Actor network, with the goal of maximizing the long-term return J(π). i ); E s,a [·] represents the sampled state s and joint motion a from the experience replay pool. i Make the expectation of a; The gradient estimate of the action of the first Critic network representing the i-th drone agent indicates the direction of the action's improvement on the Q-value; This represents the gradient of the Actor network parameters with respect to the action.

[0137] To improve stability, the Actor network is updated less frequently than the Critic network. The Critic network updates once for every two updates, and the specific update formula is as follows:

[0138] φ′ i ←τφ i +(1-τ)φ′ i

[0139]

[0140] In the formula, φ′ represents the network parameters of the Target-Actor network for the i-th UAV agent, τ is the soft update coefficient used to control the update speed, and φ i This represents the current Actor network parameters of the i-th drone agent. This represents the first Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the first Critic network for the i-th drone agent. This represents the second Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the second Critic network for the i-th drone agent.

[0141] The target network parameters are slowly updated from the current network using the above update formula through an exponential moving average.

[0142] Experimental verification:

[0143] Experimental environment:

[0144] Both the simulation and the training of the reinforcement learning model were developed in the Python 3.9.19 and PyTorch 1.13.0 environment. The specific hardware configuration of the training environment is shown in Table 1.

[0145] Table 1. Parameter Configuration Table for Simulation Environment

[0146] parameter Specific content CPU 12th Gen Intel(R)Core(TM)i7-12700 2.10GHz Onboard RAM 16 GB system Windows 11 Home Chinese Edition

[0147] During the experiment, a multi-agent reinforcement learning network was trained, and the parameters of the MATD3 algorithm were set as shown in Table 2.

[0148] Table 2 Parameter Table of MATD3 Algorithm

[0149]

[0150] Experimental results:

[0151] Simulation experiments were conducted comparing the method of this invention with existing MADDPG and DDPG algorithms. The experimental results showed a comparison of the reward values ​​of the UAV agent. Figure 4 As shown. After 10,000 rounds of learning, the reward value curves of multiple drones are as follows. Figure 4 As shown, although the algorithms exhibit some impact due to random noise during training, all three algorithms show a convergence trend. However, their convergence speed and time differ somewhat. The rewards for MATD3 and MADDPG begin to increase around episode 1000. DDPG has the slowest convergence speed, only showing a convergence trend around episode 4000. MATD3 and MADDPG converge earlier than DDPG, with MATD3 having the fastest convergence speed.

[0152] It should be noted that the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient) is a multi-agent cooperative algorithm based on deep reinforcement learning, which aims to solve the continuous control problem in multi-agent systems.

[0153] It should be noted that the DDPG (Deep Deterministic Policy Gradient) algorithm is a deep reinforcement learning-based algorithm suitable for solving problems in continuous action spaces.

[0154] Numerical experiments were conducted to compare and analyze indicators such as the reward value after algorithm convergence, average task success rate, and algorithm training efficiency. The experimental results show that the present invention is superior to the UAV task allocation and trajectory planning optimization method based on traditional optimization algorithms (such as DDPG and MADDPG algorithms) in the above indicators.

[0155] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.

[0156] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A joint optimization method for multi-UAV task allocation and trajectory planning in dynamic environments, characterized in that, include: A three-dimensional scene model for multi-UAV mission planning in a dynamic environment is established, and an optimization objective function for a single UAV is established based on the three-dimensional scene model. Based on the optimization objective of a single UAV, the joint optimization problem of multi-UAV task allocation and trajectory planning in a dynamic environment is transformed into a partially observable Markov decision model. Based on the constructed partially observable Markov decision model, the trained multi-agent reinforcement learning neural network is used to perform task allocation and trajectory planning for multiple UAVs in dynamic environments, and the optimal strategy for task allocation and trajectory planning for multiple UAVs is obtained.

2. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 1, characterized in that, The three-dimensional scene model includes multiple drones, multiple dynamic obstacles, multiple static obstacles, and multiple interception zones. The positions of the dynamic obstacles, static obstacles, and interception zones are randomly generated, and multiple drone teams need to collaborate to complete multiple tasks.

3. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 2, characterized in that, Under the aforementioned three-dimensional scene model, a kinematic model of the UAV and a radar identification and interception zone model of the UAV are established.

4. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 3, characterized in that, In the aforementioned three-dimensional scene model, an optimization objective function for a single UAV is established with the goals of avoiding collisions with dynamic and static obstacles, avoiding interception zones, and minimizing UAV energy consumption.

5. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 1 or 4, characterized in that, In the partially observable Markov decision model, the joint state space S, joint action space A, and reward function R of a single UAV system are defined; where, The state space of the i-th drone is s i Represented as: s i ={(p i ,v i ,r i ),(p t ,r t ),(p u ,v u ,r u ),(p o ,v o ,r o ),(p d ,r d ),a i } In the formula, s i Let p represent the state space of the i-th drone. i v represents the three-dimensional coordinates of the i-th drone. i Let r represent the speed of the i-th drone. i p represents the radius of the i-th drone; t r represents the relative distance from the i-th UAV to the target location. t p represents the radius from the i-th UAV to the target point; u v represents the relative distance between the i-th drone and other drones. u r represents the speed of other drones u p represents the radius of other drones. o v represents the relative distance between the i-th drone and a dynamic or static obstacle. o r represents the speed of a dynamic or static obstacle. o These represent the radii of dynamic or static obstacles, respectively; p d r represents the relative distance between the interception zone and the i-th UAV. d Represents the radius of the interception zone, en i This represents the remaining energy consumption of the i-th drone; The action space a of the i-th drone i Represented as: In the formula, Δθ represents the drone's turn angle increment, Δv represents the drone's pitch angle increment, and Δv represents the drone's speed increment. The formula for calculating the reward value function R of a single drone is: R = r1 + r2 + r3 + r4 In the formula, r1 represents the energy reward of the drone, r2 represents the reward obtained by the drone from avoiding dynamic obstacles, static obstacles and collisions between drones, r3 represents the reward obtained by the drone from avoiding the interception zone, and r4 represents the reward obtained by the drone from reaching the target point.

6. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 1, characterized in that, In the multi-agent reinforcement learning neural network, each UAV agent has one Actor network, one Target-Actor network, two Critic networks, and two Target-Critic networks.

7. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 1 or 6, characterized in that, The training process of the multi-agent reinforcement learning neural network includes: Step a: Initialize the multi-agent reinforcement learning neural network, which specifically includes initializing each network parameter of each UAV agent, initializing the experience pool size D, the number of samples B, and the training episodes; Step b: For each training round, obtain the state space s of each UAV agent through the simulation environment. i The action space a of each drone agent is obtained based on the Actor network. i The next state s is calculated based on the kinematic model of the UAV. i ', calculate the reward value r obtained by each drone agent. i . Sample i ,a i ,s i ′,r i > Store in the experience pool. If the number of samples in the experience pool is greater than the number of samples used for training, proceed to step c; otherwise, continue to step b.​ Step c: Randomly select B samples from the sample pool to form the training batch, and use the loss function... Update both Critic networks separately; Step d: Evaluate the current UAV strategy based on the updated two Critic networks, maximize the Q-value predicted by the Critic network, and optimize the Actor network based on the maximum Q-value predicted by the Critic network.

8. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 7, characterized in that, During the training process of the multi-agent reinforcement learning neural network, the loss function The calculation formula is: In the formula, E (s,a,r,s′)~D This represents the expectation of sampling (s,a,r,s′) from the priority replay buffer D, where s represents the joint state space of the UAV, a represents the joint action of the UAV, r represents the joint reward of the UAV, and s′ represents the next joint state of the UAV. This represents the Q-value estimate of the j-th Critic network for the i-th UAV agent on joint action s and joint action a; y i Represents the target Q value, r i This indicates that the i-th drone agent performs action a in the current state s. i Then, the instantaneous reward from environmental feedback; γ represents the discount factor. The Q-value represents the target network.

9. The joint optimization method for multi-UAV task allocation and trajectory planning in a dynamic environment according to claim 7, characterized in that, During the training process of the multi-agent reinforcement learning neural network, the Actor network is optimized based on the maximum Q-value predicted by the Critic network, and its specific update formula is expressed as follows: f' i ←tf i +(1-τ)φ′ i θ′ i 1 ←tth i 1 +(1-τ)θ′ i 1 θ′ i 2 ←tth i 2 +(1-τ)θ′ i 2 In the formula, φ′ represents the network parameters of the Target-Actor network for the i-th UAV agent, τ is the soft update coefficient used to control the update speed, and φ i This represents the current Actor network parameters of the i-th drone agent. This represents the first Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the first Critic network for the i-th drone agent. This represents the second Target-Critic network parameter of the i-th drone agent. This represents the network parameters of the second Critic network for the i-th drone agent.