Multi-UAV Multi-Target Cooperative Tracking Control Method Based on MAPPO Algorithm

By applying deep reinforcement learning in multi-UAV systems, multi-objective collaborative tracking and control of multi-UAVs is realized, solving the problems of large amount of computing and high collision risks in traditional methods, and improving the adaptability and robustness of the system.

CN115509251BActive Publication Date: 2025-06-10SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211017296.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-06-10
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

The traditional multi-drone multi-target tracking method has a large amount of calculation and is difficult to cope with the collisions between drones and environmental changes during multi-target tracking.

Method used

The MAPPO algorithm based on deep reinforcement learning is adopted to perform multi-drone multi-objective collaborative tracking and control through a distributed framework, reducing the requirements of the drone for communication and computing capabilities, and achieving adaptability and robustness.

Benefits of technology

It effectively solved the problems of large computing volume, high collision risk and difficult to deal with environmental changes during multi-drone tracking, and achieved efficient and stable coordinated tracking of multi-drone multi-targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115509251B_ABST
    Figure CN115509251B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm. The method includes modeling the multi-UAV target tracking process; preprocessing of environment standardization and data normalization; multi-target task allocation; designing state, action value functions, and reward return functions; designing a deep neural network structure; inputting the local observation states of each UAV into the multi-UAV multi-target cooperative tracking controller to obtain the action control quantities of each UAV, and controlling each UAV to work according to each action control quantity to complete the task of controlling multiple UAVs to perform cooperative tracking on multiple targets. The method of the present invention adopts a distributed framework, reduces the requirements of UAVs for communication and computing capabilities, effectively solves the problems of large computational amount in traditional multi-UAV multi-target tracking methods, possible mutual influence or collision between UAVs, and difficulty in coping with environmental changes that require real-time solution, and has strong adaptability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV control, and particularly relates to a multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm of deep reinforcement learning. Background Art

[0002] In recent years, the research on the cooperation between UAVs as agents has received extensive attention from scholars at home and abroad; at the same time, in the application fields of the UAV industry, the realization of many tasks is based on multi-target tracking. For example, in the police field, UAVs are used to track the vehicles driven by criminal suspects; in the military field, UAVs are used to track multiple enemy personnel and conduct strikes. Multi-UAV cooperative tracking can effectively reduce the escape probability of the tracked targets and improve the success rate of task execution. Therefore, multi-UAV cooperative tracking of multiple targets has become an important research direction.

[0003] Traditional multi-UAV multi-target tracking methods adopt a hierarchical control algorithm: the upper-layer controller is a cooperative trajectory tracking controller, which uses formation control algorithms, such as leader-follower method, artificial potential field method, virtual structure method, etc., to calculate a series of waypoints of each UAV in the tracking process based on the UAV system state information and target state information, form a flight path trajectory, and output trajectory information; the lower-layer controller is the attitude controller of the UAV, which calculates the linear velocity and yaw angular velocity of the UAV during the process of reaching the next waypoint based on the position of the next waypoint calculated by the upper-layer controller, and maintains the stability of the roll angle and pitch angle during the flight process, and outputs a speed control command. In particular, when the tracked target is in a moving state, the system needs to continuously calculate and optimize the trajectory waypoints. If the algorithm is complex, it will consume a large amount of computing resources; in addition, when tracking multiple targets, the probability of collision between UAVs will increase greatly, and the traditional control algorithm cannot well solve the cooperative problem in the multi-UAV tracking process and cannot give full play to the advantages of multi-UAV tracking. To address the above problems, recently, research has applied self-learning strategies based on intelligent algorithms to the field of multi-UAV target tracking control. Intelligent algorithms include swarm intelligence, imitation learning, deep reinforcement learning, etc., and the self-learning strategy refers to optimizing the structure or parameters of its own policy model through its own experience. Such algorithms regard each UAV in the system as an intelligent agent with independence and autonomy, interact with the environment, and exhibit the characteristics of strong adaptability and the ability to handle complex and changeable task scenarios.

[0004] Deep reinforcement learning is a branch of machine learning that combines the perception ability of deep learning and the decision-making ability of reinforcement learning, and has been widely applied in many challenging fields, such as autonomous driving, computer vision, medical diagnosis, and robot control. When dealing with a series of environmental perception and control decision-making problems, its learning process has a certain generality and can be expressed as: (1) The interaction between the agent and the environment occurs at all times, and the agent perceives and observes high-dimensional targets through deep learning methods to obtain specific state information in the current environment; (2) The value function of each action is evaluated based on the expected reward (to motivate the agent), and an adaptive strategy is obtained through reinforcement learning methods to map the current state to the corresponding action; (3) The environment makes corresponding feedback to this action, and the agent uses this to make observations at the next moment. Through the continuous cycle of the above process, the agent can finally obtain the optimal action strategy to complete the established task.

[0005] The MAPPO algorithm is a multi-agent proximal policy optimization deep reinforcement learning algorithm, which is a variant of the PPO algorithm applied to multi-agent tasks. (The PPO algorithm, also known as the proximal policy optimization algorithm, is a policy gradient optimization algorithm based on the Actor-Critic (AC) framework proposed by OpenAI in 2017. By proposing an importance sampling and gradient parameter clipping objective function, it solves the problems of difficult step size determination and excessive update differences in the policy gradient algorithm, and realizes a good solution to continuous control problems; the PPO algorithm is an on-policy algorithm. The Actor network, also called the Policy network, receives local observations (Observations) and outputs actions (Actions). The Critic network, also called the Value network, receives states (States) and outputs action values (Values) to evaluate the quality of the actions output by the Actor network. MAPPO also adopts the Actor-Critic architecture, but the difference is that it is an algorithm based on the centralized training and decentralized execution (CTDE) framework. At this time, the Critic network learns a centralized value function. That is, after training, each agent can generate an optimal action through the action policy function generated by its own Actor network based on its local observation state, and finally combine them into a multi-agent joint action to complete the task. Cai Zhihao et al. disclosed a "method for controlling a drone target tracking based on the reinforcement learning PPO algorithm" in the Chinese authorized invention patent CN111580544B, which uses an integrated controller to replace the traditional inner and outer loop controllers, and has the characteristics of good robustness and small computational complexity. However, this method can only achieve the tracking control of a single drone and cannot perform multi-drone multi-target cooperative tracking. And there is no record of a method for multi-drone multi-target cooperative tracking control using the MAPPO algorithm.) Summary of the Invention

[0006] Aiming at the control problem of multi-drone cooperative tracking of multiple targets, the present invention proposes a multi-drone multi-target cooperative tracking control method based on the multi-agent deep reinforcement learning MAPPO algorithm, which can perform multi-drone multi-target cooperative tracking, and effectively solves the problems of large computational complexity, possible mutual influence or collision between drones, and difficulty in coping with environmental changes that require real-time calculation in traditional multi-drone multi-target tracking methods. The method of the present invention adopts a distributed framework, reduces the requirements of drones for communication and computing capabilities, and has strong adaptability and robustness.

[0007] To achieve the object of the present invention, the multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm provided by the present invention includes the following steps:

[0008] Step 1: Model the multi-UAV target tracking process, including establishing a six-degree-of-freedom kinematic model of the UAV, a constant turn rate and speed model of the moving target, and a distributed partially observable Markov decision model;

[0009] Step 2: Perform environmental standardization and data normalization preprocessing;

[0010] Step 3: Allocate the multi-UAV tracking tasks, divide N UAVs into m groups to track m targets, and at the same time, based on the distances of each UAV from the initial positions of each target, calculate the sum of the Euclidean distance costs for all UAVs to track each target, where N≥2 and m≥2;

[0011] Step 4: Construct a state value function, an action value function, and a reward return function;

[0012] Step 5: Construct a deep neural network structure, including a policy network structure and a value network structure. The policy network is used to output the action control quantities of each UAV according to the UAV state quantities of each input UAV The value network is used to output the value estimate value value corresponding to the current UAV state quantity according to the global state quantity S of the multi-UAV input ; k Step 6: Perform multi-UAV multi-target cooperative tracking training based on the MAPPO algorithm to obtain a multi-UAV multi-target cooperative tracking controller;

[0013] Step 7: Input the local observation states of each UAV

[0014] into the multi-UAV multi-target cooperative tracking controller to obtain the action control quantities of each UAV, and control each UAV to work according to each action control quantity to complete the task of controlling the multi-UAV to perform cooperative tracking on multiple targets. ;

[0015] Further, in Step 1, the method for constructing the six-degree-of-freedom kinematic model of the UAV is as follows:

[0016] Assume that the UAV is a symmetric rigid body and the influence of air resistance is ignored. The movement of the UAV in space is a six-degree-of-freedom movement, which are translational movements along the X, Y, and Z axes of the ground space coordinates and rotational movements around the main axes of the body coordinate system. The position of the UAV in the geographical coordinates is The attitude angle is The position motion equation of the UAV centroid relative to the ground coordinates is The motion equation of the UAV rotating around the centroid is where vX , v Y , v Z are the velocities of the UAV relative to the ground in the X, Y, and Z directions respectively. v x , v y , v z are the velocities of the UAV relative to its own velocity coordinate system in the X, Y, and Z directions respectively. C is the transformation matrix from the UAV velocity coordinate system to the ground space coordinate system. are the angular velocities of the UAV relative to the ground coordinate in the X, Y, and Z directions respectively. , ψ are the angles of the UAV relative to the ground coordinate system in the X and Z directions respectively. are the angular velocities of the UAV in the X, Y, and Z directions in the body coordinate system respectively.

[0017] Furthermore, in step 1, the method for constructing the constant turn rate and velocity model of the moving target is as follows:

[0018] For UAV target tracking, the UAV itself and the target to be tracked are regarded as mass points relative to the entire dynamic environment. At the same time, the process of the UAV tracking the target has nothing to do with the longitudinal space. The changes in the instantaneous velocity and turn rate of the target to be tracked and the altitude change are regarded as noises, and the constant turn rate and velocity model of the moving target is constructed. where the coordinates (x m , y m ) represent the position of the target in the environment. v, σ, represent the velocity, yaw angle, and angular velocity of the target in the ground space coordinate system respectively.

[0019] Furthermore, in step 1, the method for constructing the distributed partially observable Markov decision model is as follows:

[0020] The multi-UAV multi-target cooperative tracking control process is a fully cooperative multi-agent partially observable Markov decision process. The partially observable Markov decision model of a single UAV is extended to a multi-UAV distributed partially observable Markov decision model, which is represented by a tuple G = <S, U, P, T, Z, Ο, n, γ>. Among them, γ represents the discount factor, n represents n UAV agents, s ∈ S represents the true state information of the environment, S represents the set of true state information of the environment. At each time step, for the UAV agent i ∈ N ≡ {1,..., n}, N represents the set of UAV agents, an action a needs to be selected. i∈A, where A represents the set of actions, to form a joint action u ∈ U, where U represents the set of joint actions, and then give this joint action to the environment for state transition P(s′|s,u): S×U→[0,1], where P(s′|s,u) represents the probability that s is converted to s′ under the condition of u; afterwards, the drone agent i will receive a reward r i , the total reward obtained by all drone agents T represents the set of total rewards. For the drone agent i, it receives an independent partially observable state ζ ∈ Z. Different drone agents have different observations, and all observations come from the true state information of the environment. A conditional observation transition probability function Ο(s,i): S×N→Z, where Z represents the set of partially observable states

[0021] Furthermore, the environmental standardization preprocessing in step 2 includes: defining the environmental boundary of the multi-drone multi-target cooperative tracking task within a square area with a total area of a 2 , where a is the side length of the environmental model boundary. During the training process, the drones and targets always move within the environmental boundary. Denote the center position of the area as the coordinate origin of the environmental model. At the initial training moment, each drone and each target are at any position within the area

[0022] Furthermore, the data normalization preprocessing in step 2 includes: setting the maximum and minimum values of the drone state quantity and the target state quantity, respectively clipping the drone state quantity and the target state quantity, setting the data greater than the maximum value and less than the minimum value to the maximum value and the minimum value to prevent data overflow, and then dividing the data by the maximum value to limit its value range to [-1,1]

[0023] Furthermore, the steps for multi-task allocation in step 3 include:

[0024] Step 3.1: Establish the benefit matrix M 0 (m×N). If m < N, add (N - m) rows to the benefit matrix M 0 to form a square matrix, and set the elements in the newly added (N - m) rows to 0

[0025] Step 3.2: Subtract the minimum element of each row from each row of the benefit matrix M 0 so that each row has a zero element, obtaining the benefit matrix M 1 ;

[0026] Step 3.3: Subtract the minimum element of each column from each column of the benefit matrix M 1 so that each column has a zero element, obtaining the benefit matrix M 2 ;

[0027] Step 3.4: Cover the benefit matrix M with the fewest straight lines 2 The zero elements in the benefit matrix M 3 ,If the minimum number of straight lines is equal to m, go to step 3.6: otherwise go to step 3.5;

[0028] Step 3.5: Benefit Matrix M 3 The benefit matrix M is obtained by subtracting the smallest element from all the elements not covered by the straight line and adding the smallest element at the intersection of the straight lines. 4 , let the benefit matrix M 2 Equal to the benefit matrix M 4 ;

[0029] Step 3.6: Start allocating from the row or column with the least zero elements until all drones are assigned a target, and the allocation is completed.

[0030] Furthermore, in step 4, the constructed state value function represents the state space of agent i∈{1,…,N} at time k∈{1,…,t} as where j∈{1,…,N}∩i≠i, (d x ,d y ,d z ,d ω ) represents the relative distance between the drone and the target in the X, Y, and Z directions and the difference in yaw angle. Indicates the speed and yaw rate of the drone in the X, Y, and Z directions. represents the relative distance between drone i and drone j in the X, Y, and Z directions, t is the maximum number of steps for training, and T represents vector transpose;

[0031] The speed of the drone is defined by x, y, z direction vectors and an absolute value of speed. The constructed action value function represents the continuous action space of agent i∈{1,…,N} at time k∈{1,…,t}: in, It indicates the direction vector of the drone's action in the X, Y, and Z directions. Speed ​​indicates the absolute value of the speed at which the drone performs the action. represents the yaw angular velocity of the drone performing the action, and T represents the vector transpose;

[0032] The constructed reward function represents the reward value obtained by agent i∈{1,…,N} at time k∈{1,…,t} in They represent the tracking reward and penalty, height reward and penalty, and safety reward and penalty of agent i at time k respectively.

[0033] Furthermore, the input of the policy network in step 5 is the normalized UAV state quantity and the output is the UAV action control quantity Its policy network has 7 layers. Among them, the first layer is the input layer, the second layer is a hidden layer with 256 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third layer and the fourth layer, the fourth layer and the fifth layer are both hidden layers with 256 nodes, an Elu activation function is added between the fifth layer and the sixth layer, the sixth layer is a hidden layer for calculating the mean loc and variance scale of the UAV action control quantity, and the output is a sampling value of a multivariate Gaussian distribution The seventh layer is the output layer, and the output layer includes a Tanh activation function and normalization

[0034] Furthermore, the input of the value network in step 5 is the normalized multi-UAV global state quantity S k , and the output is the value estimate value value corresponding to the current UAV state quantity. Its structure has 6 layers. Among them, the first layer is the input layer, the second layer is a hidden layer with 512 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third layer and the fourth layer, the fourth layer is a hidden layer with 512 nodes, the fifth layer is a hidden layer with 256 nodes, and the sixth layer is the output layer

[0035] Furthermore, the training process in step 6 includes:

[0036] Set the total number of training rounds and steps

[0037] In each round, randomly initialize the positions and speeds of the targets and perform multi-target allocation. Then, each UAV interacts with the environment (the environment includes other UAVs and targets), that is, simulate the process of multiple UAVs jointly tracking multiple targets in the environment. Regardless of the tracking results, after preprocessing the interaction information data, store it in the experience pool according to the time series

[0038] Whenever the experience pool is full, take out all the data and iteratively update the parameters of the policy network and the value network according to the MAPPO algorithm. Among them, the homogeneous UAVs in the system can share parameters, and the heterogeneous UAVs cannot share parameters

[0039] When the set total number of training rounds is reached, all training ends, and take out the policy network as the multi-UAV multi-target collaborative tracking controller

[0040] Compared with the prior art, the beneficial effects that the present invention can achieve are at least as follows:

[0041] 1. The present invention uses the MAPPO algorithm to train and generate a multi - UAV multi - target cooperative tracking controller, adopting a structure of centralized training and decentralized execution. Considering the partial observability under real - world environmental conditions, each agent relies only on its own observations to output control instructions, realizing the distributed control of multi - UAVs, avoiding the disadvantages of high dimensionality and non - scalability caused by centralized control when using single - agent deep reinforcement learning algorithms to handle multi - agent problems, and at the same time reducing the requirements of UAVs for communication and computing capabilities in the real - world environment.

[0042] 2. The multi - UAV multi - target cooperative tracker trained by the method proposed in the present invention is different from the traditional multi - UAV multi - target tracking method that adopts a hierarchical control idea. The controller trained by deep reinforcement learning is essentially a deep neural network, which directly takes the observation information as input and the action instruction as output, simplifying the control process of multi - UAV multi - target cooperation. When using this controller, only operations such as multiplication, addition, and some activation functions are required, and the overall computational complexity is much lower than that required by the traditional multi - UAV multi - target tracking method to use an optimization algorithm to plan the multi - UAV target tracking route.

[0043] 3. The method proposed in the present invention enables each UAV to dynamically control its flight yaw, direction, and speed during the tracking process, and cooperate with other UAVs to form a suitable spatial topology structure in the current environment, solving problems such as the possible mutual influence or collision of multi - UAVs during the process of tracking multi - targets and the difficulty in coping with environmental changes that require real - time calculation. It is applicable to tracking targets with various motion modes (such as uniform motion, variable - speed motion, linear motion, random motion, etc.), and has strong self - adaptability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0045] Figure 1 is the design diagram of the policy network structure provided by the embodiment of the present invention;

[0046] Figure 2 is the design diagram of the value network structure provided by the embodiment of the present invention;

[0047] Figure 3 is the MAPPO algorithm structure diagram in the embodiment of the present invention;

[0048] Figure 4It is the training flowchart of the multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm of the present invention;

[0049] Figure 5 It is the average episode total reward return curve graph when 5 UAVs track 3 randomly moving targets in a single training in the embodiment of the present invention;

[0050] Figure 6 It is the single simulation test trajectory graph of 5 UAVs tracking 3 randomly moving targets in the embodiment of the present invention;

[0051] Figure 7 It is the distance difference change graph of 5 UAVs tracking 3 randomly moving targets in a single simulation test in the embodiment of the present invention;

[0052] Figure 8 It is the yaw angle difference change graph of 5 UAVs tracking 3 randomly moving targets in a single simulation test in the embodiment of the present invention; Detailed implementation manners

[0053] In order to be able to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0054] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0055] In some embodiments of the present invention, it is assumed that the current task is to have 5 quadrotor UAVs that can obtain the current position information of the target in real time through the sensor devices carried by themselves, and it is required to track 3 randomly moving vehicle targets within their reconnaissance range. Now, through the method of the present invention, a multi-UAV multi-target cooperative tracking controller is designed and trained, so that 5 quadrotor UAVs can complete the tracking task without affecting or colliding with each other. The entire design, training and verification process is completed in a simulation environment.

[0056] Step 1: Model the multi-UAV target tracking process, including establishing a UAV six-degree-of-freedom kinematic model, a moving target constant turn rate and speed model, and a distributed partially observable Markov decision model.

[0057] This step specifically includes:

[0058] Step 1.1: Construct a UAV six-degree-of-freedom kinematic model;

[0059] The motion of the UAV mainly involves changes in flight attitude and spatial position. Taking a quadrotor UAV as an example, assuming that the quadrotor UAV is a symmetric rigid body and ignoring the influence of air resistance, the motion of the quadrotor UAV in space is a six-degree-of-freedom motion, which are translational motions along the X, Y, and Z axes of the ground space coordinate and rotational motions around the main axes of the body coordinate. The position of the UAV in the geographical coordinate is The attitude angles are Let p be the pitch angle, q be the roll angle, and r be the yaw angle. Establish the position motion equation of the UAV's center of mass relative to the ground coordinate And establish the motion equation of the UAV rotating around the center of mass according to the relational expression between the rotational angular velocities of the UAV relative to the ground space coordinate system Among them, v X 、v Y 、v Z are the velocities of the UAV relative to the ground in the X, Y, and Z directions respectively, v x 、v y 、v z are the velocities of the UAV relative to its own velocity coordinate system in the X, Y, and Z directions respectively. C is the transformation matrix from the UAV's velocity coordinate system to the ground space coordinate system, are the angular velocities of the UAV relative to the ground coordinate in the X, Y, and Z directions respectively, ψ are the angles of the UAV relative to the ground coordinate system in the X and Z directions respectively, are the angular velocities of the UAV in the X, Y, and Z directions in the body coordinate system respectively;

[0060] Step 1.2: Construct a motion target constant turn rate and speed model;

[0061] For UAV target tracking, regard the UAV itself and the tracked target as mass points relative to the entire dynamic environment. At the same time, the process of the UAV tracking the target has nothing to do with the longitudinal space. Regard the changes in the instantaneous velocity and turn rate of the tracked target and the altitude change as noise, and construct a motion target constant turn rate and speed model Among them, the coordinates (x m , y m ) represent the position of the target in the environment, v, σ, represent the velocity, yaw angle, and angular velocity of the target in the ground space coordinate system respectively;

[0062] Step 1.3: Construct a distributed partially observable Markov decision model;

[0063] Distributed research involves dividing a problem that requires extremely high computing power into many small parts, then distributing these parts to multiple computers for processing, and finally combining these computing results to obtain the final result. Distributed is relative to centralized. The cooperative control of multiple UAVs is not a centralized decision-making process, but each UAV makes its own decisions. In the modeling, it is based on distributed modeling, and the MAPPO algorithm adopted is a centralized training and decentralized execution framework, achieving distributed control;

[0064] The multi-UAV multi-target cooperative tracking control process is a fully cooperative multi-agent partially observable Markov decision process, which extends the single-UAV partially observable Markov decision model to a multi-UAV distributed partially observable Markov decision model (multi-UAV POMDP or Dec-POMDP). It is represented by a tuple G = <S, U, P, T, Z, Ο, n, γ>, where γ represents the discount factor, n represents n UAV agents, s ∈ S represents the true state information of the environment, S represents the set of true state information of the environment. At each time step, for UAV agent i ∈ N ≡ {1, …, n}, N represents the set of UAV agents, an action a i ∈ A, A represents the set of actions, to form a joint action u ∈ U, U represents the set of joint actions, and then this joint action is given to the environment for state transition P(s′|s, u): S × U → [0, 1], P(s′|s, u) represents the probability that s is converted to s′ under the condition of u; after that, each UAV agent i will obtain a reward r i The sum of the rewards obtained by all UAV agents T represents the set of total reward sums; for UAV agent i, it receives an independent partially observable state ζ ∈ Z, and different UAV agents have different observations. All observations come from the true state information of the environment. A set of conditional observation transition probability functions Ο(s, i): S × N → Z, Z represents the set of partially observable states.

[0065] Step 2: Environment standardization and data normalization preprocessing;

[0066] Step 2.1: Environment standardization preprocessing;

[0067] Define the environment boundary of the multi-UAV multi-target cooperative tracking task within a square area with a total area of a 2 , where a is the side length of the environment model boundary. During the training process, the UAVs and targets always move within the environment boundary. Denote the center position of the area as the coordinate origin of the environment model. At the initial training moment, each UAV and each target are at any position within the area. Among some embodiments of the present invention, a = 5;

[0068] Step 2.2: Data normalization preprocessing;

[0069] Set the maximum and minimum values of the UAV state variables and target state variables. Clip the UAV state variables and target state variables respectively, and set the data greater than the maximum value and less than the minimum value to the maximum value and the minimum value to prevent data overflow. Then divide the data by the maximum value to limit its value range to [-1, 1];

[0070] Step 3: Multi-objective task allocation;

[0071] Use the improved Hungarian algorithm (the method of adding edges and filling zeros) to allocate tasks for the multi-UAV tracking task. In the case of non-standard assignment problems where multiple UAVs (N≥2) track multiple targets (m≥2), to ensure that each target is tracked by at least 1 UAV, the number of UAVs should be greater than or equal to the number of targets (N≥m). The solution space is an m×N matrix. Here, N UAVs need to be divided into m groups to track m targets, and the number of UAVs in each group should be as equal as possible (at most 1 difference). At the same time, based on the distance of each UAV from the initial position of each target, calculate the minimum Euclidean distance cost sum for all UAVs to track each target.

[0072] In some embodiments of the present invention, in the non-standard assignment problem where 5 UAVs track 3 targets, 5 UAVs need to be divided into 3 groups to track 3 targets. The number of UAVs in each group should be as equal as possible (at most 1 difference). At the same time, based on the distance of each UAV from the initial position of each target, calculate the minimum Euclidean distance cost sum for all UAVs to track each target. (When allocating the tracking targets, we expect that on the premise of ensuring that each target can be tracked and minimizing the overall loss probability (i.e., as evenly distributed as possible), each UAV in the multi-UAV system can track the target with the shortest distance, so that each UAV can approach and track the target to be tracked with the least fuel consumption, the shortest distance, and the shortest time. Therefore, the Euclidean distance cost sum needs to be calculated.)

[0073] The way to perform task allocation in this step is as follows:

[0074] Step 3.1: Establish the benefit matrix M of the target allocation problem 0 (3×5), and add 2 rows to the benefit matrix M 0 to form a square matrix, and set the elements in the newly added 2 rows to 0;

[0075] Step 3.2: Subtract the minimum element of each row from each row in the benefit matrix M 0 so that each row has a zero element, and obtain the benefit matrix M 1 ;

[0076] Step 3.3: From the benefit matrix M 1Subtract the smallest element from each column so that each column has a zero element, and get the benefit matrix M 2 ;

[0077] Step 3.4: Cover the benefit matrix M with the fewest straight lines 2 The zero elements in the benefit matrix M 3 ,If the minimum number of straight lines is equal to m, go to step 3.6: otherwise go to step 3.5;

[0078] Step 3.5: Benefit Matrix M 3 The benefit matrix M is obtained by subtracting the smallest element from all the elements not covered by the straight line and adding the smallest element at the intersection of the straight lines. 4 , let the benefit matrix M 2 Equal to the benefit matrix M 4 ;

[0079] Step 3.6: Start allocating from the row or column with the least zero elements until all drones are assigned a target, and the allocation is completed.

[0080] Step 4: Design state value function, action value function and reward return function;

[0081] This step specifically includes:

[0082] Step 4.1: Design state value function;

[0083] The drone does not need to pitch or roll during the tracking process. It only needs to keep stable during the tracking process and ensure that there is no collision between drones. Under this premise, the shortest and best path is taken to track the target. For the problem of multi-drone multi-target collaborative tracking, the designed state value function represents the state space of agent i∈{1,…,N} at time k∈{1,…,t} (t is the maximum number of training steps) as where j∈{1,…,N}∩j≠i, Indicates the relative distance between the drone and the target in the X, Y, and Z directions and the difference in yaw angle. Indicates the speed and yaw rate of the drone in the X, Y, and Z directions. represents the relative distance between drone i and drone j in the X, Y, and Z directions, and T represents vector transpose;

[0084] Step 4.2: Design action-value function;

[0085] The speed of the drone is defined by x, y, z direction vectors and an absolute value of speed. The designed action value function represents the continuous action space of agent i∈{1,…,N} at time k∈{1,…,t}: in, Denote the direction vector of UAV i in the X, Y, and Z directions when performing an action at time k, and speed represents the absolute value of the speed of the UAV when performing the action. Denote the yaw angular velocity of UAV i when performing an action at time k, and T represents the vector transpose;

[0086] Step 4.3: Design a reward return function based on the potential plastic return function and the non-potential plastic return function;

[0087] The goal of training is to enable the UAV to move towards the target point, and collisions between UAVs are not allowed during this process. Consider the random movement trajectory of the target as a time series of position coordinate points. The UAV can track the position of the current target at each moment, that is, complete the tracking of the target point position over the entire time series; design tracking rewards and punishments based on the potential plastic return function, and design height rewards and punishments and safety rewards and punishments based on the non-potential plastic return function; the designed reward return function represents the reward return value obtained by agent i ∈ {1, …, N} at time k ∈ {1, …, t} (t is the maximum number of training steps). Where Represent the tracking rewards and punishments, height rewards and punishments, and safety rewards and punishments of agent i at time k, respectively;

[0088] In the formula,

[0089]

[0090]

[0091] Among them, μ, δ, τ, α, and β all represent discount factors, Denote the distance difference between UAV i and its target at time k, Denote the yaw angle difference between UAV i and its target at time k, H max 、H min 、D min Represent the maximum allowable flight height, minimum allowable flight height of the UAV, and the minimum safe distance between UAVs, respectively, Represent the flight height of UAV i at time k and the minimum spatial distance from the nearest other UAV at time k, respectively. Denote the coordinate position of UAV i at time k, Denote the coordinate position of its target at time k, Denote the yaw angle of UAV i at time k, Denote the yaw angle of its target at time k.

[0092] The discount factors belong to hyperparameters, and their values are determined by their impact on Based on the contribution degree (influence degree) and experimental results, in some embodiments of the present invention, μ = 1, δ = 1, τ = 15, α = 10, β = 3.

[0093] Step 5: Design a deep neural network structure, including a policy network structure and a value network structure;

[0094] Step 5.1: Design a policy network structure;

[0095] As Figure 1 shown, the policy network structure has seven layers. The input of the policy network is the normalized UAV state quantity. The first layer is the input layer, the second layer is a hidden layer with 256 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third and fourth layers, the fourth and fifth layers are both hidden layers with 256 nodes, an Elu activation function is added between the fifth and sixth layers, the sixth layer is a hidden layer for calculating the mean loc and variance scale of the UAV action control quantity, and the output is a sampled value of a multivariate Gaussian distribution. The seventh layer is the output layer, which includes a Tanh activation function and normalization. The output of the policy network structure is the UAV action control quantity.

[0096] Step 5.2: Design a value network structure;

[0097] As Figure 2 shown, the value network structure has six layers. The input of the value network is the normalized multi-UAV global state quantity S. k , the first layer is the input layer, the second layer is a hidden layer with 512 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third and fourth layers, the fourth layer is a hidden layer with 512 nodes, the fifth layer is a hidden layer with 256 nodes, and the sixth layer is the output layer. The output is the value estimate value value corresponding to the current UAV state quantity.

[0098] Step 6: Multi-UAV multi-target cooperative tracking training based on the MAPPO algorithm;

[0099] The structure of the MAPPO algorithm is as Figure 3 shown. Use the MAPPO algorithm for multi-UAV multi-target cooperative tracking training. The training process is as Figure 4As shown, the total number of training rounds is set to Rounds = 400, and the number of training steps per round is Steps = 240. In each round, the positions and velocities of the targets are randomly initialized, and multi-target allocation is performed. Then, each UAV interacts with the environment (the environment includes other UAVs and targets), that is, simulates the process of multi-UAVs collaborating to track multiple targets in the environment. Regardless of the tracking results, the interaction information data is preprocessed and stored in the experience pool according to the time series. Whenever the experience pool is full, all the data is taken out, and importance sampling is used to iteratively update the parameters of the policy network and the value network according to the MAPPO algorithm. The homogeneous UAVs in the system can share parameters, while the heterogeneous UAVs cannot. Until the set total number of training rounds is reached, all training ends, and the policy network is taken out as the multi-UAV multi-target collaborative tracking controller.

[0100] In some embodiments of the present invention, the total number of training rounds set here needs to satisfy that the average total reward return function τ(s, u) of the final training converges to a stable state or the total reported reward value is close to the reward value when the set tracking is successful, as Figure 5 shown.

[0101] Step 7: Use of the multi-UAV multi-target collaborative tracking controller: Input the local observation states of each UAV into the multi-UAV multi-target collaborative tracking controller to obtain the action control quantities of each UAV, and control each UAV to work according to each action control quantity to complete the task of controlling multi-UAVs to collaboratively track multiple targets.

[0102] Based on the characteristics of deep reinforcement learning, when using the MAPPO algorithm for multi-UAV multi-target collaborative tracking training, only train multi-UAVs to track multiple targets in one motion state (such as stationary, uniform linear motion, variable-speed curve motion, etc.). The obtained multi-UAV multi-target collaborative tracking controller can be directly applied to the tracking of randomly moving multiple targets. In addition, when applying, there is no need to introduce the randomness that needs to be added during training (the "random initialization of environmental information" in step 6). Each UAV can use its own locally observed state after normalization as the input to obtain the action control quantities of each UAV, and achieve distributed control of each UAV to track the assigned targets, thereby completing the task of controlling multi-UAVs to collaboratively track multiple targets.

[0103] In some embodiments of the present invention, 100 tracking experiments are carried out, and the environmental information is randomly initialized each time (the initial positions of 5 rotor UAVs and the initial positions, directions, and velocities of 3 vehicle targets).

[0104] Take the results of one of the tests. Initially, the initial positions and task assignments of 5 quadrotor UAVs and 3 randomly moving vehicle targets are as follows: the initial position of UAV id0 is [-0.17, 0.44, 0.5], the initial position of UAV id1 is [-1, -0.4, 0.5], the initial position of UAV id2 is [-0.71, -0.82, 0.5], the initial position of UAV id3 is [-0.63, -0.31, 0.5], the initial position of UAV id4 is [-0.21, 0.08, 0.5], the initial position of target id0 is [-0.57, 0.69, 0], the initial position of target id1 is [0.16, -0.81, 0], the initial position of target id2 is [0.24, -0.15, 0]; the details of the task assignment are that UAV id0 tracks target id0, with the minimum Euclidean distance cost of 0.69, UAV id1 tracks target id0, with the minimum Euclidean distance cost of 1.27, UAV id2 tracks target id1, with the minimum Euclidean distance cost of 1.00, UAV id3 tracks target id2, with the minimum Euclidean distance cost of 1.02, UAV id4 tracks target id2, with the minimum Euclidean distance cost of 0.71; the total minimum path cost is 4.69.

[0105] The trajectory diagram of 5 quadrotor UAVs tracking 3 randomly moving vehicle targets under the action of the multi-UAV multi-target cooperative tracking controller is as Figure 6 shown, where the dots represent the starting positions of the UAVs, and the triangles represent the ending positions of the UAVs; the diagram of the change in the distance difference is as Figure 7 shown, the distance difference rapidly shrinks in the early stage of tracking and then remains constant; the diagram of the change in the yaw angle difference is as Figure 8 shown, the yaw angle difference rapidly shrinks in the early stage of tracking and then remains constant. It can be seen from the figure that after starting from the random starting positions, the 5 quadrotor UAVs can finally stably track 3 randomly moving vehicle targets. Therefore, through the method of the present invention, it is possible to ensure that the multi-UAVs are in a reasonable spatial distribution without collision and mutual influence (such as occlusion, etc.), and at the same time, through the way of mutual cooperation, continuously and stably track the targets.

[0106] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. Multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm, characterized in that, it includes the following steps: Step 1: Model the multi-UAV target tracking process, including establishing a six-degree-of-freedom kinematic model of the UAV, a constant turn rate and speed model of the moving target, and a distributed partially observable Markov decision model; Step 2: Perform environmental standardization and data normalization preprocessing; Step 3: Allocate tasks for the multi-UAV tracking task, divide N UAVs into m groups to track m targets, and at the same time, based on the distance of each UAV from the initial position of each target, calculate the sum of the Euclidean distance costs for all UAVs to track each target, where N≥2 and m≥2; Step 4: Construct a state value function, an action value function, and a reward return function; Step 5: Construct a deep neural network structure, including a policy network and a value network. The policy network is used to output the UAV action control quantity according to the local state quantity of each input UAV The value network is used to output the value estimate value value corresponding to the current UAV state quantity according to the global state quantity S of multiple UAVs input The value network is used to output the value estimate value value corresponding to the current UAV state quantity according to the global state quantity S of multiple UAVs input k Output the value estimate value value corresponding to the current UAV state quantity; Step 6: Perform multi-UAV multi-target cooperative tracking training based on the MAPPO algorithm to obtain a multi-UAV multi-target cooperative tracking controller; Step 7: Input the local state variables of each UAV into the multi-UAV multi-target cooperative tracking controller to obtain the action control quantities of each UAV, and control each UAV to work according to each action control quantity, so as to complete the task of controlling multiple UAVs to conduct cooperative tracking of multiple targets; In step 4, the constructed state value function represents that the state space of UAV \(i\in\{1,\ldots,N\}\) at time \(k\in\{1,\ldots,t\}\) is where \(j\in\{1,\ldots,N\}\cap j\neq i\), \((d x ,d y ,d z ,d ω ) represents the relative distance between the UAV and the target and the difference angle of the yaw angle in the three directions of X, Y, and Z. represents the speed magnitudes and yaw angular velocity of the UAV in the three directions of X, Y, and Z. represents the relative distance between UAV \(i\) and UAV \(j\) in the three directions of X, Y, and Z. \(t\) is the maximum number of training steps, and \(T\) represents vector transpose. The velocity of the UAV is defined by the x, y, and z direction vectors and an absolute value of the velocity. The constructed action-value function represents the continuous action space of UAV i ∈ {1, …, N} at time k ∈ {1, …, t} as where represents the direction vectors of the UAV's action in the X, Y, and Z directions, and speed represents the absolute value of the speed of the UAV's action. represents the yaw angular velocity of the UAV's action, and T represents the vector transpose; The constructed reward function represents the reward value obtained by drone \(i\in\{1,\ldots,N\}\) at time \(k\in\{1,\ldots,t\}\). where respectively represent the tracking reward and punishment, altitude reward and punishment, and safety reward and punishment of drone \(i\) at time \(k\). The training process in Step 6 includes: Set the total number of training episodes and steps; In each episode, randomly initialize the positions and speeds of the targets and perform multi-target allocation. Then, each UAV exchanges information with other UAVs and targets, that is, simulate the process of multi-UAVs performing a cooperative tracking of multiple targets in the environment. Regardless of the tracking result, preprocess the exchanged information data and store it in the experience pool according to the time series; Whenever the data in the experience pool is full, take out all the data and iteratively update the parameters of the policy network and the value network according to the MAPPO algorithm. Among them, the homogeneous UAVs in the system can share parameters, and the heterogeneous UAVs cannot share parameters; When the set total number of training episodes is reached, all training ends, and take out the policy network as the multi-UAV multi-target cooperative tracking controller.

2. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, In Step 1, the method for constructing the six-degree-of-freedom kinematic model of the UAV is: Assume that the UAV is a symmetric rigid body and the influence of air resistance is ignored. The motion of the UAV in space is a six-degree-of-freedom motion, namely translational motions along the X, Y, and Z axes of the ground space coordinates and rotational motions around the main axes of the body coordinate system. The position of the UAV in the geographical coordinates is The attitude angles are The position motion equations of the UAV's center of mass relative to the ground coordinates are The motion equations of the UAV rotating around the center of mass are where, v X 、v Y 、v Z are the velocities of the UAV relative to the ground in the X, Y, and Z directions respectively, v x 、v y 、v z are the velocities of the UAV relative to its own velocity coordinate system in the X, Y, and Z directions respectively. C is the transformation matrix from the UAV's velocity coordinate system to the ground space coordinate system, are the angular velocities of the UAV relative to the ground coordinates in the X, Y, and Z directions respectively, ψ are the angles of the UAV relative to the ground coordinate system in the X and Z directions respectively, are the angular velocities of the UAV in the X, Y, and Z directions in the body coordinate system respectively.

3. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, In Step 1, the method for constructing the constant turn rate and speed model of the moving target is: For UAV target tracking, the UAV itself and the target to be tracked are regarded as mass points relative to the entire dynamic environment. At the same time, the process of the UAV tracking the target has nothing to do with the longitudinal space. The changes in the instantaneous speed and turning rate of the target to be tracked and the altitude change are regarded as noise, and a constant turning rate and speed model of the moving target is constructed. Among them, the coordinates (x m , y m ) represent the position of the target in the environment, and v, σ, respectively represent the speed, yaw angle and angular velocity of the target in the ground space coordinate system.

4. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, In Step 1, the method for constructing the distributed partially observable Markov decision model is: The multi-UAV multi-target cooperative tracking control process is a fully cooperative multi-UAV partially observable Markov decision process, which extends the single-UAV partially observable Markov decision model to a multi-UAV distributed partially observable Markov decision model, and is represented by a tuple G as G = <S, U, Ρ, Τ, Z, Ο, n, γ>. Among them, γ represents the discount factor, n represents n UAVs, s ∈ S represents the true state information of the environment, and S represents the set of true state information of the environment. At each time step, for UAV i ∈ N ≡ {1, …, n}, where N represents the set of UAVs, an action a i ∈ A needs to be selected. A represents the set of actions, to form a joint action u ∈ U, where U represents the set of joint actions, and then this joint action is given to the environment for state transition Ρ(s ′ |s, u): S × U → [0, 1], (Ρ(s ′ |s, u) represents the probability that s is converted to s ′ under the condition of u; afterwards, UAV i will receive a reward r i , and the total reward obtained by all UAVs Τ represents the set of total reward; for UAV i, an independent partially observable state ζ ∈ Z is received, and different UAVs have different observations. All observations come from the true state information of the environment. A set of conditional observation transition probability functions Ο(s, i): S × N → Z, where Z represents the set of partially observable states.

5. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, In step 2, environmental standardization preprocessing is performed, including: defining the environmental boundary for the multi-UAV to carry out the multi-target cooperative tracking task within a square area with a total area of a 2 , where a is the boundary side length of the environmental model. During the training process, the UAVs and targets always move within the environmental boundary. Denote the center position of the area as the coordinate origin of the environmental model. At the initial moment of training, each UAV and each target are at arbitrary positions within the area; In Step 2, the data normalization preprocessing includes: setting the maximum and minimum values of the UAV state quantity and the target state quantity, respectively clipping the UAV state quantity and the target state quantity, setting the data greater than the maximum value and less than the minimum value to the maximum value and the minimum value to prevent data overflow, and then dividing the data by the maximum value to limit its value range to [-1,1].

6. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, The steps for multi-task allocation in step 3 include: Step 3.1: Establish the benefit matrix M of the target allocation problem 0 (m×N). If m < N, in the benefit matrix M 0 add (N - m) rows to form a square matrix, and set the elements in the newly added (N - m) rows to 0; Step 3.2: From the benefit matrix M 0 Subtract the minimum element of each row from each element in that row so that each row has a zero element, obtaining the benefit matrix M 1 ; Step 3.3: From the benefit matrix M 1 Subtract the minimum element of each column from each element in the column to make each column have a zero element, obtaining the benefit matrix M 2 ; Step 3.4: Cover the zero elements in the benefit matrix M with the least number of straight lines to obtain the benefit matrix M 2 ; if the number of the least straight lines is equal to m, go to Step 3.6; otherwise go to Step 3.5 3 ​ Step 3.5: Benefit matrix M 3 Subtract the smallest element among all the elements not covered by the straight lines in 4 from all the elements not covered by the straight lines, and add this smallest element to the intersection points of the straight lines to obtain the benefit matrix M 2 , and let the benefit matrix M 4 be equal to the benefit matrix M Step 3.6: Start the allocation from the row or column with the fewest zero elements until all drones are assigned a target, and the allocation is completed at this point.

7. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, The input of the policy network in step 5 is the normalized local state quantity of the UAV. The output is the UAV action control quantity. The structure of its policy network has 7 layers. Among them, the first layer is the input layer, the second layer is a hidden layer with 256 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third layer and the fourth layer, both the fourth layer and the fifth layer are hidden layers with 256 nodes, an Elu activation function is added between the fifth layer and the sixth layer, the sixth layer is a hidden layer that calculates the mean loc and variance scale of the UAV action control quantity, and the output is a sampling value of a multivariate Gaussian distribution. The seventh layer is the output layer, and the output layer includes a Tanh activation function and normalization.

8. The multi-UAV multi-target cooperative tracking control method based on the MAPPO algorithm according to claim 1, characterized in that, The input of the value network in step 5 is the normalized multi-UAV global state quantity S k , and the output is the value estimate value value corresponding to the current UAV state quantity. Its structure has 6 layers. Among them, the first layer is the input layer, the second layer is a hidden layer with 512 nodes, the third layer is a LayerNorm normalization layer, a Tanh activation function is added between the third layer and the fourth layer, the fourth layer is a hidden layer with 512 nodes, the fifth layer is a hidden layer with 256 nodes, and the sixth layer is the output layer.

Citation Information

Patent Citations

  • A UAV Target Tracking Control Method Based on Reinforcement Learning PPO Algorithm

    CN111580544B

  • Multi-target continuous tracking system and method based on cloud deck control technology

    CN108062115A

  • Multi-unmanned aerial vehicle and user cooperative communication optimization method based on MAPPO algorithm

    CN113359480A