A multi-unmanned aerial vehicle cooperative target tracking method in a communication denial environment

By introducing attention contrastive learning and the MASAC algorithm into a multi-UAV system, the problem of global information inference in UAV cooperative target tracking under communication denial environment is solved, and more efficient target tracking and obstacle avoidance cooperation is achieved.

CN118963413BActive Publication Date: 2025-11-28SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411023598.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-11-28
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

In communication-denied environments, existing technologies struggle to effectively infer global information from local observations in multi-UAV collaborative target tracking missions, resulting in poor collaborative behavior among UAVs.

Method used

We employ a multi-agent reinforcement learning method based on attention contrastive learning. By constructing the MASAC algorithm and contrastive learning framework, we infer global information from local observation information and combine depth images and obstacle avoidance information to optimize the decision-making strategy of the UAV.

Benefits of technology

In communication-denied environments, it improves the cooperation between drones, enabling them to better complete target tracking tasks, maintain a safe distance, and effectively avoid obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118963413B_ABST
    Figure CN118963413B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-unmanned plane cooperative target tracking methods under communication denial environment, comprising the following steps: setting the relevant elements of multi-agent learning system, including agent state type and dimension, action type and dimension, reward function and algorithm related hyperparameters;Multi-agent strategy network, evaluation network and contrast learning framework are constructed;Establish the simulation environment of multi-unmanned plane target tracking, obtain agent observation information through the interaction of unmanned plane agent and environment, part of environmental information and reward information are stored in experience replay pool;Through sampling data in experience pool, the reinforcement learning target and the contrast learning target are optimized simultaneously, wherein the auxiliary information output by the contrast learning framework can help the agent obtain additional global information, after training to convergence, the strategy network obtained is used to generate the action to be executed by the unmanned plane.The application can realize the multi-unmanned plane cooperative target tracking task under communication denial environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of multi-unmanned aerial vehicle cooperative decision-making, and particularly relates to a multi-unmanned aerial vehicle cooperative target tracking method in a communication denial environment. BACKGROUND

[0002] In recent years, unmanned aerial vehicles (UAVs) are widely used in various fields such as military early warning, detection, rescue, etc., and one of their prototype tasks is mobile target tracking. When performing the tracking task, the sensor technology is used to locate the mobile target, and the control system performs action decision-making according to effective information input to achieve the approach and tracking of the target, so the decision-making and control algorithm is the key to the autonomous tracking of mobile targets by unmanned aerial vehicles. Deep reinforcement learning emphasizes the learning of the mapping from the environment to the behavior of an agent, and finds the most correct action decision by maximizing the value function, which is consistent with the decision-making and control needs of unmanned aerial vehicles. Therefore, some scholars have applied deep reinforcement learning technology to unmanned aerial vehicle target tracking tasks and have achieved certain results. A method of unmanned aerial vehicle autonomous obstacle avoidance and target tracking based on meta-reinforcement learning (Unmanned Aerial Vehicle Autonomous Obstacle Avoidance and Target Tracking Based on Meta-Reinforcement Learning, Jiang Weiyi, Wu Jun, Wang Yaonan, School of Electrical and Information Engineering, Hunan University, Changsha, Hunan) improves the training efficiency and decision-making effect of the original reinforcement learning, but it does not consider the multi-unmanned aerial vehicle task scenario; for the problem of unmanned aerial vehicle swarm obstacle avoidance and tracking, a decoupled multi-agent deep deterministic policy gradient algorithm (Wen Chao, Dong Wenhan, Jie Wu, Cai Ming, Hu Duoxiu, Air Force Engineering University, Xi'an, Shaanxi) is proposed, which has strong accuracy and real-time performance, but it assumes that there is communication between the unmanned aerial vehicle swarm, which has a certain degree of limitation; and the applicability of the reinforcement learning method in the problem of multi-unmanned aerial vehicle target search in a communication denial battlefield simulation environment (Wang Liang, Wang Wen, Wang Yuyou, Hou Songlin, Qiao Yuzhe, Wu Tianpeng, Tao Xianping, State Key Laboratory of Computer Software and New Technology, Nanjing University) verifies the applicability of reinforcement learning in the problem of multi-unmanned aerial vehicle target search in a communication denial environment, but does not consider how to enhance the cooperative ability between unmanned aerial vehicles.

[0003] Existing achievements often assume that the communication between UAVs or between UAVs and ground stations is in good condition, so that the position and other state information of other UAVs can be directly obtained, and then the target tracking task is completed. However, in reality, there are often poor communication conditions or no communication in many scenarios, so it is necessary for multiple UAVs to cooperatively complete the tracking task only through their own sensors and keep inter-vehicle collision avoidance. Multi-agent deep reinforcement learning is one of the effective tools for solving the multi-UAV target tracking task, and a centralized training and distributed execution framework is usually used to train the UAV. Agents can be guided by the same signal, such as the global state, during centralized training. However, during execution, if in a communication denial environment, the agent will lack a shared signal and can only choose an operation according to local observations. This is particularly serious in partially observable Markov decision processes, which lack common guidance, making it difficult for UAVs to form good collaborative behavior. Therefore, it is necessary to make full use of the local observations of the agent to infer global information to help the agent make better decisions. SUMMARY

[0004] The technical problem to be solved by the present application is that in view of the existing technical problems in the prior art, based on the above background, the present application proposes a multi-UAV cooperative target tracking method in a communication denial environment.

[0005] The present application is realized at least by one of the following technical solutions.

[0006] A multi-UAV cooperative target tracking method in a communication denial environment, comprising the following steps:

[0007] S1, setting related elements of the multi-agent learning system, including agent state type and dimension, action type and dimension, reward function, and related hyperparameters;

[0008] S2, constructing a multi-agent policy network, an evaluation network, and a contrastive learning framework;

[0009] S3, establishing a simulation environment for multi-UAV target tracking, obtaining agent observation information, part of the environment information, and reward information through the interaction between the UAV agent and the environment, and storing them in an experience replay pool for subsequent training;

[0010] S4, by sampling the data in the experience pool, the reinforcement learning target and the contrastive learning target are optimized through centralized training-distributed execution, wherein the auxiliary information output by the contrastive learning framework can help the agent obtain additional global information, and after training to convergence, the obtained policy network is used to generate the action to be executed by the UAV.

[0011] Further, the step S1 comprises:

[0012] Step 1-1, set the number and type of multiple UAVs and moving targets to be tracked, wherein the UAVs are homogeneous;

[0013] Step 1-2, combine the actual requirements of the multi-UAV target tracking task to set the state observation information of the UAVs Set as:

[0014]

[0015] Wherein represents the observation information in the form of a general vector, that is, and represent the velocities of the UAVs in the X and Y axis directions; and represent the coordinate difference values of the UAVs and the targets in the X and Y axis directions; represents the deep image information obtained by the on-board camera of the UAV, and the global state is the combination of the local observations of all UAVs;

[0016] Step 1-3, combine the actual requirements of the multi-UAV target tracking task to set the actions of the UAVs Set as:

[0017]

[0018] Wherein, is the velocity component of the UAV along the X and Y axes, and in addition, the velocity of the UAV is constrained : wherein , represent the minimum and maximum velocities of the UAV, respectively;

[0019] Step 1-4, during the execution of the target tracking task by the UAV, both the tracking of the target and the collision avoidance sub-tasks need to be considered, and the reward function is defined:

[0020]

[0021] The tracking reward and the collision avoidance reward of the UAV are set, and the weighted sum of the two is taken as the total reward obtained by the UAV, wherein, , represent the weight coefficients of the tracking reward and the collision avoidance reward, respectively;

[0022] ​Step 1-5, setting reinforcement learning hyperparameters and task-related parameters, including: sampling batch size, learning rate, total number of training rounds, length of each round, discount factor, initial position coordinates of multiple UAVs, initial coordinates of the target with speed, number of obstacles in the environment, maximum and minimum speed of the UAV, flight height of the UAV.

[0023] Further, the depth image of the forward-looking camera of the UAV is used as part of the state input, a non-RGB image, the depth image provides implicit information for obstacle avoidance for the UAV, which can more effectively learn the semantic information of the surrounding environment after network processing.

[0024] Further, the continuous four frames of images obtained by the depth camera are stacked, thereby helping the network to learn the ability to process time sequence information.

[0025] Further, when the distance between the UAV and the target is greater than the expected distance, if the UAV makes a movement towards the target, a certain positive reward is given, otherwise a corresponding penalty is given; when the distance between the UAV and the target is within the expected distance and greater than the minimum distance, a fixed positive reward is given to the UAV; when the distance between the UAV and the target is less than the minimum distance, a fixed negative reward is given to the UAV as a penalty.

[0026] Further, in the setting of the collision avoidance reward function, when the distance between the UAV and the surrounding obstacles is less than the set minimum safety distance, a negative reward is given to the UAV, otherwise a positive reward is given; when the UAV collides with the obstacle, a larger negative reward is given, thereby helping the UAV to learn collision avoidance behavior.

[0027] Further, the step S2 comprises:

[0028] Step 2-1, establishing a MASAC algorithm, the MASAC algorithm is a model-free off-policy multi-agent deep reinforcement learning algorithm, using a centralized training-distributed execution (CTDE) training framework, expanding a soft actor-critic (SAC) algorithm to a multi-agent continuous task, training multiple agents in a continuous action space, and performing a multi-UAV cooperative target tracking task; introducing a parameter sharing technique to improve the learning speed and convergence efficiency of the algorithm;

[0029] Step 2-2, constructing a policy network in the MASAC algorithm, i.e., an actor policy network, for an input observation , the output action of the policy network is expressed as:

[0030]

[0031] wherein, for global information of student network output in the contrastive learning framework, for parameters of the policy network, for the agent policy;

[0032] Step 2-3, build the evaluation network in the MASAC algorithm, that is, Critic (evaluation network), in order to alleviate Overestimating (overfitting), Clip Double Q learning (Clip Double Q learning) is used to learn two evaluation networks and ; according to the observation information of all unmanned aerial vehicle agents at the current moment , the action set , and the corresponding global information set = , and then the centralized state-value function Q is calculated; and are the observation and action of the n th agent, is the global information of the n th agent inferred by the contrastive learning framework;

[0033] Step 2-4, build two target networks of the Critic evaluation network and , copy the weights of the two evaluation networks and to the respective target networks, that is , ;

[0034] Step 2-5, build the contrastive learning network, including the student network and the teacher network , both have the same network architecture, that is, composed of backbone (backbone) part and projector (projector) part, wherein the backbone part includes multi-head attention layer and fully connected layer, and the projector part includes fully connected layer.

[0035] Further, the step S3 comprises:

[0036] Step 3-1, initialize the Airsim simulation environment parameters, including the initial position coordinates of multiple unmanned aerial vehicles, the target initial coordinates, the speed, the number of obstacles in the environment, the maximum and minimum speed of the unmanned aerial vehicle, and the flight height of the unmanned aerial vehicle; initialize the policy network, the evaluation network, and the contrastive learning network and the experience pool, and obtain the initial state observation of the unmanned aerial vehicle agent;

[0037] Step 3-2, calculate the global observation according to the initial observation of the agent;

[0038] Step 3-3, input the initial observation of the agent and the global observation into the policy network together to calculate the action of the agent;

[0039] Step 3-4, each multi-agent performs the action and obtains the corresponding reward;

[0040] Step 3-5, store the trajectory information obtained in the above process in the experience pool for subsequent training.

[0041] Further, a domain randomization method is used to improve the robustness of the network model, that is, the positions of the unmanned aerial vehicles and the static obstacles in the environment are randomly generated in each initialization stage of the environment, thereby helping the network model to train a more general strategy.

[0042] Further, the step S4 comprises:

[0043] Step 4-1, sample a batch of trajectory data B from the experience pool for training;

[0044] Step 4-2, use the local observation of all agents in the same trajectory sequence to infer the global information from the local observation of the agent by using contrastive learning, sample the observation information of all unmanned aerial vehicle agents at the same time step from the experience pool, and represent the sequence as follows:

[0045]

[0046] wherein is the observation of the nth unmanned aerial vehicle; n

[0047] input the sequence into the student and teacher networks respectively, and the output is and , wherein as follows:

[0048]

[0049] wherein is the global information of the nth unmanned aerial vehicle inferred by the contrastive learning framework;

[0050] contrastive learning is performed on and , and the agent can infer the global information from the local observation, thereby helping it to make better decisions;

[0051] ​The global observation information is extracted by constructing an attention integration module (AIM). The attention weight uses a self-attention mechanism to obtain global information for each agent: Given the hidden states of two agents, the attention weight uses a bilinear mapping, and then uses a softmax function to normalize it to obtain, as follows:

[0052]

[0053] In the formula, and are two learnable linear transformations, representing key and query, respectively; is the attention weight coefficient of the i-th agent, i is the transpose of is the transpose of is the hidden state of the agent To extract more diverse global information, multiple attention heads are used and their outputs are connected together; By using the relative weight of each agent, the global information is obtained through the weighted sum of all other agents: i

[0054]

[0055]

[0056] where, is the global information of the agent i is the attention weight coefficient of the agent is the hidden state of the agent i Step 4-3, the global information j and the local observation of the agent

[0057] are spliced as the new agent observation, which is input into the policy network and the evaluation network for reinforcement learning training; i Step 4-4, update the policy network and the evaluation network. The loss function of the policy network is as follows:

[0058]

[0059]

[0060] where, is the batch size of sampling, is the local observation of all agents, is the action of all agents, where​​​​​​ is a loss function of the strategy network, is an agent strategy, is an evaluation network, is a set of global information of all agents;

[0061] The loss function of the evaluation network is:

[0062]

[0063] wherein, , is a target Q value; and is an evaluation network and its loss function;

[0064] Step 4-4, in order to realize the adaptive entropy temperature coefficient, the entropy temperature coefficient is updated by using the following loss function:

[0065]

[0066] wherein, H is a target strategy entropy; and is a temperature coefficient and its loss function;

[0067] Step 4-5, the target evaluation network is updated by using the following loss function in a soft updating manner:

[0068]

[0069] wherein, , is a soft updating coefficient; and respectively are an evaluation network and a corresponding target evaluation network;

[0070] Step 4-6, while the strategy network and the evaluation network are updated, the updating of the comparison learning network is also carried out synchronously, and the design idea of the comparison loss is that the reconstructed feature should be similar to the corresponding original feature and different from other features, and the parameters of the teacher network are updated on the basis of the parameters of the student network by using the exponential moving average line EMA, i.e. momentum encoder.

[0071] Compared with the prior art, the present application has the beneficial effects that:

[0072] The application is directed to a multi-UAV cooperative target tracking task in a communication denial environment, and proposes a multi-agent reinforcement learning method ACL based on attention contrast learning assistance. The method uses contrast learning to compare the attention fusion of the local observation of a single agent with the local observation of all agents, so that the processed information can contain global information to a certain extent. By splicing the processed information and the original observation information of the agent as the input of the policy network, the partial observable problem existing during the execution of the agent is relieved, the cooperation effect between the UAVs is promoted, and the target tracking task can be better completed. BRIEF DESCRIPTION OF DRAWINGS

[0073] Figure 1 is a framework structure diagram of an embodiment of a multi-UAV cooperative target tracking method in a communication denial environment;

[0074] Figure 2 is a task scene schematic diagram of an embodiment of the application in an Airsim simulation environment;

[0075] Figure 3 is an average reward change curve diagram of a UAV agent obtained in the training process of an embodiment of the application;

[0076] Figure 4 is a distance change curve diagram between the UAV and the obstacle in the test process of an embodiment of the application;

[0077] Figure 5 is a distance change curve diagram between each UAV in the test process of an embodiment of the application;

[0078] Figure 6 is a distance change curve diagram between the UAV and the target in the test process of an embodiment of the application;

[0079] Figure 7 is a speed change curve diagram of each UAV in the test process of an embodiment of the application;

[0080] Figure 8 is a tracking trajectory schematic diagram of a UAV in the test process of an embodiment of the application;

[0081] Figure 9 is a cooperation analysis diagram between UAVs in the test process of an embodiment of the application;

[0082] Figure 10 is a flowchart of an embodiment of a multi-UAV cooperative target tracking method in a communication denial environment. DETAILED DESCRIPTION

[0083] The application will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the application are not limited to this embodiment.

[0084] As shown in Figure 1 , Figure 10 , the multi-UAV cooperative target tracking method in a communication denial environment described in this example uses multi-agent deep reinforcement learning to train a distributed target tracking strategy for multiple UAVs in a communication-denied obstacle environment. Meanwhile, to address the problem of local observation between UAVs due to the inability to communicate during strategy execution, attention mechanisms and contrastive learning are introduced to infer global information, thereby assisting UAVs in making better decisions. The method includes the following steps:

[0085] S1, set the relevant elements of the multi-agent learning system, including the state type and dimension of the agent, the action type and dimension, the reward function, and the algorithm-related hyperparameters;

[0086] Step 1-1, this example sets the number of multi-UAVs to 3, and the target to be tracked to 1; the UAVs are homogeneous; and the corresponding environment is built in the Airsim simulation environment according to the settings, as shown in Figure 2 .

[0087] Step 1-2, according to the actual requirements of the multi-UAV target tracking task, the state observation information of the UAV is set as:

[0088]

[0089] Among them represents the observation information in the form of a normal vector, i.e. . and represent the velocity of the UAV in the X and Y axis directions; and represent the coordinate difference between the UAV and the target in the X and Y axis directions; represents the depth image information obtained by the UAV-mounted camera, and the resolution is set to 84x84 in this example. Since the task is set to be communication-free, this image provides implicit information for the UAV to avoid obstacles, including the avoidance between the UAV and static obstacles in the environment, and the collision avoidance between UAVs. The global state is the combination of all local observations of the UAVs.

[0090] Step 1-3, according to the actual requirements of the multi-UAV target tracking task, the action of the UAV is set as:

[0091]

[0092] Among them, are the velocity components of the UAV along the X and Y axes, and in addition, the velocity of the UAV is constrained as follows: ; wherein , respectively represent the minimum and maximum velocities of the UAV, which are set to -3 m / s and 3 m / s respectively in this example;

[0093] Step 1-4, define the reward function:

[0094]

[0095] In the process of performing the target tracking task, the UAV needs to consider both the tracking and collision avoidance sub-tasks, based on which the tracking reward and the collision avoidance reward of the UAV are set, and the weighted sum of the two is taken as the total reward obtained by the UAV. Among them, , respectively represent the weight coefficients of the tracking reward and the collision avoidance reward, which are set to 2 and 1 respectively in this example

[0096] When the distance between the UAV and the target is greater than the expected distance, if the UAV makes a movement towards the target, a certain positive reward is given, otherwise a corresponding penalty is given; when the distance between the UAV and the target is already within the expected distance, and is greater than the minimum distance, a fixed positive reward is given to the UAV; when the distance between the UAV and the target is less than the minimum distance, a fixed negative reward is given to the UAV as a penalty.

[0097] In setting the collision avoidance reward function, it is satisfied that when the distance between the UAV and the surrounding obstacles is less than the set minimum safety distance, a negative reward is given to the UAV, and vice versa; when the UAV collides with the obstacle, a larger negative reward is given, thereby helping the UAV to learn collision avoidance behavior;

[0098] Step 1-5, the reinforcement learning hyperparameters and task-related parameters are set in this example, including: sampling batch size, learning rate, total number of training rounds, length of each round, discount factor, initial position coordinates of multiple UAVs, initial coordinates and speed of the target, number of obstacles in the environment, maximum and minimum speed of the UAV, flight height of the UAV.

[0099] As an embodiment, the sampling batch size is set to 256, the learning rate is set to 3e-4, the total number of training rounds is set to 450K, the length of each round is set to 300, the discount factor is set to 0.99, etc.; the initial position coordinates of the multiple UAVs are set to (0, 0, 0), (0, 10, 0), (0, -10, 0) respectively, the initial coordinates of the target are set to (10, 0, 0), the speed is set to (random(-1, 1), 0.5, 0), the number of obstacles in the environment is set to 11, the flight height of the UAV is set to 10 m, etc.

[0100] S2, build multi-agent policy network, value network and contrast learning framework:

[0101] Step 2-1, using centralized training-distributed execution (CTDE) training framework, expand SAC algorithm to multi-agent continuous task, propose MASAC algorithm. The algorithm is a model-free off-policy multi-agent deep reinforcement learning algorithm, which can train multiple agents in continuous action space, so it is suitable for multi-UAV cooperative target tracking task.

[0102] As an embodiment, considering that multi-UAV is homogeneous, parameter sharing technique is introduced to help improve the learning speed and convergence efficiency of the algorithm. A contrast learning framework similar to DINO is used to help agents infer part of the global information from local information and optimize decision-making effect.

[0103] Step 2-2, build the policy network in MASAC algorithm, i.e. Actor policy network. For input observation , the output action of the policy network is expressed as:

[0104]

[0105] wherein, is the global information output by the student network in the contrast learning framework, is the parameter of the policy network; is the agent policy;

[0106] Step 2-3, build the evaluation network in MASAC algorithm, i.e. Critic evaluation network. In order to alleviate Overestimating, Clip Double Q learning is used, so two evaluation networks and are needed. According to the observation of all UAV agents at the current time, the action set , and the corresponding global information set = , the centralized state-value function Q is calculated; and are the observation and action of the nth agent, is the global information of the nth agent inferred by the contrast learning framework;

[0107] Step 2-4, in order to ensure the stability of learning, build two target networks and of the Critic evaluation network,

[0108] The weights of the evaluation network and are copied to the respective target network, i.e. , ;

[0109] Step 2-5, build a contrast learning network. Specifically, a twin network framework similar to DINO is adopted, including a student network (Student) and a teacher network (Teacher), both of which have the same network architecture, i.e., composed of a backbone part and a projector part. Among them, the backbone part is composed of multi-head attention layers and fully connected layers, and the projector part is composed of fully connected layers.

[0110] S3, establish a simulation environment for multi-UAV target tracking, obtain the observation information of the agent, part of the environment information and the reward information through the interaction between the UAV agent and the environment, and store them in the experience replay pool for subsequent training:

[0111] Step 3-1, initialize the Airsim simulation environment parameters, including the initial position coordinates of the multi-UAV, the initial coordinates of the target with speed, the number of obstacles in the environment, the maximum and minimum speed of the UAV, the flight height of the UAV, etc. Initialize all neural networks and experience pools, and obtain the initial state observation of the UAV agent;

[0112] Step 3-2, calculate the global observation according to the initial observation of the agent;

[0113] Step 3-3, input the initial observation of the agent and the global observation into the policy network together. Calculate the action of the agent;

[0114] Step 3-4, each multi-agent executes the action and obtains the corresponding reward;

[0115] Step 3-5, store the trajectory information obtained in the above process in the experience pool for subsequent training.

[0116] S4: Through sampling data in the experience pool, the reinforcement learning target and the contrast learning target are optimized simultaneously through centralized training-distributed execution, and the auxiliary information output by the contrast learning framework can help the agent obtain additional global information. After training to convergence, the obtained policy network is used to generate the action to be executed by the UAV:

[0117] Step 4-1, sample a batch of trajectory data B from the experience pool for training;

[0118] Step 4-2, using the local observations of all agents in the same trajectory sequence, global information is inferred from the local observations of the agents using contrastive learning. Specifically, the observations of all agents at the same time step are sampled from the experience pool and represented as the following sequence:

[0119]

[0120] The sequence is input into the student and teacher networks, respectively, and the output is and , where as follows:

[0121]

[0122] and are contrastively learned, and the agent can infer global information from local observations, thereby helping it make better decisions. An attention integration module (AIM) is constructed to achieve the extraction of global observation information. Specifically, the attention weight uses a self-attention mechanism to obtain global information for each agent: given the hidden states of two agents, the attention weight can use a bilinear mapping, and then use the softmax function to normalize it to obtain, as follows:

[0123]

[0124] In the above formula, and are two learnable linear transformations, representing "key" and "query", respectively. is the attention weight coefficient for the i-th agent, is the transpose of , is the transpose of , is the hidden state of agent i. In addition, in order to extract more diverse global information, we use multiple attention heads and concatenate their outputs.

[0125] By using the relative weight of each agent, global information can be obtained through the weighted sum of all other agents:

[0126]

[0127] where, is the global information of agent i, is the attention weight coefficient of agent i, is the hidden state of agent j; ​

[0128] Step 4-3, global information of agent i is spliced with local observation as new agent observation, input into policy network and value network for training of reinforcement learning;

[0129] Step 4-4, update policy network and value network. Specifically, the loss function of policy network is as follows:

[0130]

[0131] wherein, is the size of the sampled batch, is the local observation of all agents, is the action of all agents, wherein is the loss function of policy network, is the agent policy, is the evaluation network, is the set of global information of all agents;

[0132] The loss function of the evaluation network is as follows:

[0133]

[0134] wherein, , is the target Q value; and is the evaluation network and its loss function;

[0135] Step 4-4, in order to realize adaptive entropy temperature coefficient, the entropy temperature coefficient is updated by using the following loss function:

[0136]

[0137] wherein, H is the target policy entropy; and is the temperature coefficient and its loss function;

[0138] Step 4-5, the target value network is updated by using the following loss function in the way of soft update:

[0139]

[0140] wherein, , is the soft update coefficient; and are the evaluation network and its corresponding target evaluation network respectively;

[0141] ​Step 4-6, while updating the strategy network and the value network, the update of the contrast learning network is also synchronized. The design idea of the contrastive loss is that the reconstructed feature should be similar to the corresponding original feature and different from other features. The parameters of the teacher network are updated based on the student network parameters through the exponential moving average (EMA), that is, the momentum encoder. Figure 3 is the average reward change curve of the UAV agent during the training process. It can be seen from the figure that the reward obtained by the UAV increases continuously and finally converges to a higher level, which shows that the UAV has learned an effective target tracking strategy.

[0142] After the training is completed, the trained strategy is deployed on the UAV to verify the learning effect of the algorithm. Figure 4 is the distance change curve between the UAV and the obstacle during the test process, which reflects that each UAV can maintain a safe distance from the obstacle during tracking and will not collide; Figure 5 is the distance change curve between each UAV during the test process, which reflects that each UAV moves close to the target during the initial tracking stage, so the distance between the UAVs will decrease, and when all UAVs reach the vicinity of the target, they will maintain a certain formation to follow the target and will not collide with each other; Figure 6 is the distance change curve between the UAV and the target during the test process, which shows that the UAV can quickly follow the target and continuously follow the target after reaching the vicinity of the target; Figure 7 is the speed change curve of each UAV during the test process, from which it can be seen that during the initial tracking stage, the UAV tracks the target at a relatively large speed, and when it reaches the vicinity of the target, it automatically adjusts its speed to continuously follow the target; Figure 8 is the tracking trajectory diagram of the UAV during the test process, in which Figure 8 (a), (b), and (c) reflect that the UAV accelerates to track the target, while Figure 8 (d), (e), and (f) reflect that the UAV continuously follows the target; Figure 9 is the cooperation analysis diagram of the UAVs during the test process, in which Figure 9 (a), (b), and (c) reflect that the UAVs accelerate to track the target from different initial positions, while Figure 9 (d), (e), and (f) reflect that the multiple UAVs adjust their speeds to form a certain spatial distribution to cooperatively track the target. From the above analysis, it can be seen that the method of the present application has high learning efficiency and superior performance in the multi-UAV target tracking task under the communication denial environment.

[0143] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations, simplifications, etc. made without departing from the spirit and principles of the present application should be equivalent replacement manners and should be included in the protection scope of the present application.

Claims

1. A multi-UAV cooperative target tracking method for communication-denied environments, characterized in that: The method comprises the following steps: S1, setting related elements of the multi-agent learning system, including agent state type and dimension, action type and dimension, reward function, and related hyperparameters; S2, constructing a multi-agent strategy network, an evaluation network, and a contrastive learning framework, comprising: Step 2-1, establishing a MASAC algorithm, which is a model-free off-policy multi-agent deep reinforcement learning algorithm, using a centralized training-distributed execution (CTDE) training framework to extend the soft actor-critic (SAC) algorithm to multi-agent continuous tasks, training multiple agents in a continuous action space for a multi-UAV cooperative target tracking task; introducing a parameter sharing technique to improve the learning speed and convergence efficiency of the algorithm; Step 2-2, build the policy network in MASAC algorithm, i.e. Actor policy network, for input observation , the output action of policy network is expressed as: wherein, is a global information of the student network output in the contrastive learning framework, is a parameter of the policy network, is an agent policy; Step 2-3, build the evaluation network in the MASAC algorithm, that is, the evaluation network adopts truncated double Q learning to alleviate overfitting, and learns two evaluation networks and ; according to the observation information of all unmanned aerial vehicle agents at the current moment , the action set , and the corresponding global information set = , and then the centralized state-value function Q is calculated; and are the observation and action of the n th agent, is the global information of the n th agent inferred by the contrastive learning framework; Step 2-4, building two target networks for Critic evaluation network and copy the weights of two evaluation networks and to the respective target networks, i.e. , ; Step 2-5, constructing a contrastive learning network, including a student network and a teacher network , both having the same network architecture, i.e. composed of a backbone part and a projection head part, wherein the backbone part includes a multi-head attention layer and a fully connected layer, and the projection head part includes a fully connected layer; S3, establishing a simulation environment for multi-UAV target tracking, obtaining agent observation information, part of the environment information, and reward information through the interaction between the UAV agent and the environment, and storing them in an experience replay pool for subsequent training; S4, optimizing the reinforcement learning target and the contrastive learning target simultaneously through centralized training-distributed execution by sampling data in the experience pool, wherein the auxiliary information output by the contrastive learning framework can help the agent obtain additional global information, and after training to convergence, the obtained strategy network is used to generate actions to be executed by the UAV.

2. The method of claim 1, wherein, The step S1 comprises: Step 1-1, setting the number and type of multiple UAVs and the target to be tracked, wherein the UAVs are homogeneous; Steps 1-2, in combination with the actual demand of the multi-UAV target tracking task, the state observation information of the UAV is collected is set to: wherein represents the observation information in the form of a common vector, i.e. , and represent the speed of the UAV itself in the X-axis and Y-axis directions; and represent the coordinate difference value of the UAV and the target in the X-axis and Y-axis directions; represents the depth image information obtained by the on-board camera of the UAV, and the global state is a combination of all local observations of the UAVs; Steps 1-3, in combination with the actual demand of the multi-UAV target tracking task, the actions of the UAVs are set as: are set as: wherein the speed of the drone is constrained : , , represent the minimum and maximum speed of the drone, respectively Step 1-4, in the process of performing target tracking task, the UAV needs to consider the two sub-tasks of tracking close to the target and avoiding collision, and define the reward function : Setting tracking reward for drone and collision avoidance reward , and a weighted sum of the two as the total reward obtained by the drone, wherein, , represent weight coefficients of the tracking reward and the collision avoidance reward, respectively. Step 1-5, setting reinforcement learning hyperparameters and task-related parameters, including: sampling batch size, learning rate, total number of training rounds, length of each round, discount factor, initial position coordinates of multiple UAVs, initial coordinates and speed of the target, number of obstacles in the environment, maximum and minimum speed of the UAV, flight height of the UAV.

3. The method of claim 2, wherein, The depth image of the forward-looking camera of the UAV is used as part of the state input, and the depth map of the non-RGB image provides implicit information for obstacle avoidance for the UAV.

4. The method of claim 2, wherein, Stacking four consecutive images obtained by the depth camera helps the network learn the ability to process time series information.

5. The method of claim 2, wherein, When the distance between the UAV and the target is greater than the desired distance, if the UAV makes a movement towards the target, a certain positive reward is given, otherwise a corresponding penalty is given; when the distance between the UAV and the target is within the desired distance and greater than the minimum distance, a fixed positive reward is given to the UAV; when the distance between the UAV and the target is less than the minimum distance, a fixed negative reward is given to the UAV as a penalty.

6. The method of claim 2, wherein, When setting the collision avoidance reward function, when the distance between the UAV and the surrounding obstacles is less than the set minimum safety distance, a negative reward is given to the UAV, otherwise a positive reward is given; when the UAV collides with the obstacle, a larger negative reward is given, thereby helping the UAV learn collision avoidance behavior.

7. The method of claim 1, wherein, The step S3 comprises: Step 3-1, initialize the Airsim simulation environment parameters, including the initial position coordinates of multiple unmanned aerial vehicles, target initial coordinates and speed, the number of obstacles in the environment, the maximum and minimum speed of the unmanned aerial vehicle, the flight height of the unmanned aerial vehicle; initialize the strategy network, evaluation network, and contrast learning network and experience pool, and obtain the initial state observation of the unmanned aerial vehicle agent; Step 3-2, calculate the global observation according to the initial observation of the agent; Step 3-3, input the initial observation and global observation of the agent into the strategy network to calculate the action of the agent; Step 3-4, each multi-agent executes the action and obtains the corresponding reward; Step 3-5, store the trajectory information obtained in the above process in the experience pool for subsequent training.

8. The method of claim 7, wherein, The Airsim simulation environment initialization method uses domain randomization to improve the robustness of the network model. In each initialization phase of the environment, the positions of the unmanned aerial vehicles and the static obstacles in the environment are randomly generated.

9. The method of claim 1, wherein, The step S4 comprises: Step 4-1, sample a batch of trajectory data B from the experience pool for training; Step 4-2, use the local observation of all agents in the same trajectory sequence to infer the global information from the local observation of the agent using contrast learning, and sample the observation information of all unmanned aerial vehicle agents at the same time step from the experience pool, denoted as the following sequence: wherein is the observation of the n th agent; The sequence is input into the student and teacher networks, respectively, and the output is and where as follows: wherein global information of the n-th agent inferred by the contrastive learning framework; Comparative learning is performed on and The agent can infer global information from local observations, which helps it make better decisions. The global observation information is extracted by constructing an attention integration module; the attention weight uses a self-attention mechanism to obtain global information for each agent: given the hidden states of two agents, the attention weight uses a bilinear mapping, and then uses a softmax function to normalize it to obtain, as follows: In the formula, and are two learnable linear transformations, representing the key and query, respectively; is the attention weight coefficient of the i th agent, is the transpose of , is the transpose of , is the hidden state of the agent i ; In order to extract more diversified global information, multiple attention heads are used and their outputs are connected together; By using the relative weight of each agent, the global information is obtained by the weighted sum of all other agents: wherein, is global information of the agent, i is attention weight coefficient of the agent, is hidden state of the agent, i is hidden state of the agent, is hidden state of the agent, j is hidden state of the agent, Step 4-3: The intelligent agent i Global information and local observation The data are spliced ​​together and used as new agent observations, which are then input into the policy network and evaluation network for reinforcement learning training. Step 4-4, update the strategy network and the evaluation network, and the loss function of the strategy network is as follows: wherein, is the batch size of the samples, is the local observation of all agents, is the action of all agents, wherein is the loss function of the policy network, is the agent policy, is the critic network, is the set of global information of all agents; The loss function of the evaluation network is: wherein, , is the target Q value; and is the evaluation network and its loss function; Step 4-4, in order to realize the adaptive entropy temperature coefficient, the entropy temperature coefficient is updated by using the following loss function: H is a target policy entropy; and is a temperature coefficient and its loss function; Step 4-5, the target evaluation network is updated by using the following loss function in a soft update manner: wherein, , is a soft update coefficient; and are an evaluation network and its corresponding target evaluation network, respectively; Step 4-6, while updating the strategy network and the evaluation network, the update of the contrast learning network is also performed simultaneously. The design idea of the contrast loss is: the reconstructed feature should be similar to the corresponding original feature and different from other features. The parameters of the teacher network are updated based on the parameters of the student network by using the exponential moving average line EMA, i.e. momentum encoder.

Citation Information

Patent Citations

  • Dynamic multi-target interference channel allocation and power decision method

    CN117580166A

  • Method for estimating dynamic target of unmanned aerial vehicle in information rejection environment

    WO2024120187A1