Cooperative tracking hunting method and system based on improved MADDPG
By improving the MADDPG algorithm and combining segment tree-based experience replay with the advantages of UAV vision, the problem of slow training convergence in UAV-UAV collaborative tracking and capture was solved, achieving more efficient collaborative target tracking and capture results.
Patent Information
- Application Number
- CN202511093672.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-12-23
AI Technical Summary
Traditional algorithms suffer from slow training convergence and difficulty in achieving convergence in collaborative tracking and capture missions involving drones and unmanned vessels. They are particularly difficult to achieve efficient collaboration in complex and ever-changing environments. Furthermore, traditional experience replay mechanisms cannot effectively distinguish key experiences, resulting in low learning efficiency.
The MADDPG algorithm is improved by introducing a segment tree data structure to prioritize experience replay mechanism, which prioritizes the processing of the state and action information of individual agents, reduces the amount of computation, accelerates the learning of key experiences through priority sampling, combines the aerial vision advantage of UAVs to transmit information to unmanned ships, and designs specific reward and punishment functions to promote collaborative cooperation.
It significantly improves the training efficiency and convergence speed of collaborative tracking and capture by UAVs and unmanned vessels, and can effectively avoid static and dynamic obstacles, achieving more efficient collaborative target tracking and capture.
Smart Images

Figure CN121187348A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-agent cooperative decision-making technology, and more specifically, to a cooperative tracking and containment method and system based on an improved MADDPG. Background Technology
[0002] With the rapid development of unmanned systems technology, unmanned aerial vehicles (UAVs) and unmanned surface vessels (USVs) are increasingly widely used in military, civilian, and scientific research fields. UAVs, with their high-altitude perspective, rapid maneuverability, and flexible deployment, excel in aerial reconnaissance, surveillance, and data collection; while USVs, with their surface navigation capabilities, long endurance, and strong environmental adaptability, are suitable for water patrol, target tracking, and mission execution. The combined system of the two enables three-dimensional, multi-dimensional target tracking and capture in complex environments, significantly improving mission efficiency and success rate. The collaborative execution of missions by UAVs and USVs, through an air-sea joint three-dimensional combat mode, can effectively cope with complex and ever-changing environments, providing real-time intelligence support and efficient mission execution capabilities, becoming an important technological means to address modern maritime challenges. Compared with traditional combat methods, this unmanned system is not only more efficient and lower in cost, but also significantly reduces personnel casualties. Collaborative tracking and capture involves air power (UAVs) and sea power (USVs) working together or coordinating operations to continuously track and encircle a target.
[0003] In 2025, Liu Lu et al. invented a maneuvering decision-making method for marine unmanned swarm encirclement based on an improved DDPG algorithm. This method includes: acquiring the previous state position of each vehicle; inputting the previous state position of each vehicle into an LSTM network to obtain the historical state position of each vehicle; inputting the historical state position of each vehicle into a corresponding pre-trained DDPG network to determine the action to be performed by each vehicle, thus obtaining the maneuvering decision for marine unmanned swarm encirclement. During training, the reward value of the pre-trained DDPG network is the sum of individual reward, team reward, and additional reward. This can further improve the success rate of encirclement. In 2025, Zhang Ya et al. invented a multi-target encirclement method and system for mobile robots based on density field and deep reinforcement learning. This method includes seven steps: perception information acquisition, target density calculation, target allocation decision, encirclement state judgment, observation state generation, action decision and execution, and encirclement process iteration. It utilizes a density field algorithm combined with relative position to calculate scores, thereby achieving reasonable real-time allocation of targets and robot grouping, ensuring a balanced encirclement force. Robots within the same group use a strategy trained by deep reinforcement learning to efficiently encircle targets. This invention significantly improves system performance in terms of mobile robots rapidly approaching targets and quickly forming encirclement formations, providing an efficient and intelligent solution for multi-target encirclement scenarios. In 2025, Tian Yongxiao et al. invented a multi-drone encirclement method and system, including: dynamic modeling of drones; constructing a 7-tuple of multiple drones in a three-dimensional dynamic environment to describe the encirclement scenario; each drone includes an Actor network and a Critic network, where the Actor network generates the drone's actions, and the Critic network generates evaluations of the corresponding actions. Extreme value search and adaptive entropy are combined to iteratively update the Actor and Critic networks, using the optimized Actor network to generate the optimal actions. Each drone executes the optimal actions, completing the encirclement of a single drone by multiple drones. The advantages of this invention are: avoiding getting trapped in local optima and improving encirclement efficiency.
[0004] The main problems in the collaborative tracking and containment of dynamic targets using UAVs and unmanned surface vessels are as follows: Traditional algorithms have many limitations in solving tracking and containment problems, requiring prior knowledge, and struggling to provide effective solutions when the objective function and constraints are complex. Furthermore, traditional algorithms are difficult to apply directly if the target is dynamic or the environment is constantly changing.
[0005] MADDPG (Multi-Agent Deep Deterministic Policy Gradient), a method in reinforcement learning, was proposed by OpenAI in 2017 and is specifically designed for multi-agent environments. Its core principle is that multiple agents interact with the environment, with each agent learning the optimal policy based on its own observations and reward functions, while also considering the behavior of other agents. It is well-suited for collaborative tracking and containment tasks involving UAVs and unmanned surface vessels. This algorithm is a deep reinforcement learning method for solving multi-agent reinforcement learning problems. It combines the powerful perceptual capabilities of deep learning with the decision-making capabilities of deterministic policy gradient methods, enabling multiple agents to collaboratively learn and optimize their respective policies without continuous supervision. It is suitable for handling complex decision-making tasks in high-dimensional state spaces and continuous action spaces. The core idea of this algorithm is centralized training and decentralized execution. During the training phase, each agent can learn using the state, action, and reward information of all other agents, thus obtaining a more stable policy update direction; while during the execution phase, each agent makes independent decisions based solely on its own observation information, achieving decentralized autonomous operation. In MADDPG, each agent has an independent Actor-Critic network structure. The Actor is responsible for generating the current policy and choosing the action, while the Critic evaluates the value of the policy in the global environment. The action policy output by the Actor network is continuously optimized through feedback from the Critic network. This architecture gives MADDPG strong coordination and adaptability, effectively addressing the non-stationarity issues caused by changes in the policies of other agents in multi-agent environments. The MADDPG algorithm, by employing a centralized training and distributed execution strategy, enhances the collaborative capabilities among multiple agents, giving it a better grasp of global information. However, precisely because of this architecture, each agent's Critic network needs to process the state and action information of all agents during training, leading to a significant increase in computational load as the number of agents increases, potentially causing problems such as decreased training efficiency.
[0006] In collaborative tracking and encirclement tasks, experience replay mechanisms are widely used to improve sample utilization and training stability. However, traditional uniform experience replay has some significant problems. First, because tracking and encirclement tasks are typically highly dynamic and complex, with frequent interactions between agents and non-stationary state transitions, only a few key experiences from a large amount of experience data are truly "valuable" for policy updates. Traditional experience replay uses an equal-probability sampling method, which cannot distinguish the importance of these experiences. This causes the model to frequently learn inefficient or redundant information during training, reducing learning efficiency and potentially slowing down convergence. Second, in multi-agent environments, the policy updates of individual agents are often affected by changes in the behavior of other agents, increasing environmental uncertainty. If the experience replay mechanism cannot capture those key experiences that bring large TD errors in a timely manner, it will be difficult to quickly correct policy biases, affecting the overall collaborative effect.
[0007] Therefore, current UAV and unmanned surface vessel (USV) collaborative target tracking technologies suffer from slow training convergence speed and difficulty in achieving convergence. Due to the complex environment and large state space of multi-agent systems, traditional reinforcement learning methods perform poorly when facing high-dimensional continuous action spaces and long-term dependencies, making it difficult for the model to quickly and stably converge to the optimal policy, thus affecting the system's real-time performance and stability. Summary of the Invention
[0008] The purpose of this invention is to provide a cooperative tracking and capture method and system based on an improved MADDPG, which can improve the efficiency of high-level cooperative tracking and capture between UAVs and unmanned vessels.
[0009] This invention provides a cooperative tracking and containment method based on an improved MADDPG, comprising the following steps: S1: Obtain the current state using each agent, and obtain the action execution set based on the current state; S2: Obtain a reward based on the current state and the set of actions to be executed; S3: Obtain the next state, and based on the current state, action execution set, reward, and next state, obtain experience samples and experience base; S4: Based on the experience base, set the priority of each experience sample in the experience base; based on the priority, construct a segment tree, and use the next state as the current state to update the state; S5: Use the segment tree to sample the experience base to obtain sampled experience samples; S6: Use the sampled experience samples to independently update each agent, and obtain the updated agent; S7: Based on the updated agent, update the priority of the experience samples in the experience base using the time difference error to obtain the updated experience base. S8: Using the updated experience base and the updated agent, perform cyclic sampling until training converges to obtain a trained agent; use the trained agent for cooperative tracking and capture.
[0010] The present invention also provides a cooperative tracking and capture system based on an improved MADDPG, the system comprising the following modules: The state acquisition and action generation module is configured to: acquire the current state using each intelligent agent, and obtain a set of actions to be executed based on the current state; The reward generation module is configured to: generate a reward based on the current state and the set of actions executed; The experience generation module is configured to: obtain the next state, and obtain experience samples and an experience library based on the current state, the action execution set, the reward, and the next state; The experience priority setting, segment tree construction, and state update module is configured as follows: based on the experience library, set the priority of each experience sample in the experience library; based on the priority, construct a segment tree, and update the state by taking the next state as the current state; The experience base sampling module is configured to: use the segment tree to sample the experience base to obtain sampled experience samples; The agent-independent update module is configured to: independently update each agent using the sampled experience data to obtain the updated agent; The experience sample priority update module is configured to: update the priority of experience samples in the experience database based on the updated agent using time difference error, so as to obtain the updated experience database. The training and execution module is configured to: perform cyclic sampling using the updated experience base and the updated agent until training converges to obtain a trained agent; and use the trained agent to perform cooperative tracking and capture.
[0011] Implementing the cooperative tracking and capture method and system based on the improved MADDPG provided by this invention has the following beneficial effects: This invention addresses the problems of long training times and difficulty in convergence in reinforcement learning-based UAV and unmanned surface vessel (USV) cooperative tracking and capture tasks due to high environmental complexity and the difficulty of multi-agent cooperation. It improves the MADDPG method by first modifying the Critic network in MADDPG, reducing computation by focusing on processing the state and action information of individual agents rather than calculating global information for all agents. It also fully leverages the aerial vision advantage of UAVs, enabling them to perceive target information in the environment that USVs cannot observe and transmit this information to USVs, achieving more efficient and stable cooperative target tracking. Finally, this invention improves the reward / penalty function design in MADDPG, guiding UAVs and USVs to cooperate in capturing dynamic targets while ensuring their own safety. This invention introduces a priority experience replay mechanism based on a segment tree data structure during training. By introducing priority, it improves the learning efficiency of the agent for important experiences. Traditional experience replay mechanisms typically use uniform sampling to randomly extract experience data from the experience replay buffer for training. While this helps break data correlation and improve sample utilization, it does not consider the differences in importance of different experiences to the learning process. Priority experience replay, on the other hand, assigns different sampling probabilities based on the "importance" of experiences, allowing experiences that bring larger TD errors to be sampled and learned more frequently, thereby accelerating convergence. Segment tree-based priority experience replay is an efficient data structure scheme for implementing priority experience replay. Its core lies in using a segment tree to maintain the priority information of experiences, thereby achieving fast priority query and update. This invention effectively solves the above problems by assigning a priority to each experience, allowing those experiences that bring greater learning benefits to be sampled and utilized more frequently. This mechanism can significantly accelerate the model convergence speed and improve the quality of policy learning. Especially in priority experience replay implemented using a segment tree structure, the insertion, updating, and sampling of experiences can all be done within a short timeframe. It completes efficiently within time complexity, making it ideal for processing large-scale, high-dimensional multi-agent experience data. It increases the sampling probability of important experiences, thereby significantly improving training efficiency and convergence speed, enabling UAVs and unmanned vessels to avoid not only static obstacles but also dynamic obstacles.
[0012] In summary, this invention can improve the convergence speed of intelligent agent training and enhance the efficiency of high-level collaborative tracking and capture between UAVs and unmanned vessels. Attached Figure Description
[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the cooperative tracking and capture method based on the improved MADDPG provided by the present invention; Figure 2This is a comparative diagram of the improved MADDPG algorithm provided by this invention; Figure 3 This is a schematic diagram of the improved MADDPG algorithm provided by the present invention; Figure 4 This is a schematic diagram of the summation tree provided by the present invention; Figure 5 This is a schematic diagram of the minimum value tree provided by the present invention; Figure 6 This is a schematic diagram of the encirclement formation provided by the present invention; Figure 7 This is a schematic diagram of the overall process of the cooperative tracking and capture method based on the improved MADDPG provided by the present invention. Detailed Implementation
[0014] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0015] Figure 1 A schematic diagram of the cooperative tracking and trapping method based on the improved MADDPG of this embodiment is shown. In this embodiment, the cooperative tracking and trapping method based on the improved MADDPG includes the following steps: S1: Obtain the current state using each agent, and obtain the action execution set based on the current state; In one exemplary embodiment, the intelligent agent includes an Actor network and a Critic network. The Actor network generates actions based on currently observed state information, and the Critic network evaluates the value of each agent's acquired state and actions. If the agent is a drone, it communicates unidirectionally with the unmanned surface vessel (USV) to transmit state information not yet observed by the USV. The current state includes the agent's own information, information of other agents, static obstacles, dynamic obstacles, target information, and missile information. The agent's own information includes speed and position. The information of other agents includes speed and position. The target information includes speed and position. The state information not yet observed by the USV includes target information, static obstacle information, dynamic obstacle information, and information of other USVs in the field of view. S2: Obtain a reward based on the current state and the set of actions to be executed; In one exemplary embodiment, step S2 specifically includes: obtaining a reward based on the current state and the set of actions performed, as shown in the formula: , , , , , , , , , , , , , , in, The total reward for each agent; The global reward reflects the collective performance of the agent team. The reward is the average distance the drone or unmanned vessel travels to the target; and These are rewards for target visibility and successful tracking, respectively. This represents the number of intelligent agents capable of observing the target; Indicates other situations; Describes the minimum value function; It is the time step being tracked; This is the average of the two-dimensional Euclidean distances from all unmanned vessels to the target; The minimum distance from any unmanned vessel to the target; The number of unmanned ships; Representation of intelligent agents Location; Indicates the location of the target; Denotes the Euclidean norm; This represents the maximum possible distance in the two-dimensional map. Map size; For the first Individual rewards for each drone agent; For the first Individual rewards for the intelligent agents of an unmanned vessel; Indicates distance penalty; This represents the rewards and punishments for being hit by a missile or evading a missile; in drones This represents coverage rewards in unmanned ships. The system rewards unmanned vessels for communicating with drones, meaning that unmanned vessels will be rewarded for being within the communication range of drones, and there will be no longer any penalties. Indicates a flight altitude bonus for the drone; The reward represents the penalty mechanism for encountering obstacles; The representative example is the penalty mechanism for collisions between intelligent agents; It is the ideal distance threshold to the target; This is the drone's current flight altitude; Indicates the minimum flight altitude of the drone; Indicates the maximum altitude at which the drone can fly; This indicates a reward for the unmanned vessel encirclement formation; It is an exponential function; It is the standard deviation of the angle difference, reflecting the uniformity of the encirclement; It is the ideal angular interval; S3: Obtain the next state, and based on the current state, action execution set, reward, and next state, obtain experience samples and experience base; S4: Based on the experience base, set the priority of each experience sample in the experience base; based on the priority, construct a segment tree, and use the next state as the current state to update the state; S5: Use the segment tree to sample the experience base to obtain sampled experience samples; In one exemplary embodiment, step S5 specifically includes: S51: Divide the sampling probability interval into a preset number of sub-intervals; In one exemplary embodiment, the preset quantity is the batch size for training; S52: Obtain random sampling samples based on the sub-intervals and the segment tree; As an exemplary embodiment, in step S52, a number is randomly selected from each sub-interval, and the summation tree is traversed from top to bottom and from left to right until a leaf node is reached to find the sample; S53: Calculate the sampling probability and weight of the random sample based on the random sample, and normalize the weight based on the sampling probability and weight and the minimum priority of the root node of the minimum tree in the segment tree to obtain the empirical sample weight. As an exemplary embodiment, in step S53, the sampling probability is calculated. and weight Then use Figure 5 Find the minimum priority at the root node of the minimum tree, and normalize the weights to obtain the final weights. ; S54: Repeat steps S51-S53 for a preset number of repetitions, and obtain the sampled experience samples according to the experience sample weights. As an exemplary embodiment, in step S54, the process is repeated Batch-size times to obtain a batch of extracted experience; Batch-size is the batch size. S6: Use the sampled experience samples to independently update each agent, and obtain the updated agent; As an exemplary embodiment, in step S6, for each agent... Perform independent updates: Calculate the target Q-value; calculate the TD error; calculate the Critic loss function; apply gradient descent to update the Critic network parameters. ; Calculate the Actor loss function; Apply gradient ascent to update the Actor network parameters. ; S7: Based on the updated agent, update the priority of the experience samples in the experience base using the time difference error to obtain the updated experience base. S8: Using the updated experience base and the updated agent, perform cyclic sampling until training converges to obtain a trained agent; use the trained agent for cooperative tracking and capture.
[0016] This embodiment provides a cooperative tracking and capture system based on an improved MADDPG. The system includes the following modules: a state acquisition and action generation module, configured to: acquire the current state using each agent, and obtain an action execution set based on the current state; a reward generation module, configured to: obtain a reward based on the current state and the action execution set; an experience generation module, configured to: acquire the next state, and obtain experience samples and an experience library based on the current state, the action execution set, the reward, and the next state; and an experience priority setting, segment tree construction, and state update module, configured to: set the priority of each experience sample in the experience library based on the experience library; construct a segment tree based on the priority, and update the next state... The current state is used to update the state; the experience base sampling module is configured to: sample the experience base using the segment tree to obtain sampled experience samples; the agent independent update module is configured to: independently update each agent using the sampled experience samples to obtain an updated agent; the experience sample priority update module is configured to: update the priority of experience samples in the experience base using time difference error based on the updated agent to obtain an updated experience base; the training and execution module is configured to: perform cyclic sampling using the updated experience base and the updated agent until training converges to obtain a trained agent; and use the trained agent to perform cooperative tracking and capture.
[0017] In one exemplary embodiment, the intelligent agent includes an Actor network and a Critic network. The Actor network generates actions based on currently observed state information, and the Critic network evaluates the value of each agent's acquired state and actions. If the agent is a drone, it communicates unidirectionally with the unmanned surface vessel (USV) to transmit state information not yet observed by the USV. The current state includes the agent's own information, information of other agents, static obstacles, dynamic obstacles, target information, and missile information. The agent's own information includes speed and position. The information of other agents includes speed and position. The target information includes speed and position. The state information not yet observed by the USV includes target information, static obstacle information, dynamic obstacle information, and information of other USVs in the field of view. In one exemplary embodiment, the reward generation module is specifically configured to: obtain a reward based on the current state and the set of actions executed, as shown in the formula: , , , , , , , , , , , , , , in, The total reward for each agent; The global reward reflects the collective performance of the agent team. The reward is the average distance the drone or unmanned vessel travels to the target; and These are rewards for target visibility and successful tracking, respectively. This represents the number of intelligent agents capable of observing the target; Indicates other situations; Describes the minimum value function; It is the time step being tracked; This is the average of the two-dimensional Euclidean distances from all unmanned vessels to the target; The minimum distance from any unmanned vessel to the target; The number of unmanned ships; Representation of intelligent agents Location; Indicates the location of the target; Denotes the Euclidean norm; This represents the maximum possible distance in the two-dimensional map. Map size; For the first Individual rewards for each drone agent; For the first Individual rewards for the intelligent agents of an unmanned vessel; Indicates distance penalty; This represents the rewards and punishments for being hit by a missile or evading a missile; in drones This represents coverage rewards in unmanned ships. The system rewards unmanned vessels for communicating with drones, meaning that unmanned vessels will be rewarded for being within the communication range of drones, and there will be no longer any penalties. Indicates a flight altitude bonus for the drone; The reward represents the penalty mechanism for encountering obstacles; The representative example is the penalty mechanism for collisions between intelligent agents; It is the ideal distance threshold to the target; This is the drone's current flight altitude; Indicates the minimum flight altitude of the drone; Indicates the maximum altitude at which the drone can fly; This indicates a reward for the unmanned vessel encirclement formation; It is an exponential function; It is the standard deviation of the angle difference, reflecting the uniformity of the encirclement; It is the ideal angular interval; In one exemplary embodiment, the experience base sampling module is specifically configured as follows: dividing the sampling probability interval into a preset number of sub-intervals; obtaining random sampling samples based on the sub-intervals and the segment tree; calculating the sampling probability and weight of the random sampling samples; normalizing the weights based on the sampling probability and weights and the minimum priority of the root node of the minimum tree in the segment tree to obtain the experience sample weights; repeating the above steps for a preset number of repetitions, and obtaining sampled experience samples based on the experience sample weights.
[0018] In some embodiments, the above-described cooperative tracking and capture method based on the improved MADDPG can also be implemented in the following ways.
[0019] In this embodiment, the cooperative tracking and capture method based on the improved MADDPG includes: (I) Design of an improved scheme based on MADDPG: like Figure 2 The diagram shows a comparison of the improved MADDPG algorithm. The left image shows the application effect of the original MADDPG algorithm, and the right image shows the application effect of the improved MADDPG algorithm. In the improved MADDPG algorithm, the state of the unmanned vessel is obtained by combining its own observation and the observation of the UAV. In this invention, the human and the machine use their visual advantage to transmit information that the unmanned vessel has not yet observed. In this embodiment, the MADDPG algorithm is first improved. The agent's Critic network focuses only on itself, reducing the computational load of the Critic network. If the agent is a drone, it sends information from the drone's field of view to the unmanned surface vessel (USV), using communication to help the USV acquire more information. This compensates for the Critic network's limitation of only processing local information, enabling better collaboration among these heterogeneous agents. Furthermore, a segment tree-based priority experience replay mechanism is introduced to improve sample efficiency and accelerate network convergence. The algorithm flow is as follows: Figure 3 As shown: The entire process can be described as follows: First, each agent has its own independent Actor network and Critic network. The Actor network is responsible for generating actions based on the currently observed state information, while each agent's Critic network only processes its own historical state and action information, no longer directly relying on all historical data of other agents. This effectively reduces the input dimensionality of a single Critic network and improves model training efficiency. The Critic network is used to evaluate the value (Q-value) of the agents' joint actions. When an agent is a drone with aerial vision advantage, it can use its observation capabilities to identify obstacles, target locations, and the positions of other unmanned surface vessels (USVs) within its field of vision, and transmit this crucial information to the USV agent. After receiving this information, the USV inputs it as part of its extended state into its Actor and Critic networks, using communication to help the USV acquire more information, compensating for the Critic network's limitation of only processing local information, and enabling better collaboration among these heterogeneous agents.
[0020] During task execution, the agent interacts with the environment, obtains the current state, performs actions, receives rewards, and enters the next state. These experiences are recorded, including the states, actions, rewards, and next states of all agents. Subsequently, the system calculates the TD error (time difference error) of this experience as a measure of its importance, sets its priority accordingly, and inserts the experience and its priority into a segment tree-based priority experience replay buffer.
[0021] To improve the efficiency of priority experience playback, this embodiment defines a summation tree and a minimum value tree based on a segment tree, such as... Figure 4 The image shows a summation tree. Figure 5 The image shows a minimum value tree. Figure 4 The definition is a summation tree used to store the priorities of empirical samples, and the actual priorities of the samples. The priorities of two leaf nodes are stored in their respective parent nodes, and this summation continues until convergence to the root node. When using empirical samples, the sampling probability interval is... Divide into Batch-size lengths For each sub-interval, a number is randomly selected. This process is repeated traversing the summation tree from top to bottom and left to right until a leaf node is reached to find the sample. Then, the sampling probability is calculated. and weight Then use Figure 5 Find the minimum priority at the root node of the minimum tree, and normalize the weights to obtain the final weights. Repeat the process Batch-size times to obtain a batch of extracted experience. Using this data structure for priority experience replay can effectively accelerate the convergence speed of the algorithm.
[0022] Next, the system samples from the experience buffer according to priority, prioritizing the experiences that contribute significantly to the TD error. This strategy ensures that the model can learn key experiences more quickly, thereby accelerating convergence and improving policy quality.
[0023] In each training iteration, the Critic network updates its parameters by minimizing the TD error, while the Actor network updates its policy using the policy gradient method based on the Critic's evaluation results to maximize the expected reward. Furthermore, the target network maintains a smooth parameter transition through a soft update mechanism, thereby improving training stability.
[0024] The entire training process is iterative, and the agent gradually learns how to coordinate efficiently in continuous interaction with the environment to achieve the perception, tracking and capture of the target.
[0025] Overall algorithm flow: (1) Initialize environment parameters and network parameters; (2) Interact with the environment to obtain the current state in Belongs to intelligent agents The observation, according to the strategy Select Action If the intelligent agent is a drone, the drone transmits target information, static obstacles, dynamic obstacles, and information about other unmanned vessels in its field of vision. To the unmanned boat in sight; (3) Set of actions to be performed ; (4) Receive a reward To obtain the next state ; (5) Empirical samples Store in the experience base. Initially set priority. Storage priority To a segment tree; (6) Update status for ; (7) Sample batch data from the experience base: Priority range ; Randomly sample from each interval; Conversion priority to index; Calculate importance sampling weights ; Selected batch-size samples ; (8) For each agent Perform independent updates: Calculate the target Q value; Calculate the TD error; Calculate the Critic loss function; Apply gradient descent to update Critic network parameters ; Calculate the Actor loss function; Apply gradient ascent to update Actor network parameters ; (9) Soft update the target network; (10) Calculate the average TD error and update the empirical priority; (II) Design of reward and punishment functions To ensure the pursuers can safely and collaboratively track and capture the target, they must possess pursuit capabilities, collaborative communication capabilities, obstacle avoidance capabilities, missile evasion capabilities, and the ability to avoid collisions with other pursuers. Furthermore, the missions of UAVs and unmanned surface vessels (USVs) differ. UAVs utilize their visual advantage to assist USVs in the capture operation, while the USVs are primarily responsible for forming the capture formation. The system employs one-way communication, meaning the UAV sends information to the USV. Therefore, the UAV must not only observe the target but also be able to observe the USV, as communication between them is only possible within the UAV's observation range. Consequently, this embodiment designs corresponding reward and penalty functions to address these requirements.
[0026] (1) Global Rewards The global reward reflects the collective performance of the agent team, including the average distance from the drone / unmanned surface vessel to the target, the reward for target visibility, and the reward for successful tracking, and is defined as: (1)
[0027]
[0028] The global rewards for drones and unmanned ships are the same. Taking unmanned ships as an example, here... A distance-based reward system penalizes the average distance the drone travels to the target, while providing additional rewards for drones that approach the target. (2)
[0029]
[0030] in For all unmanned ships (index) The average value of the two-dimensional plane Euclidean distance to the target. The minimum distance from any unmanned vessel to the target. The maximum possible distance in the 2D map is denoted as . This refers to the map size. , The threshold value is set.
[0031] In formula (1) It is a reward for target visibility. This represents the number of intelligent agents capable of observing the target. Successful tracking reward It is the time step being tracked.
[0032] (2) Total Rewards The total reward for each agent combines a portion of the global reward and the individual reward: (3) UAV rewards Represented as: (4)
[0033] The individual reward and penalty formulas for drones and unmanned surface vessels (USVs) are very similar, except that USVs are not subject to altitude constraints. We will use the reward and penalty formula for drones as an example here. This represents the rewards and punishments for being hit by a missile or evading a missile. The representative is the penalty mechanism for encountering obstacles. The representative example is the penalty mechanism for collisions between intelligent agents. These two obstacle avoidance and collision avoidance mechanisms are the most basic: if a collision occurs, a penalty is imposed. Therefore, this embodiment will not use formulas for explanation. Indicates distance penalty, It is the ideal distance threshold to the target. It's about target visibility rewards and penalties. The difference between the reward and penalty functions for drones and unmanned surface vessels (USVs) lies in the fact that USVs lack... However, there are limitations, but there is a possibility of forming an encirclement formation. The second requirement is There are some differences, in drones This represents coverage reward, meaning that the drone's observation range can cover the unmanned vessel. In the settings, the drone's observation range is also its communication range. The system uses the drone to inform the unmanned vessel in advance of information that it has not yet observed.
[0034] (5) in It is the total number of unmanned boats. This refers to the number of unmanned vessels observed by drones. The unmanned vessel reward and punishment mechanism... This represents communication rewards, which are determined by whether the user is within the drone's observation range.
[0035] Drones are subject to altitude restrictions, so The purpose is to encourage drones to fly within a suitable altitude range to optimize their visibility and mission performance, while avoiding the risks associated with flying too high or too low. The definition is shown in formula (6). This is the drone's current flight altitude: (6) The unmanned ships also need to form an encirclement formation. The definition is shown in formula (7): (7) in It is the standard deviation of the angle difference, reflecting the uniformity of the encirclement. It is the ideal angular interval.
[0036] This is to ensure a more uniform encirclement formation; for example, three unmanned boats would form a triangle around the target. Figure 6 As shown in (a), the four unmanned boats will form a square around the target, as... Figure 6 As shown in (b), the overall process of the cooperative tracking and capture method based on the improved MADDPG is as follows: Figure 7 As shown.
[0037] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A cooperative tracking and containment method based on an improved MADDPG, characterized in that, Includes the following steps: S1: Obtain the current state using each agent, and obtain the action execution set based on the current state; S2: Obtain a reward based on the current state and the set of actions to be executed; S3: Obtain the next state, and based on the current state, action execution set, reward, and next state, obtain experience samples and experience base; S4: Based on the experience base, set the priority of each experience sample in the experience base; based on the priority, construct a segment tree, and use the next state as the current state to update the state; S5: Use the segment tree to sample the experience base to obtain sampled experience samples; S6: Use the sampled experience samples to independently update each agent, and obtain the updated agent; S7: Based on the updated agent, update the priority of the experience samples in the experience base using the time difference error to obtain the updated experience base. S8: Using the updated experience base and the updated agent, perform cyclic sampling until training converges to obtain a trained agent; use the trained agent for cooperative tracking and capture.
2. The cooperative tracking and containment method based on improved MADDPG according to claim 1, characterized in that, The intelligent agent comprises an Actor network and a Critic network. The Actor network generates actions based on currently observed state information, while the Critic network evaluates the value of each agent's acquired state and actions. If the agent is a drone, it communicates unidirectionally with the unmanned surface vessel (USV) to transmit state information not yet observed by the USV. The current state includes the agent's own information, information of other agents, static obstacles, dynamic obstacles, target information, and missile information. The agent's own information includes speed and position. The information of other agents includes speed and position. The target information includes speed and position. The state information not yet observed by the USV includes target information, static obstacle information, dynamic obstacle information, and information of other USVs in its field of vision.
3. The cooperative tracking and containment method based on improved MADDPG according to claim 1, characterized in that, Step S2 specifically includes: obtaining a reward based on the current state and the set of actions to be executed, as shown in the formula: , , , , , , , , , , , , , , in, The total reward for each agent; The global reward reflects the collective performance of the agent team. The reward is the average distance the drone or unmanned vessel travels to the target; and These are rewards for target visibility and successful tracking, respectively. This represents the number of intelligent agents capable of observing the target; Indicates other situations; Describes the minimum value function; It is the time step being tracked; This is the average of the two-dimensional Euclidean distances from all unmanned vessels to the target; The minimum distance from any unmanned vessel to the target; The number of unmanned ships; Representation of intelligent agents Location; Indicates the location of the target; Denotes the Euclidean norm; This represents the maximum possible distance in the two-dimensional map. Map size; For the first Individual rewards for each drone agent; For the first Individual rewards for the intelligent agents of an unmanned vessel; Indicates distance penalty; This represents the rewards and punishments for being hit by a missile or evading a missile; in drones This represents coverage rewards in unmanned ships. The system rewards unmanned vessels for communicating with drones, meaning that unmanned vessels will be rewarded for being within the communication range of drones, and there will be no longer any penalties. Indicates a flight altitude bonus for the drone; The reward represents the penalty mechanism for encountering obstacles; The representative example is the penalty mechanism for collisions between intelligent agents; It is the ideal distance threshold to the target; This is the drone's current flight altitude; Indicates the minimum flight altitude of the drone; Indicates the maximum altitude at which the drone can fly; This indicates a reward for the unmanned vessel encirclement formation; It is an exponential function; It is the standard deviation of the angle difference, reflecting the uniformity of the encirclement; It is the ideal angular interval.
4. The cooperative tracking and containment method based on improved MADDPG according to claim 1, characterized in that, Step S5 specifically includes: S51: Divide the sampling probability interval into a preset number of sub-intervals; S52: Obtain random sampling samples based on the sub-intervals and the segment tree; S53: Calculate the sampling probability and weight of the random sample based on the random sample, and normalize the weight based on the sampling probability and weight and the minimum priority of the root node of the minimum tree in the segment tree to obtain the empirical sample weight. S54: Repeat steps S51-S53 for a preset number of repetitions, and obtain the sampled experience samples according to the experience sample weights.
5. A cooperative tracking and containment system based on an improved MADDPG, characterized in that, The system includes the following modules: The state acquisition and action generation module is configured to: acquire the current state using each intelligent agent, and obtain a set of actions to be executed based on the current state; The reward generation module is configured to: generate a reward based on the current state and the set of actions executed; The experience generation module is configured to: obtain the next state, and obtain experience samples and an experience library based on the current state, the action execution set, the reward, and the next state; The experience priority setting, segment tree construction, and state update module is configured as follows: based on the experience library, set the priority of each experience sample in the experience library; based on the priority, construct a segment tree, and update the state by taking the next state as the current state; The experience base sampling module is configured to: use the segment tree to sample the experience base to obtain sampled experience samples; The agent-independent update module is configured to: independently update each agent using the sampled experience data to obtain the updated agent; The experience sample priority update module is configured to: update the priority of experience samples in the experience database based on the updated agent using time difference error, so as to obtain the updated experience database. The training and execution module is configured to: perform cyclic sampling using the updated experience base and the updated agent until training converges to obtain a trained agent; and use the trained agent to perform cooperative tracking and capture.
6. The cooperative tracking and capture system based on the improved MADDPG according to claim 5, characterized in that, The intelligent agent comprises an Actor network and a Critic network. The Actor network generates actions based on currently observed state information, while the Critic network evaluates the value of each agent's acquired state and actions. If the agent is a drone, it communicates unidirectionally with the unmanned surface vessel (USV) to transmit state information not yet observed by the USV. The current state includes the agent's own information, information of other agents, static obstacles, dynamic obstacles, target information, and missile information. The agent's own information includes speed and position. The information of other agents includes speed and position. The target information includes speed and position. The state information not yet observed by the USV includes target information, static obstacle information, dynamic obstacle information, and information of other USVs in its field of vision.
7. The cooperative tracking and capture system based on the improved MADDPG according to claim 5, characterized in that, The reward generation module is specifically configured to: obtain a reward based on the current state and the set of actions executed, as shown in the formula: , , , , , , , , , , , , , , in, The total reward for each agent; The global reward reflects the collective performance of the agent team. The reward is the average distance the drone or unmanned vessel travels to the target; and These are rewards for target visibility and successful tracking, respectively. This represents the number of intelligent agents capable of observing the target; Indicates other situations; Describes the minimum value function; It is the time step being tracked; This is the average of the two-dimensional Euclidean distances from all unmanned vessels to the target; The minimum distance from any unmanned vessel to the target; The number of unmanned ships; Representation of intelligent agents i Location; Indicates the location of the target; Denotes the Euclidean norm; This represents the maximum possible distance in the two-dimensional map. Map size; For the first Individual rewards for each drone agent; For the first Individual rewards for the intelligent agents of an unmanned vessel; Indicates distance penalty; This represents the rewards and punishments for being hit by a missile or evading a missile; in drones This represents coverage rewards in unmanned ships. The system rewards unmanned vessels for communicating with drones, meaning that unmanned vessels will be rewarded for being within the communication range of drones, and there will be no longer any penalties. Indicates a flight altitude bonus for the drone; The reward represents the penalty mechanism for encountering obstacles; The representative example is the penalty mechanism for collisions between intelligent agents; It is the ideal distance threshold to the target; This is the drone's current flight altitude; Indicates the minimum flight altitude of the drone; Indicates the maximum altitude at which the drone can fly; This indicates a reward for the unmanned vessel encirclement formation; It is an exponential function; It is the standard deviation of the angle difference, reflecting the uniformity of the encirclement; It is the ideal angular interval.
8. The cooperative tracking and capture system based on the improved MADDPG according to claim 5, characterized in that, The specific configuration of the experience base sampling module is as follows: Divide the sampling probability interval into a preset number of sub-intervals; Based on the sub-intervals and the segment tree, random sampling samples are obtained; Based on the random sample, calculate its sampling probability and weight. Based on the sampling probability and weight and the minimum priority of the root node of the minimum tree in the segment tree, normalize the weight to obtain the empirical sample weight. Repeat the above steps for a preset number of times, and obtain the sampled experience samples according to the experience sample weights.