UAV Cooperative Pursuit Method for Multi-Degree-of-Freedom Model of Multi-Agent Reinforcement Learning
Through the multi-agent deep reinforcement learning algorithm MADDPG, a multi-degree-of-freedom drone model is built, which solves the problems of poor coordination and slow learning of the drone cluster in complex scenarios, and achieves efficient multi-drone collaborative pursuit effect.
Patent Information
- Application Number
- CN202310296946.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-24
AI Technical Summary
There are problems of poor synergy, slow learning and convergence speed in the existing drone cluster hunting methods, especially under conditions of multiple degrees of freedom and non-equal motion parameters, the pursuit scenarios are less studied, and traditional methods are difficult to apply to complex three-dimensional scenarios.
The multi-agent deep reinforcement learning algorithm MADDPG is adopted to build a multi-degree-of-freedom drone model through centralized training and decentralized execution, and use the Actor and Critic network to make collaborative decisions between agents, and optimize the pursuit strategy of the drone cluster with the reward mechanism.
The accuracy and efficiency of the coordinated pursuit of drone clusters in complex scenarios has been improved, and efficient coordinated combat between multiple drones has been achieved.
Smart Images

Figure CN116225065B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of reinforcement learning and multi-UAV confrontation, and relates to a UAV cooperative pursuit method with a multi-degree-of-freedom model for multi-agent reinforcement learning. Specifically, it relates to a UAV cooperative pursuit method with a multi-degree-of-freedom model based on multi-agent reinforcement learning. It mainly completes the research on the pursuit method of multiple low-speed pursuit UAVs for a single high-speed escaping UAV in a military combat simulation scenario using a multi-degree-of-freedom UAV model and a multi-agent reinforcement algorithm, which has very important practical significance for improving the cooperative air combat confrontation ability of multiple UAVs. Background Art
[0002] With the rapid development of modern technology, the future battlefield environment has become increasingly complex and changeable. Unmanned combat equipment with strong concealment, low cost, and high accompanyability has become increasingly important, even subverting traditional war concepts. With the increasing complexity of the unmanned equipment system, the concept of cooperative combat proposed to improve combat effectiveness has also developed rapidly. However, when designing encirclement strategies using traditional methods, a single assumption is often made about the movement strategy of the escaping target. However, in the real battlefield environment, it is difficult for one's own side to know the control strategy of the escaping target. At the same time, when the environmental model changes, it is difficult to quickly adapt the controller parameters, which has certain limitations.
[0003] In recent years, with the continuous enrichment of reinforcement learning algorithms, the problems that can be solved by artificial intelligence technology have shifted from fully informed dynamic game problems in simple environments to incomplete information dynamic game problems in complex environments. The development of multi-agent reinforcement learning provides a new method for solving the UAV swarm pursuit problem. Major military powers continue to develop UAV swarm combat capabilities, hoping to use a system of low-cost UAV swarms to harass relatively isolated high-value military targets and gain an asymmetric combat advantage. Win victory in future multi-domain and multi-dimensional system-of-systems warfare.
[0004] In future wars, UAV swarms will surely play an important role in the battlefield, and the intelligence of agent swarms will become more and more in-depth. Therefore, in the face of the UAV swarm pursuit problem with multiple degrees of freedom, using reinforcement learning algorithms to construct a set of high-efficiency training algorithms to teach agents to complete cooperative pursuit work in a continuous and dynamically changing environment, and improving the self-adaptability and cooperation of multi-agents have important guiding significance for intelligent agent cooperative combat in modern battlefields.
[0005] Solutions of the prior art:
[0006] In existing drone swarm pursuit methods based on reinforcement learning, the control of drone models is generally a single-degree-of-freedom model. Based on this model, pursuit drones are selected in a two-dimensional scenario to surround and capture an escaping drone. At the same time, the control algorithm for the pursuit drone swarm uses a single-agent algorithm for control, that is, there is no communication between units within the drone swarm.
[0007] Disadvantages of the prior art:
[0008] 1. Some problems of drone swarms based on reinforcement learning are simplified to problems of single-agent drones. Using such algorithms in multi-agent unmanned systems will result in a series of problems such as poor coordination, slow learning and convergence speeds, and even difficulty in convergence.
[0009] 2. Most of the existing combat simulation scenarios are two-dimensional scenarios, that is, the controlled drones in the algorithm are single-degree-of-freedom models. Such methods are relatively difficult to apply in practice.
[0010] 3. Currently, in most of the pursuit problem scenarios, it is set that the speed of the pursuit drones is superior to that of the escaping drone. However, there is relatively little research on the scenario where the speed of the pursuit drones is inferior to that of the escaping drone. It is necessary to study a more complex and accurate model that can handle the pursuit problem under such non-equal motion parameter conditions based on the advantages of swarm intelligence. Summary of the Invention
[0011] Technical problems to be solved
[0012] In order to avoid the deficiencies of the prior art, the present invention proposes a method for collaborative pursuit of drones with a multi-degree-of-freedom model based on multi-agent reinforcement learning, explores the confrontation strategy of using a multi-degree-of-freedom drone model to surround and capture a high-speed escaping drone with a low-speed pursuit drone swarm in a military combat scenario, and uses a multi-agent deep reinforcement learning algorithm to control the communication and cooperation between agents, which has certain practical guiding significance for modern drone swarm air combat.
[0013] Technical solution
[0014] A method for collaborative pursuit of drones with a multi-degree-of-freedom model based on multi-agent reinforcement learning, characterized in that: there are multiple homogeneous pursuit drones of the red side and a single escaping drone of the blue side in the combat area, and the red side drones cooperate with each other to successfully surround and capture the escaping target as soon as possible. The steps are as follows:
[0015] Step 1: For the intelligent agents of the two warring sides, the red side and the blue side, the red side units are controlled using a reinforcement learning algorithm, and the blue side units are based on traditional combat rules. The intelligent agent environment models of both sides are:
[0016] Let P n(n = 1, 2, …, N) represents multiple red - side drones for encirclement and capture, E represents the escaping drone, and v E represents the magnitude of the velocity of the escaping drone, represents the magnitude of the velocity of the pursuing drone, d cap represents the encirclement radius, ψ E represents the yaw angle of the escaping drone, represents the yaw angle of the pursuing drone, d t is the distance between the pursuing drone and the escaping drone, d i is the distance between the pursuing drones;
[0017] The red - side algorithm agent model includes the kinematic equations of the pursuing drones, the state space, action space, and reward function of the agents;
[0018] The blue - side rule agent model is the escape countermeasure strategy adopted by the escaping drone;
[0019] Step 2: Adopt the Multi - Agent Deep Deterministic Policy Gradient algorithm (MADDPG) as the red - side agent algorithm, where MADDPG uses the method of centralized training and decentralized execution;
[0020] Construct a value Critic network and a policy Actor network, where: The value network Critic is deployed on the global controller, and the policy network Actor is deployed on each agent. During training, the agent i transmits the observation value state i to the global value network, and the value network transmits the TD error back to the agent for the agent to train the policy network. At this time, the agents do not communicate directly, but the trained policy network makes decisions;
[0021] Use the MADDPG algorithm to train and optimize the red - side agents;
[0022] Step 3: Combine the agent environment model constructed in Step 1 and the multi - agent reinforcement learning algorithm in Step 2 to generate the final multi - drone collaborative encirclement and capture method based on reinforcement learning. The process is as follows:
[0023] Step 3 - 1: Based on the current agent, calculate the difference between the current agent and the other agents. The difference is:
[0024] Longitude difference
[0025] Latitude difference
[0026] Height difference
[0027] Distance difference
[0028] Obtain the yaw angle of the current agent Input the joint state of the agent where
[0029] Step 3-2: Input the agent joint state into the multi-agent reinforcement learning algorithm to obtain the next joint action where And execute the action in the 3D simulation combat environment;
[0030] Step 3-3: After the action is executed, obtain the next action of the agent and the reward value R of the current action n , and store the data (S n , A n , S n+1 , R n ) into the experience buffer pool, and extract a batch of data to train the algorithm;
[0031] During the entire encirclement process, step 3 operations are looped.
[0032] The successful encirclement satisfies the following conditions: 1) There exists any pursuit UAV P n (n = 1, 2,..., N) whose distance from the escape target E is less than the encirclement radius d cap ; 2) The encirclement angle between adjacent pursuit UAVs is not greater than π.
[0033] The following constraints are satisfied during the encirclement process: 1) To avoid the influence of terrain and temperature on the UAVs, the flight height of the UAVs is restricted between 1000 meters and 3000 meters; 2) The pursuit UAVs need to capture the escape UAV within the restricted area, and if the escape UAV exceeds the restricted area, the mission is judged as failed; 3) Collisions cannot occur between the pursuit UAVs.
[0034] The kinematic equation of the UAV in the red side algorithm agent model is:
[0035]
[0036] where (x i , y i ) represents the current position of the UAV, h i represents the current height of the UAV, respectively represent the track yaw angle and track pitch angle of UAV i in the nth cycle. The track yaw angle δ i and the track pitch angle ωi Constrained by: -ω max <ω i <ω max , -δ max <δ i <δ max ;
[0037] The state space of the agent is:
[0038]
[0039] Where: is the situation information of a single pursuit UAV at simulation step n;
[0040] The action space of the agent is:
[0041]
[0042] Where: is the action taken by a single pursuit UAV at simulation step n, where:
[0043]
[0044] The reward function is: The reward function design adopts a combination of continuous rewards and sparse rewards. For the UAV cooperative pursuit problem, two main factors are considered: First, the pursuit UAVs need to successfully capture the escaping UAV. In a multi-UAV pursuit scenario, as long as one UAV captures the escaping UAV, the mission is considered successful; Second, the pursuit UAVs cannot collide with each other. The specific expression is as follows:
[0045] R = r sparse + r step
[0046] Where: includes the sparse reward r sparse and the step reward r step .
[0047] The situation information of a single pursuit UAV at simulation step n is:
[0048]
[0049] Where:
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] Wherein: are respectively the relative longitude, relative latitude, and relative altitude between the pursuing UAV and the escaping UAV. and are respectively the track deviation angle and track tilt angle of the pursuing UAV. is the distance between the pursuing UAV and the escaping UAV.
[0057] The sparse reward r sparse and the step reward r step are:
[0058] The sparse reward r sparse of the pursuing UAV is divided into the following two modules: one is to give a positive reward when one UAV in the pursuing UAV cluster successfully captures the escaping UAV; the other is that when the escaping UAV successfully escapes from the area, it is regarded as a mission failure and a negative reward is given;
[0059]
[0060] Each pursuing UAV will obtain a step reward r step once according to the executed action after each simulation step, step and the UAV is guided to complete the established task through this reward. The step reward r
[0061] r step =αr1 + βr2 + γr3
[0062] Wherein: r1 is the pursuit distance reward, r2 is the pursuit height difference reward, and r3 is the UAV collision reward. α, β, and γ are weighting coefficients, and α + β + γ = 1.
[0063] The pursuit distance reward r1, the pursuit height difference reward r2, and the UAV collision reward r3 are:
[0064] r1 = -k(d t -d max )
[0065] Wherein: d t is the relative distance between UAVs, and d max is the maximum strike range of the pursuing UAV; r1 is set as a negative reward function, and when the distance between the pursuing UAV and the escaping UAV is the strike distance of the pursuing UAV, r1 = 0;
[0066] r2 = -k(h i -h E )
[0067] When the height difference h i -h E = 0 between the pursuing UAV and the escaping UAV, the height relationship between the pursuing UAV and the escaping target is locally optimal;
[0068]
[0069] A reward function r3 in the form of a negative exponent is established to describe the collision risk between pursuing UAVs, and d min represents the closest distance between the current UAV and other UAVs.
[0070] The escape countermeasure strategy adopted by the escaping UAV is as follows: when surrounded by pursuing UAVs, the escaping UAV escapes towards the midpoint of the farthest distance among the midpoints of all side lengths of the polygon formed by the pursuing UAVs; when not surrounded by escaping UAVs, the idea of the artificial potential field method is adopted. It is assumed that the pursuing UAV exerts a repulsive force in the vector direction towards the escaping UAV, and the repulsive force component between the two is an inverse function of the distance between the two: as the distance increases, the repulsive force decreases. The escaping UAV escapes in the direction of the repulsive force after synthesizing the repulsive force vectors given by all pursuing UAVs.
[0071] The Actor network structure in the MADDPG algorithm is as follows:
[0072]
[0073] The Critic network structure in the MADDPG algorithm is as follows:
[0074]
[0075] Beneficial effects
[0076] A method for cooperative pursuit of UAVs with a multi-degree-of-freedom model based on multi-agent reinforcement learning proposed by the present invention. Since the multi-agent reinforcement learning algorithm is used to study the problem of multi-UAV pursuit, it shows more intelligent autonomous decision-making than traditional mathematical model methods or single-agent reinforcement learning methods. At the same time, in the present invention, a method for deducing the multi-UAV encirclement strategy based on reinforcement learning is established, and a multi-degree-of-freedom UAV model cluster confrontation strategy is formulated. Due to the use of a multi-degree-of-freedom UAV model, a more complex and accurate model update and optimization are constructed, making up for the deficiencies of existing methods in the air combat confrontation method of multi-agent systems in complex scenarios and improving the accuracy of the air combat model. Description of the drawings
[0077] Figure 1 : Schematic diagram of UAV encirclement situation
[0078] Figure 2 : Schematic diagram of the UAV coordinate system
[0079] Figure 3 : Actor-Critic network framework diagram
[0080] Figure 4 : MADDPG algorithm flowchart
[0081] Figure 5 : Reward diagram of the UAV pursuit algorithm in a three-dimensional scenario Specific implementation manner
[0082] The present invention will be further described in combination with embodiments and drawings as follows:
[0083] The technical solution adopted by the present invention:
[0084] Step 1, using the neural network model, battlefield environment model, situation judgment and combat target allocation model in the existing system, assuming that the intelligent agents on both sides of the battle are the red side and the blue side, the red side units are controlled using the reinforcement learning algorithm, and the blue side units are constructed based on traditional combat rules. First, construct the red side algorithm intelligent agent model and the blue side rule intelligent agent model.
[0085] The task scenario of the present invention is described as follows: There are multiple homogeneous pursuit UAVs of the red side and escape UAVs of the blue side in the combat area, and both sides have opposite tactical purposes: The UAVs of the red side need to cooperate with each other to quickly surround and capture the escape target, while the escape target has to avoid and stay away from the UAV group of the red side. Existing research usually believes that when the distance between any pursuer and the escapee is less than a given threshold, the encirclement task is considered successfully completed. As Figure 1 shown.
[0086] Figure 1 In, P n (n = 1, 2,..., N) represents the UAVs of the red side, E represents the escape UAV, v E represents the speed magnitude of the escape UAV, represents the speed magnitude of the pursuit UAV, d cap represents the encirclement radius, ψ E represents the yaw angle of the escape UAV, represents the yaw angle of the pursuit UAV, d t is the distance between the pursuit UAV and the escape UAV, d i is the distance between the pursuit UAVs.
[0087] It is stipulated that the following conditions need to be met for a successful encirclement: 1) There exists any pursuit UAV P n (n = 1, 2,..., N) whose distance from the escape target E is less than the encirclement radius d cap; 2) The encirclement angle between adjacent pursuit drones is not greater than π.
[0088] The following constraints need to be satisfied during the encirclement process: 1) To avoid the influence of terrain and temperature on the drones, the flight altitude of the drones is restricted between 1000 meters and 3000 meters; 2) The pursuit drones need to capture the escaping drone within the defined area, and if the escaping drone exceeds the defined area, the mission is judged as failed; 3) Collisions cannot occur between the pursuit drones.
[0089] Step 2, adopt the MADDPG algorithm as the multi-agent deep reinforcement learning algorithm, and construct a suitable Actor network and Critic network.
[0090] Step 3, combine the intelligent agent environment model constructed in Step 1 with the multi-agent deep reinforcement learning algorithm in Step 2 to generate the final multi-agent collaborative optimization method driven by reinforcement learning in a multi-domain heterogeneous environment.
[0091] Furthermore, the specific steps for constructing the red-side algorithm intelligent agent model and the blue-side rule intelligent agent model in Step 1 are as follows:
[0092] Step 1-1: Construct the blue-side rule intelligent agent model; construct the blue-side escaping drone unit, and the escaping drone adopts the following flexible escaping and countering strategy: that is, comprehensively and simply consider the battle situation. When surrounded by pursuit drones, the escaping drone escapes towards the midpoint of the longest distance among the midpoints of all side lengths of the polygon formed by the pursuit drones; when not surrounded by escaping drones, adopt the idea of the artificial potential field method. Assume that the pursuit drone exerts a repulsive force in the vector direction towards the escaping drone, and the repulsive force component between the two is an inverse function relationship with the distance between the two: as the distance increases, the repulsive force decreases. The escaping drone escapes in the direction of the repulsive force synthesized from the repulsive force vectors given by all pursuit drones.
[0093] Step 1-2: Construct the red-side algorithm intelligent agent model; the specific steps are as follows:
[0094] Step 1-2-1: Construct the red-side intelligent agent unit, and construct the kinematic equation of the pursuit drone as:
[0095]
[0096] where (x i , y i ) represents the current position of the drone, h i represents the current altitude of the drone, respectively represent the track yaw rate and track pitch rate of drone i in the nth cycle. The track yaw rate δ i and the track pitch rate ω i are restricted by constraints: -ω max<ω i <ω max , -δ max <δ i <δ max .
[0097] Step 1-2-2: Construct the state space of the agent; For collaborative pursuit in a three-dimensional environment, the longitude, latitude, and altitude of the pursuit UAVs need to be considered. It is assumed that both UAVs carry on-board GPS devices and gyroscopes, and can obtain their own position information, altitude information, and their own orientation angles, i.e., (x i , y i , h i , φ i ); They carry on-board fire control radar equipment and can obtain the position information, altitude information, and orientation angle (x E , y E , h E , ψ E ) of the detected target (air combat target). Considering the characteristics of the multi-agent pursuit problem, a rectangular coordinate system is constructed with the escaping UAV as the origin, and the relative values of the position information of the pursuing UAV and the escaping UAV are calculated.
[0098] The designed joint state space of the UAV pursuit problem at the simulation step of n is as follows:
[0099]
[0100] Where: is the situation information of a single pursuing UAV at the simulation step of n, specifically including:
[0101]
[0102] Among them:
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109] Where: are the relative longitude, relative latitude, and relative altitude between the pursuing UAV and the escaping UAV, respectively. and They are the track yaw rate and track pitch rate of the pursuit UAV respectively. is the distance between the pursuit UAV and the escaping UAV.
[0110] Step 1-2-3: Construct the action space of the agent; This patent designs an action space applicable to the multi-degree-of-freedom UAV model pursuit problem, finds the maximum influencing factor affecting the UAV pursuit strategy in the kinematic model of the UAV, decouples the action space into the current yaw angle, current pitch angle and current roll angle of the UAV, and controls the next flight direction of the UAV through the orientation angle of the UAV. Limited by the maximum yaw angle, in each simulation step, the maximum yaw angle of the UAV cannot exceed 15°.
[0111] The designed joint action space for the UAV pursuit problem is as follows:
[0112]
[0113] In the formula: is the action taken by a single pursuit UAV at the simulation step n, specifically including:
[0114]
[0115] Step 1-2-4: Set the reward and punishment mechanism in the environment, which is the reward and punishment return given by the environment when the agents reach a certain state. The reward function design adopts a combination of continuous reward and sparse reward. For the UAV cooperative pursuit problem, two main elements are considered: one is that the pursuit UAV successfully captures the escaping UAV. In the multi-UAV pursuit scenario, as long as one UAV captures the escaping UAV, the task is considered successful; the other is that the pursuit UAVs cannot collide with each other. Therefore, the relative distance of the UAVs also needs to be considered in the design of the reward function. The specific expression is as follows:
[0116] Step 1-2-4-1 Global reward function design. During the task process, the global reward of the pursuit UAV is divided into the following two modules: one is to give a positive reward return when one UAV in the pursuit UAV cluster successfully captures the escaping UAV; the other is that when the escaping UAV successfully escapes from the area, it is considered a task failure and a negative reward return is given.
[0117]
[0118] Step 1-2-4-2 Local reward function design. For each pursuit UAV, after each simulation step, a step reward will be obtained according to the executed action, and this reward is used to guide the UAV to complete the established task. The step reward r step is composed of multiple sub-rewards weighted, and the definition of the sub-reward r k is as follows:
[0119] 1) Pursuit distance reward r1
[0120] r1 = -k(d t -d max )
[0121] Where: d t is the relative distance between the UAVs, and d max is the maximum strike range of the pursuit UAV. To ensure that the pursuit UAV can efficiently complete the pursuit task, the relative distance between the pursuit UAV and the escaping UAV is calculated at each time step. r1 is set as a negative reward function, and this distance is positively correlated with the pursuit distance reward r1. The farther the relative distance is, the smaller r1 becomes. When the distance between the pursuit UAV and the escaping UAV is the strike distance of the pursuit UAV, r1 = 0.
[0122] 2) Pursuit altitude difference reward r2
[0123] r2 = -k(h i -h E )
[0124] When the altitude difference h i -h E = 0, it can be considered that the altitude relationship between the pursuit UAV and the escaping target is locally optimal.
[0125] 3) UAV collision reward r3
[0126]
[0127] A negative exponential form of the reward function r3 is established to describe the collision risk between the pursuit UAVs, and d min represents the closest distance between the current UAV and other UAVs.
[0128] In summary, the step reward of each UAV is the weighted sum of the above two reward functions:
[0129] r step = αr1 + βr2 + γr3
[0130] Where: α, β, and γ are weighting coefficients, and α + β + γ = 1.
[0131] The step reward r step in each item is set to a negative value, and when the cooperative situation formed between the UAVs is closer to the ideal state, the value of T step is closer to 0, so as to guide the UAVs to update to a better cooperative strategy; when the encirclement task is completed, all UAVs will receive a positive reward, enabling the UAV swarm to achieve the purpose of rapid encirclement.
[0132] In step 2, the MADDPG algorithm is used as the multi-agent reinforcement learning algorithm, and its algorithm architecture is shown in the figure. MADDPG uses the method of centralized training and decentralized execution. That is, each agent obtains the action executed in the current state according to its own policy, interacts with the environment, and stores the experience in its own experience cache pool. After all agents interact with the environment, each agent randomly extracts experiences from the experience pool to train its own neural network. In this architecture, we need to obtain the states of the agents in the environment and let the agents execute their respective actions to obtain rewards and return them to the reinforcement learning algorithm for training. The value network (Critic) is deployed on the global controller, and the policy network (Actor) is deployed on each agent. During training, the agent i transmits the observation value state i to the value network, and the value network transmits the TD error back to the agent for the agent to train the policy network. At this time, the agents do not communicate with each other, and the trained policy network makes decisions. The specific steps are as Figure 3 follows:
[0133] Step 2-1: Establish the network structures of the actor module and the critic module, and initialize the network parameters. The actor module is used for decision-making actions, and the critic module is used for evaluation feedback, which is divided into the following two steps:
[0134] Step 2-1-1: The schematic diagram of the network structure of the actor module used in the present invention is shown in Table 1. Taking the state s of each motion node as the input, it passes through three fully connected layers (Inner product layer). After the first two fully connected layers, the rectified linear unit (ReLU) is used as the activation function. The output of the third layer passes through a hyperbolic tangent function tanh(). The tanh() function is a variant of the sigmoid() function, and its value range is [-1, 1], rather than [0, 1] of the sigmoid function. The output result is two values, namely the current orientation angle of the drone and the current inclination angle of the drone. In each round of the iterative process, since the parameters of the network are dynamically changing, in order to make the learning of the parameters more stable, a copy of the actor network structure is retained, and the parameters of this copy are updated only at a certain time step;
[0135] Table 1 Actor Network Structure in MADDPG Algorithm
[0136]
[0137] Step 2-1-2: The schematic diagram of the critic module network structure used in the present invention is shown in Table 2. Taking the state s of each motion node as the input, it passes through a fully connected layer and a rectified linear activation function; then the output and the action a are used as the input of the second fully connected layer. After the output result is activated by the rectified linear unit, it is input into a long short-term memory network (LSTM), and the output result is the action-value Q corresponding to the state s and the action a.
[0138] Table 2 Critic Network Structure in MADDPG Algorithm
[0139]
[0140] Step 2-2: Train and optimize the deep deterministic policy gradient algorithm. The parameter update of the critic module depends on the action a calculated by the actor module; while the parameter update of the actor module depends on the action-value gradient calculated by the critic module. The two feedback on each other to optimize the algorithm. Therefore, repeat Step 2 until the optimization termination condition for multi-agent collaborative decision-making is met or the maximum number of iterations is reached.
[0141] In the said Step 3, combine the agent environment model constructed in Step 1 and the multi-agent reinforcement learning algorithm in Step 2 to generate the final multi-UAV collaborative enclosing method based on reinforcement learning.
[0142] Step 3-1: Taking the current agent as the reference, calculate the longitude difference latitude difference altitude difference distance difference of the current agent and the other agents, and obtain the orientation angle of the current agent Input the joint state of the agents where
[0143] Step 3-2: Input the joint state of the agents into the multi-agent reinforcement learning algorithm to obtain the next joint action where and execute the action in the three-dimensional simulation combat environment.
[0144] Step 3-3: After the execution of the action, obtain the next action of the agent and the reward value R n of the current action, and save the data (S n ), A n ), S n+1 ), R n) Store it in the experience buffer pool and extract a batch of data to train the algorithm.
[0145] Step 3-4: Loop and execute the above operations.
[0146] The algorithm flow chart is as Figure 4 shown:
[0147] The effects of the present invention can be further illustrated by the following simulation experiments.
[0148] 1. Simulation conditions
[0149] The present invention is based on a central processing unit of Inter(R) Core(TM) i7-10870H 2.20GHz CPU, NVIDIA GeForce GTX1660 GPU, 32GB of memory, and a Windows 10 operating system. Using a certain military chess simulation and deduction platform as the military simulation environment, the algorithm framework uses Baidu's PaddlePaddle framework.
[0150] 2. Simulation content
[0151] The number of random explorations designed in this experiment is 100 times. It can be seen from Figure 5 that in the first 100 random exploration stages, the rewards obtained by the agent are basically -100, that is, the escaping drone can escape successfully every time. After 100 rounds, the actions trained by the algorithm are used for execution. It can be seen that the reward value of the pursuing drone has increased significantly and stabilized at about 500 points, that is, the pursuing drone can catch up at the fastest speed every time. To prevent the algorithm from falling into a local optimum, random exploration noise is added during training. Therefore, there is also a possibility of random exploration for the drone after 100 rounds. Therefore, the combat success rate reaches 99% when using this model. The following figure is the Reward curve of this algorithm.
Claims
1. A method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning, characterized in that: There are multiple homogeneous pursuit drones of the red side and a single escaping drone of the blue side in the combat area. The red drones cooperate to surround and capture the escaping target as soon as possible. The steps are as follows: Step 1: For the intelligent agents of both sides in the battle, the red side units are controlled using reinforcement learning algorithms, and the blue side units are based on traditional combat rules. The intelligent agent environment models of both sides are as follows: Let P n (n = 1, 2, …, N) represent multiple red pursuit drones, E represent the escape drone, v E represent the magnitude of the velocity of the escape drone, represent the magnitude of the velocity of the pursuit drone, d cap represent the pursuit radius, ψ E represent the yaw angle of the escape drone, represent the yaw angle of the pursuit drone, d t be the distance between the pursuit drone and the escape drone, d i be the distance between the pursuit drones; The intelligent agent model of the red side algorithm includes the kinematic equation of the pursuit drone, the state space, action space, and reward function of the intelligent agent; The intelligent agent model of the blue side rule is the escape countermeasure strategy adopted by the escaping drone; Step 2: Use the multi-agent deep deterministic policy gradient algorithm as the red side intelligent agent algorithm, where MADDPG uses the method of centralized training and decentralized execution; Construct a value critic network and a policy actor network, where: the value network critic is deployed on the global controller, and the policy network actor is deployed on each agent. During training, the agent i The observed value state i Transmitted to the global value network, the value network transmits the TD error back to the agent for the agent to train the policy network. At this time, there is no direct communication between the agents, but the trained policy network makes decisions; Use the MADDPG algorithm to train and optimize the red side intelligent agent; Step 3: Combine the intelligent agent environment model constructed in Step 1 and the multi-agent reinforcement learning algorithm in Step 2 to generate the final multi-drone cooperative surrounding method based on reinforcement learning. The process is as follows: Step 3-1: Based on the current intelligent agent, calculate the difference between the current intelligent agent and the other intelligent agents. The difference is: Longitude difference Latitude difference Height difference Distance difference Obtain the yaw angle of the current agent Input the joint state of the agent where Step 3-2: Input the joint state of the agents into the multi-agent reinforcement learning algorithm to obtain the next joint action wherein and execute the action in the three-dimensional simulation combat environment; Step 3-3: After the execution of the action, obtain the next action of the agent and the reward value R of the current action n , and store the data (S n , A n , S n+1 , R n ) into the experience buffer pool, and extract a batch of data to train the algorithm; During the entire surrounding process, Step 3 operations are looped.
2. The method for UAV cooperative pursuit of the multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 1, characterized in that: The successful encirclement satisfies the following conditions: 1) There exists any pursuit drone P n (n = 1, 2, …, N) whose distance from the escaping target E is less than the encirclement radius d cap ; 2) The encirclement angle between adjacent pursuit drones is not greater than π.
3. The UAV cooperative pursuit method for the multi-degree-of-freedom model of multi-agent reinforcement learning according to claim 1, characterized in that: The following constraints are satisfied during the surrounding process: 1) To avoid the influence of terrain and temperature on the drones, the flight altitude of the drones is restricted between 1000 meters and 3000 meters; 2) The pursuit drones need to capture the escaping drone within the defined area. If the escaping drone exceeds the defined area, the mission is judged as failed; 3) Collisions cannot occur between the pursuit drones.
4. The method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 1, characterized in that: The kinematic equation of the drone in the intelligent agent model of the red side algorithm is: where (x i , y i ) represents the current position of the UAV, and h i represents the current altitude of the UAV; respectively represent the track yaw angle and track pitch angle of UAV i in the nth period; the track yaw angle δ i and the track pitch angle ω i are restricted by constraints: -ω max < ω i < ω max , -δ max < δ i < δ max ; The state space of the intelligent agent is: Wherein: is the situation information of a single pursuit UAV at the simulation step of n; The action space of the intelligent agent is: In the formula: is the action taken by a single pursuit UAV at the simulation step n, where: The reward function is: The reward function design combines continuous rewards and sparse rewards. For the problem of multi-drone cooperative pursuit, two main elements are considered: First, the pursuit drones need to successfully capture the escaping drone. In the multi-drone pursuit scenario, as long as one drone captures the escaping drone, the mission is considered successful; Second, collisions cannot occur between the pursuit drones. The specific expression is as follows: R=r sparse +r step wherein: includes sparse reward r sparse and step reward r step .
5. The UAV cooperative pursuit method for the multi-degree-of-freedom model of multi-agent reinforcement learning according to claim 4, characterized in that: The situation information of the single pursuit UAV at the simulation step of n is as follows: Where: wherein: are respectively the relative longitude, relative latitude, and relative altitude between the pursuit UAV and the escaping UAV; and are respectively the track deflection angle and track tilt angle of the pursuit UAV; is the distance between the pursuit UAV and the escaping UAV.
6. The method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 4, characterized in that: The sparse reward r sparse and the step reward r step are as follows: Sparse reward r for the pursuit drone sparse It is divided into the following two modules: First, when a drone in the pursuit drone swarm successfully captures the escaping drone, a positive reward is given; second, when the escaping drone successfully escapes from the area, it is considered a mission failure and a negative reward is given. Each pursuit drone obtains a step reward r once per simulation step according to the executed action step , and the drone is guided to complete the established task through this reward; the step reward r step is composed of a weighted combination of multiple sub-rewards: r step = αr1 + βr2 + γr3 In the formula: r1 is the pursuit distance reward, r2 is the pursuit height difference reward, and r3 is the drone collision reward; α, β, and γ are weighting coefficients, and α + β + γ = 1.
7. The method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 6, characterized in that: The pursuit distance reward r1, the pursuit height difference reward r2, and the drone collision reward r3 are: r1 = -k(d t -d max ) where: d t is the relative distance between the UAVs, and d max is the maximum strike range of the pursuit UAV; set r1 as the negative reward function, when the distance between the pursuit UAV and the escaping UAV is the strike distance of the pursuit UAV, r1 = 0; r2 = -k(h i -h E ) When the height difference h between the pursuing drone and the escaping drone i -h E = 0, the height relationship between the pursuing drone and the escaping target is locally optimal; Establish a reward function \(r_3\) in negative exponential form to describe the collision risk between pursuit drones, \(d\) min represents the closest distance between the current drone and other drones.
8. The method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 1, characterized in that: The escape countermeasure strategy adopted by the escaping drone is: When surrounded by the pursuit drones, the escaping drone escapes towards the midpoint of the longest distance among the midpoints of all side lengths of the polygon formed by the pursuit drones; When not surrounded by the escaping drones, the idea of the artificial potential field method is adopted. Assume that the pursuit drones exert a repulsive force in the vector direction towards the escaping drone, and the repulsive force component between the two is an inverse function relationship with the distance between the two: as the distance increases, the repulsive force decreases; The escaping drone escapes in the direction of the repulsive force after synthesizing the repulsive force vectors given by all the pursuit drones.
9. The method for collaborative pursuit of drones with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 1, characterized in that: The Actor network structure in the MADDPG algorithm is:
10. The method for collaborative pursuit of unmanned aerial vehicles with a multi-degree-of-freedom model based on multi-agent reinforcement learning according to claim 1, characterized in that: The Critic network structure in the MADDPG algorithm:
Citation Information
Patent Citations
Target tracking and hunting method for unmanned aerial vehicle group adaptive environment
CN113268078A
PER-IDQN-based multi-unmanned aerial vehicle hunting tactical method
CN114815891A