Unmanned aerial vehicle target tracking method based on game theory
Through the unmanned aerial vehicle target tracking method based on game theory, the problem of collaborative tracking and obstacle avoidance of multiple drone systems in complex tasks is solved, and the rapid and intelligent target tracking and strike capabilities are achieved.
Patent Information
- Application Number
- CN202510442621.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-09
AI Technical Summary
A single drone has limited load and is difficult to complete complex and cross-distance tasks independently. Multi-drone systems lack intelligence requirements when tracking and avoiding obstacles, especially in future air combat, they lack synergistic capabilities in tracking and attacking non-cooperative drones.
The unmanned aerial vehicle target tracking method based on game theory is adopted to establish a mathematical model of the drone, define the game equilibrium point, design reward functions, and use the strategy network and value network of the Soft Actor-Critic architecture to realize intelligent decision-making and collaborative tracking of multi-unmanned aerial vehicles.
The pursuit ability of multi-UAV systems in complex scenarios has been improved, and the pursuit ability of rapid and coordinated target tracking and obstacle avoidance has been achieved, which has enhanced the anti-tracking and strike capabilities in future air combat.
Smart Images

Figure CN120295364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-end equipment manufacturing, especially the technical field of unmanned aerial vehicles, and specifically discloses a method for tracking targets of unmanned aerial vehicles based on game theory. Background Art
[0002] Unmanned aerial vehicles, also known as drones (UAV, Unmanned Aerial Vehicle), have made great development in recent years. In the civilian field, due to their diverse functions, they are widely used in agriculture, transportation, exploration and other fields. In the military aspect, drones have the advantages of being unmanned and having low costs, and can perform various more dangerous flight tasks, so they have become the focus of development for various countries. Due to the limited load of a single drone, its functions are restricted, and it is difficult to independently complete tasks with complex content and long duration. Therefore, the idea of a multi-drone system or a drone swarm came into being.
[0003] A swarm is a biological swarm similar to a bee swarm or an ant swarm. Individuals can communicate with each other and share information, so as to cooperate to complete various tasks. With the advantages of its large scale and mutual communication, the swarm has the characteristics of high intelligence, fierce attack, strong survivability and flexible use, making it a powerful weapon that can change future wars.
[0004] In future air combat, due to the increasing intelligence of unmanned systems, non-cooperative drones will have stronger anti-tracking capabilities. Multi-drone systems rely on their large quantity and strong coordination, and have greater advantages than single drones in scenarios of tracking and attacking non-cooperative drones. In actual situations, while a drone swarm is in a game with enemy aircraft, it needs to avoid obstacles in real time to avoid collisions, which requires the swarm to have a high enough intelligence. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a method for tracking targets of unmanned aerial vehicles based on game theory, including the following steps:
[0006] S1. Establish a mathematical model of the drone;
[0007] S2. Use the mathematical model of the drone in S1 to define the game equilibrium point;
[0008] S3. Based on the definition in S2, define the reward function R;
[0009] S4. Design a policy network and a value network based on the Soft Actor-Critic architecture;
[0010] S5. Deploy an agent network on the drone platform, and the drone extracts the state information s according to the environment and its own state tAs the input of the policy network, generate the UAV action decision a at each moment t , and control the UAV to track the target.
[0011] Furthermore, S1 specifically includes:
[0012] Assume the UAV has a two-dimensional motion model, and define the UAV state X including two-dimensional position and heading angle: Define the control quantity u as the speed in two dimensions: u = [ω, v], where ω is the speed in the x direction and v is the speed in the y direction; the change rate of the UAV state X is expressed as: where, is the yaw angle of the UAV, define the line-of-sight angle θ as the angle between the UAV velocity direction and the line connecting the tracker and the escapee, the relative position vector is: r = [d, θ], where d is the distance between the tracker and the escapee, define the angle to obtain the line-of-sight angle θ: Project the velocity vectors of the UAV and the target onto the line connecting the two to obtain the change rate of the relative position vector r: In the simulation, X and r are updated as:
[0013]
[0014] where t is time, the action space of the UAV is a continuous two-dimensional space, and the tracker and the escapee have the same action space: A i = [ω i , v i , the remaining energy and speed of the UAV are used as state inputs, and the state space of the UAV is:
[0015] Furthermore, S2 specifically includes:
[0016] Define the game equilibrium point as for any strategy π of the game player i i all satisfy:
[0017]
[0018] where J is the payoff function for a certain tracker or escapee, representing the cost paid during the pursuit and escape process. The UAV target tracking task is a zero-sum game for both the tracker and the escapee, that is:
[0019]
[0020] Simplify it to a two-person zero-sum game, and the equilibrium point is expressed as:
[0021]
[0022] The payment function in reinforcement learning is expressed as:
[0023]
[0024] Furthermore, in S3, the reward function R includes a tracker reward function and an escapee reward function. Both the tracker reward function and the escapee reward function include a distance reward function, an energy reward function, a time reward function, an action change rate reward function, and a boundary reward function. Among them,
[0025] The distance reward function is:
[0026]
[0027] where D is the boundary between continuous rewards and final rewards; when the distance between the drone and the target is greater than D, a reward that gradually increases as the distance decreases is given; when the distance between the drone and the target is less than D, it is initially determined that the target has been caught up with and the reward is doubled; β is a constant greater than 0 to prevent the reward function from tending to infinity when d approaches 0 and to make the reward function satisfy the bounded condition; W distance is the weight coefficient of the distance reward function;
[0028] The energy reward function is:
[0029]
[0030] where W energy is a positive constant; the threshold E threshold satisfies 0 < E threshold < 1, and when the remaining energy is less than the threshold E threshold a penalty is given;
[0031] The time reward function is:
[0032]
[0033] where W time is a positive constant; when the distance between the drone and the target is less than the threshold D threshold it is considered that the target has been basically caught up with and no time penalty is given; otherwise, it is considered that the target has not been caught up with and a time penalty is given;
[0034] The action change rate reward function is:
[0035] R uav_action = -W action (|v t-1 - v t | + |ω t-1 - ω t |) / dt,
[0036] where Waction is a positive constant;
[0037] The boundary reward function is:
[0038]
[0039] where, W boarder is a positive constant; d boarder is the shortest distance between the UAV and the boundary; ε is a small quantity used to prevent the reward function from tending to infinity. In a square area, when the UAV or the target goes out of the square area, a penalty is given.
[0040] Furthermore, S3 also includes that the reward for a single step of a single UAV is:
[0041] R uav = R uav_d + R uav_energy + R uav_time + R uav_action + R uav_boarder ,
[0042] The sum of the payoff functions of both sides in the zero-sum game is 0, that is:
[0043]
[0044] The payoff function of the UAV is equal to its long-term cumulative reward, that is:
[0045]
[0046] When a tracker is detected, the game starts, and the reward function is:
[0047]
[0048] where, d min is the shortest distance between the tracker closest to the escapee and the escapee;
[0049] The single-step reward of the escapee is:
[0050] R target = R target_pedg + R target_action .
[0051] Furthermore, in S4, both the policy network and the value network adopt a multi-layer neural fully connected network to realize the mapping relationship from the state to the action or the cumulative return value, and an intelligent agent network is constructed, which includes a policy network, two softq networks, and two target networks.
[0052] The present invention exploits the advantage of multiple unmanned aerial vehicles (UAVs) pursuing a single UAV, and combines complex scenarios, multiple targets, multiple constraints, etc. to achieve the application of multiple UAVs quickly pursuing a non-cooperative single UAV. The present invention selects the soft actor-critic algorithm (SAC) as the deep reinforcement learning algorithm for the AC architecture, and writes the entropy of the policy into the value function, enhancing the exploration of the policy and having better performance in many tasks. Description of the Drawings
[0053] Figure 1 It is a definition diagram of the UAV coordinate system of the present invention;
[0054] Figure 2 It is a distance reward function diagram;
[0055] Figure 3 It is a measurement network architecture diagram;
[0056] Figure 4 It is a value network architecture diagram;
[0057] Figure 5 It is a soft actor-critic algorithm framework diagram;
[0058] Figure 6 It is a parameter update flowchart;
[0059] Figure 7 It is a single-agent network structure diagram. Detailed Implementation Manner
[0060] The technical solution of the present invention will be further described below through the drawings and embodiments. Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs.
[0061] In a specific embodiment of the present invention, the following steps are included:
[0062] S1. Establish a UAV mathematical model, and simplify the UAV into a two-dimensional motion model through assumptions;
[0063] Define that the UAV state includes three dimensions of two-dimensional position and heading angle:
[0064]
[0065] Define the control quantity as the speed in two dimensions;
[0066] u = [ω, v]
[0067] The change rate of the UAV state can be expressed as:
[0068]
[0069] Among them is the yaw angle of the UAV, which is the angle between the orientation (velocity direction) and the north direction, with the clockwise direction being positive, as Figure 1 shown.
[0070] Next, define the relative position vector. First, define the line-of-sight angle θ as the angle between the UAV velocity direction and the line connecting the pursuer and the evader. The relative position vector is as follows:
[0071] r = [d, θ]
[0072] Define the angle The line-of-sight angle θ can be obtained as:
[0073]
[0074] Project the velocity vectors of the UAV (unmanned aerial vehicle) and the target onto the line connecting the two, and the rate of change of the relative position can be obtained:
[0075]
[0076] So far, the kinematic model between two UAVs has been obtained, and this model can be used pairwise in a multi-UAV system.
[0077] The positions and relative positions in the simulation are updated as follows:
[0078]
[0079] Among them, the relative position update can use a simpler method, that is, first update the UAV position, and then directly calculate the relative position relationship at the new moment through the geometric relationship based on the position at the new moment. Therefore, the control quantity of the UAV is u = [ω, v], and the actions are continuous. So, the action space of the UAV is a continuous two-dimensional space, and the pursuer and the evader have the same action space:
[0080] A i = [ω i , v i
[0081] The remaining energy and velocity of the UAV are used as state inputs, enabling the UAV to have a more reasonable plan for subsequent actions. Therefore, the state space of the UAV is:
[0082]
[0083] S2. Using the mathematical model in step S1, define the game equilibrium point as: for any strategy π of the game player i i , it satisfies:
[0084]
[0085] The right side of the inequality is the payoff function of all players under the optimal strategy. At the same time, this task is a zero-sum game, so we have:
[0086]
[0087] To simplify this problem, considering the pursuer as an individual, it can be simplified to a two-player zero-sum game, and its equilibrium point can be expressed as:
[0088]
[0089] In reinforcement learning, the reward function is expressed as:
[0090]
[0091] S3. Based on the definition of S2, design the reward function R, including the pursuer reward function and the evader reward function. The reward function includes distance reward, energy reward, time reward, action change rate reward, and boundary reward, etc. The main ones are as follows:
[0092] (1) Distance reward
[0093] The pursuer is required to catch up with the target. Therefore, first, a reward for catching up with the target needs to be designed. To avoid the convergence difficulty caused by sparse rewards, a continuous reward designed according to the distance between the pursuer and the target is designed. The closer the distance to the target, the greater the reward:
[0094]
[0095] Among them, D is the boundary between the continuous reward and the final reward. When the distance between the UAV and the target is greater than D, a reward that gradually increases as the distance decreases is given. When the distance is less than D, it is initially determined that the target has been caught up, and the reward is doubled. β is in the denominator to prevent the reward function from tending to infinity when d approaches 0, so that the reward function meets the bounded condition; W distance is the weight coefficient of the distance reward function. The curve of the reward function is as Figure 2 shown, where, W distance = 0.6, β = 0.05, D = 0.05.
[0096] (2) Energy reward function
[0097] In order to minimize the energy consumption during tracking as much as possible, it is hoped that the remaining energy can be higher than a certain threshold E threshold (0 < E threshold < 1), when the remaining energy is less than the threshold, a penalty is given:
[0098]
[0099] where W energy is a positive constant.
[0100] (3) Time reward function
[0101] To enable the UAV to catch up with the target as quickly as possible, a time reward function is designed. When the distance between the UAV and the target is less than a threshold D threshold , it is considered that the UAV has basically caught up and no time penalty is given. Otherwise, it is considered that the UAV has not caught up yet, and a time penalty is given:
[0102]
[0103] where W time is a positive constant.
[0104] (4) Action change rate reward function
[0105] In the simulation, the energy consumption is calculated only considering the current speed magnitude, without considering the energy consumption during acceleration and braking. Therefore, an action change rate reward is added here, which can not only consider the energy consumption during acceleration and deceleration, but also make the UAV's action change as smooth as possible, with a smoother trajectory:
[0106] R uav_action = -W action (|v t-1 - v t | + |ω t-1 - ω t |) / dt
[0107] where W action is a positive constant.
[0108] (5) Boundary reward function
[0109] This task is designed in a square area. When the UAV or the target exits this area, a penalty is given:
[0110]
[0111] where W boarder is a positive constant, d boarder is the closest distance between the UAV and the boundary; ε is a small quantity to prevent the reward function from tending to infinity.
[0112] In summary, the single-step reward of a single UAV is:
[0113] R uav = R uav_d + R uav_energy + R uav_time + R uav_action + R uav_boarder
[0114] The reward of the evader mainly depends on the pursuer. The cumulative long-term reward is the quantity that the agent needs to maximize. In differential games, the cumulative long-term reward is equivalent to the payoff function. In zero-sum games, the sum of the payoff functions of both sides is 0, that is:
[0115]
[0116] And the payoff function of the UAV is equal to its long-term cumulative reward, that is:
[0117]
[0118] Therefore, the payoff function of the evader is the opposite of the total long-term cumulative reward of the UAV. The single-step reward function of the evader (the PEDG part) is equal to the opposite of the sum of the single-step rewards of the pursuers. For the evader, when a pursuer is detected, the game starts. So the reward function is:
[0119]
[0120] where d min is the distance to the nearest pursuer among all pursuers and the evader.
[0121] At the same time, the evader also needs to consider the amplitude of its own actions. So the single-step reward of the evader is:
[0122] R target = R target_pedg + R target_action .
[0123] S4. Design the policy network and value network based on the Soft Actor-Critic (SAC) architecture. Both use multi-layer neural fully connected networks to implement the mapping relationship from state to action or cumulative return value, as shown in Figure 3 、 Figure 4 respectively.
[0124] Each UAV agent has two softq-networks, and a corresponding target-network is set for each softq-network. The relationship between these softq-networks is as shown in Figure 5 respectively.
[0125] Input the joint state and joint action into the softq-network to obtain a predicted Q value. The idea of updating the softq-network is to use the TD error as the loss function and perform backpropagation to update the network parameters. To ensure stable training, the Q value at the next moment in the TD error is predicted by the target-network. Input the joint action and state at the next moment into two target-softq-networks, and use the smaller predicted value to calculate the target Q value (target Q). Then, take the difference between the predicted Q value and the target Q value to obtain the loss function, and perform backpropagation to complete one update of the softq-network parameters. After a certain number of steps, assign the parameters of the softq-network to the target-network to complete the update of the target network parameters.
[0126] According to the loss function of the policy network:
[0127]
[0128] After improvement using the reparameterization trick, the derivative of the expectation with respect to the policy parameters can be directly obtained, resulting in a new loss function:
[0129]
[0130] The process of updating the policy network is as Figure 6 shown.
[0131] Input the current state into the policy network, which outputs an action and a log probability. This action is the one processed by the reparameterization trick, i.e., f φ (ε; s t ), and the log probability corresponds to logπ φ (f φ (ε; s t )|s t ). Input this action into two Q networks, take the smaller predicted value and substitute it into the loss function of the policy network, and perform backpropagation to update the parameters.
[0132] In summary, each agent has five neural networks, including one policy network, two softq networks, and their corresponding two target networks. The network structure is as Figure 7 shown.
[0133] S5. Deploy the trained agent network on the UAV platform. The UAV extracts the state information s according to the environment and its own state tAs the input of the policy network, generate the drone action decision a at each moment t , and control the drone actuator to fly to achieve the tracking of the target.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for tracking an unmanned aerial vehicle target based on game theory, characterized in that It includes the following steps: S1. Establish a mathematical model of the drone; S2. Define the game equilibrium point by using the mathematical model of the drone in S1; S3. Define the reward function R based on the definition in S2; S4. Design a policy network and a value network based on the Soft Actor-Critic architecture; S5. Deploy an agent network on the UAV platform. The UAV extracts the state information s based on its own state and the environment t as the input of the policy network to generate the UAV action decision a at each moment t , and control the UAV to track the target.
2. The method for tracking an unmanned aerial vehicle target based on game theory according to claim 1, characterized in that, S1 specifically includes: Assume the unmanned aerial vehicle (UAV) has a two-dimensional motion model, and define the UAV state X to include two-dimensional position and heading angle: Define the control input u as the velocities in two dimensions: u = [ω, v], where ω is the velocity in the x-direction and v is the velocity in the y-direction; the rate of change of the UAV state X is expressed as: where, is the yaw angle of the UAV. Define the line-of-sight angle θ as the angle between the UAV velocity direction and the line connecting the tracker and the evader. The relative position vector is: r = [d, θ], where d is the distance between the tracker and the evader. Define the angle to obtain the line-of-sight angle θ: Project the velocity vectors of the UAV and the target onto the line connecting the two to obtain the rate of change of the relative position vector r: In the simulation, X and r are updated as: where \(t\) is time, the action space of the UAV is a continuous two-dimensional space, and the pursuer and the evader have the same action space: \(\mathcal{A}\) i =\([\omega i , v i \), the remaining energy and speed of the UAV are used as state inputs, and the state space of the UAV is:
3. The method for tracking an unmanned aerial vehicle target based on game theory according to claim 2, wherein, S2 specifically includes: Define the game equilibrium point as any strategy π of the game player i i which satisfies: Where J is the payoff function for a certain tracker or escapee, representing the cost incurred during the pursuit and escape process. The drone target tracking task is a zero-sum game for both the tracker and the escapee, that is: Simplified to a two-player zero-sum game, the equilibrium point is expressed as: The payoff function in reinforcement learning is expressed as:
4. A method for tracking an unmanned aerial vehicle target based on game theory according to claim 3, characterized in that In S3, the reward function R includes a tracker reward function and an escapee reward function. Both the tracker reward function and the escapee reward function include a distance reward function, an energy reward function, a time reward function, an action change rate reward function, and a boundary reward function. Among them, The distance reward function is: Among them, D is the demarcation between the continuous reward and the final reward; when the distance between the UAV and the target is greater than D, a reward that gradually increases as the distance decreases is given; when the distance between the UAV and the target is less than D, it is initially determined that the target has been caught up with and the reward is doubled; β is a constant greater than 0, in order to prevent the reward function from tending to infinity when d approaches 0 and make the reward function satisfy the bounded condition; W distance is the weight coefficient of the distance reward function; The energy reward function is: Among them, W energy is a positive constant; the threshold E threshold satisfies 0 < E threshold < 1, and when the remaining energy is less than the threshold E threshold a penalty is given. The time reward function is: Among them, W time is a positive constant; when the distance between the UAV and the target is less than the threshold D threshold it is considered that the UAV has basically caught up and no time penalty is given; otherwise, it is considered that the UAV has not caught up and a time penalty is given. The action change rate reward function is: R uav_action = -W action (|v t-1 -v t | + |ω t-1 -ω t |) / dt, where W action is a positive constant; The boundary reward function is: Among them, W boarder is a positive constant; d boarder is the shortest distance between the UAV and the boundary; ε is a small quantity used to prevent the reward function from tending to infinity. In a square area, when the UAV or the target goes out of the square area, a penalty is given.
5. A method for tracking an unmanned aerial vehicle target based on game theory according to claim 4, characterized in that S3 also includes that the reward for a single step of a single drone is: R uav = R uav_d + R uav_energy + R uav_time + R uav_action + R uavbo_boarder , The sum of the payoff functions of both sides in the zero-sum game is 0, that is: The payoff function of the drone is equal to its long-term cumulative reward, that is: When a tracker is detected, the game starts, and the reward function is: where d min is the shortest distance among all trackers to the escapee; The single-step reward of the escapee is: R target = R target_pedg + R target_action .
6. A method for tracking an unmanned aerial vehicle target based on game theory according to claim 5, characterized in that In S4, both the policy network and the value network use a multi-layer neural fully connected network to implement the mapping relationship from state to action or cumulative return value, and construct an agent network, which includes a policy network, two softq networks, and two target networks.
Citation Information
Patent Citations
Unmanned aerial vehicle one-to-one pursuit game method based on M2GPI
CN116796844A
ME-DDPG-based many-for-one chasing game method for unmanned aerial vehicles
CN116976442A
Cross-domain heterogeneous unmanned cluster game confrontation strategy generation method and system
CN118885000A
Cluster collaborative hunting method based on reinforcement learning
CN119687727A
Model-based reinforcement learning
US20240320505A1