Partial observable unmanned aerial vehicle cluster collaborative pursuit game decision-making method, system and equipment
Through the combination of deep reinforcement learning and stable matching theory, the problem of long decision-making cycle and insufficient reliability of drone clusters in complex environments is solved, and efficient pursuit and escape decision-making of drone clusters in complex environments is achieved.
Patent Information
- Application Number
- CN202510273789.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
AI Technical Summary
In complex environments or obstacles, traditional reinforcement learning methods used for drone cluster pursuit strategies face the problems of long decision-making cycles and insufficient strategy reliability, and it is difficult to complete the pursuit task efficiently.
Deep reinforcement learning combined with stable matching theory is adopted to obtain local observation information of drones, build global observations, and use comprehensive preference index and multiple reward functions to generate strategies for drones to realize effective screening and fusion of perceived information among drones.
It significantly enhances the autonomous group intelligence perception ability of the drone cluster in the battlefield environment, and improves the efficiency of the pursuit mission and the effectiveness and adaptability of the strategy.
Smart Images

Figure CN120143874A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of UAV decision-making, and particularly relates to a partially observable UAV swarm collaborative pursuit-evasion game decision-making method, system, and device. Background Art
[0002] The intelligent pursuit-evasion of UAV swarms is a new and frequently applied situation that utilizes the high mobility, rapid deployment, and collaborative operation capabilities of UAVs. With the increasing maturity and popularization of unmanned system technologies (such as unmanned ships, GPS-assisted autonomous vehicles, etc.) and the continuous evolution and innovation of frontier algorithms such as deep reinforcement learning in the field of artificial intelligence, how to efficiently formulate and execute high-quality strategies has become the focus of current research and demonstrated extensive application potential in multiple fields. Under complex environments or conditions with obstacles, the pursuit-evasion strategies of UAV swarms trained using traditional reinforcement learning methods face problems such as long decision-making cycles and insufficient policy reliability, making it difficult to efficiently complete the pursuit-evasion tasks. To address this challenge, deep reinforcement learning emerged as an alternative. It combines the essence of deep learning and reinforcement learning and shows great potential in improving the performance of UAVs in pursuit-evasion tasks. Policy formulation plays a crucial role in intelligent pursuit-evasion tasks, and its quality and acquisition speed are directly related to the efficiency of task execution. However, it is not easy to achieve efficient pursuit-evasion decision-making using deep reinforcement learning because the actual pursuit-evasion scenarios usually involve the collaboration of multiple UAV swarms, and the environment faced by each UAV changes dynamically. Coupled with the limitations of observation information, global state perception becomes a difficult problem. In addition, the rapid generation and high-quality requirements of policies pose another major challenge, that is, how to ensure the effectiveness and adaptability of policies within a limited time.
[0003] There is no denying that there has been a lot of research on the pursuit and evasion decision-making of UAV swarms, which can generally be divided into three categories. (1) Methods based on mathematical models and bionic algorithms, which is a research approach that combines mathematical theories with the biological behavior mechanisms in nature to innovatively solve complex problems or optimize system performance. For example, the Voronoi diagram and Apollonius circle theory methods are used to solve the problem of UAV swarm cooperative pursuit. This method has strong theoretical support and the advantages of bionic innovation, and is particularly suitable for solving complex optimization problems, simulating natural phenomena, and providing decision support in uncertain environments. However, when dealing with complex systems, this method may face challenges such as high computational complexity, difficult parameter adjustment, and difficulty in capturing the dynamic changes in the real world. (2) Methods based on differential game theory and deterministic methods, which model the multi-agent pursuit-evasion problem and build a strategy selection model for both the pursuer and the evader to obtain a real-time strategy selection algorithm. Although the deterministic method has a strict derivation process, it is difficult to achieve good application results in the complex and highly dynamic scenarios of UAV swarm confrontation. (3) Methods based on reinforcement learning and intelligent optimization algorithms, which are important technical paradigms for realizing intelligent decision-making in air combat. This method can adaptively learn and improve strategies, but at the same time faces challenges such as high computational costs, low sample efficiency, and difficulty in dealing with continuous high-dimensional space problems. In addition, existing research mainly adopts the one-to-many confrontation scenario, while paying less attention to the more common many-to-many complex confrontation situation in reality. Summary of the Invention
[0004] The purpose of the present invention is to provide a partially observable UAV swarm cooperative pursuit and evasion game decision-making method, system, and device, which can realize the task allocation and autonomous coordination of both the pursuer and the evader, effectively screen and fuse the perception information among UAVs, and significantly enhance the autonomous swarm intelligence perception ability of the swarm for the battlefield environment;
[0005] To achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A partially observable UAV swarm cooperative pursuit and evasion game decision-making method includes the following steps:
[0007] Obtain the local observation information of the UAV to be decided.
[0008] According to the local observation information, obtain the global observation value within the UAV swarm to be decided.
[0009] Communicate and transmit the motion information in the local observation information to the stable matching model to obtain the target information for the UAV to pursue or evade preferentially.
[0010] According to the global observation value and the target information for the UAV to pursue or evade preferentially, construct a strategy model and generate the strategy for the UAV for the preferential pursuit or evasion target.
[0011] Preferably, the motion information in the local observation information is communicated and transmitted to the stable matching model to obtain the target information for the UAV to preferentially pursue or evade, including the following steps:
[0012] According to the motion information in the local observation information, the situation of the UAV is evaluated using the comprehensive preference index, and potential pursuit-evasion pairs are screened out from the pursuing UAV cluster and the evading UAV cluster, where the comprehensive preference index is the linear weighted sum of the angle preference index, the distance preference index, and the speed preference index.
[0013] Preferably, the motion information in the local observation information is communicated and transmitted to the stable matching model to obtain the target information for the UAV to preferentially pursue or evade, and the following steps are also included:
[0014] Traverse all possible combinations of the pursuing and evading parties, and evaluate the initial interactive situation between each pair;
[0015] Through multiple iterative calculations, determine the optimal game strategy at the first time step through the comprehensive preference index and the preset UAV objective function;
[0016] Before the start of each new time step, use the stable strategy obtained in the previous time step as the initial strategy for the current time step;
[0017] Iterate for each pursuing UAV to obtain the optimal game target of the current pursuing UAV;
[0018] Based on the above iterative results, update the strategy matrix of the entire pursuit-evasion game;
[0019] According to the obtained pursuit-evasion game strategy matrix, establish a targeted pursuit-evasion relationship between the pursuer and the evader.
[0020] Preferably, the strategy model is the Actor model. According to the global observation value and the target information for the UAV to preferentially pursue or evade, the Actor model generates the strategy of the UAV for the preferential pursuit or evasion target. After the UAV executes the decision, the current observation value is obtained and stored in the buffer pool. The Actor model is updated with the maximization of the UAV objective function as the guide, and the Critic model is used to evaluate the observation values in the buffer pool to guide the update iteration of the Actor model.
[0021] Preferably, the UAV objective function can be expressed as:
[0022]
[0023] where, π i represents the strategy adopted by the i-th UAV, o i represents the local observation information of the UAV, ai Denotes the strategy adopted by UAV i.
[0024] Preferably, the update of the Actor model involves the comprehensive reward function of the UAV, where the comprehensive reward function of the UAV is a linear weighted sum of the capture reward function and the guidance reward function, and the capture reward function is:
[0025]
[0026] The guidance reward function is:
[0027]
[0028] where d ij Denotes the distance between the pursuing UAV i and the escaping UAV j, and r i Represents the action radius of the capture action of the pursuing UAV, and r j Represents the action radius of the evasion action of the escaping UAV, Represents the observation radius of the pursuing UAV, Represents the observation radius of the escaping UAV, H represents the pursuing UAV cluster, and E represents the escaping UAV cluster.
[0029] Preferably, it further includes the following steps:
[0030] Based on the kinematic model, realize the real-time update of its own position and speed;
[0031] Judge whether the pursuit and escape are over. If not, repeat the steps and continue to loop; otherwise, stop the loop.
[0032] The present invention also provides a partially observable UAV cluster cooperative pursuit and escape game decision-making system, including:
[0033] An acquisition module for acquiring the local observation information of the UAV to be decision-making;
[0034] A processing module for obtaining the global observation value within the UAV cluster to be decision-making according to the local observation information;
[0035] A pursuit and escape game strategy formulation module. The processing module communicates and transmits the motion information in the local observation information to the stable matching model of the pursuit and escape game strategy formulation module to obtain the target information of the UAV to preferentially pursue or preferentially escape, and then returns the target information of the UAV to preferentially pursue or preferentially escape to the processing module;
[0036] A strategy module. The processing module sends the global observation value and the target information of the UAV to preferentially pursue or preferentially escape to the strategy module, constructs a strategy model, and generates the strategy of the UAV for the target of preferentially pursuing or preferentially escaping.
[0037] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the steps in the above-mentioned partially observable UAV swarm cooperative pursuit-evasion game decision-making method.
[0038] The present invention also provides a partially observable UAV swarm cooperative pursuit-evasion game decision-making device, including:
[0039] A memory for storing software application programs,
[0040] A processor for executing the software application programs, and each program of the software application programs correspondingly executes the steps in the above-mentioned partially observable UAV swarm cooperative pursuit-evasion game decision-making method.
[0041] The present invention utilizes a lightweight explicit communication mechanism to share the UAV individual state information and local perception information within the UAV swarm, realizing the effective screening and fusion of the perception information among UAVs, enabling each UAV to make optimal decisions based on the current environmental state and the behaviors of teammates and opponents. In addition, a UAV confrontation game strategy based on the stable matching theory is introduced, effectively simplifying the cluster confrontation decision-making process, providing a flexible but highly coordinated task allocation mode for UAVs, and further enhancing the robustness and reaction speed of the entire system. Through multiple reward functions including guiding rewards and capture rewards, the pursuit-evasion efficiency of the UAV swarm is further optimized. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the pursuit-evasion game.
[0043] Figure 2 It is a schematic diagram of the algorithm flow of the present invention.
[0044] Figure 3 It is a schematic diagram of the communication model flow.
[0045] Figure 4 It is a schematic diagram of the stable game model flow.
[0046] Figure 5 It is a complete capture trajectory diagram of 2-on-2 UAVs for the algorithm of the present invention.
[0047] Figure 6 It is a complete capture trajectory diagram of 5-on-5 UAVs for the algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments.
[0049] The collaborative pursuit-evasion game decision-making of the proposed unmanned aerial vehicle (UAV) swarm in the present invention is about the pursuit-evasion confrontation situation among UAV swarms within a limited combat airspace. The tactical objectives of the two opposing UAV swarms are opposite, that is, the pursuing UAVs aim to capture the escaping UAVs, while the escaping UAVs aim to avoid or stay away from the pursuing UAVs. In the two-dimensional plane area of the pursuit-evasion game between the two sides, as Figure 1 shown, H represents the pursuing UAV swarm, E represents the escaping UAV swarm, v i represents the speed of the pursuing UAV, v j represents the speed of the escaping UAV, α i represents the heading angle of the pursuing UAV, α j represents the heading angle of the escaping UAV, r i represents the action radius of the capture action of the pursuing UAV, r j represents the action radius of the evasion action of the escaping UAV, represents the observation radius of the pursuing UAV, represents the observation radius of the escaping UAV.
[0050] β i is the included angle - line-of-sight angle of the line of sight (LOS) of UAV i, and the line of sight refers to the ray from the pursuing UAV h with a pursuit-evasion relationship to the escaping UAV e it needs to pursue.
[0051] The pursuing UAVs are action-oriented to maximize the number of captured targets, while the escaping UAVs are action-oriented to increase the distance from the pursuing UAVs.
[0052] Within a specified period, the pursuing UAV swarm aims to maximize the number of captured opponent UAVs, that is:
[0053] max G h = f(v i , β i , d ij , v j , β j )
[0054] The escaping UAV swarm aims to minimize the number of captured UAVs, that is:
[0055] min G e = f(v i , β i , d ij , v j , β j )
[0056] Among them, d ij represents the distance between the pursuing UAV i and the escaping UAV j.
[0057] A method for collaborative pursuit-evasion game decision-making of a partially observable unmanned aerial vehicle (UAV) swarm, comprising the following steps:
[0058] S1. Obtain the local observation information of the UAV to be decided;
[0059] S2. Obtain the global observation value within the UAV swarm to be decided according to the local observation information;
[0060] Specifically, as Figure 2 shown, the algorithm of the present invention consists of a communication mechanism, a UAV swarm game strategy, and an independent Actor-Critic network framework. Among them, the communication mechanism is: at time t, first collect the local observation information of each UAV to be decided i.e., the state of the UAV action (a 1 , a 2 ,..., a N ), reward (r 1 , r 2 ,..., r N ) and the expected subsequent state
[0061]
[0062] After the above information is fused, screened, and calculated, the global observation value within the UAV swarm to be decided is obtained:
[0063]
[0064] And, as Figure 3 shown, the key motion parameters of each UAV at time t are screened from the local observation information:
[0065]
[0066] Specifically, it includes the position p t , speed v t and angle α t information. This information is used to help the UAV determine which targets should be the priority pursuit objects, so as to formulate an effective UAV pursuit-evasion game strategy, that is, to determine the targets that the UAV gives priority to pursue or evade;
[0067] S3. Communicate and transmit the motion information in the local observation information to the stable matching model to obtain the target information of the UAV to give priority to pursue or evade;
[0068] Specifically, communicate and transmit the key motion parameters of each UAV to the stable matching model, and use the comprehensive preference index to evaluate the UAV situation, as Figure 4As shown, potential pursuer-evader pairs are screened from the pursuing UAV cluster and the evading UAV cluster, where the comprehensive preference index includes the angle preference index distance preference index and speed preference index
[0069]
[0070] where β ij is the included angle-line of sight angle of the UAV, that is, the included angle between the target line of sight of the UAV and the UAV heading;
[0071] The comprehensive preference index is the linear weighted sum of various preference indices:
[0072]
[0073] In the formula, ω 1 , ω 2 and ω 3 are the weight values of the angle, distance and speed preference indices respectively, and satisfy
[0074] More specifically, in the pursuer-evader game process, the pursuer-evader game pairs will change in real time according to the preference index during the game process. The present invention divides the entire game process into many discrete time steps, and each time step contains several key stages to ensure the real-time optimization and effective execution of the strategy. The specific steps are as follows:
[0075] (1) Initialization stage (the first time step): Initially, there is no pre-set pursuer-evader pair, that is, the set of pursuer-evader game pairs is empty. Subsequently, based on the comprehensive preference index, potential pursuer-evader pairs are screened from the pursuing UAV set H and the evading UAV set E.
[0076] S311. The initial game strategy is empty:
[0077] S312. Blocking pair traversal: First, traverse all possible combinations of pursuer and evader (h i , e j ), and evaluate the initial interactive situation (distance, speed, angle) between each pair;
[0078] S313. Strategy iterative optimization: Through multiple iterative calculations, the optimal game strategy for the first time step is finally determined through the comprehensive preference index and the preset UAV objective function
[0079] (2) Strategy iteration stage (subsequent time steps):
[0080] S321. Policy inheritance and update: Before the start of each new time step, the stable policy obtained in the previous time step is used as the initial policy for the current time step.
[0081] S322. Policy iteration for pursuer UAVs: Iterate for each pursuer UAV p i to obtain the optimal game objective e i of the current pursuer p j ;
[0082] S323. Policy matrix adjustment: Based on the above iteration results, update the policy matrix of the entire pursuit-evasion game to reflect the latest tactical layout;
[0083] (3) Targeted pursuit-evasion establishment stage:
[0084] S33. Establishment of pursuit-evasion relationship: Based on the obtained pursuit-evasion game policy matrix in each time slot establish the targeted pursuit-evasion relationship between the pursuer and the evader.
[0085] When the algorithm converges, that is, when a stable pursuit-evasion game policy is found, output the pursuit-evasion game policy matrix
[0086] S4. Generate the policy of the UAV for the target of preferential pursuit or preferential evasion according to the global observation value and the target information of the UAV for preferential pursuit or preferential evasion.
[0087] Specifically, the state, action, and observed environmental information of each UAV change continuously over time. The continuous-domain partially observable Markov decision process (POMDP) is used to construct the UAV objective function, denoted as <S, A, O, P, R i , γ>, and its specific meanings are as follows:
[0088] S represents the set of states in the environment. The state refers to the information useful for decision-making that the agent can obtain. In reinforcement learning, the state of the environment needs to be artificially abstracted and selected to extract those signals that are valuable to the agent and can reflect the interaction results as the state. Assume that in the current state s, the agent has N actions to choose from. When the agent selects one of the actions i, a new state si will be generated. Then, in the current state s, the state space that the agent can reach is S = {s1, s2,..., sN}, where si represents the new state that the agent may reach by taking different actions.
[0089] A represents the set of actions of the agent. An action refers to various behaviors that the agent can choose in the current state s. Suppose there are M actions that the current agent can choose, then the action set can be expressed as A = {a1, a2,..., aM}, where ai represents the i-th action that the agent can choose. The action space can be discrete or continuous. In a discrete action space, the agent can only choose from a finite set of actions, such as {move up, move down, turn left, turn right}. In a continuous action space, the agent can choose any value within a range, such as speed or angle.
[0090] P represents the state transition probability. It represents the probability of transferring to another state s' (s' belongs to S) after the action a acts in the current state s (s ∈ S). The specific mathematical expression is Given a policy π and a Markov decision process (MDP): M = <S, A, P, R, γ>, then when executing the policy π, the probability that the state transfers from s to s' is equal to the sum of a series of probabilities. This series of probabilities refers to the probability π(a|s) of executing a certain action a when executing the current policy π and the probability that this action can make the state transfer from s to s′ The product of. The specific mathematical expression is:
[0091] R is the reward function. It represents the reward obtained after taking the action a (a ∈ A) in the current state s (s ∈ S). The immediate reward obtained by executing the specified policy π in the current state s is the reward obtained from all possible actions under this policy π The sum of the product of the probability π(a|s) of this behavior occurring: The reward function provides the basis for the agent to evaluate the current state and select the optimal action, and is the foundation for the effective operation of the reinforcement learning algorithm. Through the reward function, the agent can learn and optimize its decision-making strategy in the interaction with the environment to maximize the long-term cumulative reward.
[0092] r is the discount factor, γ ∈ [0, 1]. The discount factor can balance the cumulative reward of the current state and the immediate reward of future moments.
[0093] Based on the above analysis, the objective function of the UAV i can be expressed as:
[0094]
[0095] where, π i represents the policy adopted by the i-th UAV, o i represents the local observation information of the UAV, ai Denote the strategy adopted by UAV i
[0096]
[0097] Among them, the target optimal strategy of UAV i Satisfy:
[0098]
[0099] Specifically, send the global observation value and the target information of the UAV's priority pursuit or priority evasion to the Actor model. Based on this, the Actor model formulates a control strategy. After the UAV executes the decision, it obtains the current observation value, stores the current observation value in the buffer pool, updates the Actor model with the maximization of the UAV's objective function as the guide, and uses the Critic model to evaluate the observation values in the buffer pool to guide the update iteration of the Actor model. The Critic model calculates the Q value and compares it with TargetQ, and feeds back the difference information to guide the update iteration of the Actor network. When the difference between the Q value and TargetQ approaches zero, it indicates that the policy optimization degree has improved and the iteration process tends to be perfect.
[0100] More specifically, update the Actor model with the maximization of the UAV's objective function as the guide, which involves the calculation of the UAV's reward function. The UAV's reward function mainly includes: capture reward and guidance reward. Among them, the guidance reward encourages its active movement by shortening the distance between the UAV and the target, and the capture reward prompts the UAV swarm to capture more targets, thereby accelerating and optimizing the convergence process of the algorithm:
[0101] The capture reward function is:
[0102]
[0103] The guidance reward function is:
[0104]
[0105] Among them, d ij Denote the distance between the pursuing UAV i and the evading UAV j, r i Represents the action radius of the capture action of the pursuing UAV, r j Represents the action radius of the evasion action of the evading UAV, Represents the observation radius of the pursuing UAV, Represents the observation radius of the evading UAV, H represents the pursuing UAV swarm, E represents the evading UAV swarm;
[0106] The comprehensive reward function of the UAV is R hit ,R guideLinear weighted sum of two parts:
[0107] R total = λ 1 ·R hit + λ 2 ·R guide
[0108] Where λ 1 , λ 2 are coefficients for adjusting the weights of the reward functions of each part (λ 1 > 0, λ 2 > 0, (λ 1 + λ 2 ) = 1)
[0109] Specifically, in step S4, when the UAV makes a pursuit or evasion decision for the target, the pursuit UAV cluster aims to maximize the number of opponent UAVs captured, while the evasion UAV cluster aims to minimize the number of UAVs captured.
[0110] It also includes the following steps:
[0111] S5. Based on the kinematic model, realize real-time update of its own position and speed;
[0112] Specifically, the movement of the UAV follows the following kinematic equation:
[0113]
[0114] Where v t is the speed of the UAV in the heading direction at time t, and its maximum is v max , a is the acceleration of the UAV movement, v x and v y represent the components of the speed v in the x-axis direction and y-axis direction respectively, and (x, y) is the position coordinate of the UAV.
[0115] S6. Judge whether the pursuit and evasion are over. If not, return to step S1 and continue to loop; otherwise, stop the loop.
[0116] Specifically, the condition for the end of the pursuit and evasion is shown by the following formula, that is, when the straight-line distance between the two sides of the pursuit and evasion is less than the sum of the action radius of the capture action of the pursuer and the action radius of the evasion action of the evading UAV, the pursuit and evasion between the UAVs end.
[0117]
[0118] In the formula, Loc i and Loc j are the positions of the pursuing UAV i and the evading UAV j.
[0119] The following is a comparative experiment between the algorithm of the present invention and the MADDPG algorithm:
[0120] Create a pursuit-evasion game scenario for a UAV swarm in a finite area on a two-dimensional plane, where N pursuit UAVs confront N evasion UAVs in an adversarial pursuit-evasion game. The number of rounds is 8000, and the time step for each round is 300. During the execution of the strategy, the UAVs are in a partially observable state. The combat parameters of the UAVs are different. The speed of the pursuit UAVs is greater than that of the evasion UAVs. At the same time, to make the strengths of both sides of the game have their own advantages, the acceleration of the evasion UAVs is set higher than that of the pursuit UAVs. The specific parameters are shown in Table 1, and the UAV training parameters are shown in Table 2.
[0121] Table 1 UAV parameter list
[0122]
[0123] Table 2 Training parameter list
[0124]
[0125] The experimental results are shown in Table 3. The complete capture rate of the algorithm of the present invention is greater than that of the MADDPG method in most cases;
[0126] Table 3 Performance comparison of different algorithms under different numbers of UAVs
[0127]
[0128]
[0129] As Figure 5 、 Figure 6 shown, the partial dynamic trajectories of the 2-on-2 and 5-on-5 UAV pursuit-evasion of the algorithm of the present invention are shown.
[0130] The present invention also provides a partially observable UAV swarm cooperative pursuit-evasion game decision-making system, including:
[0131] An acquisition module for acquiring the local observation information of the UAV to be decided;
[0132] A processing module for obtaining the global observation value within the UAV swarm to be decided according to the local observation information;
[0133] A pursuit-evasion game strategy formulation module. The processing module communicates and transmits the motion information in the local observation information to the stable matching model of the pursuit-evasion game strategy formulation module to obtain the target information of the UAV to give priority to pursuit or priority to evasion, and then returns the target information of the UAV to give priority to pursuit or priority to evasion to the processing module;
[0134] The strategy module. The processing module sends the global observation value and the target information of the UAV's priority pursuit or priority evasion to the strategy module to generate the strategy of the UAV for the priority pursuit or priority evasion target.
[0135] Specifically, the local observation information obtained by the acquisition module is screened and processed by the processing module to extract key information such as the position, speed, heading, and state of the UAV, and part of the information is transmitted to the pursuit-evasion game strategy formulation module; the pursuit-evasion game strategy formulation module calculates and matches the received information to determine the action target of the UAV and returns this target to the processing module for subsequent processing. Then, the processing module integrates the global observation information and the result of the pursuit-evasion game strategy and sends it to the strategy module. The Actor model formulates a control strategy based on this and stores the current observation value in the buffer pool. The Critic model evaluates the observation value in the buffer pool to guide the update and iteration of the Actor model.
[0136] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the steps in the above-mentioned partially observable UAV cluster cooperative pursuit-evasion game decision-making method.
[0137] The present invention also provides a partially observable UAV cluster cooperative pursuit-evasion game decision-making device, including:
[0138] A memory for storing software application programs,
[0139] A processor for executing the software application programs, and each program of the software application programs correspondingly executes the steps in the above-mentioned partially observable UAV cluster cooperative pursuit-evasion game decision-making method.
Claims
1. A partially observable UAV swarm collaborative pursuit and escape game decision-making method, characterized in that: The following steps are involved: Obtain local observation information of the drone to be decided; Based on the local observation information, the global observation value within the UAV cluster to be decided is obtained; The motion information in the local observation information is communicated to the stable matching model to obtain the target information that the UAV should preferentially pursue or evade; According to the global observation value and the target information that the UAV prioritizes to pursue or escape, a strategy model is constructed to generate the UAV's strategy for the target that it prioritizes to pursue or escape.
2. The partially observable UAV swarm collaborative pursuit and escape game decision-making method according to claim 1 is characterized in that: The motion information in the local observation information is communicated to the stable matching model to obtain the target information of the drone's priority pursuit or priority escape, including the following steps: According to the motion information in the local observation information, the comprehensive preference index is used to evaluate the UAV situation, and potential pursuit and escape pairs are screened out from the pursuit UAV cluster and the escaping UAV cluster. The comprehensive preference index is the linear weighted sum of the angle preference index, distance preference index and speed preference index.
3. The method for cooperative pursuit and escape game decision-making of partially observable drone swarm according to claim 1 or 2, characterized in that: The motion information in the local observation information is transmitted to the stable matching model to obtain the target information of the drone's priority pursuit or priority escape, and the following steps are also included: Traverse all possible combinations of the chasing and fleeing parties and evaluate the initial interaction between each pair; By calculating the comprehensive preference index, the game strategy matrix of the first time step is determined; Before each new time step begins, the game strategy matrix obtained in the previous time step is used as the game strategy matrix of the current time step; At each time step, iterate each pursuing drone to obtain the optimal game goal of the current pursuing drone; Based on the above iteration results, the game strategy matrix of the entire pursuit and escape game is updated; According to the obtained pursuit and escape game strategy matrix, a targeted pursuit and escape relationship is established between the pursuer and the escapee.
4. The partially observable UAV swarm collaborative pursuit and escape game decision-making method according to claim 1 is characterized in that: The strategy model is an Actor model. According to the global observation value and the target information of the drone's priority pursuit or priority escape, the Actor model generates the drone's strategy for the priority pursuit or priority escape target. After the drone executes the decision, it obtains the current observation value, stores the current observation value in the buffer pool, updates the Actor model guided by the maximization of the drone's objective function, and uses the Critic model to evaluate the observation value in the buffer pool to guide the update iteration of the Actor model.
5. The method for cooperative pursuit and escape game decision-making of partially observable drone swarm according to claim 3 or 4, characterized in that: The UAV objective function can be expressed as: Among them, π i represents the strategy adopted by the i-th drone, o i represents the local observation information of the UAV, a i Represents the strategy adopted by drone i.
6. The partially observable UAV swarm collaborative pursuit and escape game decision-making method according to claim 4 is characterized in that: The update of the Actor model involves the drone comprehensive reward function, where the drone comprehensive reward function is a linear weighted sum of the capture reward function and the guidance reward function, and the capture reward function is: The guided reward function is: Among them, d ij represents the distance between the chasing drone i and the fleeing drone j, r i Represents the capture action radius of the pursuit drone, r j Represents the evasion action radius of the evasion drone. Represents the observation radius of the hunting drone, represents the observation radius of the evading drone, H represents the pursuit drone cluster, and E represents the evasion drone cluster.
7. The partially observable UAV swarm collaborative pursuit and escape game decision-making method according to claim 1 is characterized in that: The following steps are also included: Based on the kinematic model, the real-time update of its own position and speed is realized; Determine whether the pursuit is over. If not, repeat the steps and continue the cycle; Otherwise stop the loop.
8. A partially observable UAV swarm collaborative pursuit and escape game decision-making system, characterized in that: include: The acquisition module is used to obtain the local observation information of the drone to be decided; A processing module is used to obtain the global observation value within the UAV cluster to be decided based on the local observation information; The pursuit-escape game strategy formulation module, the processing module communicates the motion information in the local observation information to the stable matching model of the pursuit-escape game strategy formulation module, obtains the target information that the drone prioritizes to pursue or prioritizes to escape, and then returns the target information that the drone prioritizes to pursue or prioritizes to escape to the processing module; Strategy Module,The processing module sends the global observation value and the target information of the UAV’s priority pursuit or priority escape to the strategy module, builds a strategy model, and generates the UAV’s strategy for the priority pursuit or priority escape target.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the partially observable drone cluster collaborative pursuit and escape game decision-making method described in any one of claims 1 to 7 are performed.
10. A partially observable drone cluster collaborative pursuit and escape game decision-making device, characterized in that: include: Memory, for storing software applications, A processor is used to execute the software application, and each program of the software application correspondingly executes the steps in the partially observable drone cluster collaborative pursuit and escape game decision-making method described in any one of claims 1 to 7.
Citation Information
Cited By
Unmanned aerial vehicle cooperation-based driving carrying platform adaptive adjustment method and system
CN121680465A
A method and system for adaptive adjustment of a driving carrier platform under cooperation of a UAV
CN121680465B