Simulation method for coordinated defense of UAV group under electronic interference and wave reinforcement
By employing a multi-agent proximal policy optimization algorithm and an adaptive reward function, the problem of collaborative defense simulation of UAV swarms under dynamic uncertainty and electronic interference was solved, achieving stable collaborative decision-making in wave-based reinforcement and deceptive target environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2025-11-28
- Publication Date
- 2026-06-26
AI Technical Summary
Existing collaborative defense simulation methods for drone swarms struggle to maintain robustness in the face of dynamic uncertainties and electronic interference. In particular, in environments involving wave-based reinforcement and deceptive targets, the policy network of multi-agent reinforcement learning algorithms struggles to extract key interaction information from high-dimensional situations, affecting stability and generalization ability.
The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, combined with the Multi-Agent Partially Observable Markov Decision Method (MA-POMDP) and an adaptive reward function, is used to establish the state space and the optimization objective function. The optimization objective function is then solved using the MAPPO algorithm to achieve adversarial simulation of UAV swarms.
Reliable UAV swarm cooperative combat simulation was achieved in complex operating environments, improving simulation stability and the effectiveness of cooperative decision-making. It can maintain the stability and generalization ability of the strategy under electronic interference and wave reinforcement.
Smart Images

Figure CN121680112B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control technology, and in particular to a simulation method for cooperative defense against UAV swarms under electronic interference and wave reinforcement. Background Technology
[0002] In modern aerial operations, UAV swarms offer greater mission adaptability and system redundancy compared to single platforms, enabling them to address complex enemy situations through formation and strategic division of labor. However, real-world operational scenarios are often accompanied by dynamic uncertainties: enemy UAVs employ wave-based reinforcements, decoy deception, and electronic interference, leading to incomplete perception and high-noise collaborative decision-making. These factors collectively expand the state space and induce multi-agent non-stationarity, making it difficult for rule-based and greedy heuristic methods to maintain robustness under dynamic changes in task density and interference intensity.
[0003] In recent years, multi-agent reinforcement learning (MARL) has demonstrated its potential for collaborative decision-making within the centralized training-distributed execution (CTDE) paradigm. However, existing research is mostly based on static or semi-static environment settings, neglecting observation noise and policy drift caused by decoys and disturbances. Furthermore, the policy networks are mostly shallow MLPs, making it difficult to extract key interaction information from high-dimensional situations, which affects stability and generalization ability. Summary of the Invention
[0004] This invention discloses a simulation method for cooperative defense against unmanned aerial vehicle (UAV) swarms under electronic jamming and wave reinforcement, in order to overcome the above-mentioned technical problems.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] A simulation method for cooperative defense against unmanned aerial vehicle (UAV) swarms under electronic jamming and wave-reinforcement includes the following steps:
[0007] S1: Establish the state space of a multi-agent partially observable Markov decision-making method based on wave reinforcement;
[0008] S2: Establish a reward function that takes into account the presence of false drone targets and electronic interference;
[0009] S3: Based on the state space and reward function, establish an optimization objective function to maximize the total reward of our drone swarm in performing the task;
[0010] S4: Employ a multi-agent near-end strategy optimization algorithm to solve the optimization objective function, thereby obtaining the target selection of our UAVs and realizing the adversarial simulation of UAV swarms.
[0011] Furthermore, the reward function is expressed as follows:
[0012]
[0013] in:
[0014]
[0015] In the formula: The reward function value is given by the multi-agent partially observable Markov decision-making method. Rewards for completing the task; Rewards are imposed to ensure safety; Rewards for collaborative efficiency; For robust rewards; Reward for efficiency; All are weights.
[0016] Furthermore, the task completion reward is represented as follows:
[0017]
[0018]
[0019]
[0020] In the formula: , , These are all hyperparameters used to balance rewards; In the first The number of real targets newly destroyed in each execution step; The number of false targets that were mistakenly hit; This is a task completion indicator function that is triggered when the percentage of real targets destroyed reaches a preset threshold. For the first The total number of real targets destroyed in each execution step; For the first The total number of false targets mishit in each execution step.
[0021] Furthermore, the security constraint reward is represented as follows:
[0022]
[0023] In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; , These are all hyperparameters used to balance the safety constraints of our drones; For indicator functions; Indicates the first Loss per execution step They deployed our drones.
[0024] Furthermore, the collaboration efficiency reward is represented as follows:
[0025]
[0026] In the formula: , , These are all weighting parameters in the collaboration efficiency reward; Represents the count of coordinated actions; Represents the count of conflicting actions; This represents the count of information-sharing actions.
[0027] Furthermore, the robustness reward is represented as follows:
[0028]
[0029] In the formula: , , These are all weight parameters in robust reward; This indicates the number of times a target was correctly identified in a jammed environment; This indicates the number of drones affected by electronic interference; This indicates the number of drones that have recovered from the interference.
[0030] Furthermore, the efficiency reward is represented as follows:
[0031]
[0032] In the formula: These are all weighting parameters in efficiency rewards; This indicates that the task is complete. and They represent the first The cumulative energy consumption and total energy of our drones.
[0033] Furthermore, the state space includes the state space of the enemy drone and the state space of our drone;
[0034] The state space representation of our UAV is as follows:
[0035]
[0036] In the formula: For the first The state space of our drones; They represent the first The horizontal and vertical coordinates of the position of our drone; Indicates the first The remaining fuel of our drone; For the first Target selection for our drones; For the first Decision parameters for determining whether our drones are subject to electronic interference; This represents the total number of our drones;
[0037] The state space representation of the enemy drone is as follows:
[0038]
[0039] In the formula: For the first The state space of the enemy drone; They represent the first The horizontal and vertical coordinates of the enemy drone's position; Indicates the first The movement speed of the enemy drone; Indicates the type of enemy entity. Indicates the actual number of enemy drones. K Indicates the number of decoy targets. L Indicates the number of dynamic interference sources; An index for real drone targets.
[0040] Furthermore, the objective function for maximizing the total reward of our drone swarm during mission execution is expressed as follows:
[0041]
[0042] In the formula: This represents the total reward for our drone swarm during mission execution; The reward function value is given by the multi-agent partially observable Markov decision-making method. This represents the maximum number of steps the task can take. Index for the number of execution steps; Indicates the expected reward; Parameters representing the joint strategy; This represents the expectation of the trajectory distribution; This indicates the trajectory generated by the interaction between our drone and the environment; Discount factor;
[0043] The constraints of the objective function include: UAV survival constraints, mission completion constraints, and time constraints;
[0044] Among them, the drone survival constraints are:
[0045]
[0046] In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; This is an index for our drones.
[0047] Task completion constraints:
[0048]
[0049] In the formula: Indicates the actual number of enemy drones; An index for real drone targets; It is an indicator function, if the first... j The indicator function takes a value of 1 if a real target has been destroyed, and 0 otherwise. Indicates the first j Whether the actual target has been destroyed; Indicates the first j A real goal; Indicates the percentage of mission objectives destroyed; Time constraint:
[0050] .
[0051] Furthermore, the loss function of the multi-agent proximal policy optimization algorithm is expressed as follows:
[0052]
[0053]
[0054] In the formula: Optimize the target for the cropping strategy; This represents the ratio of the current strategy to the old strategy. The dominant function; These are the trimming parameters; This is the current strategy; This is the old strategy; The expected loss of the strategy; KL penalty coefficient; The Kullback-Leibler divergence between policy distributions; For the first Target selection with a number of execution steps; For the first The state of each execution step.
[0055] Beneficial Effects: This invention provides a cooperative defense drone swarm simulation method under electronic jamming and wave-reinforcement conditions. Based on a constructed reward function considering the presence of false drone targets and electronic jamming, and a state space based on a multi-agent partially observable Markov decision method using wave-reinforcement, an optimization objective function is established to maximize the total reward of the friendly drone swarm during mission execution. A multi-agent proximal policy optimization algorithm is then used to solve this objective function. This invention models scenarios with false targets and electronic jamming, and by employing a multi-agent proximal policy optimization algorithm, it can extract key interaction information from high-dimensional situations, exhibiting high simulation stability and enabling reliable cooperative drone swarm simulation in complex operational environments. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart of the drone swarm confrontation simulation method of the present invention.
[0058] Figure 2 This is a schematic diagram of a simulation scenario in an embodiment of the present invention.
[0059] Figure 3 This is a comparison chart of the algorithm reward curves in the embodiments of the present invention.
[0060] Figure 4 This is a comparison chart of the success rate curves of the algorithms in the embodiments of the present invention.
[0061] Figure 5 This is a comparison chart of the overall performance of the algorithms in the embodiments of the present invention.
[0062] Figure 6 This is an ablation diagram of the improved MAPPO algorithm technology in an embodiment of the present invention. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] This embodiment introduces a simulation method for cooperative defense against unmanned aerial vehicle (UAV) swarms under electronic jamming and wave reinforcement, including the following steps: Figure 1 As shown:
[0065] S1: Establish the state space, observation space, action space, and wave-based state transition function based on the multi-agent partially observable Markov decision-making method;
[0066] This embodiment focuses on multi-UAV collaborative operations in a dynamic asymmetric environment. Specifically, the research scenario is a continuous two-dimensional space (1000×1000), where our UAV swarm performs collaborative strike and suppression missions in an asymmetric confrontation environment, while the opposing side increases the complexity of the confrontation through wave-based reinforcements, deception of false targets, and electronic jamming.
[0067] Among them, dynamic asymmetric environment operation scenarios have the following core challenges:
[0068] (1) Asymmetric confrontation and deception identification: The opponent creates uncertainty in target identification by using false targets and electronic interference. Real / false targets are mixed and the tag noise changes dynamically with the interference intensity, which directly impacts the reliability of our identification, locking and engagement strategies.
[0069] (2) Wave-based reinforcement leads to non-stationarity: the enemy UAVs enter in batches based on Poisson distribution, and the instantaneous target density and threat topology change, which induces formation congestion and task queue accumulation, requiring the strategy to have the ability to quickly replan and divide tasks under sudden load.
[0070] (3) Resource-constrained collaborative allocation: Constraints such as fuel / energy, communication bandwidth, and attack window coexist with collaborative tasks. Dynamic trade-offs must be made between attack efficiency and survival security to avoid excessive concentration or dispersion leading to overall utility loss.
[0071] Preferably, the state space includes the state space of the enemy drone and the state space of our drone;
[0072] This embodiment incorporates an optimization problem design, modeling the multi-UAV cooperative adversarial task as a multi-agent partially observable Markov decision process (MA-POMDP). It uses a quintuple... This indicates the process. Among them, For state space, For observation space, For action space; This is the state transition function; For the reward function;
[0073] (1) State space This includes the status of all drones (including our drones and enemy drones), the status of targets, the status of enemy interference sources, and overall environmental information.
[0074] In a continuous two-dimensional plane, global information such as the coordinates and remaining fuel of each friendly UAV, the position and speed of the enemy UAV (both real and false targets), and the current time step are collected.
[0075] The state space representation of our UAV is as follows:
[0076] (1)
[0077] In the formula: For the first The state space of our drones; They represent the first The horizontal and vertical coordinates of the position of our drone; Indicates the first The remaining fuel of our drone; For the first Target selection for our drones; For the first Decision parameters for determining whether our drones are subject to electronic interference; This represents the total number of our drones;
[0078] The state space representation of the enemy drone is as follows:
[0079]
[0080] In the formula: For the first The state space of the enemy drone; They represent the first The horizontal and vertical coordinates of the enemy drone's position; Indicates the first The movement speed of the enemy drone; Indicates the type of enemy entity. Indicates the actual number of enemy drones. K Indicates the number of decoy targets. L Indicates the number of dynamic interference sources; An index for real drone targets.
[0081] in,
[0082] .
[0083] (2) Observation space : No. This refers to the local observation information that our UAVs can acquire. This includes their own status information (position coordinates, flight speed, remaining fuel) and information on enemy targets within their detection range (position, type identification, etc.). Due to electronic interference and limited field of view, each UAV's observations are local and may contain noise.
[0084] (3) Action space Each of our drones can execute a set of actions, including target attack commands and non-attack maneuver commands. All drones simultaneously select actions at each time step.
[0085] .
[0086] (4) State transition function This describes the dynamic evolution of the environment. Based on the current state and the joint actions of all agents, the system transitions to the next state. ,in, For the first The state of each execution step. Indicates the first Our drone was deployed in the 1st The target selection involves a number of execution steps. Based on the selected action, our drones execute the corresponding maneuver strategy; simultaneously, the enemy drones move dynamically according to a preset behavior pattern. The enemy employs a wave-based reinforcement strategy, dynamically generating new combat units at specific time points. Each wave contains 2-5 drones, entering the combat area from random locations at the battlefield boundary. Environmental constraints limit the total number of enemy drones to a preset upper limit of 20; if this limit is exceeded, no more drones will be generated.
[0087] (5) Reward function The design of the reward function is crucial to the success of the CTDE framework. This embodiment employs an adaptive reward mechanism that combines dense and sparse rewards. In the early stages of training, dense rewards (such as target approach rewards and effective detection rewards) guide the agent to learn basic behavioral patterns, encouraging it to actively approach the target. As training progresses, the reward mechanism gradually shifts to sparse rewards centered on task completion, primarily based on key events such as target destruction and task success.
[0088] S2: Establish a reward function based on the multi-agent partially observable Markov decision-making method that takes into account the presence of false UAV targets, the survivability of our UAVs, and the impact of electronic interference, so as to achieve formal modeling of MA-POMDP.
[0089] Specifically, during the intensive training phase, the reward function needs to simultaneously satisfy two objectives: first, to provide a clear task-oriented signal for the centralized value network; and second, to provide intensive feedback shaping for the local behaviors of each agent. This design helps alleviate the non-stationarity of multi-agent environments and promotes effective collaborative cooperation. Based on this, this embodiment proposes the following reward function representation:
[0090] (2)
[0091] The weights satisfy the following:
[0092] (3)
[0093] In the formula: The reward function value is given by the multi-agent partially observable Markov decision-making method. Rewards for completing the task; Rewards are imposed to ensure safety; Rewards for collaborative efficiency; For robust rewards; Reward for efficiency; All are weights.
[0094] Specifically, task completion rewards The aim is to encourage our drones to accurately identify and destroy real targets while avoiding attacks on false targets.
[0095] (4)
[0096]
[0097]
[0098] In the formula: , , These are all hyperparameters used to balance rewards; In the first The number of real targets newly destroyed in each execution step; The number of false targets that were mistakenly hit; This is a task completion indicator function that is triggered when the percentage of real targets destroyed reaches a preset threshold. For the first The total number of real targets destroyed in each execution step; For the first The total number of false targets mishit in each execution step.
[0099] Specifically, ensuring the effective execution of the mission is the prerequisite for the stable operation of the system, and the safety of our drones is a prerequisite for the system's stable operation.
[0100] Specifically, safety constraint rewards Ensuring the survivability of drones during mission execution is defined as:
[0101] (5)
[0102] In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; , These are all hyperparameters used to balance the safety constraints of our drones; For indicator functions; Indicates the first Loss per execution step Our drones were deployed;
[0103] Specifically, the first aspect of the safety constraint rewards encourages drones to remain alive, while the second aspect penalizes drone losses.
[0104] Specifically, besides individual safety, the overall collaborative efficiency of a multi-agent system directly determines the quality of task completion. Collaboration efficiency rewards. The formula for promoting effective coordination and information sharing among drones is as follows:
[0105] (6)
[0106] In the formula: , , These are all weighting parameters in the collaboration efficiency reward; Represents the count of coordinated actions; Represents the count of conflicting actions; This represents the count of information-sharing actions.
[0107] Specifically, considering the complexity of actual combat environments, the system also needs to possess robustness against various forms of interference. Robustness bonuses The ability of a drone to adapt to electronic interference and environmental noise is defined as follows:
[0108] (7)
[0109] In the formula: , , These are all weight parameters in robust reward; This indicates the number of times a target was correctly identified in a jammed environment; This indicates the number of drones affected by electronic interference; This indicates the number of drones that have recovered from the interference.
[0110] Finally, to achieve optimal resource allocation, the system needs to pursue execution efficiency while ensuring task quality. Efficiency rewards. The expression for incentivizing drones to complete tasks quickly and efficiently while controlling resource consumption is as follows:
[0111] (8)
[0112] In the formula: These are all weighting parameters in efficiency rewards; This indicates that the task is complete. and They represent the first The cumulative energy consumption and total energy of our drones.
[0113] Specifically, the efficiency rewards consist of three components: a first-time reward for completing the task for the first time, a time efficiency penalty, and an energy consumption penalty. The weights of each parameter are adjusted based on task priority and the training phase to achieve a balance among the reward components.
[0114] S3: Based on the reward function, establish an optimization objective function and constraints to maximize the total reward of our drone swarm in performing the mission;
[0115] Specifically, the task is viewed as an optimization problem, with the goal of maximizing the total reward of our drone swarm in performing the task.
[0116] Preferably, the objective function for maximizing the total reward of our drone swarm in performing the mission is expressed as follows:
[0117] (9)
[0118] In the formula: This represents the total reward for our drone swarm during mission execution; The reward function value is given by the multi-agent partially observable Markov decision-making method. This represents the maximum number of steps the task can take. Index for the number of execution steps; Indicates the expected reward; Parameters representing the joint strategy; This represents the expectation of the trajectory distribution; This indicates the trajectory generated by the interaction between our drone and the environment; Discount factor;
[0119] Specifically, the optimization objective function is determined by a combination of factors, including the destruction of enemy targets, drone collaboration efficiency, and fuel consumption.
[0120] The constraints include:
[0121] Drone survival constraint: The fuel of each drone cannot fall below 0 at any given moment, that is:
[0122] (10)
[0123] In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; This is an index for our drones.
[0124] Mission completion constraint: The mission succeeds by destroying a certain percentage of real targets without mistakenly attacking fake targets. The mission objective is to destroy a certain percentage of real targets. ,
[0125] (11)
[0126] In the formula: Indicates the actual number of enemy drones; An index for real drone targets; It is an indicator function, if the first... j The indicator function takes a value of 1 if a real target has been destroyed, and 0 otherwise. Indicates the first j Whether the actual target has been destroyed; Indicates the first j A real goal; This indicates the percentage of mission objectives that were destroyed; in this embodiment, =80%.
[0127] Time constraint: The task must be completed within the maximum number of steps. Completed internally:
[0128] (12).
[0129] S4: The Multi-Agent Proximity Policy Optimization (MAPPO) algorithm is used to solve the objective function to obtain the target selection of our UAV.
[0130] This embodiment employs a multi-agent reinforcement learning framework of Centralized Training with Decentralized Execution (CTDE). The core idea of this framework is that the training phase utilizes global information for centralized learning, while the execution phase involves decentralized decision-making based on local observations. This mechanism effectively avoids non-stationarity issues in multi-agent environments, thereby improving the stability of collaborative decision-making and the scalability of the policy.
[0131] Multi-Agent Proximal Policy Optimization (MAPPO) is an extension of PPO in multi-agent scenarios, exhibiting good stability and efficiency. This algorithm employs a typical Actor-Critic architecture: each agent is equipped with an independent Actor network to generate action policies, while sharing a centralized Critic network to estimate the global state value. Centralized value assessment alleviates the credit allocation problem, while distributed actors ensure scalability during execution. It can provide more stable policy updates in partially observable, cooperatively coupled, and strongly non-stationary scenarios.
[0132] In actual optimization, each of our drones has an independent Actor, which generates a local policy and performs updates based on the policy gradient estimate. For example, for policy parameters... :
[0133] (13)
[0134] In the formula: For the first Strategy parameters for deploying our drones The partial derivatives; For strategy parameters The expected reward below; The expected outcome of the strategy; For strategy parameters The following strategy For the first Local observations conducted by our drones; The advantage function can be calculated from the global value function estimated by the ensemble Critic.
[0135] Specifically, through gradient ascent updates, the strategies of each agent are continuously improved in the direction of increasing the team's cumulative reward.
[0136] Specifically, driven by asymmetric adversarial training and wavering load peaks, simple MAPPO is prone to early training instability and mid-term policy jitter. To address this, this embodiment introduces adaptive curriculum learning, generalized advantage estimation (GAE), dual-constraint PPO loss, and a multi-head attention mechanism.
[0137] Specifically, unlike the existing MAPPO, this example makes targeted improvements in key aspects such as observation modeling, advantage estimation and credit allocation, policy update loss, attention network structure, and course learning and scheduling, solving problems such as training instability under wavelet reinforcement and strong interference conditions, and insufficient representation of policy mutation and cooperative coupling.
[0138] 1) Observational modeling and action masking for interfering sources and perception of real and false targets: Explicitly model the target set (including real / false targets) and the interfering source set (denoted as ) in the overall state. L The system adapts the observation noise to the interference power and distance to improve the robustness of the interference scenario. It uses a mask to shield invisible or heavily interfered entities and limits the neighborhood radius or Top-K important entities on the Actor side. It also uses a mask for discrete actions and maps continuous actions to the legal range with tanh to ensure the stability of execution constraints and training.
[0139] 2) Dominance Estimation and Unified Credit Allocation: Generalized Dominance Estimation (GAE) is used to reduce the variance of dominance estimation under partially observable and strongly disturbed conditions, and to mitigate the jitter in policy updates when distribution abruptly occurs due to wave peaks. Dominance calculation follows trajectory-value assessment, according to the centralized Critic... and structure And recursively obtain :
[0140] (14)
[0141]
[0142] In the formula: For generalized advantage estimation; Discount factor; For GAE parameters; For time step backoff; This refers to timing difference error; The state value function for a centralized Critic;
[0143] Specifically, in the implementation, each trajectory is recursively analyzed from back to front, and the advantages are standardized to prevent numerical instability; all agents within a batch share the same... By unifying credit allocation, variance and jitter are significantly reduced in some observable and strongly disturbed conditions.
[0144] 3) Dual-Constraint PPO Loss for Policy Update Stabilization: The dual-constraint PPO loss enhances the stability of near-end updates through "pruning + KL constraint". Waveforming and disturbances can cause instantaneous changes in policy distribution, and a single pruning is insufficient to avoid the risk of large-step updates. Therefore, in addition to the standard pruning loss... A KL penalty term is added to the basic structure, and a KL threshold is set to stop the current update early or reduce the learning rate, in order to suppress the instantaneous changes in policy distribution and the risk of "large step updates"; during the policy update phase, the Actor loss adopts the following pruning of the PPO loss:
[0145] (15)
[0146]
[0147] In the formula: Optimize the target for the cropping strategy; This represents the ratio of the current strategy to the old strategy. The dominant function; These are the trimming parameters; This is the current strategy; This is the old strategy; The expected value of the strategy loss; KL penalty coefficient; The Kullback-Leibler divergence between policy distributions; For the first Target selection for each execution step; For the first The state of each execution step.
[0148] Specifically, the Critic side uses the squared error value loss.
[0149] 4) Collaborative coupling modeling of multi-head attention.
[0150] Multi-head attention mechanisms are used to explicitly characterize the importance of collaborative coupling and neighborhood interactions among agents, highlighting key neighbors and entities under peak interference and strong disturbance conditions, thereby improving the quality of division of labor and conflict avoidance. In the network structure, a centralized Critic performs multi-head attention convergence on all agent embeddings to generate global semantic features for value estimation; the centralized Critic also performs multi-head attention convergence on all agent embeddings and global environmental features to generate global semantic features, which are then processed by an MLP to obtain... The Actor side applies local multi-head attention to the agent or target entity within its neighborhood, forming a context-enhanced local representation. The attention calculation is represented as follows:
[0151] (16)
[0152] in: Let be the dimension of the key vector in the attention mechanism; Q, K, and V are obtained by linear mapping and concatenated after parallel computation using H heads. To balance computational efficiency and robustness, a mask and neighborhood pruning are uniformly used to shield invisible or heavily distorted entities.
[0153] 5) Adaptive learning and phased hyperparameter scheduling.
[0154] To mitigate the severe non-stationarity and load spikes caused by wave-based reinforcements, this example uses a three-stage progressive task difficulty approach, dynamically adjusting hyperparameters (learning rate, entropy coefficient, GAE λ, and KL weights, etc.) based on training success rate and stability metrics. From basic to intermediate to advanced stages, the number of real targets, enemy aircraft generation rate and cluster size, interference intensity, false target ratio, observation radius, and noise are gradually increased. Stage switching is triggered by "the task completion rate of the most recent N rounds reaching a threshold and round loss stabilizing." If the failure rate increases, KL becomes abnormal, or value drifts, the upgrade is paused or regressed, and the learning rate is reduced. This mechanism follows cognitive load and stochastic gradient convergence theory, achieving orderly learning from easy to difficult and fine-grained convergence in the later stages, significantly reducing training oscillations.
[0155] Specifically, the introduction of the adaptive learning mechanism aims to alleviate the severe non-stationarity and load peaks caused by wave-based reinforcement, and to avoid "sudden instability" in the early stages of training. This mechanism simultaneously considers the difficulty of identifying deception and interference, and promotes the gradual advancement of recognition and collaborative capabilities by increasing the proportion of observation noise and false targets in stages.
[0156] This embodiment designs an adaptive course learning mechanism based on training success rate. This mechanism achieves an ordered learning process from simple to complex through progressive adjustment of environmental difficulty and dynamic optimization of hyperparameters. The former controls the number of real targets, enemy aircraft generation rate and cluster size, and interference intensity. The proportion of false targets and the observation radius and noise; the latter adjusts the learning rate, entropy coefficient, and GAE in segments at different stages. Equal to KL weights.
[0157] In practice, a three-stage progressive strategy is adopted, guiding the agent's capabilities to improve by gradually increasing task complexity. The basic learning stage (rounds 1-300) sets 8 targets and an enemy generation rate of 0.05 to establish a basic cooperative mode; the capability enhancement stage (rounds 301-700) increases to 12 targets and a generation rate of 0.10 to strengthen cooperation and recognition; the strategy refinement stage (rounds 701-1000) reaches 16 targets and a generation rate of 0.15 to achieve stable decision-making under complex loads and disturbances.
[0158] To ensure optimal algorithm performance at different training stages, this embodiment designs a phased hyperparameter dynamic optimization mechanism. This strategy fine-tunes the algorithm's default parameter configuration, primarily involving the dynamic optimization of four key hyperparameters. The learning rate employs a decreasing strategy, starting from 3×10⁻⁶ in the basic stage. -4 Gradually reduce to 2×10 in the intermediate stage -4 It eventually converges to 1×10 in the high-level stage. -4 This approach aims to balance rapid learning in the early stages with fine-tuning in the later stages. The entropy coefficient follows a similar decreasing pattern, decreasing from an initial 0.02 to 0.01 and then 0.005, ensuring sufficient exploration of the environment's state space in the early stages of training, while focusing on policy convergence in the later stages. Stage switching is triggered when "the task completion rate in the most recent N rounds reaches a threshold and the round loss is stable"; if the failure rate increases, KL becomes abnormal, or value drifts, the upgrade is paused or rolled back, and LR is reduced. The reward weight also adapts to the stage: initially increased... and Mid-term improvement and In the later stages, under the premise of stable identification, the level should be appropriately increased. .
[0159] Specifically, the adaptive learning mechanism is designed based on the following theoretical considerations: First, the progressive difficulty adjustment follows cognitive load theory, avoiding cognitive overload caused by complex environments in the early stages of learning and effectively preventing "sudden change instability." Second, the phased hyperparameter adjustment strategy balances exploration and utilization; a larger entropy coefficient and learning rate in the early stages promote sufficient exploration, while the tightening of parameters in the later stages is conducive to policy convergence and performance stability. Finally, the decreasing learning rate adjustment conforms to the convergence theory of stochastic gradient descent, which helps the algorithm achieve refined policy optimization in the later stages and reduces training oscillations. Through this systematic learning design, the MAPPO algorithm can achieve more stable and efficient policy learning in complex multi-agent adversarial environments.
[0160] Specifically, to balance computational efficiency and robustness, masks are used to shield invisible or heavily distorted entities, and neighborhood radii or Top-K important entities are limited on the Actor side. The Critic output aggregates features and then passes them through an MLP to obtain... ,in, To estimate the global state value of the centralized Critic output, the Actor fuses context-enhanced features with its own observations, outputting... .
[0161] Specifically, the optimized MAPPO consists of distributed Actors and a centralized Critic. The Actor network takes each agent's local observations as input, including its own state and features of neighboring targets / enemy aircraft, and can access the shared task context. Feature extraction uses MLP (ReLU) and optional LayerNorm, followed by weighted aggregation of neighboring entities through local multi-head attention, concatenated, and then generated as policy parameters by MLP. The output layer uses a Softmax distribution for discrete actions and tanh mapping to legal ranges for continuous actions, strictly satisfying constraints with action masks. The centralized Critic receives all agent embeddings and global environment features (including wave scheduler state, interference intensity, and task progress), and performs global aggregation through multi-head attention to form a semantic representation, ultimately outputting the state value V(s) for GAE and value loss calculation. To improve generalization ability and robustness, lightweight regularization (such as Dropout and weight decay) is introduced into the Critic and Actors, and masks and neighborhood pruning are uniformly used in the attention module.
[0162] The training stream consists of five stages: First, trajectory acquisition is performed in an environment with wavelet reinforcements and interference using the current joint strategy, caching the local observations, actions, rewards, and global states of each agent; second, in the advantage estimation stage, a temporal difference error is constructed using the V(s) of the centralized Critic. The advantages of GAE are derived by recursion from back to front. Third, the strategy update stage uses "dual-constraint PPO loss" to achieve near-end stability optimization, that is, adding KL constraints to the pruning loss to control the update step size, while updating the centralized Critic with value regression loss; Fourth, the course scheduling stage adaptively increases the task difficulty and adjusts the hyperparameters simultaneously based on the training success rate and stability indicators to achieve progressive learning from easy to difficult; Fifth, attention converges at the network level to perform multi-head attention integration on multi-agent embedding, explicitly modeling "important neighbors / key entities" to enhance the effectiveness of collaborative behavior and local interaction.
[0163] To verify the performance of the improved MAPPO algorithm in multi-agent cooperative decision-making under dynamic asymmetric environments, this embodiment constructs a high-fidelity UAV swarm cooperative adversarial simulation environment. The environment follows the MA-POMDP modeling and CTDE training paradigm: during execution, each agent makes independent decisions based on local observations, and during training, a centralized value network estimates the state value under global information.
[0164] The simulation scenario is a 2D battlefield with a 1000×1000 element size (time step 1 s, maximum number of steps 500). Blue circles represent 8 friendly drones, green squares represent enemy drones, yellow circles represent enemy decoy targets, and red crosses represent dynamic electronic jamming sources. There are 8 friendly drones and 18–25 enemy drone targets, of which 20%–30% are decoy targets to simulate deception and confusion. 3–6 dynamic electronic jamming sources are introduced into the environment. Initial deployment uses a partitioned uniform random or Poisson distribution with hard boundary conditions (outbound penalty). Observation employs local field of view plus neighborhood pruning, and masking preventable actions to ensure strategy feasibility. Figure 2 This is a simulation scenario illustration.
[0165] The dynamic nature of the environment is reflected in several aspects, mainly including target movement, the deceptive characteristics of false targets, the impact of electronic interference, and environmental noise. Target movement is modeled using a random walk, with targets moving randomly across the battlefield at a speed of 25 units per step, simulating the uncertainty and dynamic changes of targets in a real-world scenario. False targets possess deceptive characteristics; by setting them to comprise 20% to 30% of the total number of targets and enabling them to disguise themselves as real targets, the accuracy and robustness of the UAV swarm in decision-making are tested.
[0166] The impact of electronic jamming sources was a major challenge in the experiment. The jamming range was 30 units, causing degradation in communication and observation. Sensor noise was also introduced in this environment, with a standard deviation of 0.1, to further simulate the impact of uncertainties in the battlefield. The sensor noise was zero-mean Gaussian noise (standard deviation 0.1), and the communication delay was 1–3 steps. To simulate the “non-stationarity” of real air situations, wave-based reinforcements were introduced: every 80–120 steps, a small number of enemy reinforcements (3–5 aircraft) were generated according to a non-homogeneous Poisson process or a predetermined schedule, and the targets and jamming sources were slightly reset or their intensity adjusted, making the decision-making difficulty change over time.
[0167] A series of physical constraints are set in the environment to ensure the realism and operability of the drones' movement and mission execution. The maximum speed of each drone is limited to 30 units / step, simulating the energy consumption and operational capabilities of a real drone in high-speed flight. Each drone has an initial fuel supply of 100 units, with fuel consumption related to flight and attack actions. The attack range is set to 60 units to ensure that drones can engage targets within a certain distance. To avoid collisions between drones, a minimum safe distance of 50 units is set to ensure that each agent can maintain a safe distance during mission execution in a multi-agent collaborative environment. Energy / loss and risk constraints are incorporated into the reward and constraint items, consistent with the reward design of this embodiment (mistakenly hitting decoys, exposure to areas with strong interference, close-range collisions, boundary crossings, and invalid actions all incur penalties).
[0168] This experiment compares the performance of three deep reinforcement learning-based algorithms (improved MAPPO, MADDPG, and QMIX) with a traditional greedy algorithm. The main objective of the experiment is to evaluate the collaborative ability, target recognition accuracy, and task completion efficiency of different algorithms in complex dynamic environments. To ensure the fairness of the experimental results, all algorithms were trained in the same environment and on the same hardware platform. The hardware configuration consisted of an Intel i7 CPU and an NVIDIA RTX 4060 GPU to ensure consistent computational resources for each experiment.
[0169] like Figure 3 The graph shows a comparison of the reward learning curves of four algorithms, including MAPPO. It can be seen that the MAPPO algorithm exhibits the best performance, ultimately achieving an average reward of 79.5 points, which is 103% higher than the greedy strategy's 39 points. MADDPG performs second best, ultimately achieving an average reward of 76.7 points. Figure 4 This is a comparison chart of the success rates of four algorithms including MAPPO. The MAPPO algorithm has a task success rate of 89.5%, which is 130% higher than the traditional greedy algorithm's 38.8%. MADDPG has a success rate of 82.3%. Figure 5 This chart provides a comprehensive performance comparison of four algorithms, including MAPPO. Performance metrics include final average reward, task success rate, training stability, and training time. Training stability is evaluated by examining the fluctuations in reward at the end of convergence. Traditional greedy strategies, lacking a learning mechanism, show "stability" but not ideal performance. The MAPPO algorithm exhibits exploratory fluctuations in the early stages of training, but the overall fluctuation amplitude is minimal, resulting in the best convergence performance and outstanding stability. Overall, MAPPO achieves the highest reward while maintaining good convergence stability and training efficiency: under the same training environment, its convergence time is approximately 241 seconds, significantly faster than QMIX's 378 seconds and MADDPG's 402 seconds. Figure 6The ablation experiment diagrams for the improved MAPPO algorithm are shown, with "final success rate" as the main evaluation metric. The success rate of the complete MAPPO is 0.892. After removing course learning, GAE, double pruning, and multi-head attention respectively, the success rate drops to 0.743, a relative decrease of 16.7%, 0.698 (21.8%), 0.721 (19.2%), and 0.675 (24.3%).
[0170] In the experiment, all algorithms were trained using the same reward function and environment settings. At the start of each training round, targets and decoys were randomly generated within the battlefield area and moved along random trajectories. The location and intensity of ground jamming sources were also dynamically changed to increase the environmental complexity of the experiment. The objective of the UAV swarm was to maximize the number of real targets destroyed within a limited time, while minimizing losses due to misjudging false targets and encountering jamming sources. In this experiment, the time step was set to 1 second to match the time scale of real air combat environments, with a maximum of 500 steps. Each algorithm was trained for 1000 rounds, and each algorithm was tested five times repeatedly to average the results, in order to evaluate its stability and reliability.
[0171] Experimental results show that the MAPPO algorithm exhibits the best overall performance across multiple key performance indicators, with MADDPG and QMIX ranking second and third respectively, while the traditional greedy algorithm performs the worst. The MAPPO algorithm significantly outperforms other methods in core metrics: the task success rate reaches 89.5%, a 130% improvement compared to the traditional greedy algorithm's 38.8%; the average reward is 79.5 points, a 103% improvement compared to the greedy strategy's 39 points. The improved MADDPG performs second best, with a success rate of 82.3% and an average reward of 76.7 points, close to MAPPO's performance. The improved QMIX has a success rate of 72% and an average reward of 61.2 points, lower than the former two but still significantly better than the greedy algorithm. These results validate the significant advantages of advanced multi-agent reinforcement learning methods in complex collaborative tasks.
[0172] Training stability was evaluated by examining the fluctuations in rewards at the end of convergence for each algorithm: the traditional greedy strategy, lacking a learning mechanism, showed "stability" but not ideal performance; the MAPPO algorithm exhibited exploratory fluctuations in the early stages of training, but the overall fluctuation amplitude was minimal, resulting in the best convergence effect and outstanding stability. MADDPG showed moderate fluctuations during training and tended to stabilize in the later stages. QMIX occasionally exhibited instability in the value function under complex environments, leading to relatively significant performance fluctuations and the lowest stability score.
[0173] In summary, MAPPO achieves the highest return while maintaining good convergence stability and training efficiency: under the same training environment, its convergence time is approximately 241 seconds, significantly faster than QMIX's 378 seconds and MADDPG's 402 seconds. This means that MAPPO utilizes the sample more fully for policy updates per round, resulting in a training efficiency approximately 66% higher than MADDPG, demonstrating its advantage in algorithmic efficiency. MAPPO's significant advantage stems from its global state modeling capability and stable policy optimization method.
[0174] Overall, the experimental results demonstrate that the MAPPO algorithm exhibits superior adaptability and robustness in complex multi-agent adversarial environments. Benefiting from global state information input and the strong generalization ability of the policy network, MAPPO can quickly adapt to environmental changes, such as accurately identifying newly appearing obstacles and effectively distinguishing and ignoring decoy targets, maintaining a leading position in both task success rate and execution efficiency. In contrast, MADDPG and QMIX each have their limitations: MADDPG is prone to training instability or getting trapped in local optima in complex cooperative scenarios; while QMIX has a relatively stable training process, it is constrained in terms of policy flexibility and optimal performance. The significant performance gap with traditional greedy algorithms further validates the superiority of reinforcement learning methods in complex cooperative decision-making tasks.
[0175] To verify the effectiveness of the core components in the improved MAPPO algorithm proposed in this embodiment, we conducted systematic ablation experiments. The "final success rate" was used as the primary evaluation metric (e.g., Figure 5 The success rate of the complete MAPPO was 0.892. After removing course learning, GAE, double pruning, and multi-head attention, the success rate dropped to 0.743, representing relative decreases of 16.7%, 0.698 (21.8%), 0.721 (19.2%), and 0.675 (24.3%), respectively. The results indicate that the absence of any core component leads to significant performance degradation, with multi-head attention and GAE having the greatest impact. This suggests that the ability to represent multi-agent interactions and the quality of advantage estimation are key to improving collaborative performance. Double pruning effectively suppresses excessive policy updates and improves optimization stability. Course learning significantly improves the overall success rate through progressive difficulty and parameter scheduling. Overall, the modules work together to create a synergistic effect, forming an efficient improvement framework for complex collaborative tasks.
[0176] This embodiment focuses on cooperative adversarial tasks involving UAV swarms. A highly complex simulation environment incorporating decoy targets and electronic interference was constructed. The improved MAPPO algorithm and several optimized multi-agent reinforcement learning algorithms were systematically evaluated and compared with a greedy baseline. Experimental results show that in partially observable, highly adversarial dynamic scenarios, multi-agent reinforcement learning has significant advantages over traditional methods: it can achieve more reasonable target allocation and decoy avoidance through self-learning, significantly improving the task success rate. The improved MAPPO algorithm exhibits the best overall performance, combining a higher final success rate with a more stable training process. Through network structure improvement and ablation analysis, the important role of attention mechanisms and other techniques in improving policy performance and stability is demonstrated. Future work will consider introducing more real-world factors and combining them with transfer learning, autonomous game-playing adversarial methods, etc., to continuously improve the robustness and generalization ability of intelligent decision-making in UAV swarms.
[0177] In summary, this embodiment addresses the challenges of target identification and allocation in dynamic asymmetric aerial operations environments, where UAV swarms must perform tasks under conditions of incomplete observability and strong interference. Facing the challenges of wave-like reinforcements, decoy deception, and high-noise decision-making caused by electronic interference, a two-dimensional operational environment simulation platform is constructed. This platform includes 18–25 moving targets, 20–30% decoys, and 3–6 interference sources. The platform speed is limited to 30 units / step, with enemy reinforcements forming waves from the boundary using a Poisson process. The task is modeled as a multi-agent partially observable Markov decision process (MA-POMDP), employing a centralized training-distributed execution (CTDE) framework. An improved MAPPO method is proposed: multi-head attention is introduced into the policy network to aggregate key situational information, and residual connections are used to stabilize the deep network. The training stage combines adaptive curriculum learning and dynamic rewards to progressively increase the task difficulty. The optimization stage employs generalized advantage estimation (GAE) and a dual-constraint PPO loss to enhance stable updates under high noise conditions. Experimental comparisons with greedy strategies and standard baselines such as MAPPO / QMIX / MADDPG show that the improved method in this embodiment outperforms in success rate, average reward, and robustness against interference, while maintaining good sample efficiency and training stability. This provides a scalable decision-making method reference for cooperative combat in UAV swarms in complex operational environments. It not only provides theoretical and experimental support for improving the cooperative decision-making capabilities of agents in complex battlefields but also offers a practical path for the engineering implementation of multi-agent algorithms.
[0178] The main contributions of this invention include:
[0179] 1) Construct a dynamic simulation platform that includes wave-like moving targets, decoy deception, and electronic jamming to formally characterize the key elements of combat in complex operational environments;
[0180] 2) Based on the MAPPO policy network that integrates multi-head attention and residual connections, the objective function is solved to significantly improve the representation of high-dimensional situational information and the ability to train stably.
[0181] 3) Introduce a joint optimization mechanism of adaptive course learning, dynamic reward and GAE+ dual-constraint PPO to balance stability and sample efficiency;
[0182] 4) Establish a unified evaluation framework and systematically compare and verify the effectiveness and robustness of the method through ablation analysis.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A simulation method for cooperative defense against unmanned aerial vehicle (UAV) swarms under electronic jamming and wave-reinforcement, characterized in that, Includes the following steps: S1: Establish the state space of a multi-agent partially observable Markov decision-making method based on wave reinforcement; S2: Establish a reward function that takes into account the presence of false drone targets and electronic interference; The reward function is expressed as follows: in: In the formula: The reward function value is given by the multi-agent partially observable Markov decision-making method. Rewards for completing the task; Rewards are imposed to ensure safety; Rewards for collaborative efficiency; For robust rewards; Reward for efficiency; All are weights; The reward for completing the task is as follows: In the formula: , , These are all hyperparameters used to balance rewards; In the first The number of real targets newly destroyed in each execution step; The number of false targets that were mistakenly hit; This is a task completion indicator function that is triggered when the percentage of real targets destroyed reaches a preset threshold. For the first The total number of real targets destroyed in each execution step; For the first The total number of false targets misclicked in each execution step; The robustness reward is represented as follows: In the formula: , , These are all weight parameters in robust reward; This indicates the number of times a target was correctly identified in a jammed environment; This indicates the number of drones affected by electronic interference; Indicates the number of drones that recovered from the interference; The efficiency reward is represented as follows: In the formula: These are all weighting parameters in efficiency rewards; This indicates that the task is complete. and They represent the first The cumulative energy consumption and total energy of our drones; S3: Based on the state space and reward function, establish an optimization objective function to maximize the total reward of our drone swarm in performing the task; The objective function for maximizing the total reward of our drone swarm during mission execution is expressed as follows: In the formula: This represents the total reward for our drone swarm during mission execution; The reward function value is given by the multi-agent partially observable Markov decision-making method. This represents the maximum number of steps the task can take. Index for the number of execution steps; Indicates the expected reward; Parameters representing the joint strategy; This represents the expectation of the trajectory distribution; This indicates the trajectory generated by the interaction between our drone and the environment; Discount factor; The constraints of the objective function include: UAV survival constraints, mission completion constraints, and time constraints; Among them, the drone survival constraints are: In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; This serves as an index for our drones; Task completion constraints: In the formula: Indicates the actual number of enemy drones; An index for real drone targets; It is an indicator function, if the first... j The indicator function takes a value of 1 if a real target has been destroyed, and 0 otherwise. Indicates the first j Whether the actual target has been destroyed; Indicates the first j A real goal; Indicates the percentage of mission objectives destroyed; Time constraint: In the formula: This represents the maximum number of steps the task can take. S4: Employ a multi-agent near-end strategy optimization algorithm to solve the optimization objective function, thereby obtaining the target selection of our UAVs and realizing the adversarial simulation of UAV swarms.
2. The simulation method for cooperative defense against unmanned aerial vehicle swarms under electronic interference and wave reinforcement as described in claim 1, characterized in that, The security constraint reward is represented as follows: In the formula: Indicates the first Our drone was deployed in the 1st Remaining fuel for each execution step; This represents the total number of our drones; , These are all hyperparameters used to balance the safety constraints of our drones; For indicator functions; Indicates the first Loss per execution step They deployed our drones.
3. The simulation method for cooperative defense against unmanned aerial vehicle swarms under electronic interference and wave reinforcement as described in claim 1, characterized in that, The collaboration efficiency reward is represented as follows: In the formula: , , These are all weighting parameters in the collaboration efficiency reward; Represents the count of coordinated actions; Represents the count of conflicting actions; This represents the count of information-sharing actions.
4. The simulation method for cooperative defense against unmanned aerial vehicle swarms under electronic interference and wave reinforcement as described in claim 1, characterized in that, The state space includes the state space of the enemy drone and the state space of our drone; The state space representation of our UAV is as follows: In the formula: For the first The state space of our drone; They represent the first The horizontal and vertical coordinates of the position of our drone; Indicates the first The remaining fuel of our drone; For the first Target selection for our drones; For the first Decision parameters for determining whether our drones are subject to electronic interference; This represents the total number of our drones; The state space representation of the enemy drone is as follows: In the formula: For the first The state space of the enemy drone; They represent the first The horizontal and vertical coordinates of the enemy drone's position; Indicates the first The movement speed of the enemy drone; Indicates the type of enemy entity. Indicates the actual number of enemy drones. K Indicates the number of decoy targets. L Indicates the number of dynamic interference sources; An index for real drone targets.
5. The simulation method for cooperative defense against unmanned aerial vehicle swarms under electronic interference and wave reinforcement as described in claim 1, characterized in that, The loss function of the multi-agent proximal policy optimization algorithm is expressed as follows: In the formula: Optimize the target for the cropping strategy; This represents the ratio of the current strategy to the old strategy. The dominant function; These are the trimming parameters; This is the current strategy; This is the old strategy; The expected value of the strategy loss; KL penalty coefficient; The Kullback-Leibler divergence between policy distributions; For the first Target selection for each execution step; For the first The state of each execution step.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster collaborative confrontation decision-making method based on reinforcement learning
CN119002521A
Decentralized policy gradient descent and ascent for safe multi-agent reinforcement learning
US20230113168A1