Layered hybrid action decision method for unmanned aerial vehicle cluster in complex electromagnetic environment
By constructing a predictive hierarchical hybrid action control architecture, the problems of decision delay and communication interference in the complex electromagnetic environment of UAV swarms were solved, and the coordinated control of active interception and continuous maneuvering was realized, thereby improving the tactical effectiveness and survivability of UAV swarms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-21
AI Technical Summary
Existing drone swarms suffer from problems such as purely reactive decision-making delays, difficulty in adapting to dynamic node changes, paralysis of coordination due to high-intensity communication interference, and difficulty in coupling discrete tactics with continuous maneuvering in complex electromagnetic environments.
A predictive hierarchical hybrid action control architecture is constructed, which adopts a cluster-based many-to-many dynamic target allocation module and a multinomial fitting trajectory prediction mechanism, combined with a hybrid action near-end strategy optimization algorithm, to achieve active interception and continuous maneuver control of UAV swarms, and reconstructs the global battlefield model using the trajectory prediction module when communication is interrupted.
It enables active interception of UAV swarms in complex electromagnetic environments, improves target kill rate and stealth penetration capability, has strong communication robustness and zero-sample generalization capability for dynamic node changes, and is suitable for large-scale dynamic air combat battlefields.
Smart Images

Figure CN122431400A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of UAV swarm collaborative control and artificial intelligence, specifically involving a hierarchical hybrid action decision-making method for UAV swarms under complex electromagnetic environments. Background Technology
[0002] In technical solutions for UAV cooperative decision-making based on traditional multi-agent reinforcement learning (MARL), similar standard reinforcement learning methods (such as MAPPO, MADDPG, etc.) mostly rely on the purely reactive Markov decision process (MDP) paradigm when dealing with UAV swarm air combat. These methods typically only utilize the observation information at the current moment (such as local position, attitude, etc.) to directly generate corresponding action commands through neural networks.
[0003] Technical defects: (1) Severe delay in passive decision-making: When facing highly maneuverable targets and dynamically changing radar threats, this purely reactive decision-making mode inevitably produces decision lag, causing the UAV swarm to be in a passive "tail-chasing" state for a long time, making it difficult to achieve active interception. (2) Poor ability to balance the mixed action space: High-fidelity fixed-wing air combat requires simultaneous optimization of three-degree-of-freedom (3-DOF) continuous motion control (including throttle, pitch, roll) and discrete tactical decision-making (covering radar silence, weapon lock-on, and firing). Existing standard reinforcement learning algorithms have obvious defects in balancing such mixed action spaces, which not only exacerbates the non-stationarity problem, but also ultimately leads to low swarm survival rate and penetration rate.
[0004] In technical solutions for cluster collaboration and communication methods based on graph neural networks or attention mechanisms: Existing methods address the issue of cluster scaling by introducing graph neural networks (GNNs) and attention mechanisms (such as TarMAC algorithms) into multi-agent reinforcement learning. These methods enable agents to handle information from a variable number of neighboring nodes, achieving global situational awareness and collaboration through real-time state sharing and feature aggregation among nodes.
[0005] Technical shortcomings: (1) Weak generalization ability for dynamic node changes: The real battlefield environment is constantly changing, and the cluster size of the red and blue sides often fluctuates dynamically due to battle losses or enemy reinforcements. Although the existing methods perform well in static scenarios, they are difficult to adapt to highly asymmetric confrontations where the size of both sides changes dynamically. They lack zero-shot generalization ability and often need to be retrained for specific input dimensions. (2) Extremely vulnerable to high-intensity communication interference: In real electronic warfare (EW) environments, strong communication interference can lead to severe network packet loss, thereby cutting off real-time state sharing. These mainstream collaborative networks are highly dependent on continuous data links. Once communication is interrupted, the global attention mechanism will collapse, causing the agent to lose global situational awareness and the collaborative combat system to be quickly paralyzed.
[0006] To address the aforementioned shortcomings, a predictive hierarchical hybrid action control architecture is urgently needed, capable of overcoming the latency of traditional reactive decision-making while jointly achieving the coordination of discrete tactical planning and continuous attitude control. Furthermore, it is necessary to compensate for the loss of global vision caused by communication interruptions, ensuring that the swarm maintains stable interception and survivability performance even under extreme communication packet loss and dynamic node changes. Therefore, this invention provides a hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments. Summary of the Invention
[0007] The purpose of this invention is to provide a hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments. It addresses the problems of pure reactive decision-making lag, difficulty in adapting to dynamic node changes, high-intensity communication interference leading to coordination paralysis, and difficulty in coupling discrete tactics and continuous maneuvers in complex electromagnetic environments such as anti-access / area denial (A2 / AD) for fixed-wing UAV swarms. At the same time, it can solve the problem of how to overcome the delay of pure reactive decision-making, achieve seamless adaptation of large-scale dynamic nodes, and hybrid control of continuous maneuvers and discrete tactics when UAV swarms face strong radar threats and high-intensity communication interference.
[0008] The specific technical solution adopted by this invention is as follows: A hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments. Constructing a predictive hierarchical hybrid action control architecture (H-TP-MASNet) Key technology: A predictive hierarchical hybrid action architecture is proposed to address the problems of reactive decision-making latency, communication vulnerability, and poor scalability in large-scale combat in existing methods; Implementation: A two-layer collaborative structure is constructed; the upper layer adopts a clustering-based many-to-many dynamic target assignment (CM-DTA) module to decouple the global strategy into local sub-team adversarial; the lower layer integrates the kinematic principle-based multinomial fitting trajectory prediction (PF-TP) mechanism into the multi-agent scalable network (MASNet) and trains it through the hybrid action proximal policy optimization (HAPPO) algorithm. Advantages: It gives the drone swarm proactive spatiotemporal prediction capabilities, overcomes the decision delay of pure responsive strategies, and can generate high-fidelity continuous maneuver commands (such as throttle, pitch and roll) and discrete tactical commands (such as radar silence, weapon lock and fire). Constructing a spatiotemporal prediction mechanism for proactive interception Core technology: A prediction-based hybrid action architecture was designed, enabling swarm tactics to transform from passive tailing and pursuit attacks to active deflection and interception. Implementation method: Utilize a lightweight polynomial fitting trajectory prediction (PF-TP) module to infer the short-term spatiotemporal motion trends of all agents by fitting recent trajectory history based on the physical inertia of the UAV. Advantages: By predicting the future trajectory of targets in advance, the cluster can not only achieve a superior target kill rate in a highly asymmetric combat environment, but also effectively reduce the risk of radar exposure, achieving a balance between covert penetration and firepower strike. Construct a time caching mechanism to ensure extremely strong communication anti-interference capabilities Core technology: The trajectory prediction (TP) module's time information caching and reconstruction functions under extreme electronic warfare interference were thoroughly explored and verified; Implementation method: When the communication link is broken due to high-intensity interference, the agent uses the trajectory prediction module to extrapolate the future trajectories of the lost teammates and enemies, and reconstructs a coherent global battlefield model through fragmented data. Advantages: Even in extreme cases with a communication packet loss rate as high as 80%, it can still achieve smooth tactical degradation, maintain stable strike capability, and demonstrate unprecedented communication robustness. Constructing a zero-shot generalization architecture that adapts to dynamic node changes Core technology: To address the issue of dynamic changes in drone swarm nodes due to battle damage or reinforcements in actual deployments, a mechanism is proposed that can adapt to battlefields of different scales without fine-tuning. Implementation: The trajectory prediction function is integrated with the Multi-Agent Scalable Network (MASNet) aggregator. A permutation-invariant aggregation architecture is used to process the dynamic number of neighbor information, which completely decouples the policy from the fixed input dimension. Advantages: It has extremely strong structural flexibility, can directly generalize from 5 vs 8 training scenarios with zero-shot, and seamlessly adapt to large-scale dynamic air combat battlefields of 10 vs 15 and 20 vs 30, showing great potential for real-world deployment. Specifically, the following steps are included: Step 1: Modeling the kinematics and radar threat environment of unmanned aerial vehicles; Step 2: Upper-level strategy planning: Cluster-based many-to-many dynamic target assignment (CM-DTA); Step 3: Lower-level situational awareness: trajectory prediction based on polynomial fitting (PF-TP); Step 4: Lower-level feature fusion: Construct a multi-agent scalable network (MASNet); Step 5: Design and optimization of the hybrid action strategy controller (HAPPO); Step 6: Network training and parameter update.
[0009] The technical effects achieved by this invention are as follows: This invention significantly enhances tactical strike effectiveness and stealthy penetration capabilities; Experimental Comparison: In a highly asymmetric 5v8 air combat simulation environment, the H-TP-MASNet architecture proposed in this invention demonstrates a dominant tactical advantage. Experimental data shows that this invention achieves a target kill rate as high as 76.0% (an average of 6.08 kills), while suppressing the radar exposure rate to an extremely low 18.1%. In contrast, the fully distributed baseline IPPO algorithm has a kill rate and survival rate of 0.0%, while the ablation model without trajectory prediction (H-MASNetw / oTP), although highly aggressive, suffers from a radar exposure rate soaring to 39.2% due to blind pursuit, resulting in a kill rate of only 0.4%. Theoretical Support: This effect is attributed to the introduction of the polynomial trajectory prediction (PF-TP) module, which transforms the UAV swarm from a passive "tailgating" mode to an active "predictive interception" mode. By predicting the future position of highly maneuverable enemy targets and dynamic radar envelopes in advance, the agent can perfectly balance stealthy penetration and lethal firepower output.
[0010] The algorithm of this invention greatly improves training convergence stability and global optimization capability; Experimental Comparison: On the Episodic Reward curve during the training phase, the hybrid action architecture of this invention exhibits exceptionally stable convergence characteristics. In contrast, the baseline IPPO algorithm fails to converge at all, causing the cluster to fall into a suboptimal strategy with extreme negative rewards; the ablation model without a prediction module exhibits extremely high variance and instability during training.
[0011] Theoretical Support: The trajectory prediction module endows the neural network with spatiotemporal predictive capabilities, providing the model with a smooth and forward-looking state representation. This effectively overcomes the hysteresis problem of pure reactive networks when facing high-speed dynamic targets, guiding the policy network to converge quickly and stably to the globally optimal tactical strategy.
[0012] This invention provides extremely strong system resilience (robustness) under extreme communication interference. Experimental Comparison: In high-intensity communication packet loss scenarios simulating real electronic warfare (EW), this invention demonstrates superior graceful degradation capabilities. Experimental data shows that even under extremely harsh conditions with a random packet loss rate as high as 80%, the cluster's kill rate remains an astonishing 74.0%. Furthermore, at a 40% packet loss rate, the survival rate exhibits a counterintuitive peak increase (approximately 30%).
[0013] Theoretical Support: This exceptionally strong anti-interference capability stems from the trajectory prediction module's unique role as a "time information buffer." When a communication link is partially interrupted, the agent can use this module to extrapolate the future trajectories of lost teammates and enemies, thereby reconstructing a coherent global battlefield model from fragmented data. When global coordination is hindered, the agent will adaptively prioritize local evasion, thus improving its temporary survival rate.
[0014] This invention: Large-scale zero-shot generalization capability adaptable to dynamic node changes. Experimental Comparison: To verify the flexibility of the architecture, models trained only in 5 vs. 8 scenarios were directly deployed without any fine-tuning (Zero-Shot) to ultra-large-scale dynamic adversarial environments of 10 vs. 15 and even 20 vs. 30. The 3D projection results show that the cluster can still easily maintain complex spatial formations and continuous coordinated strike capabilities, exhibiting highly consistent tactical performance.
[0015] Theoretical support: This advantage stems from MASNet's internal permutation-invariant attention aggregation mechanism, combined with predictive filtering, which completely decouples the policy network from fixed input dimensions and specific cluster sizes. This mechanism allows agents to dynamically process information about any number of neighbors and enemies, endowing the model with enormous potential and seamless scalability in dynamic deployments of troops on real-world battlefields. Attached Figure Description
[0016] Figure 1 This is a schematic diagram illustrating the highly asymmetric countermeasures against environmental and radar threats of this invention; Figure 2 This is a schematic diagram comparing the training reward convergence curves of different algorithms in this invention; Figure 3 This is a bar chart comparing the tactical performance indicators of this invention in typical confrontation scenarios; Figure 4 This is a line graph showing the changes in target kill rate and survival rate under different communication packet loss rates according to the present invention; Figure 5 This invention provides a 10-to-15 cluster adversarial zero-shot generalization 3D spatial trajectory map. Figure 6 This invention provides a zero-sample generalization 3D spatial trajectory map for 20-to-30 extreme-scale cluster adversarial scenarios. Figure 7 This is a system block diagram of the hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments, as described in this invention. Detailed Implementation
[0017] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.
[0018] like Figure 1 As shown, a hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments is presented. Constructing a predictive hierarchical hybrid action control architecture (H-TP-MASNet) Key technology: A predictive hierarchical hybrid action architecture is proposed to address the problems of reactive decision-making latency, communication vulnerability, and poor scalability in large-scale combat in existing methods; Implementation: A two-layer collaborative structure is constructed; the upper layer adopts a clustering-based many-to-many dynamic target assignment (CM-DTA) module to decouple the global strategy into local sub-team adversarial; the lower layer integrates the kinematic principle-based multinomial fitting trajectory prediction (PF-TP) mechanism into the multi-agent scalable network (MASNet) and trains it through the hybrid action proximal policy optimization (HAPPO) algorithm. Advantages: It gives the drone swarm proactive spatiotemporal prediction capabilities, overcomes the decision delay of pure responsive strategies, and can generate high-fidelity continuous maneuver commands (such as throttle, pitch and roll) and discrete tactical commands (such as radar silence, weapon lock and fire). Constructing a spatiotemporal prediction mechanism for proactive interception Core technology: A prediction-based hybrid action architecture was designed, enabling swarm tactics to transform from passive tailing and pursuit attacks to active deflection and interception. Implementation method: Utilize a lightweight polynomial fitting trajectory prediction (PF-TP) module to infer the short-term spatiotemporal motion trends of all agents by fitting recent trajectory history based on the physical inertia of the UAV. Advantages: By predicting the future trajectory of targets in advance, the cluster can not only achieve a superior target kill rate in a highly asymmetric combat environment, but also effectively reduce the risk of radar exposure, achieving a balance between covert penetration and firepower strike. Construct a time caching mechanism to ensure extremely strong communication anti-interference capabilities Core technology: The trajectory prediction (TP) module's time information caching and reconstruction functions under extreme electronic warfare interference were thoroughly explored and verified; Implementation method: When the communication link is broken due to high-intensity interference, the agent uses the trajectory prediction module to extrapolate the future trajectories of the lost teammates and enemies, and reconstructs a coherent global battlefield model through fragmented data. Advantages: Even in extreme cases with a communication packet loss rate as high as 80%, it can still achieve smooth tactical degradation, maintain stable strike capability, and demonstrate unprecedented communication robustness. Constructing a zero-shot generalization architecture that adapts to dynamic node changes Core technology: To address the issue of dynamic changes in drone swarm nodes due to battle damage or reinforcements in actual deployments, a mechanism is proposed that can adapt to battlefields of different scales without fine-tuning. Implementation: The trajectory prediction function is integrated with the Multi-Agent Scalable Network (MASNet) aggregator. A permutation-invariant aggregation architecture is used to process the dynamic number of neighbor information, which completely decouples the policy from the fixed input dimension. Advantages: It has extremely strong structural flexibility, can directly generalize from 5 vs 8 training scenarios with zero-shot, and seamlessly adapt to large-scale dynamic air combat battlefields of 10 vs 15 and 20 vs 30, showing great potential for real-world deployment. Specifically, the following steps are included: Step 1: Modeling the kinematics and radar threat environment of unmanned aerial vehicles; Step 1 includes the following steps: Step 101: UAV Kinematic Modeling: Establish a three-degree-of-freedom (3-DOF) kinematic model of a fixed-wing UAV; for the motion of the UAV in three-dimensional space, its state space is represented as follows: The kinematic differential equations are defined as follows: ; in, The three-dimensional spatial coordinates of the UAV in the inertial reference frame; Indicates the magnitude of speed; The inclination angle is the flight path angle. This is the heading angle; in addition, Representing gravitational acceleration; this set of kinematic differential equations is driven by three coupled control inputs: throttle overload coefficient. Pitch overload coefficient and roll angle ; Step 102: Radar Detection Probability and Threat Field Model: For models containing... In a threat environment with a fixed radar station, the blue team's drone is located... Cumulative detection probability at location The model is as follows: ; in, For drones to the first The Euclidean distance of the radar unit; The radar cross section (RCS) is the relative azimuth angle between the UAV and the radar line of sight (LOS). The function; in addition, and These are empirically determined radar constants, representing the shape coefficient of the detection probability curve and the generalized radar performance coefficient, respectively.
[0019] Step 2: Upper-level strategy planning: Cluster-based many-to-many dynamic target assignment (CM-DTA); Step 2 includes the following specific steps: Adversarial clustering; using spatial proximity to group red team enemy aircraft clusters Divided into tactical clusters Each cluster This represents a localized threat focus, thereby decoupling the overall problem into localized squad-level confrontations; Define Boolean decision variables Indicates whether the blue team's drone i is assigned to the red team's cluster. Its cluster suppression effect Combining spatial distance advantage, angular attack advantage, and radar exposure avoidance: ; in, Represents a cluster The number of drones used by the Chinese Red team; and These are the range-based and angle-based tactical advantage functions of the blue team's UAV i against the red team's UAV j, respectively; the weight coefficients w1, w2, and w3 are positive weights used to balance the offensive geometric advantage and radar evasion performance. Global task optimization can be described as follows: ; Where Y represents all The constructed Boolean decision matrix; Assemble the blue team's drones; The total number of blue team drones; the constraint of equation (5) ensures that a sufficient red-blue force ratio is maintained in each local battle, and is determined by a pre-set minimum. and maximum Threshold constraints; Constraining the troop ratio in local battle situations (e.g., ensuring a minimum threshold) and maximum threshold Under the premise of ), maximize the global allocation utility; given the determined cluster, Then, based on the distance and the target's quantified threat level, the specific primary target is dynamically selected. : ; in, Indicates the blue team's drone i With Red Team's drone j The Euclidean distance between them; ω is the quantified tactical threat value of target j; ω is the weighting coefficient; finally, the selected target As a key interface identifier, it provides guidance for the execution of lower-level tactics.
[0020] Step 3: Lower-level situational awareness: trajectory prediction based on polynomial fitting (PF-TP); Step 3 includes the following specific steps: To achieve early interception response against highly maneuverable targets with low airborne computational latency, a quadratic polynomial is used to fit the recent historical trajectory of the enemy target to extract dynamic features. In the X-axis coordinate system, the fitted model is: ; in, As a relative time variable, Corresponding to the current time t; These are the polynomial coefficients corresponding to the acceleration, velocity, and position terms fitted along the X-axis, respectively. The corresponding acceleration, velocity, and position polynomial coefficients are extracted using the finite difference method. ; Based on these coefficients, the predicted position and predicted velocity for the next moment are calculated as follows: ; By applying them independently to the y-axis and z-axis, the complete predicted state vector can be obtained: ; in, and They represent the first j The predicted three-dimensional position vector and predicted three-dimensional velocity vector of the red team's UAV.
[0021] Step 4: Lower-level feature fusion: Construct a multi-agent scalable network (MASNet); Step 4 includes the following specific steps: Step 401: State Feature Embedding; For friendly UAVs Its kinematic state Through a friendly multilayer perceptron (MLP) Encode as feature vector : ; For enemy targets its original state Polynomial trajectory prediction and allocation indicator The parts are spliced together and then passed through the enemy's MLP network. Encoding as features : ; Here, Represents vector concatenation operation; indicator function The value is 1 when target j is the primary target assigned to the local UAV, and 0 otherwise. Step 402: Dual-stream multi-head attention and final state aggregation; to handle dynamically changing cluster node numbers, based on self-machine characteristics For querying, a variable amount of local environment information is processed through parallel friendly and adversary multi-head attention networks. and The goal is to output aggregated features with invariant permutations, which are then concatenated into a comprehensive spatiotemporal state representation. .
[0022] Step 5: Design and optimization of the hybrid action strategy controller (HAPPO); Step 5 includes the following specific steps: Step 501: Specific formulas for the hybrid action space and policy output: Discrete tactical instructions With continuous flight control commands The specific space is defined as: ; Mode 0 is silent / cruising; Mode 1 is target lock; Mode 2 is firing. Corresponding to throttle / longitudinal overload, Corresponding to pitch overload, Corresponding roll angle; Dual-head network structure: Policy network Shared MASNet encoder parameters It then branches into two independent heads: The discrete head outputs the output classification distribution. ; The continuous output head outputs a diagonal Gaussian distribution. ,in and The mean and standard deviation vectors generated for the network are state-related, and I is the identity matrix; Under the conditional independence assumption given the state characteristics, the joint policy distribution can be decomposed as follows: ; Step 502: Adversarial Reward Mechanism ; in, and For predefined weighting coefficients; and These represent the rewards for attacking targets, penalties for radar evasion, and penalties for maneuvering smoothness, respectively.
[0023] Step 6: Network training and parameter update.
[0024] In step 6, the Hybrid Proximal Policy Optimization (HybridPPO) algorithm is used to optimize the policy parameters. Update, among which This represents the network parameters used for state evaluation (Critic); it should be noted that the trajectory prediction module is deterministic and does not perform gradient updates at this stage. The specific update process includes the following sub-steps: Step 601: Calculate the importance sampling ratio of the mixed PPO: Since the action space contains both discrete and continuous components, the importance of sampling at time step t is higher than that of continuous components. The calculation is performed using the joint likelihood, and the formula is as follows: ; Step 602: Construct the surrogate objective function for pruning: To limit excessive policy updates, a pruned proxy objective function is used. : ; in, This represents the expected experience on the experience batch. It is a generalized advantage estimate (GAE) obtained from the value (Critic) network evaluation. These are the pruning hyperparameters of the PPO algorithm; Step 603: Calculate the entropy regularization term: To prevent premature model convergence and encourage policy exploration, entropy regularization terms are introduced for both discrete and continuous action heads. : ; in, Shannon entropy represents the distribution. and These represent the entropy coefficients of the discrete action space and the continuous action space, respectively. Step 604: Calculate the total reinforcement learning loss and perform end-to-end updates: The total reinforcement learning loss is obtained by weighted summation of the policy loss, value function loss, and entropy reward. : ; ; in, It is the mean squared error (MSE) loss of the value predicted by the Critic network. This is the weighting coefficient for value loss; End-to-end update backpropagation mechanism: shared MASNet encoder parameters Received by total reinforcement learning loss The backpropagation gradient is used to update the parameters; during this process, although the parameters of the trajectory prediction module remain fixed, the encoder can adapt through gradient updates, thereby effectively utilizing multinomial prediction features. To identify states where a specific tactical mode needs to be adopted.
[0025] In actual simulation experiments, this invention: First: Experimental scenario initialization and parameter configuration. The simulation experiment is set in a high-fidelity 3D highly asymmetric air combat scenario, where the airspace is composed of a boundary-constrained three-dimensional space, such as... Figure 1 As shown: Confrontational formation deployment: The initial setting is that the blue team (our control cluster) has N=5 drones and the red team (enemy cluster) has M=8 drones, forming a 5 vs 8 disadvantageous opening.
[0026] Radar Threat Deployment: Deploy a high-power early warning radar in the central battlefield area. Its detection probability field is calculated by coupling physical-driven dynamic RCS with range and depth. The radar's detection failure threshold radius is set at 30.0 km (i.e., within this radius, changes in RCS may trigger a detection alarm). Any behavior that causes the detection probability to continuously exceed the threshold is considered a mission failure or a deduction of a high survival bonus.
[0027] Communication environment settings: To simulate a real electronic warfare (EW) environment, the instance was set with a communication packet loss rate, and random communication masking of up to 80% was artificially applied in some test rounds.
[0028] Then, the performance and algorithm parameters of the UAV actuator are set: The fixed-wing UAV selected in this example follows a high-fidelity 3-DOF point mass dynamics model, and its state variables are constrained by the following physical safety envelope: Speed constraint: The actual airspeed V must be maintained within the range of [50, 200] m / s.
[0029] Maneuvering constraints: roll angle ϕLimited to ±80°; maximum longitudinal overload factor nx is limited to [-2g, 2g], and normal overload factor n z Limited to [-3g, 6g].
[0030] Trajectory prediction parameters: Quadratic polynomial fitting (PF-TP) is used, and the historical trajectory observation window size is set to H=10 time steps to predict the spatiotemporal trend of the target.
[0031] Next, the algorithm's training performance and convergence are analyzed. This example demonstrates a comparative training experiment of 30,000 episodes between the proposed H-TP-MASNet algorithm and two baseline algorithms (the fully distributed IPPO algorithm and the H-MASNet algorithm without a trajectory prediction module). The training reward curves are shown below. Figure 2 As shown; Convergence efficiency: The algorithm of this invention exhibits extremely high exploration efficiency and stability during the training period, with the average fragment reward rapidly increasing and converging.
[0032] Final reward performance: Compared to IPPO, which fails to converge at all (falling into extreme negative rewards) and the high variance oscillation of the ablation version without a prediction module, H-TP-MASNet, with its spatiotemporal predictive capabilities, successfully stabilizes the final reward at the global optimum level, achieving smooth tactical convergence.
[0033] This example reconstructs and quantitatively analyzes the evaluation results of 100 test rounds using a co-simulation platform, validating the model's tactical decision-making capabilities. Key metrics, such as... Figure 3 As shown; Proactive Interception: Faced with an absolute numerical disadvantage of 5 to 8, the Blue team completely abandoned the inefficient strategy of "passive tailing and pursuit" in traditional reinforcement learning. By predicting the Red team's maneuver trajectory through the PF-TP module, the Blue team executed the "deflection firing" tactic of preemptive positioning, ultimately achieving an overwhelming target kill rate of 76.0% (an average of 6.08 kills).
[0034] RadarEvasion: While delivering a lethal strike, the blue team successfully suppressed its radar exposure rate to an extremely low 18.1% through continuous attitude optimization and path planning. In contrast, the baseline algorithm without a prediction module blindly pursued the enemy, causing its radar exposure rate to soar to 39.2%, ultimately leading to its easy annihilation by the defense system.
[0035] Next, extreme interference and dynamic expansion capabilities were verified; real-time communication interference robustness was tested; and survival and kill indicators were monitored through artificially simulated packet loss interference, with results as follows: Figure 4As shown; Anti-interference time buffer: In the test round with an extreme communication packet loss rate of 80%, the blue team's intelligent agent used the trajectory prediction module as a "time information buffer" and relied on historical slices to extrapolate the positions of teammates and enemies, still maintaining a high kill rate of 74.0%, achieving a graceful tactical downgrade.
[0036] Zero-shot generalization expansion testing involved directly deploying the model trained in a 5-on-8 scenario to a large-scale battlefield for validation. The trajectory spatial relationships were as follows: Figure 5 , Figure 6 As shown; Zero-Shot generalization: Without any network fine-tuning, thanks to the aggregation property of MASNet permutation invariance, the cluster automatically adapts to the surge in the number of nodes, perfectly maintaining the complex space penetration formation and cooperative attack capability, proving its excellent dynamic scalability.
[0037] This invention features an innovative architectural paradigm: for the first time, it deeply integrates macroscopic dynamic target assignment (CM-DTA), microscopic multinomial trajectory prediction (PF-TP), and multi-agent hybrid action reinforcement learning (HAPPO) to construct a unified intelligent cooperative adversarial architecture of "global decoupled assignment + predictive situational awareness + continuous / discrete hybrid control".
[0038] This invention represents a breakthrough in tactical decision-making: it overcomes the severe lag bottleneck caused by the purely reactive decision-making of traditional reinforcement learning, endows UAV swarms with forward-looking spatiotemporal interception capabilities, and achieves a tactical leap from "passive tailing and pursuit" to "active prediction and interception," perfectly balancing high kill rate and low radar exposure rate in highly asymmetric environments.
[0039] This invention represents a breakthrough in communication resilience: it innovatively explores the potential of trajectory prediction models as "time information buffers," enabling self-reconstruction of the global battlefield state and smooth tactical degradation under extreme communication packet loss rates of up to 80% without the need for additional complex and massive network deployments, thus significantly improving combat survivability under strong electromagnetic interference.
[0040] This invention achieves a breakthrough in dynamic scaling generalization: relying on the permutation-invariant attention aggregation mechanism, it completely decouples the model from the fixed input dimension, breaks the limitation of traditional reinforcement learning models that need to be repeatedly trained for specific cluster sizes, and realizes seamless zero-shot scaling deployment of dynamic node adversarial systems from small-scale (5 vs. 8) to ultra-large-scale (20 vs. 30).
[0041] This invention significantly enhances tactical strike effectiveness and stealthy penetration capabilities; Experimental Comparison: In a highly asymmetric 5v8 air combat simulation environment, the H-TP-MASNet architecture proposed in this invention demonstrates a dominant tactical advantage. Experimental data shows that this invention achieves a target kill rate as high as 76.0% (an average of 6.08 kills), while suppressing the radar exposure rate to an extremely low 18.1%. In contrast, the fully distributed baseline IPPO algorithm has a kill rate and survival rate of 0.0%, while the ablation model without trajectory prediction (H-MASNetw / oTP), although highly aggressive, suffers from a radar exposure rate soaring to 39.2% due to blind pursuit, resulting in a kill rate of only 0.4%. Theoretical Support: This effect is attributed to the introduction of the polynomial trajectory prediction (PF-TP) module, which transforms the UAV swarm from a passive "tailgating" mode to an active "predictive interception" mode. By predicting the future position of highly maneuverable enemy targets and dynamic radar envelopes in advance, the agent can perfectly balance stealthy penetration and lethal firepower output.
[0042] The algorithm of this invention greatly improves training convergence stability and global optimization capability; Experimental Comparison: On the Episodic Reward curve during the training phase, the hybrid action architecture of this invention exhibits exceptionally stable convergence characteristics. In contrast, the baseline IPPO algorithm fails to converge at all, causing the cluster to fall into a suboptimal strategy with extreme negative rewards; the ablation model without a prediction module exhibits extremely high variance and instability during training.
[0043] Theoretical Support: The trajectory prediction module endows the neural network with spatiotemporal predictive capabilities, providing the model with a smooth and forward-looking state representation. This effectively overcomes the hysteresis problem of pure reactive networks when facing high-speed dynamic targets, guiding the policy network to converge quickly and stably to the globally optimal tactical strategy.
[0044] This invention provides extremely strong system resilience (robustness) under extreme communication interference. Experimental Comparison: In high-intensity communication packet loss scenarios simulating real electronic warfare (EW), this invention demonstrates superior graceful degradation capabilities. Experimental data shows that even under extremely harsh conditions with a random packet loss rate as high as 80%, the cluster's kill rate remains an astonishing 74.0%. Furthermore, at a 40% packet loss rate, the survival rate exhibits a counterintuitive peak increase (approximately 30%).
[0045] Theoretical Support: This exceptionally strong anti-interference capability stems from the trajectory prediction module's unique role as a "time information buffer." When a communication link is partially interrupted, the agent can use this module to extrapolate the future trajectories of lost teammates and enemies, thereby reconstructing a coherent global battlefield model from fragmented data. When global coordination is hindered, the agent will adaptively prioritize local evasion, thus improving its temporary survival rate.
[0046] This invention: Large-scale zero-shot generalization capability adaptable to dynamic node changes. Experimental Comparison: To verify the flexibility of the architecture, models trained only in 5 vs. 8 scenarios were directly deployed without any fine-tuning (Zero-Shot) to ultra-large-scale dynamic adversarial environments of 10 vs. 15 and even 20 vs. 30. The 3D projection results show that the cluster can still easily maintain complex spatial formations and continuous coordinated strike capabilities, exhibiting highly consistent tactical performance.
[0047] Theoretical support: This advantage stems from MASNet's internal permutation-invariant attention aggregation mechanism, combined with predictive filtering, which completely decouples the policy network from fixed input dimensions and specific cluster sizes. This mechanism allows agents to dynamically process information about any number of neighbors and enemies, endowing the model with enormous potential and seamless scalability in dynamic deployments of troops on real-world battlefields.
[0048] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.
Claims
1. A hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments, characterized by: A predictive hierarchical hybrid action control architecture (H-TP-MASNet) is proposed. A two-layer collaborative structure is constructed. The upper layer adopts a clustering-based many-to-many dynamic target assignment (CM-DTA) module to decouple the global strategy into local sub-team adversarial. The lower layer integrates the kinematic principle-based multinomial fitting trajectory prediction (PF-TP) mechanism into the multi-agent scalable network (MASNet) and is trained by the hybrid action proximal policy optimization (HAPPO) algorithm. A spatiotemporal prediction mechanism for proactive interception is constructed; a prediction-based hybrid action architecture is designed to enable swarm tactics to transform from passive tailing and pursuit attacks to proactive deflection and shooting interception; a lightweight polynomial fitting trajectory prediction (PF-TP) module is used to infer the short-term spatiotemporal motion trends of all agents by fitting recent trajectory history based on the physical inertia of the UAV. Construct a time caching mechanism to ensure extremely strong communication anti-interference capabilities; When the communication link is broken due to high-intensity interference, the agent uses the trajectory prediction module to extrapolate the future trajectories of the lost teammates and the enemy, and reconstructs a coherent global battlefield model through fragmented data. We construct a zero-shot generalization architecture that adapts to dynamic node changes; we integrate trajectory prediction with the Multi-Agent Scalable Network (MASNet) aggregator, and use a permutation-invariant aggregation architecture to handle the dynamic number of neighbor information, thus completely decoupling the policy from the fixed input dimension. Specifically, the following steps are included: Step 1: Modeling the kinematics and radar threat environment of unmanned aerial vehicles; Step 2: Upper-level strategy planning: Cluster-based many-to-many dynamic target assignment (CM-DTA); Step 3: Lower-level situational awareness: trajectory prediction based on polynomial fitting (PF-TP); Step 4: Lower-level feature fusion: Construct a multi-agent scalable network (MASNet); Step 5: Design and optimization of the hybrid action strategy controller (HAPPO); Step 6: Network training and parameter update.
2. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: Step 1 includes the following steps: Step 101: UAV Kinematic Modeling: Establish a three-degree-of-freedom kinematic model of a fixed-wing UAV; for the motion of the UAV in three-dimensional space, its state space is represented as follows: The kinematic differential equations are defined as follows: ; in, The three-dimensional spatial coordinates of the UAV in the inertial reference frame; Indicates the magnitude of speed; The inclination angle is the flight path angle. This is the heading angle; in addition, Representing gravitational acceleration; this set of kinematic differential equations is driven by three coupled control inputs: throttle overload coefficient. Pitch overload coefficient and roll angle ; Step 102: Radar Detection Probability and Threat Field Model: For models containing... In a threat environment with a fixed radar station, the blue team's drone is located... Cumulative detection probability at location The model is as follows: ; in, For drones to the first The Euclidean distance of the radar unit; The radar cross section (RCS) is the relative azimuth angle between the UAV and the radar line of sight (LOS). The function; in addition, and These are empirically determined radar constants, representing the shape coefficient of the detection probability curve and the generalized radar performance coefficient, respectively.
3. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: Step 2 includes the following specific steps: Adversarial clustering; using spatial proximity to group red team enemy aircraft clusters Divided into tactical clusters Each cluster This represents a localized threat focus, thereby decoupling the overall problem into localized squad-level confrontations; Define Boolean decision variables Indicates whether the blue team's drone i is assigned to the red team's cluster. Its cluster suppression effect Combining spatial distance advantage, angular attack advantage, and radar exposure avoidance: ; in, Represents a cluster The number of drones used by the Chinese Red team; and These are the range-based and angle-based tactical advantage functions of the blue team's UAV i against the red team's UAV j, respectively; the weight coefficients w1, w2, and w3 are positive weights used to balance the offensive geometric advantage and radar evasion performance. Global task optimization can be described as follows: ; Where Y represents all The constructed Boolean decision matrix; Assemble the blue team's drones; The total number of blue team drones; the constraint of equation (5) ensures that a sufficient red-blue force ratio is maintained in each local battle, and is determined by a pre-set minimum. and maximum Threshold constraints; Maximize the overall allocation effectiveness while constraining the troop ratio in local battles; determine the assigned clusters. Then, based on the distance and the target's quantified threat level, the specific primary target is dynamically selected. : ; in, Indicates the blue team's drone i With Red Team's drone j The Euclidean distance between them; ω is the quantified tactical threat value of target j; ω is the weighting coefficient; finally, the selected target As a key interface identifier, it provides guidance for the execution of lower-level tactics.
4. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: Step 3 includes the following specific steps: A quadratic polynomial is used to fit the recent historical trajectory of the enemy target to extract dynamic features; in the X-axis coordinate, the fitting model is... ; in, As a relative time variable, Corresponding to the current time t; These are the polynomial coefficients corresponding to the acceleration, velocity, and position terms fitted along the X-axis, respectively. The corresponding acceleration, velocity, and position polynomial coefficients are extracted using the finite difference method. ; Based on these coefficients, the predicted position and predicted velocity for the next moment are calculated as follows: ; By applying them independently to the y-axis and z-axis, the complete predicted state vector can be obtained: ; in, and They represent the first j The predicted three-dimensional position vector and predicted three-dimensional velocity vector of the red team's UAV.
5. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: Step 4 includes the following specific steps: Step 401: State Feature Embedding; For friendly UAVs Its kinematic state Through a friendly multilayer perceptron (MLP) Encode as feature vector : ; For enemy targets its original state Polynomial trajectory prediction and allocation indicator The parts are spliced together and then passed through the enemy's MLP network. Encoding as features : ; Here, Represents vector concatenation operation; indicator function The value is 1 when target j is the primary target assigned to the local UAV, and 0 otherwise. Step 402: Dual-stream multi-head attention and final state aggregation, based on machine features. For queries, a variable amount of local environment information is processed through parallel friendly and adversary multi-head attention networks. and The goal is to output aggregated features with invariant permutations, which are then concatenated into a comprehensive spatiotemporal state representation. .
6. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: Step 5 includes the following specific steps: Step 501: Specific formulas for the hybrid action space and policy output: Discrete tactical instructions With continuous flight control commands The specific space is defined as: ; Mode 0 is silent / cruising; Mode 1 is target lock; Mode 2 is firing. Corresponding to throttle / longitudinal overload, Corresponding to pitch overload, Corresponding roll angle; Dual-head network structure: Policy network Shared MASNet encoder parameters It then branches into two independent heads: The discrete head outputs the output classification distribution. ; The continuous output head outputs a diagonal Gaussian distribution. ,in and The mean and standard deviation vectors generated for the network are state-related, and I is the identity matrix; Under the conditional independence assumption given the state characteristics, the joint policy distribution can be decomposed as follows: ; Step 502: Adversarial Reward Mechanism ; in, and For predefined weighting coefficients; and These represent the rewards for attacking targets, penalties for radar evasion, and penalties for maneuvering smoothness, respectively.
7. The hierarchical hybrid action decision-making method for UAV swarms in complex electromagnetic environments according to claim 1, characterized in that: In step 6, the Hybrid Proximal Policy Optimization (HybridPPO) algorithm is used to optimize the policy parameters. Update, among which This represents the network parameters used for state evaluation (Critic); it should be noted that the trajectory prediction module is deterministic and does not perform gradient updates at this stage. The specific update process includes the following sub-steps: Step 601: Calculate the importance sampling ratio of the mixed PPO: Since the action space contains both discrete and continuous components, the importance of sampling at time step t is higher than that of continuous components. The calculation is performed using the joint likelihood, and the formula is as follows: ; Step 602: Construct the surrogate objective function for pruning: To limit excessive policy updates, a pruned proxy objective function is used. : ; in, This represents the expected experience on the experience batch. It is a generalized advantage estimate (GAE) obtained from the value (Critic) network evaluation. These are the pruning hyperparameters of the PPO algorithm; Step 603: Calculate the entropy regularization term: To prevent premature model convergence and encourage policy exploration, entropy regularization terms are introduced for both discrete and continuous action heads. : ; in, Shannon entropy represents the distribution. and These represent the entropy coefficients of the discrete action space and the continuous action space, respectively. Step 604: Calculate the total reinforcement learning loss and perform end-to-end updates: The total reinforcement learning loss is obtained by weighted summation of the policy loss, value function loss, and entropy reward. : ; ; in, It is the mean squared error (MSE) loss of the value predicted by the Critic network. This is the weighting coefficient for value loss; End-to-end update backpropagation mechanism: shared MASNet encoder parameters Received by total reinforcement learning loss The backpropagation gradient is used to update the parameters; during this process, although the parameters of the trajectory prediction module remain fixed, the encoder can adapt through gradient updates, thereby effectively utilizing multinomial prediction features. To identify states where a specific tactical mode needs to be adopted.