Unmanned aerial vehicle cluster confrontation game method based on clustering average field reinforcement learning

By using a clustered mean field reinforcement learning method, the UAV swarm is divided into several clusters for state and action representation, which solves the problems of dimensionality curse, non-stationarity and incomplete information in large-scale UAV swarm adversarial operations, and realizes efficient and robust online decision-making and engineering applications.

CN122064093APending Publication Date: 2026-05-19CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning methods face bottlenecks in large-scale UAV swarm adversarial scenarios, such as the curse of dimensionality, environmental non-stationarity, incomplete information, and homogeneous/heterogeneous contradictions, leading to high computational complexity, difficulty in real-time decision-making, policy mismatch, and lack of engineering scalability.

Method used

A clustered mean-field reinforcement learning method is adopted to divide the UAV swarm into several clusters. Through cluster-level state and action representations, a cluster-level value function is constructed, and parameter learning is carried out under a centralized training and distributed execution framework. Combined with adversary strategy prediction and payoff function design, cluster-level online decision-making and replanning are realized, which can adapt to battlefield changes and make robust decisions using incomplete information.

Benefits of technology

It effectively reduces complexity, enables efficient and robust decision-making for large-scale UAV swarms in complex environments, adapts to battlefield changes, meets real-time requirements, and possesses scalability and stability for engineering implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064093A_ABST
    Figure CN122064093A_ABST
Patent Text Reader

Abstract

The invention relates to the field of unmanned aerial vehicle cluster control, and discloses an adversarial game method based on clustering average field reinforcement learning. The method comprises the following steps: firstly, performing dynamic clustering according to spatial proximity, task similarity and link quality to form a cluster-level state and an average action; under a centralized training and distributed execution system, average field modeling is introduced, probability prediction is carried out on opponent intentions in combination with hidden Markov and other time sequence models, and deep reinforcement learning is driven to solve a cluster-level strategy; and mapping the cluster-level macroscopic instruction into intra-cluster individuals for collaborative execution through hierarchical control. The system comprises a clustering module, an opponent prediction module, a learning module, an income and constraint fusion module, a decision issuing module, a communication self-adaption module, a safety verification module and an execution feedback module. According to the scheme, under the conditions of incomplete information and limited communication, the calculation complexity is reduced, non-stability is relieved, the decision real-time performance and robustness are improved, and the method is suitable for defense penetration, interception, suppression and other confrontation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) swarm control technology, and in particular to a method and simulation verification system for intelligent game decision-making in complex, dynamic, and non-stationary adversarial environments for UAV swarms. Background Technology

[0002] Unmanned aerial vehicle (UAV) swarms, as a crucial carrier of swarm intelligence in missions such as airspace warfare, reconnaissance and surveillance, electronic warfare, disaster relief, and emergency communications, exhibit comprehensive characteristics including multi-agent capabilities, intense competition, high maneuverability, imperfect information, and strong constraints. Compared to single-unit intelligence, swarms must complete multi-objective collaborative decision-making in uncertain and highly perturbed environments, and maintain stable and efficient group behavior under constraints such as limited communication, energy, airspace, and strict real-time requirements. In high-intensity competitive scenarios such as military confrontations, swarm operation inevitably involves interaction and game theory with enemy agents, essentially constituting a complex multi-agent dynamic game problem.

[0003] To enhance the autonomous combat effectiveness of swarms, mainstream technical approaches typically model them as multi-agent systems (MAS) and introduce methods such as multi-agent reinforcement learning (MARL) to train individual UAVs (agents) within the swarm with strategies executable in adversarial scenarios. However, existing MARL theories generally face the following key bottlenecks when dealing with large-scale, highly adversarial, strongly constrained, and partially observable real-world scenarios:

[0004] (1) Dimensional disaster (exponential growth of state-action space).

[0005] Traditional MARL often treats each UAV as an independent learning entity. As the cluster size N increases, the joint state and joint action space of the system expands exponentially. This leads to a sharp increase in the demand for computing and storage resources, a surge in the complexity of training samples and exploration costs, difficulty in convergence, and even failure, making it difficult to meet the real-time requirements of second-level online decision windows.

[0006] (2) Environmental non-stationarity.

[0007] In adversarial games, the opponent is also continuously learning and adapting; from the perspective of any friendly drone, the combined actions of teammates and opponents constitute its "environment." As the opponent's strategy evolves, this environment becomes unstable for the individual drone, breaking the fundamental assumptions of many single-agent reinforcement learning algorithms. The direct consequence is that learned strategies are prone to rapid mismatch with changes in the opponent, leading to sudden performance drops or misjudgments.

[0008] (3) Incomplete information and partially observable (POMDP).

[0009] Real-world battlefields are characterized by "fog of war," limitations in communication distance and bandwidth, electromagnetic interference and deception, and the limitations of sensor detection capabilities, resulting in incomplete, noisy, and time-delayed state information. Game theory and decision-making models based on the assumption of perfect information are difficult to implement directly; and a unified and effective engineering solution for achieving robust perception, reasoning, and decision-making under imperfect information conditions still lacks a clear framework.

[0010] (4) The contradiction between homogeneity and heterogeneity of models (the balance between scalability and individual differences).

[0011] If a homogeneous strategy network is adopted for all drones, it will be convenient for training and deployment, but it will be difficult to reflect the strategy heterogeneity caused by different tactical positions and task divisions, and it will be easy to "one-size-fits-all"; on the contrary, if a heterogeneous network is designed for each individual, the design, training and maintenance costs will increase dramatically with the scale, and the project will not be scalable.

[0012] In summary, existing MARL-based UAV swarm adversarial decision-making technologies still suffer from fundamental limitations in areas such as the curse of dimensionality, nonstationarity, incomplete information, and homogeneous / heterogeneous tradeoffs, hindering their technological maturity and practical effectiveness. The industry urgently needs a new theoretical and engineering framework that can effectively reduce dimensionality and decompose collaborative complexity structurally, mitigate nonstationarity and strategy mismatch during adversarial operations, and achieve robust online decision-making and reproducible experimental evaluation loops for partially observable and constrained communication. This framework will drive substantial progress and engineering applications of UAV swarm intelligent adversarial technology. Summary of the Invention

[0013] This invention addresses several bottlenecks in adversarial scenarios involving large-scale UAV swarms, including the curse of dimensionality (the exponential growth of the joint state and action space as the swarm size increases, making learning and planning difficult to complete within a real-time window); environmental non-stationarity (adversary strategies change with game evolution, leading to mismatch of learned strategies); incomplete information (limited and noisy observations, latency, and packet loss, making stable decision-making difficult under partially observable conditions); and the homogeneity / heterogeneity contradiction (homogeneous networks struggle to reflect task differences, while heterogeneous networks impose unscalable engineering burdens). This invention aims to provide an adversarial game method and system based on clustered mean-field reinforcement learning, achieving structured dimensionality reduction with clusters as the basic decision-making unit, mitigating the dimensionality explosion and non-stationarity of multi-agent interactions with mean-field approximation, robust online decision-making and replanning under limited communication and partially observable conditions, and a unified framework for engineering deployment and reproducible experimental evaluation. This invention provides a UAV swarm adversarial game method based on clustered mean-field reinforcement learning, comprising the following steps:

[0014] S1: Scene modeling and cluster generation. The set of drones is denoted as U = {u...} iBased on at least one or a combination of spatial proximity, task similarity, and communication link quality, U is divided into several clusters C = {C...}. k}, obtain cluster-level state S k Local observation of individuals within a cluster i With available action set A i Clustering employs threshold-based or density-based criteria, and allows reclustering during operation according to preset periods or trigger conditions;

[0015] S2: Cluster-level state and average action representation. For cluster C k Internal individual action a i We perform weighted aggregation to obtain the average action representation within the cluster:

[0016]

[0017] Where w i Z represents a non-negative weight related to individual confidence, link quality, or task priority. k The normalization factor is used. Cluster-level actions are denoted as A. k It can be taken from the cluster instruction template set or generated by a continuous parameterization strategy;

[0018] S3: Mean-field reinforcement learning modeling. Constructing a cluster-level value function with cluster-level states, cluster-level actions, and average action representations as inputs:

[0019] Parameter learning is performed under a centralized training, distributed execution (CTDE) architecture. The training sample quadruple is (S k A k ,r k ,S k The loss function can take the following values:

[0020]

[0021] in Here, γ represents the target network parameters, and γ is the discount factor. Experience playback and target network updates are performed at set intervals.

[0022] S4: Adversary Cluster-Level Policy Probability Prediction. Based on historical interactions and current observations, estimate the opponent's policy distribution π at the cluster level. opp (A -k |S -k The prediction model can employ at least one of prior instance library screening and temporal probability inference to output a cluster-level action probability vector for the next decision slot, which can be used for target value calculation and policy improvement in S3.

[0023] S5: Profit Function Design and Constraint Fusion. Constructing the profit function:

[0024] r k =αJ task -βJ cons -γJ risk -δJ fair (4)

[0025] J task Characterizing task achievement (such as at least one of penetration rate, interception rate, and coverage), J cons Characterizing resource consumption (such as at least one of energy, mileage, and payload consumption), J risk Characterizing safety risks (such as minimum separation violations or no-fly zone violations), J fair Characterizes resource fairness among clusters; α, β, γ, δ are non-negative weights. Constraints are incorporated using either soft penalties or feasibility projection.

[0026] S6: Cluster-level adversarial action sequence generation and delivery. Based on the current Q... k With π opp Generate cluster-level action sequences using greedy or temperature-controlled strategies. Under communication constraints, these instructions are mapped to executable commands for individual cluster members. Command delivery employs at least one of the following methods: fragmentation, compression, and acknowledgment retransmission, to ensure timely delivery.

[0027] S7: Online replanning and triggering mechanism. Replanning is triggered when any of the following conditions occur:

[0028] 1) The difference between the observed and predicted distributions of opponent cluster-level strategies exceeds a threshold;

[0029] 2) Link quality indicators deteriorate beyond the threshold;

[0030] 3) The mission situation assessment is below the threshold.

[0031] Replanning is completed within a given real-time window, updating cluster-level policies and instruction sequences;

[0032] S8: Safety Constraint Verification and Degradation Handling. The generated action sequence undergoes safety verification, including minimum safe distance, no-fly zones, and collision risks; actions that do not meet the constraints are downgraded using at least one of the following methods: replacement, speed / altitude limiting, or delayed execution.

[0033] S9: Execution Feedback and Parameter Adaptation. Collect execution results and link statistics to form cluster-level feedback quantities, which are used to update experience replay, opponent prediction models, and communication adaptive parameters (such as compression ratio and synchronization frequency).

[0034] Compared with existing technologies, the method proposed in this invention has significant advantages. First, by using clustering and mean-field theory, it successfully decouples problem complexity from cluster size, effectively avoiding the curse of dimensionality. This allows the algorithm to be efficiently applied to large-scale UAV swarms and exhibits good scalability. Second, mean-field theory smooths out the non-stationarity of the environment. Combined with intention prediction based on Hidden Markov Models and dynamic selection using targeted neural networks, the decision-making process can adapt to battlefield changes and opponent strategy adjustments in real time, demonstrating strong robustness and situational adaptability. Furthermore, this method does not rely on perfect global information but effectively utilizes incomplete and uncertain battlefield information through probabilistic models to achieve high-quality decision-making under near-real combat conditions. Finally, the entire method has a clear flow and well-defined module functions. Its hierarchical control architecture facilitates design, implementation, testing, and upgrades, meeting the requirements for the engineering implementation of complex systems. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below:

[0036] Figure 1 This is a schematic diagram of the system functional architecture and data flow of a drone swarm adversarial game method according to an embodiment of the present invention.

[0037] Figure 2 This is a conceptual diagram of the UAV swarm clustered adversarial game model in this invention.

[0038] Figure 3 This is a schematic diagram of the deep Q-learning network architecture used in the cluster-level game decision-making module of this invention.

[0039] Figure 4 This is a schematic diagram of the distributed information fusion process used for predicting adversary intentions in this invention. Detailed Implementation

[0040] This invention presents a method for adversarial game theory involving drone swarms based on clustered mean-field reinforcement learning (MFRL). The overall process and module interactions are as follows: Figure 1 As shown. The process includes:

[0041] S1. Scene modeling and constraint setting;

[0042] Establish a partitioned model of the mission area (which may include observation area / defense area / penetration area, etc.), set hard constraints such as no-fly zone, minimum safe interval, speed / turn rate / altitude, and define communication constraints (throughput, latency, packet loss rate threshold, etc.).

[0043] S2. Cluster generation;

[0044] Upon mission initiation, the Red Team system immediately performs dynamic clustering. Based on the initial positions of the 30 drones, the K-Means clustering algorithm is used to divide them into 5 clusters (C1 to C5), with approximately 6 drones in each cluster. Within each cluster, the drone with the strongest communication capabilities is elected as the cluster leader, responsible for external communication and internal coordination. This process reduces the original complex 30 vs. 20 game to a macroscopic game of 5 pairs of "several objectives," the concept of which is as follows: Figure 2 As shown.

[0045] Clustering of UAV assemblies is performed based on at least one or a combination of spatial proximity, task similarity, and link quality to obtain clusters. Clustering can be performed using a threshold method (such as a distance threshold ε). d Link quality threshold ε q (or density clustering.) Output cluster-level state S k (Including intra-cluster statistics, neighborhood summaries, and constraint summaries). For individual actions a within the cluster... i After weighted aggregation, we get:

[0046]

[0047] Where the weight w i It can be related to individual confidence, link quality, or task priority, and updated at fixed intervals or triggered by events;

[0048] Cluster-level value learning in S3.CTDE training;

[0049] Within the centralized training, distributed execution (CTDE) framework, a cluster-level Q-function is constructed, and the training sample quadruples are (S k A k ,r k ,S k Using empirical replay with the target network, the loss can be taken as the squared TD error; the soft update coefficient τ of the target network is preferably between 0.90 and 0.999. The actors in the execution phase rely only on the information available to their own cluster (S). k , (and short-term historical summary);

[0050] S4. Opponent cluster-level strategy prediction;

[0051] Based on historical interactions and current observations, output the opponent cluster-level policy probability vector. It can be implemented using a prior instance library and filtering, or using a time-series probabilistic model (such as HMM / Bayesian filtering / RNN);

[0052] S7. Action Generation and Command Issuance:

[0053] Based on the current Q k and Cluster-level action sequences are generated using greedy algorithms and temperature sampling strategies. Actions are sent to individuals within the cluster within a given instruction cycle through adaptive communication mechanisms (compression, fragmentation, retransmission, rate control). Figure 3 The cluster-level Q-network structure is illustrated: the upper box is... The computational logic, with inputs including: cluster-level states s t (including neighborhood summary), average action within cluster Optional prior features (such as task stage identifiers); output is the Q-value of multiple candidate cluster-level actions; the lower network consists of an input layer, hidden layers, and an output layer, with 2–4 hidden layers and 128–512 units per layer, using the target network and experience replay in parallel for stable training; TD loss is used during training with a learning rate of 1×10⁻⁶. -4 ~5×10 -4 Batch size 32–256, experience pool 10 4 ~10 6 Samples; construct global Q tot =f(Q1,...,Q) K Monotonic mixing of clusters is used to improve inter-cluster credit allocation.

[0054] S8. Strategic replanning;

[0055] Replanning is triggered when any of the following conditions occur:

[0056] 1) The difference between the observed and predicted distributions of opponent cluster-level strategies exceeds a threshold;

[0057] 2) Link quality indicators deteriorate beyond the threshold;

[0058] 3) The mission situation assessment is below the threshold.

[0059] The strategy is then updated within the real-time window (e.g., ≤3s) and returned to the "Decision Issuance Module" via a dashed arrow.

[0060] Figure 4 The evaluation process of "multi-cluster-multi-target" is presented; the upper part is the set of M×S solution targets in the blue team's UAV swarm; the middle part is the multiple local evaluation clusters (1...L) of the red team, each cluster based on its own observation X. l Output the posterior probability P(H) of the hypothesis / target. m |X l The lower "Posterior Probability Fusion" unit fuses the posterior probabilities of each cluster to obtain a comprehensive evaluation Y (such as the suppression / interception ratio, mission success rate range, etc.); this Y serves as... Figure 1 J, which integrates "rewards and constraints" task It is a component and one of the triggers for "strategy replanning" (dashed loop identifier information feedback).

Claims

1. A method for adversarial game theory involving drone swarms based on clustered average field reinforcement learning, characterized in that, include: S1: The drone swarm is divided into clusters based on spatial proximity, task similarity, and link quality, resulting in multiple clusters and their corresponding cluster-level states; S2: Construct a cluster-level value function with the average representation of cluster-level state, cluster-level action, and individual actions within the cluster as input, and train it using mean-field reinforcement learning to obtain a cluster-level adversarial strategy; S3: For each cluster, construct a cluster-level reinforcement learning model based on the average representation of individual actions within the cluster to obtain a value function with cluster state, cluster action, and average action as inputs; S4: Generate adversarial action sequences at the cluster level according to the profit function that includes task success rate and resource consumption constraints, and distribute them to the corresponding individuals within the cluster for execution; S5: Dynamic synchronization control module, which adaptively adjusts node coupling weights based on Hamiltonian minimization algorithm to minimize the overall network collaborative potential energy.

2. Under the conditions of limited communication and real-time constraints, the cluster-level adversarial strategy is reprogrammed online to cope with changes in the environment and the opponent's strategy.

3. The method according to claim 1 or 2, wherein, The average is represented as a weighted average of individual actions within the cluster, with the weighting coefficients related to individual confidence / link quality.

4. The method according to any one of claims 1 to 3, wherein, The opponent policy prediction uses a hidden Markov model / Bayes update / deep network to output the cluster-level policy probability for the next time period.

5. The method according to any one of claims 1 to 4, wherein, The revenue function also includes a fairness coefficient / penalty term to constrain inter-cluster resource consumption.

6. The method according to any one of claims 1 to 5, wherein, The real-time constraints are: cluster-level solution delay ≤ 2 seconds, online replanning cycle ≤ 3 seconds.

7. The method according to claim 1, characterized in that, The revenue function includes a task success rate term, a resource consumption penalty term, and an inter-cluster fairness coefficient term. The task success rate term is measured by the proportion of suppressed or intercepted targets, the resource consumption penalty term is measured by the proportion of energy consumption or path length, and the inter-cluster fairness coefficient term is used to limit the excessive use of resources by individual clusters for a long time.

8. The method according to claim 1, characterized in that, The triggering conditions for online replanning include: the observed difference between the opponent cluster-level policy and the predicted distribution exceeds the third threshold, or the change in network link quality exceeds the fourth threshold, or the task situation assessment is below the fifth threshold.

9. The method according to claim 1, characterized in that, The method further includes a safety constraint execution step, which is used to perform constraint verification on minimum safety interval, no-fly zone crossing and collision risk after the adversarial action sequence is generated, and to replace or downgrade actions that do not meet the constraints.

10. The method according to claim 1, characterized in that, The method further includes a cluster-level resource allocation step, used to constrain the allocation of the number of UAVs and payload types according to the payoff function and the current policy confidence when multiple clusters compete for the same task objective.