Multi-target-based unmanned aerial vehicle group collaborative confrontation method, system and application

By improving the encoding and decoding mechanism and the dynamic task weighting strategy, the problems of unpredictable state information and credit allocation in UAV swarm cooperative combat are solved, the situational reasoning ability of UAV swarms and the stability of task allocation are improved, and more efficient cooperative combat effects are achieved.

CN121806982APending Publication Date: 2026-04-07ZHEJIANG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for cooperative combat against unmanned aerial vehicle (UAV) swarms suffer from problems in predicting state information, credit allocation, and task assignment, which affect decision-making capabilities and training effectiveness.

Method used

Based on the QMIX reinforcement learning algorithm, an improved encoding and decoding mechanism is designed, and a network architecture with multi-objective reward allocation and dynamic task weights is constructed. Key situational information in complex and incomplete scenarios is extracted through encoder and decoder networks, and the cooperative combat process of UAV swarms is optimized by combining multi-objective reward allocation and dynamic task weight strategies.

Benefits of technology

It enhances the situational reasoning capabilities of drones, alleviates the problem of credit allocation imbalance, reduces the limitations of algorithmic decision-making, and improves learning efficiency and task completion stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121806982A_ABST
    Figure CN121806982A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-target-based unmanned aerial vehicle group collaborative confrontation method and system and application, and the method comprises the steps: constructing a beyond-visual-range collaborative confrontation scene based on an unmanned aerial vehicle group, obtaining a collaborative confrontation task, and defining a confrontation evaluation index; key situation information in a complex and incomplete scene is extracted based on an improved network; designing a strategy of multi-target reward allocation and dynamic task weight, deploying the strategy to a control end of the unmanned aerial vehicle, executing the task, evaluating and optimizing the strategy, and continuously executing unmanned aerial vehicle group collaborative confrontation until the task is completed; the system obtains observation information and positioning information based on a first unmanned aerial vehicle group and a second unmanned aerial vehicle group through a PC terminal, and a processor outputs a strategy action instruction to a control terminal of a corresponding unmanned aerial vehicle; the method is applied to multi-unmanned aerial vehicle cooperative confrontation in a multi-target complex environment. The situation reasoning capability of the unmanned aerial vehicle based on local observation is enhanced, the problem of credit distribution imbalance of a traditional fixed weight method is relieved, and the problem of limitation of an algorithm in target decision making is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of calculation, estimation or counting, and in particular to a multi-target-based UAV swarm cooperative confrontation method and system and application. BACKGROUND

[0002] UAV swarm cooperative confrontation refers to an intelligent combat system constructed by multiple UAVs relying on cooperative confrontation algorithms. Through information sharing, task division and dynamic cooperation, the system can break through multiple defense systems such as enemy radar, electronic jamming and missile interception, and complete tasks such as reconnaissance, attack and interference. Compared with traditional single-machine combat, the advantage of UAV swarm lies in real-time decision-making ability based on algorithms and highly flexible cooperative strategies. In the process of cooperative confrontation, algorithms play a core role. Through distributed perception fusion algorithms, each UAV can share target information, environmental data and battlefield situation detected by itself, achieve overall information advantage, rely on task allocation and path planning algorithms to dynamically coordinate each UAV to perform different roles such as main attack, decoy attack, interference and induction, thereby maximizing combat effectiveness in a complex electromagnetic environment; based on game theory and reinforcement learning decision-making algorithms, the UAV swarm can autonomously select the optimal penetration path and combat strategy. In recent years, reinforcement learning decision-making algorithms can optimize strategies in real time in highly dynamic and strongly confrontational battlefields through continuous interaction with the environment and adaptive learning, and are widely applied in the field of UAV swarm cooperative confrontation.

[0003] Chinese Patent No. CN120456119A discloses a multi-task-oriented UAV swarm group assembly method, which adopts a hierarchical information collection and diffusion mechanism to divide the nearest unit information from the enemy into granular size, and then based on the ally information, the method can cope with complex environments with intensive tasks; however, it often needs to be based on the case where global situation information can be observed.

[0004] Chinese Patent No. CN115755949A discloses a multi-UAV formation cluster control method based on multi-agent deep reinforcement learning, which introduces a graph attention mechanism based on the MADDPG algorithm to optimize the learning process of multiple UAVs for the task of autonomous aggregation of heterogeneous UAV swarms. However, this method needs to be implemented through unified modeling and does not have strong generalization, and cannot solve problems such as task allocation.

[0005] A Chinese patent with publication number CN119882773A discloses a multi-agent based unmanned aerial vehicle formation penetration method and its application. By decomposing the penetration task into multiple subtasks and adopting a hierarchical reinforcement learning strategy, the training effect of multiple unmanned aerial vehicles in a multi-target environment is effectively improved. Although this method can better solve the task allocation problem, credit allocation may occur during the training process, causing some unmanned aerial vehicles to fail to achieve good training results and affecting the overall decision-making ability. SUMMARY

[0006] To solve the problems of state information being difficult to predict, credit allocation and task allocation in the process of unmanned aerial vehicle group cooperative confrontation, an improved encoding and decoding mechanism is designed based on the QMIX reinforcement learning algorithm, a network architecture integrating multi-target reward allocation and dynamic task weight is constructed, and the learning efficiency and application value of the algorithm are improved.

[0007] The technical scheme adopted by the present application is a multi-target based unmanned aerial vehicle group cooperative confrontation method, which constructs a beyond-visual-range cooperative confrontation scene based on a unmanned aerial vehicle group, obtains a cooperative confrontation task, and defines confrontation evaluation indexes; based on an improved network, key situation information in a complex incomplete scene is extracted;

[0008] A strategy of multi-target reward allocation and dynamic task weight is designed, which is deployed to the control end of each unmanned aerial vehicle of the own side, executes the task, evaluates and optimizes the strategy, and continuously executes the unmanned aerial vehicle group cooperative confrontation until the task is completed.

[0009] Preferably, the cooperative confrontation task is a Markov game tuple of unmanned aerial vehicle group confrontation :

[0010]

[0011] wherein, is the number of unmanned aerial vehicles of the own side, is battlefield situation information, is the action of each unmanned aerial vehicle, is a state transition function, is a discount factor, is a reward function of each unmanned aerial vehicle.

[0012] Preferably, based on the complex incomplete battlefield situation information , a state information set s i = [ s s e l f i , s a l l y i , s e n e m y i , s a t t a c k i ] is obtained, wherein, is part of the observation information of the unmanned aerial vehicle , This is part of the observation information from our own drones. This is part of the observation information obtained from detecting enemy drones. The means of attack by enemy drones, such as missiles launched by the enemy;

[0013] The state information set of all UAVs is normalized and input into the improved network to extract state information parameters in the UAV cooperative adversarial environment, thereby obtaining high-quality state information.

[0014] Preferably, the improved network includes an encoder and a decoder connected in sequence;

[0015] The encoder network includes a sequentially connected fully connected layer and several LSTM layers, which perform time-step feature extraction and information encoding on the state information.

[0016] The decoder includes a self-attention module, an LSTM layer, and two fully connected layers arranged in sequence.

[0017] Preferably, the strategy for multi-objective reward allocation and dynamic task weighting is as follows: For time t, the drone The observation vector is input, and the task weights for the UAV's offensive and defensive missions are dynamically adjusted. For time t, the drone The normalized observation vector, For time t, the drone The time-series reward characteristics.

[0018] Preferably, obtain Offensive and defensive reward sequences at each time step The network adjusts the task weights for offensive and defensive tasks through dynamic task weight evaluation. and ;

[0019] based on Step reward sequence, calculate the average reward for corresponding offensive and defensive tasks. ,variance and slope Let n be 1 or 2 corresponding to offense or defense, and obtain f r , i t = [ μ o b j 1 , i t , σ o b j 1 , i t , θ o b j 1 , i t , μ o b j 2 , i t , σ o b j 2 , i t , θ o b j 2 , i t ] ;

[0020]

[0021] in, Time step index within the time window Time Index The average value, This represents the average reward.

[0022] Preferably, drones exist Attack reward is

[0023]

[0024] in, This is the reward range control coefficient. This is the distance attenuation coefficient. For drones The Euclidean distance to the nearest enemy drone. For the enemy's nearest drone, A collection of enemy drones;

[0025] drones exist The reward for constant defense is

[0026] r o b j 2 , i t = γ ⋅ [ 1 − t a n h ( 1 N − 1 ∑ a j ϵ A \ { a i } ‖ p a i t − p a j t ‖ 2 ) ]

[0027] in, This represents the total number of friendly drones. As the reward coefficient, For normalization function, To extract from drone set A, except All drones except those mentioned above;

[0028] Offensive and defensive rewards are redistributed, and joint rewards are updated based on the mission weights of drone offensive and defensive missions.

[0029] Preferably, a multi-objective local Q-value network is constructed, and the strategy is evaluated and optimized based on the dynamically adjusted task weights of the UAV attack and defense tasks.

[0030] A multi-target-based UAV swarm cooperative combat system, the system comprising:

[0031] The first and second drone swarms each include several drones. Each drone is equipped with a radar module, a positioning module, and an attack module. The radar module collects observation information, the positioning module obtains the positioning information of the corresponding drone, and the attack module is used for attack. A control terminal is provided in conjunction with the radar module, positioning module, and attack module.

[0032] On the PC side, in conjunction with the first and second UAV swarms, a system is set up including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the aforementioned multi-target UAV swarm cooperative combat method.

[0033] The PC-side obtains observation and positioning information based on the first and second UAV swarms, and the processor outputs policy action instructions to the corresponding UAV control terminal.

[0034] An application of the aforementioned multi-target-based UAV swarm cooperative combat method is demonstrated in multi-UAV cooperative combat under complex multi-target environments.

[0035] This invention relates to a method, system, and application of multi-target cooperative combat by a drone swarm. It constructs a beyond-line-of-sight cooperative combat scenario based on a drone swarm, acquires cooperative combat tasks, and defines combat evaluation indicators. It extracts key situational information in complex and incomplete scenarios using an improved network. It designs a strategy for multi-target reward allocation and dynamic task weighting, deploys the strategy to the control terminal of each drone, executes the task, evaluates and optimizes the strategy, and continuously executes drone swarm cooperative combat until the task is completed. The system uses a PC to obtain observation and positioning information based on the first and second drone swarms, and the processor outputs strategy action instructions to the control terminals of the corresponding drones. The method is applied to multi-drone cooperative combat in complex multi-target environments.

[0036] The beneficial effects of this invention are as follows:

[0037] (1) Enhance the situational reasoning capabilities of UAVs based on local observations;

[0038] The temporal features of the observation sequence are extracted by the encoder network, and the intermediate state sequence is decoded into the output sequence after the attention is calculated by the decoder network, which provides high-level semantic representation for subsequent multi-UAV reinforcement learning.

[0039] (2) Alleviating the credit allocation imbalance problem of the traditional fixed-weight method;

[0040] By decomposing global rewards into individual task dimensions and combining overall contribution assessment, the dynamic contribution of drones in different tasks is quantified. Furthermore, a dynamic task weight network is constructed to adaptively adjust task weight allocation based on local observations and historical rewards.

[0041] (3) Reduce the limitations of the algorithm in target decision-making;

[0042] By designing a target-oriented local Q-value network, the motion value function of each UAV under different targets is estimated. At the same time, a parallel Q-value fusion network is combined to realize the efficient calculation of the joint motion value function of multiple targets. After being weighted by task weights, the fusion network and policy network are jointly guided to update, effectively expanding the applicability of the algorithm in multi-UAV multi-task scenarios. Attached Figure Description

[0043] Figure 1 This is a flowchart of the method of the present invention;

[0044] Figure 2 The diagram shows the network structure of the encoder and decoder of the present invention, where (a) is the encoder network and (b) is the decoder network.

[0045] Figure 3 This is a framework diagram of the dynamic multi-task weight algorithm of the present invention;

[0046] Figure 4 This is a diagram showing the experimental application effect of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0048] This invention relates to a multi-target-based UAV swarm cooperative adversarial method, which constructs a beyond-line-of-sight cooperative adversarial scenario based on UAV swarms and obtains cooperative adversarial tasks and defines adversarial evaluation indicators; and extracts key situational information in complex and incomplete scenarios based on an improved network.

[0049] Here, "beyond visual range" refers to the situation where the enemy drone is outside the strike range from the initial position. In this case, the friendly drone explores the environment by interacting with it through reinforcement learning, finds the enemy drone's position, and approaches it in the shortest time to attack it. The so-called "beyond visual range" is an objectively existing distance that can be directly set and obtained.

[0050] "Complex and incomplete" means that the information dimension of the environment is very large. The observation dimension of a single UAV is 33 dimensions, which is the "complexity" of the environment. "Incomplete" means that the initial position of our UAV has a limited observation distance from the enemy UAV, and we cannot obtain all the information of the enemy aircraft. Therefore, the information obtained is "incomplete".

[0051] Design a strategy for multi-objective reward allocation and dynamic task weighting. After the strategy is deployed to the control terminal of each UAV, the task is executed, the strategy is evaluated and optimized, and the UAV swarm cooperative combat is continuously executed until the task is completed.

[0052] This invention extracts features from observation sequences through an encoder network and uses a decoder network combined with an attention mechanism to focus on and interactively model key information in the UAV's observation situation, thereby enhancing the UAV's reasoning ability using local observation sequences.

[0053] By decomposing the global reward into individual task dimensions and combining time-series contribution evaluation to quantify the dynamic contribution of drones in different tasks, we can solve the problems of reward sparsity and allocation bias in traditional methods and promote the stability of the strategy optimization process.

[0054] The method of the present invention includes the following steps:

[0055] (1) Based on the beyond-line-of-sight cooperative combat scenario of UAV swarms, obtain cooperative combat tasks and define combat evaluation indicators;

[0056] (2) Extract key situational information in complex and incomplete scenarios based on improved networks;

[0057] (3) Design a strategy for multi-objective reward allocation and dynamic task weighting, and then deploy the strategy to the control terminal of each UAV to execute the task;

[0058] (4) Evaluate and optimize the strategy, and continue to execute the drone swarm cooperative confrontation until the mission is completed.

[0059] The method will be described in detail below with reference to the embodiments.

[0060] In this embodiment, a two-dimensional scenario of red and blue forces confrontation is constructed using the Micius combat platform. Each side is equipped with four isomorphic UAVs, each UAV having radar detection capabilities, capable of detecting enemy UAVs and missiles within its range and obtaining corresponding coordinate information. Each aircraft is equipped with 10 missiles, which can attack targets within its attack range. The initial orientation of the red and blue forces is towards each other, and their flight altitudes are consistent.

[0061] (1) Based on the beyond-line-of-sight cooperative combat scenario of UAV swarms, obtain cooperative combat tasks and define combat evaluation indicators;

[0062] First, a mathematical model of the adversarial task is constructed to obtain the distance vectors between the red and blue teams during the cooperative combat process involving multiple drones. Euclidean distance vector Attack angle and the angle of attack :

[0063]

[0064] in, and The two-dimensional coordinates representing the red and blue drones, and The drones representing the red and blue teams respectively were in shaft and The speed on the axis can be controlled by adjusting the speed increment of each drone. and heading angle increment To perform real-time control, in this example Δ v i ∈ [ − 5 , 5 ] , Δ ω i ∈ [ − 1 0 ∘ , 1 0 ∘ ] .

[0065] The cooperative adversarial task is a Markov game tuple of drone swarm adversarial combat. :

[0066]

[0067] in, This refers to the number of friendly drones; in this example, it is 4. This is battlefield situational information, presented in this embodiment as 33 dimensions. For the actions of each drone, This is the state transition function. As a discount factor, it generally satisfies In this embodiment, 0.95 is used. The reward function for each drone.

[0068] In this embodiment, a multi-level scoring system is constructed. Firstly, the number of surviving aircraft at the end of each round is used as the primary scoring criterion; the side with more surviving aircraft wins. If both sides have the same number of surviving aircraft, the number of remaining missiles is used; the side with more remaining missiles wins. If the number of missiles is also the same, the advantage in air combat angle is used for scoring.

[0069] { e ( n ) = [ a l o o k t ( n ) − a l o o k ( n ) + 1 8 0 ] / 3 6 0 E = ( ∑ n = 1 N f r a m e e ( n ) ) / N f r a m e

[0070] in, This refers to the time step in a single-round simulation, which is 400 steps in this example. This represents the total number of time steps in a single batch simulation, which in this example is... step, and The first The line-of-sight angle of the aircraft relative to the enemy aircraft and the line-of-sight angle of the enemy aircraft relative to the aircraft. This represents "security." The higher the value (closer to 180), the less vulnerable the user is to enemy attacks, meaning the safer the user is. Represents "aggressiveness." The smaller the value (closer to 0), the more accurate the attack direction, indicating the need to position oneself in the enemy's blind spot (larger area of ​​attack). And more precise targeting of the enemy (small) ).

[0071] (2) Extract key situational information in complex and incomplete scenarios based on improved networks;

[0072] Based on the aforementioned complex and incomplete battlefield situation information To obtain a set of state information s i = [ s s e l f i , s a l l y i , s e n e m y i , s a t t a c k i ] ,in, For drones Partial observation information, such as location coordinates ,speed and direction of action , , This is part of the observation information from our own drones. , This is part of the observation information obtained from detecting enemy drones. This is an attack method used by enemy drones;

[0073] The state information set of all UAVs (with the corresponding value set to 0 if no enemy UAV is observed) is normalized to obtain standardized state information. This standardized state information is then input into the improved network to extract state information parameters in the UAV cooperative combat environment, thereby obtaining high-quality state information. State information parameters include, but are not limited to, the position, speed, and orientation of friendly UAVs, as well as the position, speed, and orientation of detected enemy UAVs, and information on detected enemy missiles.

[0074] The improved network includes an encoder and a decoder connected in sequence;

[0075] The encoder network includes a sequentially connected fully connected layer and several LSTM layers, which perform time-step feature extraction and information encoding on the state information.

[0076] The decoder includes a self-attention module, an LSTM layer, and two fully connected layers arranged in sequence; it decodes the encoded information, focusing more on key information in the state information sequence, such as the state and actions of the drone closest to us.

[0077] By training the state information parameters in a UAV cooperative combat environment using an improved network, high-quality state information can be obtained. For example, here, the self-attention mechanism module will prioritize the results on enemy position information and missile information, so the enemy information and enemy missile information have higher weights.

[0078] The dynamic multi-task weight algorithm framework designed in this invention is as follows: Figure 3 As shown.

[0079] based on Figure 3 (3) Design a strategy for multi-objective reward allocation and dynamic task weighting, and deploy the strategy to the control terminal of each UAV to execute the task;

[0080] The strategy for multi-objective reward allocation and dynamic task weighting is as follows: For time t, the drone The observation vector is input, and the task weights for the UAV's offensive and defensive missions are dynamically adjusted. For time t, the drone The normalized observation vector, For time t, the drone The temporal reward characteristics. The adjustment of dynamic task weights here is achieved through a dedicated dynamic task weight evaluation network; specifically, this network simultaneously inputs the current local observation status of the UAV and the reward statistics over a period of time, such as the mean, variance, and growth trend of the reward, to judge recent performance. After fusing this information through a gating mechanism, it automatically assigns greater weights to tasks that are more beneficial to the current situation or have higher recent gains, thereby allowing the agent to flexibly switch its focus between offense and defense.

[0081] Get Offensive and defensive reward sequences at each time step Capture mission trends and adjust the mission weights for offensive and defensive missions through a dynamic mission weight evaluation network. and In this example It takes 20 steps; here Represents a single drone at a certain moment For specific subtasks Such as the immediate reward value gained from offense or defense;

[0082] based on Step reward sequence, calculate the average reward for corresponding offensive and defensive tasks. ,variance and slope Let n be 1 or 2 corresponding to offense or defense, and obtain f r , i t = [ μ o b j 1 , i t , σ o b j 1 , i t , θ o b j 1 , i t , μ o b j 2 , i t , σ o b j 2 , i t , θ o b j 2 , i t ] ;

[0083]

[0084] in, Time step index within the time window Time Index The average value, This represents the average reward.

[0085] drones exist Attack reward is

[0086]

[0087] in, The reward magnitude control coefficient is 0.02 in this example. This is the distance attenuation coefficient, which is 0.01 in this example. For drones The Euclidean distance to the nearest enemy drone. For the enemy's nearest drone, A collection of enemy drones;

[0088] drones exist The reward for constant defense is

[0089] r o b j 2 , i t = γ ⋅ [ 1 − t a n h ( 1 N − 1 ∑ a j ϵ A \ { a i } ‖ p a i t − p a j t ‖ 2 ) ]

[0090] in, This represents the total number of friendly drones. This is the reward coefficient, which is 0.6 in this example, used to control the maximum value of the defensive reward. This is a normalization function used to control the average distance range between the current drone and friendly drones. To extract from drone set A, except All drones except those mentioned above;

[0091] In this invention, "keeping the formation compact" is defined as the core indicator of the defense mission. If the drones are too scattered, they are easy to be defeated one by one. Therefore, it is also necessary to maintain the formation during the defense process. By calculating the average distance with all friendly forces, the agents are encouraged to keep the formation compact.

[0092] Offense and defense rewards are redistributed, and the joint rewards are updated based on the mission weights of drone offense and defense tasks. Specifically, the offense and defense rewards are redistributed.

[0093]

[0094] in, The symbol represents the mission number, with 1 for attack and 2 for defense. and Represents the adjustment coefficient. The magnitude of this factor influences the variability in the distribution of drone rewards, which is 0.6 in this example. The strength of different reward modifiers is controlled to influence the strength of contributor rewards and penalties; in this example, it is 0.4. A swarm of drones that makes a positive contribution refers to a swarm of drones that, at time t, receive a positive reward through interaction with the environment. For the task The average reward for drones, This indicates that all contributing drones are in the mission. Total contribution For the task The total drone reward; after redistribution, calculate the combined reward after redistribution. ,in, and These are the task weights for drone offensive and defensive missions, respectively.

[0095] In the reallocation of attack and defense rewards, the immediate performance of the drones for the two sub-tasks at the current moment is first calculated. For the attack task, the reward is based on the Euclidean distance between the drone and the nearest enemy target; the closer the distance, the higher the reward. For the defense task, the reward is based on the average distance between the drone and all friendly drones, encouraging tight formations and moderate distances to achieve high scores (using the tanh function for normalization). Then, the original rewards of all drones are collected, and the reward for each agent is adjusted using a formula. Finally, the adjusted attack and defense rewards are weighted according to the weights generated by the dynamic task weight network. Perform a weighted summation to obtain the final reward value used to update the strategy;

[0096] Redistribution is based on differentiated competition and contribution percentage criteria; differentiated competition corresponds to... The second item calculates the average reward for all drones on this task. If a drone outperforms the average, it receives an additional positive incentive; otherwise, it is penalized negatively. This rejects egalitarianism and forces drones to strive to surpass their teammates by widening the reward gap between high-performing and low-performing drones. The contribution percentage corresponds to... The third step is to calculate the total reward pool for all drones that make positive contributions (rewards greater than 0), and to calculate the proportion of each drone's reward in this total reward pool. The higher the proportion, the larger the share of the team's total revenue will be allocated to rewarding the best-performing drones.

[0097] (4) Evaluate and optimize the strategy, and continue to execute the drone swarm cooperative combat until the mission is completed;

[0098] A multi-objective local Q-value network is constructed to evaluate and optimize strategies based on dynamically adjusted task weights for UAV offensive and defensive missions.

[0099] Specifically, a multi-target local Q-value network is constructed using multiple multilayer perceptrons and a gated recurrent unit, such as... Figure 3 As shown in the middle section, the output is a multi-objective Q-vector for each optional action.

[0100] Q i ( τ t i , ⋅ ) = [ Q o b j 1 i ( τ t i , ⋅ ) , Q o b j 2 i ( τ t i , ⋅ ) ]

[0101] By using the task weight coefficients obtained in (3) to weight the multi-objective Q-values ​​and employing a greedy strategy, we can obtain...

[0102]

[0103] The QMIX hybrid network is used to further fuse the Q-vectors of multiple objectives, which can output a joint action value vector. ,like Figure 3 As shown in the left part;

[0104] During the network update phase, a batch of data is sampled from the experience replay pool, and the evaluation network is used. With the target network To increase algorithm stability, the target network updates more slowly than the evaluation network and is periodically synchronized with it. Training data is sampled from the experience replay pool, and the target network is used to compute and optimize the target and loss. TD target. The calculation formula is as follows:

[0105]

[0106] in, For the task The reward value, in this example, is 5. Discount factor In this example, it is 0.99. For the output of the target hybrid network value;

[0107] Calculate the loss for each task separately, then sum them up to evaluate the network. The loss function is as follows:

[0108] L ( θ ) =  τ , u [ ∑ j = 1 n | y j t o t − Q j t o t ( τ , a , s ; θ ) | ]

[0109] Update Q-value fusion network parameters With target network parameters The network model is trained until the end of the round.

[0110] The QMIX algorithm employs a parallel Q-value fusion network. Traditional QMIX typically uses a single-channel structure, aiming to fit the local Q-values ​​of all agents into a single global scalar. This invention designs two channels, each corresponding to a different task objective. For example, channel 1 handles offense and channel 2 handles defense. Each channel has an independent MLP layer that processes the Q-value of its corresponding objective and fuses the data based on dynamically generated weights to ensure global accuracy. The optimization direction is consistent with the current tactical objectives. Since the multi-task QMIX algorithm is adopted, offensive and defensive tasks are considered and weighted. The weights are obtained through dynamic learning of the network and can be adjusted in a timely manner.

[0111] In this embodiment, experimental verification and comparative experiments are used, such as... Figure 4As shown, QMIX is the classic framework, MD is the multi-task allocation mechanism, ED is the encoding / decoding architecture, and RD is the reward allocation mechanism. Through ablation experiments, the encoding / decoding module and the reward allocation module are integrated into the multi-task QMIX algorithm. The experimental comparisons are the basic QMIX algorithm, the multi-task QMIX algorithm, the QMIX algorithm with the added encoding / decoding module, and the QMIX algorithm that integrates encoding / decoding and task allocation.

[0112] Through comparative verification, each module improves the method to a certain extent, demonstrating the superiority of the method of the present invention.

[0113] The present invention also relates to a computer-readable storage medium storing a multi-target-based UAV swarm cooperative combat program, which, when executed by a processor, implements the aforementioned multi-target-based UAV swarm cooperative combat method.

[0114] The present invention also relates to a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the aforementioned multi-target-based UAV swarm cooperative combat method.

[0115] This invention also relates to a multi-target-based unmanned aerial vehicle (UAV) swarm cooperative combat system, the system comprising:

[0116] The first and second drone swarms each include several drones. Each drone is equipped with a radar module, a positioning module, and an attack module. The radar module collects observation information, the positioning module obtains the positioning information of the corresponding drone, and the attack module is used for attack. A control terminal is provided in conjunction with the radar module, positioning module, and attack module.

[0117] On the PC side, in conjunction with the first and second UAV swarms, a system is set up including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the aforementioned multi-target UAV swarm cooperative combat method.

[0118] The PC-side obtains observation and positioning information based on the first and second UAV swarms, and the processor outputs policy action instructions to the corresponding UAV control terminal.

[0119] The present invention also relates to an application of the aforementioned multi-target-based UAV swarm cooperative combat method, which is applied to multi-UAV cooperative combat in complex multi-target environments.

[0120] The simulated combat environment can be constructed based on the Micius Joint Operations Simulation System. This system integrates an equipment parameter library, a tactical rule library, and deep reinforcement learning model components, enabling high-fidelity modeling of the kinematic characteristics and combat behavior of air combat units. This embodiment supplements the key parameters of the UAV platform based on the interfaces provided by the system to ensure the physical consistency and repeatability of the simulation process.

[0121] In this embodiment, both drones are assumed to be J-20 fighters, with a default cruising speed of approximately 648.2 km / h, which can be dynamically adjusted within the range of 0–750 km / h under system constraints. Speed ​​changes are linked to a fuel consumption model; the higher the flight speed, the greater the fuel consumption per unit time. The drone's flight altitude can be continuously adjusted within a preset altitude layer according to mission requirements, with altitude changes limited by platform performance and environmental constraints. All drones are equipped with an onboard radar system, with a maximum air-to-air detection range set at approximately 300 km, capable of acquiring the relative azimuth and coordinates of enemy drones or incoming missiles within radar coverage. Simultaneously, to simulate air-to-air combat capabilities, each drone carries ten PL-15 missiles, with a maximum effective range of approximately 185 km. The system simulates the missile's flight trajectory, interception process, and hit determination based on a missile dynamics model.

[0122] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0124] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0126] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0127] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for cooperative warfare among unmanned aerial vehicle (UAV) swarms based on multiple targets, characterized in that: The method constructs a beyond-line-of-sight cooperative adversarial scenario based on UAV swarms, obtains cooperative adversarial tasks, and defines adversarial evaluation indicators; it also extracts key situational information in complex and incomplete scenarios based on an improved network. Design a strategy for multi-objective reward allocation and dynamic task weighting. After the strategy is deployed to the control terminal of each UAV, the task is executed, the strategy is evaluated and optimized, and the UAV swarm cooperative combat is continuously executed until the task is completed.

2. The method for cooperative warfare between multiple UAV swarms based on claim 1, characterized in that: The cooperative adversarial task is a Markov game tuple of drone swarm adversarial combat. : , in, The number of your own drones. For battlefield situation information, For the actions of each drone, This is the state transition function. As a discount factor, The reward function for each drone.

3. The method for cooperative combat of unmanned aerial vehicle swarms based on multiple targets as described in claim 2, characterized in that: Based on the aforementioned complex and incomplete battlefield situation information To obtain a set of state information ,in, For drones Partial observation information, This is part of the observation information from our own drones. This is part of the observation information obtained from detecting enemy drones. This is an attack method used by enemy drones; The state information set of all UAVs is normalized and input into the improved network to extract state information parameters in the UAV cooperative adversarial environment, thereby obtaining high-quality state information.

4. The method for cooperative combat of unmanned aerial vehicle swarms based on multiple targets as described in claim 3, characterized in that: The improved network includes an encoder and a decoder connected in sequence; The encoder network includes a sequentially connected fully connected layer and several LSTM layers, which perform time-step feature extraction and information encoding on the state information. The decoder includes a self-attention module, an LSTM layer, and two fully connected layers arranged in sequence.

5. The method for cooperative warfare between multiple UAV swarms based on claim 1, characterized in that: The strategy for multi-objective reward allocation and dynamic task weighting is as follows: For time t, the drone The observation vector is input, and the task weights for the UAV's offensive and defensive missions are dynamically adjusted. For time t, the drone The normalized observation vector, For time t, the drone The time-series reward characteristics.

6. The method for cooperative warfare between multiple UAV swarms based on claim 5, characterized in that: Get Offensive and defensive reward sequences at each time step The network adjusts the task weights for offensive and defensive tasks through dynamic task weight evaluation. and ; based on Step reward sequence, calculate the average reward for corresponding offensive and defensive tasks. ,variance and slope Let n be 1 or 2 corresponding to offense or defense, and obtain ; , in, Time step index within the time window Time Index The average value, This represents the average reward.

7. A method for cooperative warfare between multiple UAV swarms based on claim 6, characterized in that: drones exist Attack reward is , in, This is the reward range control coefficient. This is the distance attenuation coefficient. For drones The Euclidean distance to the nearest enemy drone. For the enemy's nearest drone, A collection of enemy drones; drones exist The reward for constant defense is , in, This represents the total number of friendly drones. As the reward coefficient, For normalization function, To extract from drone set A, except All drones except those mentioned above; Offensive and defensive rewards are redistributed, and joint rewards are updated based on the mission weights of drone offensive and defensive missions.

8. A method for cooperative warfare between multiple UAV swarms based on claim 5, characterized in that: A multi-objective local Q-value network is constructed to evaluate and optimize strategies based on dynamically adjusted task weights for UAV offensive and defensive missions.

9. A multi-target-based unmanned aerial vehicle (UAV) swarm cooperative combat system, characterized in that: The system includes: The first and second drone swarms each include several drones. Each drone is equipped with a radar module, a positioning module, and an attack module. The radar module collects observation information, the positioning module obtains the positioning information of the corresponding drone, and the attack module is used for attack. A control terminal is provided in conjunction with the radar module, positioning module, and attack module. On the PC side, in conjunction with the first and second UAV swarms, a computer program is set up, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the multi-target-based UAV swarm cooperative combat method as described in any one of claims 1 to 8. The PC-side obtains observation and positioning information based on the first and second UAV swarms, and the processor outputs policy action instructions to the corresponding UAV control terminal.

10. An application of the multi-target-based UAV swarm cooperative combat method according to any one of claims 1 to 8, characterized in that: It is applied to multi-UAV cooperative combat in complex environments with multiple targets.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle formation cluster control method based on multi-agent deep reinforcement learning

    CN115755949A

  • Multi-agent-based unmanned aerial vehicle formation defense penetration method and application thereof

    CN119882773A

  • Multi-task-oriented unmanned aerial vehicle group convening method

    CN120456119A